Unstructured Raises $25M to Revolutionize Data Preparation for LLMs with CEO Brian Raymond

14 Mar 2024 · 31 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today Podcast Episode Notes

Episode Title

Unstructured Raises $25M to Revolutionize Data Preparation for LLMs with CEO Brian Raymond

Episode Summary

In this episode, host Brian Raymond discusses Unstructured's recent accomplishment of securing $25 million in funding aimed at transforming data preparation for Large Language Models (LLMs). The conversation covers the company's vision, the challenges they are addressing, and their innovative approach to making natural language data more accessible.

---

Key Participants

  • Brian Raymond: CEO of Unstructured
  • Podcast Host: Interviewer (name not specified)

---

Key Discussions

Overview of Unstructured

  • Mission: To make data preparation "cheap, fast, and easy" for natural language files (e.g., PDFs, HTML).
  • Goal: Convert these files into a machine-readable format (JSON) to ease the data science process.

Founding Experience

  • Brian's background includes roles in the U.S. intelligence community, including the CIA and White House during the Obama administration.
  • The company was created to solve a specific problem in the data processing landscape, rather than simply building on existing technologies.

Technical Solutions

  • Data Processing: Focus on automation over manual processes (like regex and OCR) to improve efficiency and reduce costs.
  • File Types Supported: Over 25 different formats (e.g., text files, PDFs).
  • Integration: Capable of linking with various platforms (e.g., SharePoint, Google Drive) to automate data ingestion.

Market Need

  • The data preparation process typically consumes 80-90% of data scientists' time.
  • Addressing this challenge can significantly improve the success rate of machine learning initiatives, which historically see a failure rate of 80-90%.

Funding and Growth

  • Recent funding will be directed towards productionizing their platform for enterprise use, focusing on secure and compliant solutions.
  • Plans to expand the engineering team to meet growing demand and enhance service capabilities.

Future Developments

  • Upcoming features will include support for various data types (e.g., audio, video) and improved processing speeds.
  • Aim to provide a robust platform for organizations, allowing for continuous data integration and workflow automation.

Ethical Considerations

  • Discussion on the evolving relationship between tech firms like Silicon Valley and defense agencies.
  • Emphasis on responsible AI usage, compliance, and addressing biases in AI systems.

Investor Relations

  • Insight into the investor concerns regarding market viability and the transition from open-source to commercial products.
  • The importance of having a clear and structured business plan to gain investor confidence.

---

Key Takeaways

  • Focus on Data Preparation: Unstructured aims to streamline the data preprocessing phase, enabling data scientists to focus more on modeling.
  • Community Engagement: A strong GitHub presence and an open-source model have helped establish user trust and engagement.
  • Growth Potential: With the recent funding, Unstructured is positioned to expand its offerings and respond to the needs of various industries, particularly in LLM applications.

---

Resources Mentioned

  • [AI Box Investment](https://republic.com/ai-box)
  • [AI Box Waitlist](https://AIBox.ai/)
  • [AI Facebook Community](https://www.facebook.com/groups/739308654562189)
  • [AI in Music](https://musicalai.pro/)
  • [AI Models Info](https://aimodelspro.com/)

---

Closing Notes

  • Brian invites listeners interested in Unstructured to reach out via GitHub or directly through email for inquiries about the platform.
  • The episode concludes with acknowledgment of the importance of adapting AI technologies to real-world applications and fostering a collaborative tech ecosystem.

---

Disclaimer

This summary includes insights and key points from the podcast episode and is intended for informational purposes only.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome to the AI Chat Podcast. businesses. His professional journey is not limited, just in private sector. He also served in the U.S. intelligence community. He worked in the White House during the Obama administration, and he had a run at the CIA. So leveraging his unique experiences, Brian is leading the charge in making AI technology more accessible and efficient by addressing critical data processes in this entire AI field. So Brian, thank you so much for coming on the show today. Thanks for having me. So, you know, it's super exciting, I'm imagining, for you to be working on this company and super excited and congratulations on this new round of funding that you were able to recently raise.

1:10Can you give us a brief overview of Unstructured and kind of tell us a little bit about what led you to co-found this company? you? Yeah, absolutely. So at Unstructured, we're just trying to make it cheap, fast, and easy to take any files containing natural language data and to render them in a format that's ready for an embedding endpoint, a vector database, or a machine learning pipeline. And so what that means in practice is, and we'll get into this later on, is taking things like PDFs and PowerPoints and XML and HTML and then normalizing that into JSON that you can clean and curate really easily, and then focus on the data science, not the data preprocessing.

1:48What really drove us to launch Unstructured was practical experience. So Prime, we're a fantastic company, really at the leading edge over the last several years in the natural language processing space. There we were fine-tuning, you know, BERT-based models, transformer models more broadly, orchestrating them into pipelines, building apps on top of them and deploying against customer data. and we didn't need another data labeling platform, a model serving or model monitoring platform, really good solutions available for those. It really, everything slowed down in terms of the time to value and the cost to value on taking customer data, the natural language data that was important to them and being able to transform that into a machine readable format.

2:35Okay, very cool. So I imagine, right, you're working at, prior to this primary AI, And this is, like you said, a really kind of bleeding edge AI company. I imagine that was really exciting and awesome. What was that like, you know, thinking about taking that next step and starting your own company? What made you like decide you could or should make that step? Well, generally, this is an oversimplification, but you have companies that emerge from technology. You worked at some cool, developing some new cool new techniques or methods at Uber or Meta or wherever. and now you want to take that and build a company around that.

3:15We were kind of the opposite. We were a problem in search of technology. And so when I was there at Primer, we looked around. We would talk to almost all of the data integration players in the space, and none of them were really focused on this problem space. We looked at some of the intelligent document processing providers that have fabulous solutions, but those weren't really tailored for this. and then um talk to others and interested how are all of you doing this and everyone said we're doing it the same way we're doing it through completely manually and and specific to every customer we're just doing regexes and python scripts and ocr and it's slow and it's ugly and it ruins our our margins on these contracts because we have huge professional services associated with it and so at that point i was like this is this is an area that urgently needs technology, right?

4:05This problem. And so that was really the light bulb moment to say, hey, this is a big enough problem that we can go build a company around it. And that's super cool. So I know you have some co-founders with you on the company. When you made that move, was this something where you guys were all sitting around one day talking about what kind of solutions should we have? Or was this one of your specific ideas and you brought the other players in? How did that conversation go, building the team there? Sure. Yeah. So, so I left Primer and been working on this for a few months and had been doing lots of user interviews and others and sort of thinking about, okay, what sort of the founding team need to look like?

4:43And I needed someone who's fabulous on infrastructure and architecture and another one who's fabulous on, on data science. And I thought to myself, okay, who knows these problems better than anyone? And Matt Robinson, who I think the world of came on to lead our data science. He worked at CIA with me. He worked at Capital One, Primer, another company called Rebellion Defense, had really deep and intimate knowledge of this problem from a data science side. And then Craig Wolf, who leads our infrastructure, he deep expertise from Red Hat and then five years of Primer on building enterprise solutions around natural language processing.

5:24And so, one, I just love working with them, but two, world-class expertise for this particular problem in particular. Yeah, that sounds amazing. And I know the team makes a huge difference. So given a lot of people say that data processing and kind of the prep time on that takes about 80 % of data science time, how over-end structured are you really focusing on optimizing that process? Yeah, absolutely. I think credit the Alex Ratner and the folks at Snorkel have focused a lot on data-centric AI. I think that kind of whiplash back around to model-centric AI over the last eight months, and now it's coming back around to refocusing on data-centric AI.

6:08and um and and you know they're primarily focused on um on annotated data and being able to have really high quality training data for the models um what we're focused on is is really to the left of that and it's so it's um suppose you want to utilize llama 2 um and put it on top of your internal company data so that you can accelerate workflows how do you get that into Pinecone or Weaviate or Chroma or these vector databases? The way that we're doing it is we're consolidating that down to a single API. So you can have, say you have 100 ,000 files in an S3 bucket, you can just point it at our API.

6:51We detect what type of file it is, route it to the appropriate strategy, and render into JSON and rejoin it with all of the other data. And so really collapsing a lot of the complexity around that. And so get data scientists out of the business of deep data engineering and give that time back to them to focus on modeling. Okay, very cool. Could you maybe give us some examples of the types of documents that you, when you talk about this, you know, that unstructured, I think their file transformation NLP model, what kind of documents is it really capable of handling? standard do you focus on? Yeah, no, great question.

7:31So right now we support more than 25 different file types. And as files come in, we detect the file extension and then route it to the appropriate strategy. And so let me unpack that. For, say, a text file, really easy to get access to that natural language data. And so the task there is extracting it from that text file, but then we also clean and curate it so we make sure that if they're as we extracting it make sure there's no weird white spaces or unicode characters or something on there but then also in the curation side we're annotating that each chunk of text with the what we call a document element that is is tied to so if it's a title we tag it as a title we're able to automatically detect that using nlp models and more body text lists etc and the idea there is that you don't want to send everything into a vector database and you may not want to send everything into a model pipeline.

8:28That'd be thoughtful and data centric about what you put in front of that model. And so for a text file, it's relatively simple. We have blended approaches for other file types all the way up to models that we'll be talking more about in September that we've pre-trained that treat every file as an image. And then we have around 20 different categories of document elements and we're able to extract that data with an extremely high level of precision without any any particular training so you don't need to like provide input on the document layout um we retrain this model on millions and millions of documents and and dozens of languages and it just works it's just that's amazing oh my god very cool very cool we'll look forward to uh hearing more about that in september when we make that announcement i'm sure um when you kind of talk Talk about interacting with other things.

9:22How does Unstructured interact with things like customer relationship management software and other data sources? Totally. So we're doing two things there. So on the left side of what we're building, we're building upstream data connectors. And we're inspired by Fivetran. These are maintained by us. So we're continuously testing them. But we have our own abstraction for them. And so that they're resilient to interruptions, that they're easy to parallelize, and they're very sensitive to versioning. And so you can grab net new data very easily. And so we have that for like SharePoint and Azure Blob and Google Drive and Notion.

9:59And I think we have, as of yesterday, 18. And this will grow to about 40 into the fall. Very cool. And then on the right-hand side, we have a bunch of what we're just calling staging utilities. So if you want to chunk it according to a particular attention window size, tokenize, vectorize, map the JSON schema to a specific vendor downstream, we have a lot of these utilities out of the box. And so you don't need to do any additional data munging before you go engage with whoever's to the right of us. Okay. Very cool. Very cool. How do you integrate? Great. So, you know, again, talking about integrations, how do you kind of handle the whole and how does Unstructured kind of handle integrations with people like LangChain and Vector Databases and, you know, MongoDB, Atlassian Vector Search, all that kind of stuff?

10:47Yeah, no, great, great question. Probably the best way to describe it is native. And so with LangChain and Llama Index on the orchestration side, with MongoDB Atlas, and with WeV8 and others, extremely simple plugin to the right of us, and you can leave with data that's immediately ingestible. We were actually there and sponsored Harrison and Langchain's kickoff party back in February. We were the first company to tie in with them. Oh, that's awesome. WeV8 and Mongo are both invested in us. And so there's a really special ecosystem of open source solutions that's emerged to help power this new LLM stack.

11:36and we're doing everything we can to make it effortless for our users to plug into that ecosystem and realize the value of those those solutions that sit downstream with us okay very cool that's awesome so correct me if i'm wrong but you transformed from offering an open source suite of data processing tools to launch a commercial api is that correct yeah and the api is actually free right now. And so in terms of like the way that we're delivering value, we start off with the Python library and we heard user feedback that's a pain in the ass to install because we have a ton of dependencies to handle the long tail of all these different file types.

12:17And so we rushed and we got some containers out and we've continued to add new containers for different hardware types. And a lot of the folks that are building on top of LMs and front-end developers said, hey, we want something even easier. And so we were excited to announce last week the rollout of our actually free API right now. And so you can go to our website and grab an API key and start hammering that if you want. What we'll be introducing over the coming weeks are tiered APIs so you can have access to GPUs and dedicated instances. And so we can bring down the latency there. And then later on in the fall and in the winter, we'll be talking more about our enterprise platform that will have everything that you'll need in order to move these LM solutions into production.

13:04So we imagine a world where you're continuously grabbing data, say, every minute or 15 seconds from Slack and from Google Drive and from email, et cetera, and wanting to move that into a vector database. And so it can accelerate and power workflows across the enterprise. And we want to be the backbone of that, making that a reality. Okay, very cool. So given your size, and I'm not 100 % sure, roughly how many users or companies do you have currently interfacing with your product, your technology? That's a great question and a really hard one to answer given that we're open source. Right. Some of the things that we're looking at in our Slack, we have like almost 600 people now and at least 100 companies represented there.

13:53We're about 2 ,800 GitHub repos use us. Okay. Those are open source repos, and we're quickly closing in on around a million downloads in terms of pip installs. And so moving very quickly kind of across the board there. That's very cool. So I guess given this, you know, a million installs and all these different people, different places, different corporations using you, how do you measure the success and impact of your product on businesses? Yeah, I think what matters at the end of the day is that the individual data scientist at Company X that's prototyping something cool with Langchain and with Weviate, that they're successful and that they're able to move that into production.

14:37That's what we care about. And it may be a different architecture that they're using, but we are laser focused on enterprise adoption of LMs. And that's going to come by data scientists inside those organizations being successful. look, historically about 80%, 80 to 90 % of machine learning initiatives within businesses fail. Right. And so it's like, how are we going to change the numbers on that? One way is to change the economics and the time required to get the data that's important to you and the knowledge that's important to the organization in front of the model. Right. And that's for us, like where we're really focused.

15:16But at the end of the day, comes down to us as an ecosystem player of enabling technologies, helping these businesses realize the power, the productivity gains, and the performance gains that LLMs can drive within their businesses. That's super cool. That's awesome. So of course, you talked a little bit about starting out as open source and you have a commercial API. Talk to the audience a little bit and I guess explain a little bit what your monetization mechanism and what that overall strategy looks like for you. Because, you know, there's different sides of it. There's the open source side and all of that.

15:52And so, yeah, I think a lot of people are curious about that. So right now the objective of open source is to remove any barriers to entry for folks that want to prototype. So we want a fabulous solution for file transformation that provides a good foundation for successful prototypes. um our you know vision here hypothesis is that as these prototypes are successful you're going to want to continuously grab that data and then push it into vector databases and what you'll want there fine-grained user permissioning you want scheduling you want a ui you want sock 2 compliance you want hipaa compliance uh you'll need a whole bunch of functionality you want premium supported connectors that we support over the long term and we fix them if they break at 2 a.m and we envision wrapping around that core file transformation technology, all of that additional functionality so that this can be running in serial perpetually into this living architecture and LM-enabled architecture that services HR and sales and R &D and every different aspect of an organization.

17:03And so that will have a license component and a usage-based component, But hopefully they'll like very similar to most of the other data integration providers. But we'll have that open source entry point that you can prototype around. Okay, very cool. So, you know, as I mentioned earlier when I was kind of giving your intro or whatever, you have had some experience working with the government, working with the CIA. I wonder, like, what influence, if any, do you feel like that your background there in the government would have into kind of your focus on what you're doing today with your current company?

17:39So this really, I mean, it's really elemental to everything that we're doing. And so when I came into the CIA, I spent most of my time with the director of analysis. It was during kind of the move to cloud and the move in big data. And big data was the mantra at that time, but you didn't have analytics to match the data. Palantir was doing some really cool stuff, mostly around structured data. NLP at the time was still pretty sleepy. So you had enormous volumes of data being collected and being produced, but you didn't have any real mechanism in order to exploit that. At Primer, we did some really interesting work on how to put analytic engines against it.

18:25But again, because the ecosystem was still so immature, it was difficult to expand that to lots of different use cases and different user personas within an enterprise. I think now with the advent of LLMs, it's changed the fundamental economics and the time to value where you don't need 30 different fine-tuned models on the ML side. If we're right and we help solve a lot of the problems on the preprocessing side, then you're going to be able to realize a lot more value of this big data that's been sitting around for years. Okay. That's very cool. So I think this is something I've just noticed in tech recently.

19:06There is kind of this move in the past away from, you know, tech companies like Google, for example, wanting to work with the Department of Defense or any kind of military thing. There was a lot of employees at a lot of companies that were kind of had an aversion to working with anything in defense or technology. And I feel like we've kind of seen this trend where the pendulum may be swinging back. There's a bunch of new, you know, interesting startups. There's Palantir, of course. There's Palmer Luckey that has his startup really focusing on defense and whatnot. And I feel like I kind of see a little bit more innovation coming into this.

19:40How do you guys view that and where do you guys stand? Because I know you have some partnerships with, I believe, the U.S. Special Operations Command and defense agencies and whatnot. So I guess what's the nature of your partnerships there and how do you kind of view that? Yeah, I've looked at the last 10 years as kind of the left coast figuring out how to work with the right coast. And there is some growing pains there. And I think that that's maturing in a really productive way and an exciting way right now. You saw CIA on stage at AWS reInvent last year, which was bizarre, right? Thinking about where we were a few years back.

20:18And a huge emphasis on responsible AI, on taking bias seriously, and on taking automation really seriously from an ethical consideration. But also a recognition that we in Silicon Valley have a responsibility as well. And that there's a lot of good to be done through our technology as well. And so it's been pretty incredible to see Silicon Valley more broadly rally behind Ukraine and then also probably around Europe over the last year and a half. I'll say this. Silicon Valley has made it, I think it's definitely could come back around and that it's a lot easier. I think there's a lot more space now for conversation around this on how to do it.

21:10from, in terms of like the Pentagon and CIA and others, they've made it a lot easier to do business with them too. And I've put a lot of structures in place that have provided a lot more confidence, right, to folks on the left coast. And so I really look at it in terms of a maturing relationship that's moving in a really positive direction. Okay, very cool. Yeah, that's very cool. In what ways would you say this, you just raised$25 million in funding, How would you say this round of funding is going to impact Unstructured's business and growth plans? And what do you plan on kind of doing with that?

21:46Yeah, absolutely. It's going towards productionizing the platform for enterprises. There is, I think that, you know, as an ecosystem of LLM-enabling technologies, Unstructured and others have gotten some great tooling out there to prototype. However, if you're at a large financial services organization, at a Walmart, at General Motors, these are really cool for demos. These are not production ready. And so myself and my peers are racing ahead as quickly as we can in order to mature the solutions that we're working on into actually production-grade platforms that they can deploy with confidence.

22:36there's technical challenges associated with this in terms of like how does all this work and you know how do you handle drift and in-context learning and rag-based approaches and all this where it's changing every week but then there's things that just don't change on the cyber security side on the compliance side on um on you know basic functionality here that needs to be established in order to uh to credibly go into these organizations and and meet their their needs. And so, you know, from where we sit, we more than doubled the size of our engineering team in the last six weeks. We'll continue to make investments there in order to build as fast as we can to support those folks that I mentioned earlier in this interview on those individual data scientists that are doing really, really cool prototypes within their organization that are trying to find a pathway to production.

23:27Okay. Very cool. So I know, you know, in the process of raising a round of funding, you obviously have to go through some pretty rigorous documentation and background checks and all that kind of stuff. Can you talk a little bit about maybe what were some of the concerns your investors had going into this and what ultimately helped you to overcome those concerns for them? Speaking of unstructured in the business. Yes. Yeah, totally. I think a few different things. One is uncertainty on what this emerging LM tech stack is going to look like. Everyone's asking questions on whether or not the service area that they're touching is sufficiently broad to capture enough value, but also not so broad that you never go deep enough to to actually generate a huge amount of value.

24:20so that that's one just because this is moving so quickly two um this is you know an old problem but everyone um that's in my fuse has to address it like how do you turn the corner from open source to a commercial to a commercial product um and then and and then three i think just um from a technical standpoint you got a great idea um how are you actually going to build it how are you actually going to deliver it and what's the thing that's going to be required and i think that I got a lot of really great questions around that. And beyond that, just normal stuff. But this is a unique moment that we're in right now where the world's changing so, so quickly.

25:03And so to earn the trust of investors to make a bet on you, you got to have a really tight plan and a lot of conviction in terms of those different pillars that I talked about. Okay, very cool. How would you say your relationship with your board members? I know you have Michael Grone and Mike Brown on there. How's that kind of shaped the trajectory of Unstructured or kind of impacted that? Yeah, absolutely. So we have some fantastic advisory board. There's some General Grown, and then Mike Brown, our fantastic individuals to have in our corner to help navigate the Department of Defense and the public sector more broadly and to advise on where we're making AI investments.

25:48uh kharan maandru who is our board member um and he led the investment in our our a-roud from madrona and then enrique selim from uh bain capital ventures who's a second board member and he led our seed and also invested in our a um world-class individuals um know how to build great companies um fantastic um vcs but i think most importantly are um our builder advocates um And I think we have deep alignment on strategy. And then also just the way that they tend to engage is so practical and so helpful that we're able to, you know, like I have both of them in Slack, for example. Okay. We're constantly bouncing stuff off each other.

26:42And so really intimate in a positive way. That's awesome. That's amazing. so talk to me about when you first kind of got this thing uh started you first started working on this what what did that look like at the very beginning the early stages for you as far as kind of funding it and getting this thing off the ground was this something that you began self funding were you immediately looking for angles or you know seed round pre-seed round what did that process look like for you yeah it was a it was a it was a that's called a non-linear process Yes. So the first thing I started to do was actually doing user interviews.

27:18I actually called and reached out and talked to, I still have the spreadsheet, more than 70 data scientists and talked to them about how they were grappling with this. And I had dozens of pages of notes in terms of validating the problem and talking about what type of solution that they'd want. And so really kind of put my product hat on for that. I did submit to for Y Combinator and was admitted to that. But at that point, I had a decent amount of institutional capital lined up for a pre-seed. And, you know, as we were about as I was about to button up the pre-seed, had an opportunity to do a larger seed round with Enrique and the team there at Bain.

28:01And so I pulled back on the pre-seed and decided to raise a larger amount from the get-go so that we'd have more runway in order to actually figure out how the market wanted to consume this and then also grow the engineering team more rapidly. Okay, yeah, that makes sense. And what was your process? What did that look like as far as reaching out and meeting these people at Bain and the institutional investor? Was that like previous connections? You were going for cold emails? What did that kind of look like for you? It was too funny. I was previous connections and cold emails. I got a, uh, the actually Enrique got ahold of me.

28:37I was at a movie with my wife and I got a text message. Hey, this is Enrique from Bain. Can we chat? And, um, and I go out of the movie column and he's like, Hey, let's make something happen. And, um, honestly, um, you know, this is, um, this is probably the single most helpful thing that, uh, if someone's in, you know, considering launching something, um, don't just change on your LinkedIn to working on something. new and still startup and then everyone that's scraping LinkedIn will automatically reach out to you. I should have done that earlier instead of hustling so hard with the pitch deck around the contacts and contacts and contacts.

29:13But lesson learned at least. That's funny. So that's what happened? You changed it there and that's how we contacted you? It was everybody else. It was too funny. About a week before we closed, I changed it. And then within 24 hours, I started getting a flood of inbounds. And I was like, they must be scraping LinkedIn because now they can see they'll start. Never went about it beforehand. Yeah. That's so funny. That's awesome. So what future developments or I guess expansions can users expect and potential investors and other people look forward to from unstructured going into the future? Yeah, I would say right now we're focused on building the equivalent of a Toyota Corolla.

Read the full transcript

29:58Okay. You just want it to be cheap to operate. You can't break it. It's a good daily commuter, and it meets all of your needs. From there, however, they can expect multimodal, so we're going to be able to ingest audio and video. You're going to see huge speed gains, algorithmic speed gains, as we continue to invest in the underlying architectures of the models and pre-training our own models from scratch. and then also much broader upstream and downstream integrations. And so at the end of the day, it's going to be easier, it's going to be faster, and it's going to be more performable. Very cool.

30:40Very exciting. So first off, it's been amazing having you on the show today, amazing picking your brain and learning all of your insights about what's going on in the industry. If companies are interested in using Unstructured, where is the best place for them to find out more about it, to get, you know, access it to kind of start learning if that's a good fit for their company. Absolutely. Three options. So one, you can go to GitHub, get everything off of GitHub. Two, off our website, you can go get your API key. It'll be automatically generated and sent to you. Or three, just feel free to email me, brian.unstructure.io.

31:14I promise I'll respond. That's amazing. Well, thank you so much for coming on today, Brian. It's been amazing to talk to you. For the listeners, thank you so much for listening to the AI Chat Podcast. make sure to rate us wherever you get your podcasts and we will see you next time

From the publisher

In this episode, we explore the groundbreaking efforts of Unstructured as they secure $25M in funding to revolutionize data preparation for Large Language Models (LLMs), featuring insights from CEO Brian Raymond on the company's vision and impact.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Unstructured Raises $25M to Revolutionize Data Preparation for LLMs with CEO Brian RaymondAI Today · 31 min
Listen in VO