Towards high-quality (maybe synthetic) datasets

9 Oct 2024 · 57 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast Episode Summary

Episode Title

Towards High-Quality (Maybe Synthetic) Datasets

Episode Description In this episode, hosts David Berenstein and Ben Burtenshaw from Argilla discuss the importance of data quality in AI and how collaboration between AI teams and domain experts can enhance this quality. The conversation delves into synthetic data generation and AI-generated labeling and feedback.

---

Key Concepts

Importance of Data Quality

  • Data Quality: Fundamental for successful AI implementation, as noted by Argilla, "Data quality is what makes or breaks AI."
  • Collaboration: Emphasizes the need for collaboration between domain experts (who provide context and knowledge) and AI engineers (who understand technical data requirements).

Data Collaboration

  • Definition: Collaboration between domain experts and AI engineers is crucial for creating effective datasets.
  • Domain Experts: Have extensive knowledge about the particular field, guiding the labeling and curation process.
  • AI Engineers: Provide technical expertise on data management, models, and outputs.

AI Feedback and Synthetic Data

  • AI Feedback: Utilizing AI models to generate feedback or labels for datasets, enhancing the quality and breadth of data.
  • Synthetic Data Generation: The process of creating artificial datasets that mimic real-world data to train models without compromising privacy.

---

Workflow for New AI Teams

  1. Understanding Domain Data:
  2. Identify and describe the specific problem domain and the associated data.
  3. Collaborate with domain experts to model the problem simply, deriving questions that AI systems should answer based on the data.
  1. Baseline Establishment:
  2. Create a benchmark for evaluating how well AI models can handle tasks with the provided domain data.
  3. Ensure proper indexing and chunking of documents for effective retrieval.
  1. Iterative Refinement:
  2. Use feedback loops to refine the dataset based on initial model outputs and expert feedback.
  3. Introduce retrieval techniques and incorporate AI tools to assess the relevance of documents.
  1. Leveraging AI Models:
  2. Use smaller, domain-specific models for better performance and less resource dependency.
  3. Utilize generative models for expanding datasets while maintaining quality through iterative evaluation.

Practical Steps for Implementation

  • Develop a List of Questions: Gather relevant questions from domain experts to align dataset construction.
  • Document Review: Associate documents with questions and verify their applicability in real-world scenarios.
  • Iterative Testing: Test how well AI models can answer these questions using the generated or collected documents to gauge effectiveness.

---

Tools and Technologies Discussed

  • Argilla: A data annotation and collaboration tool designed to streamline the feedback process between domain experts and AI engineers.
  • Distilabel: A framework for synthetic data generation and AI feedback, facilitating the creation and evaluation of datasets for training AI models.

Use Cases

  • Example of a Healthcare Company: Utilized classification and generation pipelines to manage customer query responses effectively.
  • OpenHermes and Intel Orca Datasets: Demonstrated the effectiveness of using Distilabel for cleaning and optimizing datasets based on AI feedback.

---

Insights and Future Directions

  • Multimodal Data: The future development of tools like Argilla and Distilabel will focus on handling various data types (text, image, audio, video).
  • Real-Time Feedback Loops: Enhancing the integration of feedback from domain experts into the dataset generation and model training process.

Conclusion The episode underscores the significance of collaboration between domain knowledge and AI technology to improve data quality and the efficacy of AI models. Synthetic data generation and AI feedback offer innovative paths to enhancing datasets, paving the way for more robust AI applications across various industries.

---

Sponsors

  • Fly.io
  • WorkOS
  • Eight Sleep

Visit their respective websites for more information on how they can assist in deploying and optimizing applications.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.

0:44okay friends i'm here with annie sexton over at fly and you know we use fly here at changelow we love fly it is such an awesome platform and we love building on it but for those who don't know much about fly what's special about building on fly fly gives you a lot of flexibility like a lot lot of flexibility on multiple fronts. And on top of that, you get, so I've talked a lot about the networking and that's obviously one thing, but there's various data stores that we partner with that are really easy to use. Actually, one of my favorite partners is Tigris. I can't say enough good things about them when it comes to object storage.

1:24I've never in my life thought I would have so many opinions about object storage, but I do now. Tigris is a partner of Fly and it's S3 compatible object storage that basically seems like it's a CDN, but is not. It's basically object storage that's globally distributed without needing to actually set up a CDN at all. It's like automatically distributed around the world. And it's also incredibly easy to use and set up, like creating a bucket is literally one command. So it's partners like that, that I think are this sort of extra icing on top of Fly that really makes it sort of the platform that has everything that you need.

2:00So we use Tigris here at Changelog. Are they built on top of Fly? Is this one of those examples of being able to build on Fly? Yeah. So Tigris is built on top of Fly's infrastructure, and that's what allows it to be globally distributed. I do have a video on this, but basically the way it works is whenever, like, let's say a user uploads an asset to a particular bucket. Well, that gets uploaded directly to the region closest to the user. Whereas with a CDN, there's sort of like a centralized place where assets need to get copied to. And then eventually they get sort of trickled out to all of the different global locations.

2:33Whereas with Tigris, the moment you upload something, it's available in that region instantly. And then it's eventually cached in all the other regions as well as it's requested. In fact, with Tigris, you don't even have to select which regions things are stored in. You just get these regions for free. And then on top of that, it is so much easier to work with. I feel like the way they manage permissions, the way they handle bucket creation, making things public or private is just so much simpler than other solutions. And the good news is that you don't actually need to change your code if you're already using S3.

3:07It's S3 compatible. So like whatever SDK you're using is probably just fine. And all you got to do is update the credentials. So it's super easy. Very cool. Thanks, Annie. So Fly has everything you need. Over 3 million applications, including ours here at Changelog, multiple applications have launched on Fly. Boosted by global anycast load balancing, zero configuration private networking, hardware isolation, instant wire guard VPN connections, push button deployments that scale to thousands of instances. It's all there for you right now. Deploy your app in five minutes. Go to fly.io. Again, fly.io.

3:55Welcome to another episode of the Practical AI Podcast. This is Daniel Whitenack. I am CEO at Prediction Guard, where we're building a private Securigen AI platform. And I'm joined, as always, by Chris Benson, who is a Principal AI Research Engineer at Lockheed Martin. How are you doing, Chris? Great today, Daniel. Daniel, how are you? It's a beautiful, beautiful fall day and a good day to take a walk around the block and think about interesting AI things and clear your mind before getting back into some data collaboration, which is what we're going to talk about today. Chris, I don't know if you remember our conversation.

4:37It was just me on that one, but with Bing Soon Chua, who talked about broccoli AI, the type of AI that's healthy for organizations. And in that episode, he made a call out to Argeala, which was a big part of his solution that he was developing in a particular vertical. I'm really happy today that we have with us Ben Bertenshaw, who is a machine learning engineer at Argeala, and also David Berenstein, who is a developer advocate engineer working on building Argeela and Distilled Label at Hugging Face. Welcome, David and Ben. Thank you. Great to be here. Hi. Thanks for having us. Yeah, so like I was saying, I think for some time, maybe if you're coming from a data science perspective, there's been tooling maybe around data that manages training data sets or evaluation sets or maybe MLOps tooling and this sort of thing.

5:40And part of that has to do with preparation and curation of data sets. But I found interesting, I mentioned the previous conversation with Bing Soon, he talked a lot about collaborating with his sort of subject matter experts in his company around the data sets he was creating for text classification. And that's where Arjila came up. So I'm wondering if maybe one of you could talk a little bit at a at a higher level when when you're talking about data collaboration in the context of the current kind of ai environment what what does that mean generally and how would you maybe distinguish that from previous generations of of tooling and maybe similar or different ways so data collaboration at least from from our point of view is kind of the collaboration between both the domain level experts that really have high domain knowledge actually know what they're talking about in terms of the data, the inputs and the outputs that the models are supposed to give within their domain.

6:43And then you have the data scientists or the AI engineers and this side of the coin that are more technical. They know from a technical point of view what the models expect and what the model should output. And then the collaboration between them is now even higher because nowadays you can actually prompt other lamps with natural language and you actually need to ensure that both the models actually perform well and also the prompts and these kind of things. So the collaboration is even more important nowadays. And that's also the case for still the case for text cap models and these kind of things which we also support within our job.

7:17I guess maybe in the context of let's say there's a new team that's exploring the adoption of AI technology maybe for the first time. Maybe they're not coming from that data science background, the sort of heavy ML ops stuff, but maybe they've been excited by this latest wave of AI technologies. How would you go about helping them understand how their own data, the data that they would curate, the data that they would maybe collaborate on is relevant to and where that fits into the certain workflow? So So yeah, imagine someone may be familiar with what you can do with chat GPT or pasting in certain documents or other things.

8:02And now they're kind of wrestling through how to set up their own domain specific AI workflows in their organization. What would you kind of describe about how their own domain data and how collaborating around that fits into common AI workflows? Yeah, so something that I like to think about a lot around this subject is machine learning textbooks. And they often talk about modeling a problem as well as building a model, right? There's a famous mama and matter cycle. And in that, when you model a problem, you're basically trying to explain and define the problem. So I have articles and I need to know whether they are a positive or negative rating.

8:45and I'm describing that problem and then I'm going to need to describe that problem to a domain expert or an annotator through guidelines and when I can describe that problem in such a way that the annotator or the domain expert answers that question clearly enough then I know that that's a modelled and clear problem and it's something I could then take on to build a model around in simple terms it makes sense and so I think when you're going into a new space like generous AI, and you're trying to understand your business context around these tools, you can start off by modeling the problem in simple terms by looking at the data and saying, okay, does this label make sense to this articles?

9:27If I sort all these articles down by these labels or by this ranking, are these the kinds of things I'm expecting? Starting off at quite low numbers, right? Like single articles and kind of building up to tens of hundreds. And as you do that, you begin to understand and also iterate on the problem and kind of change it and adapt it as you go. And once you've got up to a reasonable scale of the problem, you can then say, all right, this is something that a machine learning model could learn. I guess on that front, maybe one of the big confusions that I've seen floating around these days is the kind of data that's relevant to some of these workflows.

10:08So it might be easy for people to think about a labeled data set for a text classification problem, right? Like here's this text coming in. I'm going to label it spam or not spam or in some categories. But I think sometimes a sentiment that I've got very often is, hey, our company has this big file store, right, of documents. and somehow I'm going to fine-tune quote-unquote a generative model with just this blob of documents and then it will perform better for me. And there's two elements of that that are kind of mushy. One is like, well, to what end for what task, right? What are you trying to do?

10:52And then also how you curate that data then really matters. Is this a sentiment that you all are seeing or how for this latest wave of models, like how would you describe if a company has a bunch of documents and they're in this situation, they're like, hey, we know we have data and we know that these models can get better and maybe we could even create our own private model with our own domain of data. What would you walk them through to explain where to start with that process and how to start curating their data maybe in a less general way, but towards some end? I think in these scenarios, it's always good to first establish a baseline or a benchmark, because what we often see is that people come to us or come to the open source space.

11:40They say, OK, we really want to fine tune a model. We really want to do a super extensive RAC pipeline with all of the bells and whistles included, and then kind of start working on these documents. But what we often see is that they don't even have a benchmark. They don't have a benchmark to actually start with. So that's normally what we recommend. Also, whenever you work with a RAC pipeline, ensure that all of the documents that you index are actually properly indexed, properly chunked. Whenever you actually execute a pipeline and you would store these retrieved documents or these based on the question and the queries in RGLR or any other data annotation tool, you can actually have a look at the documents, see if they make sense, see if the retrieval makes sense, but also if the generated output makes sense.

12:24And then whenever you have that baseline set up, from there, actually start iterating and then kind of making additions to your pipeline. Shall I add re-ranking potentially to the retrieval if the retrieval isn't functioning properly? Shall I add a fine-tuned version of the model? Should I switch from the latest Lama model of 3 billion to 7 billion or these kind of things? And then from there on, you can actually consider maybe either fine-tuning a model if that's actually needed or fine-tuning one of the retrievers or these kind of things. As you're saying that, as you're speaking from this kind of profound expertise you have, and I think a lot of folks really have trouble just getting started.

13:01And like you asked some great questions there. But I think some of those are really tough for someone who's just getting into it, like which way to go and some of the selections that you would go with that. could you talk a little bit about the kind of go back over the same thing but um kind of make up a little workflow you know that's kind of hands-on on just like you might see this and this is how i would decide that just for a moment just so people can kind of grasp kind of the thought process you're going because you kind of described a process but uh if you could be a little bit more descriptive about that um i think when i talk to people once they get going they kind of go to the next step and go to the next step and go to the next step.

13:38But the first four or five big question marks at the beginning, they don't know which one to handle. So I can add some practical steps onto that that I've worked with in the past. That'd be fantastic. Yeah. So one thing that you can do that is really straightforward is actually to write down a list of the kinds of questions that you're expecting your system to answer. And you can get that list by speaking to domain experts, or if you are a domain expert, you can write it yourself, right? And it doesn't need to be an extensive, exhaustive list. It can be quite a small starting set. You can then take those questions away and start to look at documents or pools and sections of documents from this lake that you potentially have and associate those documents with those questions and then start to look if a model can answer those questions with those documents.

14:31in fact by not even building anything by starting to use say chat tpt or hugging chat or any of these kind of interfaces and just seeing this very very low simple benchmark see is that feasible whilst at the same time starting to ask yourself can i as a domain expert answer this and that's kind of where argilla comes in at the very first step so you start to put these documents in front of people with those questions and you start to search through those documents and say to people, can you answer this question? Or here's an answer from a model to this question in a very small setting. And you start to get basic early signals of quality.

15:13And from there, you would start to introduce proper retrieval. So you would scale up your doc, you would take all of your documents, Say you had 100 documents associated with your 10 questions. You put all those 100 documents in an index and iterate over your 10 questions and see, okay, are the right documents aligning with the right questions here? Then you start to scale up your documents and make it more and more of a real world situation. You would start to scale up your questions. You could do both of these synthetically. And then if you still started to see positive signals, you could start to scale.

15:49and if you start to see negative signals, I'm no longer getting the right documents associated with the right questions. I personally would always start from the simplest levers in a RAG setup. And what I mean there is that you have a number of different things that you can optimize. So you have retrieval, you can optimize it semantically or you can optimize it in a rule-based retrieval. You can optimize the generative model, you can optimize the prompt. And the simplest movers, the simplest levers, are the rule-based retrieval, the word search, and then the semantic search. So I would first of all add like a hybrid search.

16:30What happens if I make sure that there's an exact match in that document for the word in my query? Does that improve or improve my results? And then I would just move through that process, basically.

16:55what's up friends i'm here with a friend of mine a good friend of mine michael greenwich ceo and founder of work os work os is the all-in-one enterprise sso and a whole lot more solution for everyone from a brand new startup to a enterprise and all the ai apps in between. So Michael, when is too early or too late to begin to think about being enterprise ready? It's not just a single point in time where people make this transition. It occurs at many steps of the business. Enterprise single sign on like SAML, auth, you usually don't need that until you have users. You're not going to need that when you're getting started.

17:32And we call it an enterprise feature. But I think what you'll find is there's companies when you sell to like a 50 person company, they might want this. They actually, especially if they care about security, they might want that capability in it. So it's more of like SMB features even if they're tech forward. At WorkOS, we provide a ton of other stuff that we give away for free for people earlier in their lifecycle. We just don't charge you for it. So that AuthKit stuff I mentioned, that identity service, we give that away for free up to a million users, one million users. And this competes with Auth0 and other platforms that have much, much lower free plans.

18:06I'm talking like 10 ,000, 50 ,000, like we give you a million free because we really want to give developers the best tools and capabilities to build their products faster, you know, and to go to market much, much faster. And where we charge people money for the service is on these enterprise things. If you end up being successful and grow and scale up market, that's where we monetize. And that's also when you're making money as a business. So we really like to align, you know, our incentives across that. So we have people using AuthKit that are brand new apps, just getting started, companies in Y Combinator, side projects, hackathon things, you know, things that are not necessarily commercial focus, but could be someday.

18:42They're kind of future-proofing their tech stack by using WorkOS. On the other side, we have companies much, much later that are really big who typically don't like us talking about them. They're logos, you know, because they're big, big customers. But they say, hey, we tried to build this stuff or we have some existing technology, but we're sort of unhappy with it. The developer that built it maybe has left. I was talking last week with a company that does over a billion in revenue each year and their skim connection, the user provisioning was written last summer by an intern who's no longer obviously at the company and the thing doesn't really work.

19:14And so they're looking for a solution for that. So there's a really wide spectrum. We'll serve companies that are in a, you know, their offices in a coffee shop or their living room all the way through. They have a, you know, their own building in downtown San Francisco or New York or something. And it's the same platform, same technology, same tools on both sides. The volume is obviously different. And sometimes the way we support them from a kind of customer support perspective is a little bit different. Their needs are different, but same technology, same platform, just like AWS, right? You can use AWS and pay them$10 a month.

19:41You can also pay them$10 million a month, same product. Or more, for sure. Or more. Well, no matter where you're at on your enterprise ready journey, WorkOS has a solution for you. They're trusted by Perplexity, Copy.ai, Loom, Vercel, Indeed, and so many more. You can learn more and check them out at workos.com. That's w-o-r-k-o-s.com. Again, workos.com.

20:23I'm guessing that you all, you know, the fact that you're supporting all of these use cases on top of Argyla on the data side makes me think, like you say, there's so many things to optimize in terms of that rag process. But there's also so many AI workflows that are being thought of, whether that be code generation or assistance or content generation, information extraction. But then you kind of go beyond that. David, you mentioned text classification and of course there's image use cases. So I'm wondering from you all, at this point, one of the things Chris and I have talked about on the show a bit is we're still big proponents and believe that in enterprises, a lot of times there is a lot of mixing of rule-based systems and more kind of traditional, I guess if you want to think about it that way, machine learning and smaller models, and then bringing in these larger Gen AI models as kind of orchestrators or query layer things.

21:29And that's a story we've been kind of telling, But I think it's interesting that we have both of you here in the sense that like you really, I'm sure there's certain things that you don't or can't track about what you're doing. But just even anecdotally, out of the users that you're supporting on Argeala, what have you seen in terms of what is the mix between those using Argeala for this sort of maybe what people would consider traditional data science type of models like text classification or image classification type of things and these maybe newer workflows like RAG and other things? How do you see that balance?

22:09And do you see people using both or one or the other? Yeah, any insights there? I think we recently had this company from Germany, Elamind, over at one of our meetups that we host. And they had an interesting use case where they collaborated with this healthcare insurance platform in Germany. And one of the things that you see with large language models is that these large language models can't really produce German language properly. They're mostly trained on English text. And that was also one of their issues. And what they did was actually a huge classification and generation pipeline, combining a lot of these techniques where they would initially get an email in that they would classify to a certain category.

22:57Then based on the category, they would kind of define what kind of email template, what kind of prompt template they would use. Then based on the prompt template, they would kind of start generating and composing one of these response emails that you would expect for like a customer query request coming in for the healthcare insurance companies. And then in order to actually ensure that the formatting and phrasing and the German language was applied properly, they would then, based on that prompt, regenerate the email once more. So prompt an LLM to kind of improve the quality of the initial proposed output.

23:32And then after all of these different steps of classification, of retrieval, augmented generation, of an initial generation and a regeneration, they would then end up with their eventual output. So what we see is that all of these techniques are normally combined. And also a thing that we are strong believers in is that whenever there is a smaller model or an easier approach applicable, why not go for that instead of using one of these super big large language models so if you can just classify is this relevant or is this not relevant and based on that actually decide what to do that makes a lot of sense and also one of the interesting things that i've seen one of these open source platforms haystack out there using is also this query classification pipeline where they would classify incoming queries as either a key terminology search, a question query, or actually a phrase for an LLM to actually start prompting an LLM.

24:36And based on that, actually redirect all of their queries to the correct model. And that's also an interesting approach that we've seen. Quick follow-up on that. And it's just something I wanted to draw out because we've drawn it out across some other episodes a bit. You were just making a recommendation, kind of go for the smaller model versus the larger model. Could you, for people trying to follow and there's the divergent mindsets, could you take just a second and say why you would advocate for that, what the benefit, what the virtue is in the context of everything else? I would say smaller models are generally hostable by yourself, so it's more private.

25:15Smaller models, they are more cost-efficient. Smaller models can also be fine-tuned easier to your specific use case. So even what we see a lot of people coming to us about is actually fine-tuning LLMs. But even the big companies out there with huge amounts of money and resources and dedicated research teams still have difficulties on fine-tuning LLMs. So whenever you, instead of within your retrieval augmented generation pipeline, fine-tune like the NLM for the generation part, you can actually choose to fine-tune one of these retrieval models that you can actually fine-tune on consumer-grade hardware.

25:54You can actually fine-tune it very easily on any arbitrary data scientist developer device. And then instead of having to deploy anything on one of the cloud providers, you can start with that. And in a similar reasoning for a RAC pipeline, whenever you provide an LLM with garbage within such a retrieved augmented generation pipeline, you actually also ensure that there's less relevant content and the output of the LLM is also going to be worse. Yeah, I've seen a lot of cases where I think it was Travis Fisher who was on the show. He advocated for this hierarchy of how you should approach these problems.

26:33And there's like, you know, maybe seven things on his hierarchy that you should try before fine tuning. And I think in a lot of cases I've seen people maybe jump to that. They're like, oh, I forget which one of you said this, but, you know, this naive rag approach didn't get me quite there. So now I need to fine tune when in reality, there's sort of a huge number of things in between those two places. And you might end up just getting a worse performing model, depending on how you go about the fine tune. One of the things, David, you kind of walk through these different the example of the specific company that had these workflows that involve a variety of different operations, which I assume, you know, to Ben, you mentioning earlier, starting with a test set.

27:20and that sort of thing and how to think about the tasks. I'm wondering if you can specifically now talk just a little bit about Arjila specifically, people might, might be familiar generally with like data annotation. They might be familiar, you know, maybe even with how to upload some data to quote, fine tune some of these models in an API sense, or maybe even in a more advanced way with Q Laura or, or something like that. But could you take a minute and just talk through kind of Arjila's approach to data annotation and data collaboration? And like, it's kind of hard on a podcast because we don't have a visual to show for people.

Read the full transcript

28:01But as best you can, you know, help people to imagine, you know, if I'm using Arjila to do data collaboration, what does that look like in terms of what I would set up and who's involved? What actions are they doing? That sort of thing. Aguila, there's two sides to it, right? So there's a Python SDK, which is intended for the AI machine learning engineer. And there's a UI, which is intended for your domain expert. In reality, the engineers often also use the UI and you kind of iterate on that as you would, because it gives you a representation of your task. but there's these two sides. The UI is kind of lightweight.

28:42It can be deployed in a Docker container or on Hugging Face spaces. It's really easy to spin up. And the SDK is really about describing a feedback task and describing the kind of information that you want. So you use Python classes to construct your dataset settings. You'll say, okay, my fields are a piece of text, a chat or an image. and the questions are a text question, so like some kind of feedback, a comment, for example, a label question, so positive or negative labels, for example, a rating, let's say between one and five, or a ranking, so example one is better than example two, and you can rank a set of examples.

29:30And with that definition of a feedback task, you can create that on your server, in your UI, And then you can push what we call records, your samples into that data set. And then they'll be shown within the UI and your annotator can see all of the questions. They'll have nice descriptions that were defined in the SDK. They can tweak and kind of change those as well if you need in the UI, because that's a little bit easier. You can distribute the task between a team. So you can say, OK, this record will be accepted once we have at least two reviews of it. you can say that some questions are required and some aren't and they can skip through some of the questions the ui has a loads of keyboard shortcuts like with numbers and arrows and return so you can move through it like really fast it's kind of optimized for that and different sort of screen sizes one thing we're starting to see is that as llms get really good at quite long documents some of the stuff that they're dealing with is like a multi-page document or a really detailed image and then a chat conversation and then we want like a comment and a ranking question so it's like a lot of information in the screen so the ui kind of scales a bit like an ide like so you can drag it around to give yourself enough width to see all this stuff and then you can move through it in a reasonably efficient way with the keyboard shortcuts and stuff interesting and what do you see as kind of the the backgrounds of the roles of people that are using this tool because one of the interesting things from my perspective especially with this kind of latest wave is there there's maybe less data scientists kind of ai people that that their background and more software engineers and and just you know non-technical domain experts so how do you kind of think about the roles within that and what are you seeing in terms of who's using the system for us i think it's yeah Yeah, from the SDK Python side, it's really still developers.

31:35And then from the UI side, it's like anyone in the team that needs to have some data labeled with domain knowledge. Often these are also going to be like the AI experts. And one of the cool things is that whenever an AI expert actually sets up a data set, besides these fields and questions, they can actually come up with some interesting features that they can add on top of the data set. They are also able to add like semantic search, like attach records or semantic representation of the records to one of the records, which actually enables the users within the UI to label way more efficiently.

32:11So, for example, if someone sees a very representative example of something that's bad within their data set, they can do a semantic search, find the most similar documents, and then continue with the labeling on top of that. And besides that, you can also, for example, filter based on model certainties. So let's say that your model is very uncertain about an initial prediction that you have within your UI. And it's really interesting for the domain expert or for the data scientist to go and have a look at that specific record or range of uncertainties. And then based on that, the labeling or the data curation or whatever you would like to call it becomes way more engaging and way more interesting.

32:55And on top of that, another thing that we are starting to explore is actually using this AI feedback and synthetic data within our JLA as well. And that's actually one of the other products that we're working on, and it's called this C-label. So nowadays, what you can do with LLMs is also actually use LLMs to evaluate questions, for example, to evaluate whether something is labeled A, B, or C, or whether something is a good or bad response. and you see all kinds of like tools, open source tools out there. And that's also a thing that we are looking at for integrating with the UI, where instead of doing this more from a data science SDK perspective, users without any technical knowledge would actually be able to tweak these guidelines that Ben highlighted earlier and then say, okay, maybe instead of taking this into account, you should focus a bit more on like the harm that potentially is within your data or the risks that are within your data.

33:53And then you would be able to prompt an LLM once again to kind of label your data. And then you wouldn't directly need the Python SDK anymore. I was thinking about, as you were describing that, I work at a large organization and we certainly have a lot of domain experts in the organization I work at that are either non-technical or semi-technical. And as users, they will sometimes find it intimidating, you know, kind of getting into all this as they're starting a project. Could you talk a little bit about what it's like for a non-technical person to sit down with Argeala and start to work in a productive way?

34:29What is that experience like for them? Because it's one thing like the technical people kind of just know. They dive into it. They're going to use the SDK. They've used other SDKs. But there can be a bit of hand-holding for people who are not used to that. Could you describe the user experience for that non-technical subject matter expert coming in and what labeling is like and just kind of paint a picture of words on what their experience might be like? Yeah, I mean, one thing I guess I'd start off by saying is that Argyla is kind of the latest iteration of a problem that has existed for a long time in machine learning and data science, right?

35:06about collecting feedback from domain experts. And it's kind of gone through spreadsheets and various other tools that were substandard and really bad user experiences where domain experts were asked for information. That information was extracted and then models have been trained really poorly on that information. So as a field, we kind of know that it's something that we have to take really seriously. And that's kind of what Argyla is built on top of, right? That's part of our DNA as a product. It's like optimizing the feedback process as a user experience problem. And so when the user sits down to use Argyla, the intention is that all of the information should be right there in front of them inside their single record view.

35:56So what that means is they've got a set of guidelines that are edited in Markdown. they can contain images links to various pages or other external documents if they need and they can just kind of scroll through that it's always there it's always available they've then also got like basic metrics so they'll know how many records they've got left how many they've labeled they can view their kind of team status and see what's going on and then on the left they have their fields which they can scroll through and on the right they'll have a set of questions as I said they can move through these in keyboard shortcuts and they can switch the view so that they can scroll kind of infinitely or they can move into a kind of page swiping which yeah if you're looking at really small records for like a couple of lines and you're just assigning a symbol label to you can do that in bulk so as we said you could use a semantic search give me all the records that are similar to this one and I'll bulk label those or you could search for terms inside those records and you can both label those.

36:58And then once you're finished, you'll know about it. And one of the interesting things that I've done personally quite often is sit together with the domain experts and their AI engineers to kind of walk them through how to configure our JLA most usefully for both of them. And then the domain experts come with a lot of things to the table, like I want to see this specific representation. What if we could do this? What if we could do that? Then the AI engineers think about the data side of things. Is this possible from our point of view, from our side? And then me as a mediator, so to say, can I make the most out of the Archela configuration?

37:37And that's also how we see this collaboration process going, where domain experts really work together also with AI engineers, because AI engineers or machine learning engineers actually know what's possible from the data, what it means to get high quality data for fine tuning model. because whenever a domain engineer comes up with something that's useful for them in terms of labeling, it doesn't mean necessarily that it's actually proper data that's going to come out of there in terms of fine-tuning a model. And that's also a part of, I guess, the collaboration that we're talking about.

38:25What's up, friends? I've got something exciting to share with you today. A sleep technology that's pushing the boundaries of what's possible in our bedrooms. Let me introduce you to Eight Sleep and their cutting-edge Pod 4 Ultra. I haven't gotten mine yet, but it's on its way. I'm literally counting the days. So what exactly is the Pod 4 Ultra? Imagine a high-tech mattress cover that you can easily add to any bed. But this isn't just any cover. It is packed with sensors, heating and cooling elements, and it's all controlled by sophisticated AI algorithms. It's like having a sleep lab, a smart thermostat and a personal sleep coach all rolled into a single device.

39:06It uses a network of sensors to track a wide array of biometrics while you sleep, sleep stages, heart rate variability, respiratory rate, temperature and more. It uses precision temperature control to regulate your body's sleep cycles. It can cool you down to a chilly 55 degrees Fahrenheit or warm you up to a good, nice solid temperature of 110 Fahrenheit. And it does this separately for each side of the bed. This means you and your partner can have your own ideal sleep temperatures. But the really cool part is that the pod uses AI and it uses machine learning to learn your sleep patterns over time.

39:47And it uses this data to automatically adjust the temperature of your bed throughout the night according to your body's preferences. Instead of just giving you some stats, it understands them and it does something about it. Your bed literally gets smarter as you sleep over time. And all this functionality is accessible through a comprehensive mobile app. You get sleep analytics, trends over time, and you even get a daily sleep fitness score. Now, I don't have mine yet. it is on its way thanks to our friends over at 8sleep and I'm literally counting the days I get it because I love this stuff but if you're ready to take your sleep and your recovery to the next level head over to 8sleep.com slash practical ai and use our code practical ai to get 350 bucks off your very own pod 4 ultra and you can try it free for 30 days I don't think you want to send it back but you can if you want to.

40:40They're currently shipping to the US, Canada, United Kingdom, Europe, and Australia. Again, 8sleep.com slash practical AI.

41:09I want to maybe double click on something that David, you just said and sort of passing, which I think is quite significant. And I don't know if some people might have caught it, but when you were talking about the still label, you also talked about AI feedback. So AI feedback and synthetic data. So I'd love to get into those topics a little bit. But maybe first coming from the AI feedback side, I think this is super interesting because, you know, Ben, you talked about how this is a kind of more general problem that people have been have been looking at in various ways from various perspectives for a long time in terms of this data collaboration labeling piece.

41:48But there is this kind of very interesting element now where we have the ability to utilize these very powerful, maybe general purpose instruction following type of models to actually act as labelers within the system or at least generate drafts of labels or feedback or even preferences and scores and all of those sorts of things. So I'm wondering if one of you could speak to that. Some people might find this kind of strange that we're kind of giving feedback to AI systems with AI systems, which seems circular. And maybe like, why would that work? Or just sort of maybe that's kind of produces some weird feelings for people.

42:38But I think it is a significant thing that is happening. And so, yeah, either of you would want to kind of dive into that. What does it specifically mean in AI feedback? How are you seeing that being used most productively? So when we create a data set, either manually or with AI feedback or AI generation, we have all the information there to understand the problem. We have a set of guidelines. We have a set of labels, definitions of those labels, with documents and definitions of those documents. We give those to a manual annotator, or we'll go out and collect those documents and we'll give those documents to the manual annotator.

43:14And we're trying to describe that problem so that the person understands it to create the data. We can essentially take all of the same resources and give those to an LLM and get the LLM to perform the same steps. So there's two parts to that. There's a generative part where the LLM can generate documents. So let's say we've got 100 documents in our data set, but we want 10 ,000. We can say, generate a document like this one, but, and add variation on top of that. And we can fan out our data set, our documents from 100 to 10 ,000. We could then take those same documents or a pool of documents from elsewhere, and we could get feedback on that.

43:57So that could be qualitative feedback. Tell me which of these documents are relevant to this task. Tell me which of these documents are of a high quality, are concise, are detailed, these kind of attributes. So we could filter down our large data set or our generated data set to the best documents. We could also add labels. So we could say, tell me which of these documents relates to my business use case or not. These kind of things. Apply topics to these documents. And then we can, in doing so, create a classification data set, right, from those labels. Or we could, in one example, take a set of documents and use a generative model to generate questions or queries.

44:40about those documents. And we could use that to create a Q &A dataset or a retrieval dataset where we generate search queries based on documents. When you're doing that and you're generating the datasets with another model, how much do you have to worry about hallucination playing into that? It sounds like you have a good process for kind of trying to catch it there, but is that a small issue? Is that a larger issue? Any guidance on that? That's one of the main issues, definitely. like it is probably the main issue. And so really it's about both sides of that process that I described, that generating side and that evaluating side.

45:21So you get the large language models to do as much as possible to expose hallucination by evaluating themselves. And typically you're getting larger models to evaluate so that they're a more performant model and they should hallucinate less. The task of identifying hallucinations is not the same as generating a document. So typically LLMs are better at identifying hallucinations and nonsense if you give them the context than they are not generating it. And so you combine that within a pipeline and then you would take that to a domain expert in a tool like Argyla. And so that's really why we have these two tools, right?

46:01DistalAble and Argyla, because kind of without Argyla, DistalAble would suffer from a lot of those problems, right? Yeah, and I guess that brings us to the second tool, the distill label, which I know has some to do with this synthetic data piece as well. And I'm really intrigued to hear about this because I also see some of what you have on the documentation about what are people building with distill label. I do note a couple of data sets like the OpenHermes data set, the Intel Orca DPO data set. These are data sets that have been part of the lineage of models that I've found very, very useful.

46:45So first off, thanks for building tooling that's created really useful models in my own life. But beyond that, yeah, David, do you want to go into a little bit about what DistillLabel is? and maybe even tie into some of those things and how it's proven to be a useful piece of the process in creating some of those models. I think, yeah, the idea of DC Label kind of started to done half a year ago, more or less, or maybe a year ago, where we saw these initial new models coming out, like Alpaca and Dolly from Databricks, Alpaca from Stanford, where there were like data sets being generated with OpenAI frontier models being evaluated with OpenAI frontier models and then published and actually used for fine-tuning one of these models.

47:39So apparently there were research groups or companies kind of investing time in this. But what we also saw is when we would kind of upload these data sets into Agila, actually start looking at the data that there were a lot of flaws within there. And then whenever like ultra feedback, which is one of these specific papers that really started to scale the synthetic data and AI feedback concept, came out. We thought, okay, maybe it's worth to look into a package that can actually help us facilitate creating data sets that we can then eventually fine-tune within Arjella. And that's when we started to work on the initial version of this label.

48:16So it's kind of like application framework, like Lama Index or our LengChain, if you're familiar with those. but then specifically focused on synthetic data generation and AI feedback. So what we try to do is organize everything into this pipelining structure where you have either steps that are about basic data operations, tasks that are about prompt templates or prompting. And prompt templates you can think about either providing feedback, maybe rewriting some initial input that you provide to that prompt template, or maybe like ranking or like generating from scratch or these kind of things.

48:55And then these tasks are actually executed by LLMs and these are then all like fit together within a pipelining structure. The thing for these tasks is that nowadays we actually look at all of the most recent research implementations or most recent research papers and we try to implement them whenever they come out and are actually relevant for synthetic data generation. So you really go from the kind of finicky prompt engineering, so to say, to well-evaluated prompts that we've implemented. And the nice thing about our pipeline infrastructure is also that we run everything asynchronously. So there's multiple LLM executions being done at once, which will really speed up your pipeline.

49:40And on top of that, we also cache all of the intermediary results. So as you can imagine, calling the OpenAI API can be quite costly. And whenever you run a pipeline, a lot of things can go wrong. But whenever you actually rerun our pipelines within the C-label, you actually have these cache results already there. So you would avoid incurring additional costs whenever something within the pipeline breaks. Yeah, that's awesome. And I know that one element of this is the kind of creation of synthetic data for further fine-tuning LLMs to increase performance or maybe to some sort of alignment goal or something like that.

50:22But also, I know from working with a lot of healthcare companies, manufacturers, others that are more security privacy conscious in my day job, part of the pitch around synthetic data is maybe also creating data sets that might not kind of poison LLMs with a bunch of your own sort of private information that could be sort of exposed as part of an answer. answer that someone prompts the model in some way and this data is embedded in the data set and all of that. So yeah, I would definitely encourage people to check out the still label. And you said it's been around for half a year. So how have you seen the kind of usage and adoption so far?

51:11The usage and adoption has been quite good in terms of the number of data sets that have been released. So you mentioned the Intel Orca DPO dataset, which was an example use case of how we were initially using it, where we had this original dataset that had been labeled by Intel employees with references of what would be like the preferred response to a given prompt. And we actually used this label to kind of clean that based on prompting LLMs ourselves to reevaluate these chosen rejected pairs within the original dataset, filtering out all of the ambiguity. So sometimes the LLM wouldn't align with the original chosen rejected pair.

51:56And based on that, we were actually able to scale down the dataset by 50%, leading to less training time and also leading to a higher performing model. And that was one of the really famous examples that kind of inspired some people within the open source community to actually start looking at the distal label, to start using distal label to generate data sets. There's some hugging face teams that actually have been generating millions and millions of rows of synthetic data using distal label. And that's pretty cool to see that people are actually using it at scale. And besides that, there's also these smaller companies, so to say, but like LMI, the German consultancy, the German startup that I mentioned before, using it to also rewrite and resynthesize emails within like actual production use cases.

52:48It's really fascinating. You guys are pushing the state of the art in a big way. With the work that you've done in Distalable and Argila, where do you think things are going? Like when you're kind of end of whatever your task is of the day and you're kind of just letting your mind wander and thinking about the future, where do each of y 'all go in terms of what you think is going to happen what you're excited about what you're hoping will happen um what you might be working on in a few months or maybe a year or two what are your thoughts i suppose for me it's about two main things and the first would be modalities so moving out of text and into image and audio and video and also kind of ux environments so that maybe in Argyllar, but also in DistalAble, that we can generate synthetic data sets in different modalities and that we can review those.

53:43And that's a necessity and something that we're already working on and we've already got features around, but we've got kind of more coming. And then the second one, which I suppose is a bit more far-fetched, and that's a bit more about kind of tightening the loop between the various applications. So between DistalAble, Argyllar, and the application that you're building, So that you can deal with feedback as it's coming from your domain expert that's using your application and potentially our Giller at the same time. So we can kind of synthesize on top of that to evaluate that feedback that we're getting and generate based on that feedback.

54:18So we can add that into our Giller and then we can respond to that synthetic generation, that synthetic data. And then we can use that to train our model, this kind of tight loop between the end user, the application and our feedback. Yeah, and for me, it kind of aligns with what you mentioned before, Ben, like the multimodality, smaller, more efficient models, things that can actually run on a device. I've been playing around with this app this morning that you can actually load local LLM into, like a smaller QN or Lama model from Meta. And it actually runs on an iPhone 13, which is really cool.

54:56It's private. It runs quite quickly. And the thing that I've been wanting to play around with these speech-to-speech models where you can actually have real-time speech-to-speech. I'm currently learning Spanish at the moment. And one of the difficult things there is not being secure enough to actually talk to people out on the streets and these kind of things. So whenever you would be able to kind of practice that at home privately on your device, kind of talk some Spanish into an LM, get some Spanish back, maybe some corrections in English, these kind of scenarios are super cool for me whenever they would be able to...

55:31Muy bueno. Yeah, this is muy bueno. And yeah, I've been really, really excited to talk to you both and would love to have you both back on the show sometime to update on those things. Thank you for what you all are doing, both in terms of tooling and, you know, our GLN hugging face more broadly in terms of how you're driving things forward in the community and especially the open source side. So thank you both. thank you for taking time to talk with us and hope to talk again soon. Yeah. Thank you. And thanks for having us. Thank you.

56:15All right. That is Practical AI for this week. Subscribe now. If you haven't already, head to practicalai.fm for all the ways and join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our Beat Freaking Residence, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.

56:59Game on!

From the publisher

As Argilla puts it: “Data quality is what makes or breaks AI.” However, what exactly does this mean and how can AI team probably collaborate with domain experts towards improved data quality? David Berenstein & Ben Burtenshaw, who are building Argilla & Distilabel at Hugging Face, join us to dig into these topics along with synthetic data generation & AI-generated labeling / feedback.

Join the discussion

Changelog++ members save 11 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Fly.io – The home of Changelog.com — Deploy your apps close to your users — global Anycast load-balancing, zero-configuration private networking, hardware isolation, and instant WireGuard VPN connections. Push-button deployments that scale to thousands of instances. Check out the speedrun to get started in minutes. 
  • WorkOS – A platform that gives developers a set of building blocks for quickly adding enterprise-ready features to their application. Add Single Sign-On (Okta, Azure, Google, Microsoft OAuth), sync users from any SCIM directory, HRIS integration, audit trails (SIEM), free magic link sign-in. WorkOS is designed for developers and offers a single, elegant interface that abstracts dozens of enterprise integrations. Learn more and get started at WorkOS.com
  • Eight Sleep – Take your sleep and recovery to the next level. Go to eightsleep.com/PRACTICALAI and use the code PRACTICALAI to get $350 off your very own Pod 4 Ultra. You can try it for free for 30 days - but we’re confident you will not want to return it. Once you experience AI-optimized sleep, you’ll wonder how you ever slept without it. Currently shipping to: United States, Canada, United Kingdom, Europe, and Australia. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Towards high-quality (maybe synthetic) datasetsPractical AI · 57 min
Listen in VO