#192 Lukas Biewald: How Weights and Biases Supercharges Machine Learning

9 Jun 2024 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Notes

Episode #192

Lukas Biewald: How Weights and Biases Supercharges Machine Learning

Podcast Overview

  • Host: Craig S. Smith, former New York Times correspondent
  • Focus: Conversations with influential figures in AI, contextualizing technological advances and their global implications.

Episode Summary In this episode, Craig interviews Lukas Biewald, CEO and co-founder of Weights & Biases, a leading AI developer platform. They discuss Lukas's background, the evolution of machine learning, and how Weights & Biases provides essential tools for AI practitioners throughout the machine learning workflow.

Key Topics Discussed

  1. Lukas Biewald's Background
  2. Early education at Stanford focusing on natural language processing (NLP).
  3. Work experience at Yahoo and a startup (PowerSet) that became part of Microsoft Bing.
  4. Founded CrowdFlower (later Figure 8) to address data annotation needs.
  1. The Evolution of Machine Learning
  2. Transition from basic machine learning to modern applications requiring high-quality data.
  3. Growth of image labeling in the industry.
  4. Increasing importance of annotated data in the context of self-supervised learning.
  1. Weights and Biases: Origin and Mission
  2. Founded to enhance the machine learning toolchain, moving beyond just data annotation.
  3. Focus on providing comprehensive tools and support for model training, visualization, and compliance.
  1. Importance of Data Annotation
  2. High-quality data annotation is critical for effective machine learning models.
  3. Discussion on the shift in AI applications and the role of reinforcement learning with human feedback (RLHF).
  1. Compliance and Experiment Tracking
  2. The role of Weights and Biases in ensuring data lineage and compliance within AI, especially in light of regulatory frameworks in the EU vs. the US.
  3. Emphasis on the significance of tracking experiments to capture intellectual property and improve model performance.

Discussion Highlights

Monitoring AI Models

  • Data Drift: Addressed how models can silently fail due to upstream changes in input data.
  • Importance of monitoring input data and output distributions.
  • Suggests that monitoring practices should be tailored to the specific needs of different applications.

Evolving Products to Meet Industry Needs

  • The necessity for Weights and Biases to adapt their tools to the changing landscape of machine learning, especially with the rise of large language models and prompt engineering.
  • Current trends in machine learning that necessitate updates in tools and methodologies.

Future of Weights and Biases

  • Lukas expresses a desire to keep Weights and Biases independent, potentially through an IPO, to maintain flexibility and integration capabilities with various platforms.

Key Takeaways

  • Data Quality is Crucial: High-quality annotated data remains vital for the success of machine learning applications, despite advancements in self-supervised learning.
  • Regulatory Compliance: Tools like Weights and Biases are essential in ensuring compliance with emerging AI regulations, particularly regarding data lineage and transparency.
  • Experiment Tracking: Keeping track of experiments is crucial for retaining knowledge and improving model performance over time, especially as team members change.
  • Continuous Evolution: The landscape of machine learning is rapidly changing, and tools must evolve to meet practitioners' needs, particularly around monitoring and model performance in production.

Conclusion In summary, the episode provides insight into the challenges and opportunities within the AI field, emphasizing the critical role of data, compliance, and effective monitoring in developing reliable machine learning models. Lukas Biewald’s experiences and vision for Weights and Biases reflect the dynamic nature of the AI landscape and the tools necessary for success in this rapidly evolving domain.

Stay Connected

  • Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
  • Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Why would you want to go back and look at the data your model is trained? So in industry, there are real compliance reasons. So, you know, if you're in the EU, a lot of our customers will get data from their users, but then the users have the right to take their data out. Yeah, a great use case for weights and biases that we can produce the data that your model is trained on. I think the U.S. regulations right now are not clear enough to know exactly what you should produce. And so I worry a little bit. I worry a lot, honestly, that the U.S. regulations were just vague and scary. AI might be the most important new computer technology ever.

0:32It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds.

1:22If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. Oracle.com slash IonAI. That's oracle.com slash IonAI. Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I talk to Lucas B. Wald, co-founder of Weights and Biases, a company at the forefront of MLOps with innovative tools designed to streamline the machine learning workflow. Lucas talks about some of the under-the-hood complexities that his company addresses, from tracking changes in training data to monitoring large model behavior in production.

2:20I hope you find the conversation as informative as I did. Okay, so why don't you start, Lucas, by introducing yourself, giving some of your background, where you went to school, what you were doing before Weights and Biases, and how you came to start the company and then we'll talk about what the company does and talk more generally about neural nets and large language models and where the research is going. Sure, sure. So I was a Stanford undergrad and then a master's student where I expected to get a PhD and was doing research with Daphne Kohler back in 2004, 2005, and did work on natural language processing, which I think is all probably irrelevant at this point, but working on translation and kind of support vector machines and base nets and kind of older models.

3:22And from there, I went to work on search at Yahoo and then search at a NLP startup called PowerSet that actually became Microsoft Bing. And through that process, through actually all three of those stints that I did working on natural language, got convinced that the most important problem in machine learning was access to high quality training data. So I started a company called Crowdflower that rebranded as Figure 8 towards the end. It was a kind of early version of Scale AI. And we collected data sets for almost everyone doing natural language processing back then. And then image processing kind of came online with, you know, kind of modern methods and convolutional neural networks.

4:08What year was Crowdflower? Crowdflower was probably started, I think it was 2007. Oh. You know, back then, actually, you know, machine learning was not hot. You know, investors did not want to hear the word machine learning. It was a really different time. and we ran I think until 2018 or 2019 when it sold to Appen which is a big public company in Australia that no one's ever heard of but you know did quite a lot of natural language processing and from there I started a company Wasted Biases with the idea being we wanted to help with the rest of the tool chain in machine learning so we'd gotten training data for thousands of companies and teams and researchers.

4:48And we started to see lots of new problems with getting those models into production and making them useful. And we wanted to help with the end to end solution versus just the training data problem. I think training data now, there's lots of ways to do it, but still there's many issues making machine learning actually really work for real applications. Just going back, Daphne Collar, I've had her on the podcast talking about her drug discovery, but also about her career. And I've had Fei-Fei Li on talking about the creation of ImageNet and then what she's working on. So your crowd player was prior to ImageNet?

5:32In fact, I remember Fei-Fei Li reached out and asked me if I could help her do ImageNet. And I really, really regret not doing it. At the time, you know, it was like we were really trying to make money and academics are such horrible customers that she had all these specific requirements. So I remember turning her down, but I regret that to this day. I think ImageNet was such an amazing contribution. I'm just curious, how did Crowdflower aggregate data? I mean, the word crowd makes me think it's like Mechanical Turk or something. Yeah, so at that time, Mechanical Turk was kind of starting to become popular.

6:07And the real challenge was that getting high quality data was the biggest issue. So at first we were an aggregator on top of Mechanical Turk. But then as we kind of moved into more industrial applications, we started crowdsourcing people all over the world. So we kind of became our own Mechanical Turk in a way, but more oriented towards large scale industrial data collection. and text textual data well it started with text and that was the big thing at the time right so the applications back then our big customers were the big search companies and then e-commerce you know wanted search so at that time those are the big applications and also text extraction for financial companies like you know Bloomberg and banks and others but then over time as image applications started working, we saw an explosion in image labeling.

6:59So, you know, by the end, image labeling had overtaken text. And now with my new company, we see the opposite trend where it started off. It was mostly vision applications. And I think text has actually come back and overtaken vision, at least from what we see on our new platform at Weights and Vises. Yeah, I've worked with LabelBox. And I know it's at Alex Ratner at scale. And I've talked to some of the other snorkel. No, Alex is at snorkel. Yeah. Right. And is that business still strong? I was surprised here at the conference to see people talking about annotation. is is annotated data still given the the self-supervised learning explosion is annotated data still as important as it was yeah it's a huge huge business still it's evolved so um you know i i know all these guys and you know talk to them a fair amount um i think you see less vendors here than at cvpr or more vision oriented conference i think in vision it's still really unsolved problem.

8:18And now there's a lot of tooling to help the annotators be more efficient, more effective. And in text, what I hear is that you need more and more high quality annotation. So it used to be very, very simple annotation that you could almost do in the blink of an eye, like instant feedback from humans. And now sometimes the type of labeling they're doing takes hours. But I think RHLF is booming. I mean, it's absolutely here. Well, here to stay is a dangerous thing to say in this industry, but I think that it's not going anywhere anytime soon. And the demand is only increasing. Although I've been talking to people about RLAIF, which, you know, sounds like it may overtake RLHF.

9:07RLHF to me sounds, I mean, for listeners that don't know what it is, it's reinforcement learning with human feedback. and it's being heavily used by the large model companies to kind of nudge the models away from hallucination or other bad behavior. But just intuitively, that sounds like a very crude way to go about it. But what do you think? When you say it's here to stay, that surprises me. It sounds to me like a very short-term solution. Interesting. Well, I'm not an expert on this anymore. So this is not my area at this point. But I guess I would say broadly that as long as humans can do things better than AI algorithms, I think inputting that human data in some form is valuable.

10:03I think as AI gets more sophisticated, you know, you continue to automate away various tasks that humans used to do because it's painful and expensive to get humans involved. So I think that the tasks become more sophisticated and complicated and specific that humans are doing. But my feeling is that until we hit singularity or whatever you want to call it, why not have human input to make the algorithm better? Yeah. And we'll get to weights and biases. I'm sorry, but this is that I know this is a previous live for you, but it's interesting to me. Annotating text, just I've never quite I've never looked into it.

10:44So I don't really know what people mean by annotating text. Are you talking about annotate of sort of identifying, you know, subject, verb? uh or are you talking about uh larger concepts or what yeah what what are you annotating when you annotate text um all kinds of things i mean it's changed over time so you know the big applications used to be you know 10 or 15 years ago it would be um something like text extraction right where like you know i have this document maybe it's a medical record i want to pull out the name of the person. I want to pull out, you know, what diseases maybe they have.

11:30Maybe I want to like put that in some ontology that I have almost any unstructured document. There's some way that you want to structure it to feed it into a computer. And so it's literally having human beings do some of the tasks that you want the algorithm to do, and then using that as training data. So every NLP task you could think of would have an annotation component. Now that's a little bit out of date right now people use um you know large language models for quite a lot of tasks and so the annotations have changed right so now the things that people want is really high quality input for the algorithms right so if the algorithm is having trouble with reasoning putting in really high quality reasoning if the algorithm has trouble telling compelling stories okay you want someone to input compelling stories if the algorithm is you know saying things that are rude you want to guide it to say things that are not rude right and so um that's the sort of modern type of annotation yeah okay so on to weights and biases uh and and ml apps more generally give us uh i know what weights and biases are within a neural network but explain what weights and biases are and why you chose that as the is the is the name of the company yeah and so so we devices are kind of the coefficients the numbers aside of the neural network so a neural that basically consists of weights and biases.

12:52And that's all the numbers inside Transformer, any algorithm that you might use. And the name is like historic. I think it comes from linear regression where the coefficients are weights and then the final term was multiplied by one. So it's kind of a constant. It's called the bias. And we call the company Weights and Biases because our thesis with Weights and Biases when we started it six years ago was we really wanted to be on the side of the ML practitioner, the researcher, or the engineer that's responsible for building general network. And we felt like the software that was available at the time was very clearly being sold to the CIO or the CTO or some kind of high-level boss who has compliance requirements which are important, but he's often was neglected.

13:40At the time, the software that was available was really clearly designed for CTOs, CIOs, folks like that. And we wanted to make software that really served the ML researchers themselves, the ML practitioners. And so we wanted a name that would show that, right? So we still, to this day, have people come up to our booth at every conference we go to, except NeurIPS, where they say, oh, weights and balances, what does that do? And then they tell me what a dumb name, you know, I made the company. But I kind of love it because I think that people that are in our target market, they understand it. That's actually changing a little bit, you know, with this evolution to LLMs and prompt engineering and all that.

14:23But, you know, at this point, the name is stuck. Yeah, yeah, yeah. And it's a tool of, well, you describe it. Sure, sure. I like to think of it as a set of tools that take care of the needs that you have from getting a model from first conception or first specced out to deployed and doing something useful in production. And that's kind of our North Star is, OK, what does the industry need to get these models running reliably in production? What do practitioners need to make their life easier is another way of saying it. So some of our like guiding principles are we try to integrate with everything, including competitive products.

15:04So like, you know, because, again, I think if you had sort of like a top down MBA strategy, you might say, hey, let's block out, you know, all the competitors. Like you just have to live in weights and biases world. For us, we know that people are using lots of different stuff already. And so we want to integrate into that. And we try to pick off pieces of the solution that we know we can do really well. So, for example, things we don't do that I think are a big problem, just like ML infrastructure, right? Like everyone's GPUs. Our customers always come to us. Hey, can you get me GPUs? We don't have GPUs.

15:35We can't help you with that, right? We can't, you know, we can't help you with low-level infrastructure stuff. We're just not good at that. A lot of people are working on it. The problem is not solved. But we can help you. And I think we do a really good job. Keep track of what's happening as you build your models. I think almost everyone at NURPS uses us for that these days. things that people don't know we can do, but I think we do really well, is something called data lineage, where you want to know what's upstream of your model. So usually, often these days, your model is trained on a foundation model.

16:07Maybe that foundation model is actually fine too, not a different foundation model. So we can keep track of that lineage for you. And we can keep track of all the data that goes in. And there might be like transforms of the data along that way, data augmentation. So we can keep track of that whole pipeline for you in an efficient way and keep track of all the evaluations that you do on your model. These days, people typically have a fast set of evaluations and then slower evaluation sets that they might run asynchronously. And of course, these evaluation sets and these training data sets are always changing all the time.

16:36And so in a perfect world, they'd always be the same and everything would be organized and beautiful. But we want to help in the messy world that people actually want. Yeah. And you were showing me the interface yesterday. Then you produce these visualizations, for example, for evaluations where it's very easy to see where things are doing better, where they're doing worse. Talk about those visualizations and how people use them. And then we were talking about integrating sort of conversational models. I mean, my perspective is that we hardly ever do original research. We try to take the research and make it easy to use.

17:24I think there's like hundreds of papers out here in this conference on interesting visualizations. And we try to look at what is real, like what are people really doing and using. And then we try to make it easy to surface those in our app. So like a simple loss curve is kind of table stakes. But, you know, we try to make it easy to generate that. And, you know, these days people have lots of different loss curves and the data gets high dimensional and then you want different visualizations there to try to see patterns and higher dimensional data. Also, you would be shocked by how many of our customers have trouble producing an example of an image that they're labeling or producing the actual text, literally the text that was like fed into their system because it gets tokenized, all these weird things happen to it.

18:10And then if you're just looking at the number, that doesn't tell you what's really going on. So just surfacing literally, okay, what was this text here? And what's my LLM predict is like the next set of tokens in a way that you can read. I think in a way that's the most important thing. Yeah. And that's for evaluation primarily? Evaluation. Yeah. I mean, yeah. Evaluation. Yeah, because I was asking again yesterday why somebody would need to go back and track some of this stuff. And you were saying for all kinds of compliance reasons. So why would you want to go back and look at the data your models trained on?

18:54So in industry, there are real compliance reasons. So, you know, if you're in the EU, a lot of our customers will get data from their users, but then the users have the right to take their data out. Right. So, you know, we'll have a customer that's trained on a billion images minus three. And they have to be able to show that those three are out of their data set that the model's trained on. And when there's fine tuning and like upstream models, as you'll really see in production, that actually is kind of a complicated thing to really prove that these three data points are removed from your set.

19:23That's like an industrial application. But in research, it's really important too, right? Like it's very common, I think, for test data to leak into training data. These days, like I often wonder with some of these models performances on standard benchmarks, there's so many ways that it could potentially leak into, you know, what the model was trained on. I think it's just best practice. Like, I mean, even if no regulator is telling you to check what data your model is trained on, it'll make you more efficient to know right and and you could write it down you know it's like not i'm not saying we're doing rocket science here you can write it down once you can write it down twice but you know the thousandth time you write it down you might forget to do it or you might get it wrong and so we're finding the 500th time out of a thousand hundredth time and i also think it um you know these files get really big and that also makes people a little bit lazy right Like, you know, like, you know, it's pretty frustrating, right?

20:19When you, you know, you have like, you know, like a petabyte file and then a few of those files change. So like having kind of an efficient way to version that, we built our own system because we didn't see a good system out there. We tried to use some of the open source stuff that I think works in different cases, but we really felt our customers were asking us to design something for the case where files are so big, they don't fit into a file system. So it could live on like an object store. So that's, that's, those are some of the things that we worked on to, to make data lineage work. And I do think people don't really understand the significance of it until you get in a situation where you can't produce the training data or run the GI.

21:00Right. You talked about, you know, excluding certain things from training data or taking them out. Yeah. That's something I've heard people talking about, you know, forgetting how to teach a language model to forget something that it's been trained on. How do you, I mean, I can understand excluding images from a supervised learning model, but from these large models that internalize the training data, is it possible to take that out once the model's been trained? This is what all these copyright cases are pointing to. I think that's a research topic that I'm not 100 % of the speed on, but I would be suspicious of anyone that says they can really make a model forget something.

22:01I mean, there's so many ways that it can kind of creep out. And we've seen so many ways to do that that I would really recommend starting with a model that didn't have it in there. I mean, I think one of my takeaways from this NeurIPS though, and kind of recent discussions with folks, is that there are so many great open source models coming out right now. You know, most of them are built on weights and biases, and most of them actually produce the weights and biases reports and lineages, the open source ones do. And so I think one service that we can offer is if you really want to go and dig in and see exactly the data that models were trained on, often there's a weights and biases reportable.

22:40Yeah. You mentioned the regulation in the EU. I mean, that's coming in the U.S. Do you think that you guys are poised to be like one of the primary tools for compliance if, you know, the regulation comes in and you have to produce your training data set and show that there's no copyrighted material, for example? in the training data set or no synthetic data in the training data set or if that's something you're worried about. Yeah, a great use case for weights and biases that we can produce the data that your model is trained on. I think the U.S. regulations right now are not clear enough to know exactly what you should produce.

23:28And so I worry a little bit, I worry a lot, honestly, that the U.S. regulations were just vague and scary and not prescriptive in terms of like what you could do to comply with it. So I wish we could say, hey, this is the report that'll make you in compliance. Right now, what I tell customers is these are the reports you could produce that are best practice. And there's no tension here. The stuff that you would want as a company to operate well is the same stuff I think that the government is going to eventually want. So right now we're just squarely on the side of what's useful for our customers.

24:07But I think that if the regulators are smart, that should be the same thing that they ask for. Yeah. You were also showing me the tracking experiments, which I mentioned to you, Determined AI. I did a podcast series with them a couple of years ago. I think they were bought by somebody. But we talked about tracking experiments and why that's important. The way I understood it is you're sort of absorbed in all this stuff, and you run a test, and you think, oh, I should tweak that, and it'll be better, and then you do that, and then five times later you realize, oh, it was better back then, but then you've got to go and figure out what everything, all the settings and everything.

25:11Is that what tracking experiments is for? And can you talk about how you guys do that and how you visualize it for users? Yeah. I mean, again, I think tracking experiments is exactly what you said. I think anyone that's done machine learning realizes that, unlike software, the key difference between doing machine learning and doing software is that in this non-deterministic world, you're really doing experiments, right? Like most of the stuff that you do is throwaway code, whereas most of the software there, right, it's throwaway code, maybe you could go and use it, but it's not like, you know, it didn't actually produce the intended result.

25:46And so when you, you know, when you do all these experiments, your real IP is these experiments, like it's things that you're like learning, right? If you're going to write a paper, that's like what you go back to to figure out like what to put in that table and a big thing for me like when i was starting the company was like yeah you look at these tables inside um research papers and you just want more information you know it's like they have like limited like ink space right and so you know it's like you know that they collected more metrics you know they tried more things and you really want to see all that so for research purposes i love it when people make their weights and biases experiments available so you can see like every single thing that they tried inside of companies.

26:24What I always say is, look, if you're not tracking your experiments, when somebody leaves, their IP mostly walks out the door. Like you might have the model they produced in the end, but that model is like a byproduct of a lot of learning. And for someone to improve that model, they're going to have to catch up on the previous work. So yeah, I think it's pretty much. Yeah. And we talked about, you said that you are sort of researching or experimenting with large language model interface or layer, and there's so much being done in using large models to analyze complex systems. Could you have a model that rather than having a human sort of go through the archive or whatever you call it of your experiments and sort of looking at what changed and that sort of thing.

27:23Could you have a large model sort of flash through all that and say, oh, you know, your best performance came from this model because you had this. I mean, I'm obviously not a machine learning engineer, but yeah. Yeah, potentially. I mean, in a way, hyperparameter sweeps are sort of like a simple mathy version of that. But I do think that today, one key part of being a machine learning engineer is researching what's going on in your models. And so I think outsourcing that completely to an LLM is probably dangerous because that is like your job. So you're really automating away this core thing that you're doing, which could be effective.

28:14But I don't think that the models are good enough today to do that. But certainly we're experimenting with ways to do that. Yeah. So as a company, you built these products. How is it evolving? I mean, are you adopting the products as the research evolves so that they can fit more closely with what people are doing? Or are you creating new products? Well, it's evolving in that I think the biggest effect, there's two big things happening. So one major thing is that this conference, NeurIPS, used to be totally research. There really were not applications. There are not a lot of companies coming here.

29:02Now you see tons of executives coming here. And the reason is that this research is just incredibly applicable to so many applications. You would just not believe how many industries are touched by this today. so most big companies at this point have a machine learning team that's trying to get models into production and so that you know makes us you know maybe orient a little bit more towards compliance a little bit more towards high-scale development a little bit more towards things like monitoring in production you know because because research is sort of just get the paper out and i'm like what i'm passionate about is actually making these things work so i love you know working with like you know procter and gamble to make the medicine taste better or you know working with like a farming company to make sure that, you know, they're using less pesticides because they can see exactly where the weeds are.

29:47I mean, I love these applications. I get really excited about it. And so, you know, as you go into bigger businesses, you know, you get more complicated security controls and all these things, right, that companies care about. That's one big effect. And that's, you know, been an evolution of our business. The other thing is more like a shock, which is that the way people do machine learning has changed a lot in the last year or two, because now many things can be done with what they call prompt engineering. I think at this point, it's like a little bit of fine tuning and prompt engineering, a little bit of agents, you know?

30:18But that's like a really new workflow, still experimental. But what is the unit of experiment when you're trying to build an agent? I think it's actually a little different than modeling. So that's a place where we've been kind of rapidly updating our software to accommodate these cool new use cases. Yeah, those are the two big things. And on prompt and engineering, does each time you tweak the prompt count as an experiment? Well, it's a good question, but yeah, I think it does. I mean, I think your model essentially with these prompts is, yeah, each time you tweak the prompt. But in the real world, people don't usually use one prompt.

Read the full transcript

30:57It's like multiple prompts chained together or an agent that's even like automatically creating prompts depending on what the results are. So what's an experiment is a complicated question. What's a model even is a complicated question in this world. But from weights and biases point of view, if someone is running different prompts to try and get a certain result, you're tracking that? Yeah, we track it. Yeah. And you mentioned monitoring. So weights and biases, is it attached to different models or do you run it kind of in the background and for monitoring? And I'd like to hear a little bit about why monitoring is important.

31:53Again, I know on a high level why, but I don't, I've never really understood why model drift happens and things like that. is the product something that's embedded then in a model or in a productized model so that it's always monitoring performance and that sort of thing? Yeah. So the way it works is you instrument your code with a few lines of weights and biases code. So you import a library into the code that runs your model or trains your model. And then what happens is that model is saving data locally on your computer and files, right? Because we never want to crash your training or your model running, right?

32:39So file system is very stable, right? So we really are saving everything and caching it locally. But the amount that you're monitoring often is bigger than a file system can handle. So in the background, it's also streaming that to a central server. And for most academic use cases, they use Dolby Vita AI that you can just log into. Simple, right? Everyone can go to our website, you know, and see what's going on. Many companies want to run, you know, in their own kind of private VPN. So then they have their own URL where the stuff is streaming to. And you ask about like, well, why does data drift happen?

33:15And I actually think data drift is fun for academics to talk about. I think in the real world, the problems are even more basic, but you see them all the time. And it's because models are very sensitive to inputs and inputs can change in a lot of ways that would be surprising. Right. So, you know, models are typically trained on data inputs upstream of them. And, you know, if something just switches where like all the output is zero, you typically don't. The problem with models is they don't give you like an error message or warning. Typically, they just start behaving weird. And so I think that in the real world, the most common source of problems in production is something upstream changed.

34:01Even a lot of times models are trained on upstream models. That's really common. When I worked even 20 years ago at Yahoo, we had a search ranking algorithm. It was like, okay, is this web page relevant to this query? And I tried to test that. And then we figured out, okay, we also have this team that's working on spam. Because there's spammy websites trying to sell you a bunch of junk. And it turned out, okay, that spam score is useful to know if something's relevant, even if it's not like Mark's spam, if it's like a little spammy, that's relevant. So one day the spam team updates their model, right?

34:30They make it better. And then boom, you know, suddenly these spam scores are like higher than they used to be. And our model silently freaks out and thinks that all these results are spammy and the ranking goes haywire. And it's nothing to do with the ranking model that we had. It just was sensitive to that. and the spam team didn't know that we were consumers of that. That happens all the time in production, like every day, right? And so that's kind of one failure mode, but it's often almost always kind of human error versus these sort of more sophisticated things that academics like to study.

35:06But it is true that language does change more than we realize over time. And, you know, like a few years ago, there'd be a lot less emojis in a typical text corpus. and I think our eyes just sort of scan over that and we don't really notice the emojis. But I think from a computer's perspective, might be like, what is going on here? You know what I mean? There's all these new characters that I've never been trained on and haven't seen before. And, you know, new slang can freak things out. I don't know if like LLMs are more robust to it, but I sort of suspect not. I think our brains are just so good at like synthesizing kind of new information and adapting to it.

35:36And remember these models are typically static, right? I mean, they almost never, you know, retrain themselves. So that can be counterintuitive, I think, to downstream users. Yeah, that's interesting. And so on the monitoring side, what are you monitoring? I mean, basic stuff, right? Just like, you know, is it producing a similar output distribution to what it had in the past? Like sometimes, you know, smart companies will have like a test set that they just keep running, you know, the model on. Just make sure like it keeps, you know, working at a certain level. Also making sure that the input data, you know, isn't doing something weird.

36:12Again, there's so many different cases that, you know, weights and biases tends to not have a super strong point of view on how to do things. Like we'll sort of suggest like here's what we think are best practices and try to make the common paths easier for you. But the truth is that I think everyone in different applications should be doing monitoring differently. Like different kinds of errors are different levels of bad. I mean, it's even a kind of a complicated question, you know, for a autonomous vehicle company, like, you know, mistaking like a bicyclist versus mistaking a dog versus a baby carriage versus a sidewalk.

36:45You know, we probably have an intuitive ranking of like, you know, how we think about those errors. But you have to ultimately kind of put that in a loss function and put coefficients against how much pain you want to assign to each one. It's tough to do. Yeah. So the monitoring function of weights and biases, is that a separate model? You were saying that you have kind of a suite of products. So is there like a monitoring product and how is it all bundled or do people pick and choose what they want? And then how are they? Is it a subscription model? Yeah. So subscription model, we give a lot of way free to academics.

37:29So most academics don't pay us. You have to really be hammering our servers before you get to a state where we charge academics, but some do. So, you know, occasionally that happens. We, you know, today we have a model where we charge on a bundle. We charge for the whole bundle and we say, you know, you can pick and choose what parts are useful to you because it's so painful to sort of like gate people against the paywall each time they want to add more of a functionality. But, you know, that could change. I mean, we just try to match our pricing as best we can to the value that we're providing for companies.

38:04Yeah. And on the monitoring specifically, so you suggest these different ways to monitor, different things to monitor, and then the user tailors it to their model? Or does weights and biases sort of look at the problem and decide how to monitor the model? It's super simple. I mean, when I say suggesting to monitor, I mean, like literally in our documentation, we have suggestions on how to hook things up. So you're basically setting this up, doing the logging, you know, setting up alerts or graphs based on your needs. Is there something I haven't covered that you want to talk about? No, I don't think so.

38:47Yeah. Well, let me ask this. You've been through at least one exit, right? Yeah. Obviously, you're passionate about weights and biases, but I mentioned Determined and I know some of their other MLOps companies that I've talked to and they end up getting acquired.

39:16How do you feel about this? And there's so much happening and you do you have like, I'm occupied with this, but what I see this opportunity and, you know, I'd really like to go after that. Well, look, I've never wanted to be a serial entrepreneur. So I really, really hope that this is my last company. I mean, I think through the decisions we make with an eye towards the long term. I mean, I've done this long enough now that I know that it's dangerous to say we'll never be acquired. You know, we've raised VC funding and ultimately they will want an exit and the return. But my strong hope is that that comes from an IPO or a way that we can stay independent.

40:03I think that it's a big enough idea that it should be an independent company. And I care a lot about our ability to integrate with all the different things that our customers are using. So it'd be sad if I think one of the big benefits of us is, look, we work with your AWS models. We work with your GCP models. If you use Snowflake, we integrate with that. If you use Databricks, we integrate with that. We work with whatever you're doing. And so I would worry if we were acquired that we might lose that ability to operate independently. AI might be the most important new computer technology ever.

40:36It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds.

41:26If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, Take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. Oracle.com slash IonAI. That's oracle.com slash IonAI. That's it for this episode. I want to thank Lucas for his time. If you want to read a transcript of today's conversation, you can find one on our website, IonAI. That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but A-I is changing our world. So pay attention.

From the publisher

This episode is sponsored by Oracle. AI is revolutionizing industries, but needs power without breaking the bank. Enter Oracle Cloud Infrastructure (OCI): the one-stop platform for all your AI needs, with 4-8x the bandwidth of other clouds. Train AI models faster and at half the cost. Be ahead like Uber and Cohere.

If you want to do more and spend less like Uber, 8x8, and Databricks Mosaic - take a free test drive of OCI at https://oracle.com/eyeonai



In this episode of the Eye on AI podcast, join us as we sit down with Lukas Biewald, CEO & co-founder of Weights & Biases, the AI developer platform with tools for training models, fine-tuning models, and leveraging foundation models. 

Lukas takes us through his journey, from his early days at Stanford and his work in natural language processing, to the founding of CrowdFlower and its evolution into a major player in data annotation. He shares the insights that led him to start Weights and Biases, aiming to provide comprehensive tools for the entire machine learning workflow.

Lukas discusses the importance of high-quality data annotation, the shift in AI applications, and the role of reinforcement learning with human feedback (RLHF) in refining large models.

Discover how Weights and Biases helps ML practitioners with data lineage and compliance, ensuring that models are trained on the right data and adhere to regulatory standards. Lukas also highlights the significance of tracking and visualizing experiments, retaining intellectual property, and evolving the company's products to meet industry needs.

Tune in to gain valuable insights into the world of ML Ops, data annotation, and the critical tools that support machine learning practitioners in deploying reliable models.

Don't forget to like, subscribe, and hit the notification bell for more on groundbreaking AI technologies.



Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI



(00:00) Preview and Intro

(01:39) Lukas's Background and Career 

(04:09) Founding CrowdFlower and Early Machine Learning 

(06:59) Current Trends in Machine Learning

(08:46) Reinforcement Learning with Human Feedback (RLHF) 

(12:43) Weights and Biases: Origin and Mission  

(16:44) Visualizations and Compliance in AI 

(22:43) US vs. EU AI Regulations

(25:20) Importance of Experiment Tracking in ML 

(28:47) Evolving Products to Meet Industry Needs 

(30:38) Prompt Engineering in Modern AI 

(33:34) Challenges in Monitoring AI Models 

(37:25) Monitoring Functions of Weights and Biases

(39:33) Future of Weights and Biases

More from Eye On A.I.

All 266 episodes
#192 Lukas Biewald: How Weights and Biases Supercharges Machine LearningEye On A.I. · 42 min
Listen in VO