In short
Eye On A.I. Podcast: Episode #261 - Jonathan Frankle: How Databricks is Disrupting AI Model Training
Podcast Overview Host: Craig S. Smith Description: Eye on A.I. features discussions with individuals making significant advancements in artificial intelligence, contextualizing incremental progress against broader global implications.
Episode Description In this episode, Jonathan Frankle, the Chief Scientist at Databricks and co-founder of MosaicML, discusses TAO (Test-time Adaptive Optimization), Databricks' innovative tuning method that enhances how enterprises develop and scale large language models (LLMs). Jonathan shares insights on how TAO leverages reinforcement learning and synthetic data to optimize models without relying on extensive labeled data.
---
Key Discussion Points
Introduction to Jonathan Frankle
- Chief Scientist at Databricks.
- Co-founded MosaicML, focusing on AI model training.
- Background in making neural network training more efficient during his PhD studies.
Databricks and Data Management
- Databricks is referred to as a "data lake house," combining data lakes and data warehouses.
- Emphasizes the intrinsic link between data and AI, where effective data management is crucial for understanding and utilizing data.
- Focus on "data intelligence," enabling users to interact with unstructured and structured data without needing extensive technical knowledge.
TAO
Test-time Adaptive Optimization
- A synthetic data generation technique allowing for model tuning without the need for labeled data.
- Users provide prompts from which synthetic data is generated, integrating this back into the model iteratively.
- Addresses a significant pain point in AI deployment: the challenge of obtaining high-quality labeled data.
Comparison to Traditional Fine-tuning
- Traditional supervised fine-tuning requires meticulously labeled data, which can be resource-intensive.
- TAO simplifies this by reducing the need for labeled datasets while still achieving competitive performance.
- Results showed TAO could outperform traditional methods even without traditional labeling.
Reinforcement Learning and DBRM
- TAO employs reinforcement learning through a model called DBRM (Databricks Reward Model) that evaluates output quality.
- Focus on obtaining signals regarding output effectiveness rather than strict right or wrong answers, making it adaptable for fuzzy tasks (e.g., summarization, question-answering).
Continuous Improvement of Models
- Models can be continually improved through user interactions and prompts gathered during deployment.
- Users can refine their models based on real-world input, allowing for responsive and adaptive AI systems.
Cost and Efficiency
- TAO reduces the computational and human costs associated with traditional model training.
- Provides a framework for organizations to deploy AI solutions more quickly and efficiently.
---
Key Takeaways
- Innovative Approach: TAO represents a significant shift in how organizations can train and deploy AI models without the burdens of extensive data labeling.
- Reinforcement Learning: Utilizing reinforcement learning techniques can enhance AI performance in a way that is both effective and adaptable.
- User-Centric Design: The focus on enabling users to provide inputs reflects a broader trend in AI toward democratizing technology access and application.
- Future Applications: Ongoing developments aim to enhance reasoning capabilities of models, further improving their applicability across various domains.
Conclusion The conversation sheds light on the evolving landscape of AI model training, advocating for approaches that prioritize efficiency and accessibility. The insights shared by Jonathan Frankle indicate a promising future for enterprise AI capabilities, particularly in leveraging user-generated data and advanced modeling techniques.
---
Additional Resources
- Explore Databricks innovations: [Data + AI Summit 2025](https://www.databricks.com/events/dataaisummit-2025-announcements)
- Follow Craig Smith on [X](https://x.com/craigss) and Eye on A.I. on [X](https://x.com/EyeOn_AI).
- Try Oracle Cloud Infrastructure for AI projects: [Oracle](http://oracle.com/eyeonai).
---
This episode is a must-listen for AI researchers, enterprise leaders, and anyone interested in the future of model customization and deployment in artificial intelligence.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00The idea is that if you really want to be reductive about it, this is a synthetic data generation and usage technique. where you give us the prompts, we generate synthetic data based on those prompts, and then we integrate that synthetic data back into the model. And then you can repeat this process many times with more prompts as you collect them over time. And the nice thing about synthetic data generation at training time is like, I usually make recommendations to customers to be kind of as conservative as possible when they're doing something with AI. They're already doing something with AI and that's an aggressive choice to make.
0:30We're still learning how to do that. So, you know, don't push your luck in some sense. So what I typically tell people is don't deploy something that you haven't tested. And so my usual recommendation is deploy the model, you know, collect the data, and then redo TAU on all the data you've collected so far. Test that model. And then, you know, even consider A-B testing it to make sure that, you know, for whatever metrics you're looking at in production, this model isn't going to make things worse, and then deploy it. In business, they say you can have better, cheaper, or faster. You only get to pick two.
1:00What if you could have all three at the same time? That's exactly what Coher, Thompson Reuters, and Specialized Bikes have, since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 % less for compute, 70 % less for storage, and 80 % less for networking.
1:48Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com slash IonAI. IonAI, all run together, E-Y-E-O-N-A-I. That's oracle.com slash IonAI. But why don't you begin by introducing yourself? Go ahead. Definitely. So I'm Jonathan Frankel. I'm chief AI scientist here at Databricks. I oversee a lot of our research efforts on trying to make AI more useful for our customers and understanding how AI and data interact in a productive way.
2:39I got here. Databricks acquired my startup Mosaic ML, I guess, about a year and a half ago now. We were kind of dedicated to helping everybody train their own AI models from scratch. Founded kind of out of my PhD along with, you know, co-founders Naveen Rao and Hanlon Tang and Michael Carbon. And before that, I was a PhD student studying how to make neural network training more efficient. So it's been a wild journey over the past few years. But, you know, if you're looking for one through line, I care a lot about making sure everybody can use and control this technology and do with it what they want.
3:14I think a lot about kind of the internet where everybody could have a website. And that led to just such an explosion of creativity, a lot of really bad ideas and a few absolutely brilliant ideas. And, you know, if I told you from the beginning, oh, why don't we create an encyclopedia online that anyone can edit and contribute to? You'd say definitely one of the bad ideas. And it turns out to change the world. So the idea that AI is similar and everybody kind of has the chance to do creative things with it to solve their problems, it'll change the world again. Yeah, absolutely. I agree, which is why I have this podcast.
3:52So I was reading, and at Databricks, what's your, I mean, describe for listeners what Databricks does, first of all. I mean, I've talked, I think, before on the podcast about the data lakes and or data. What's it called? Yeah, data lakes. Isn't that the term? We're the data lake house. Right. Exactly. Data lake house. That's right. And what that means and how what you're doing on tuning models fits in that. Yeah, so I think the Databricks journey actually from the very beginning, you know, the stories have lived in legend for me as someone who's watched Databricks from afar. And now I get to be a part of it and hear all the inside stories.
4:42But the story began kind of where it is today. People have lots of data. They need a way to manage it and they need a way to understand it. And kind of data and AI are intrinsically linked to each other. Because what do you do with a bunch of data? You understand it. And Databricks is, as far as I can tell, looking through the history books, always been about using whatever the latest AI was to understand your data. Now, that AI wasn't always, you know, large language models or neural networks or anything like that. But and machine learning have been fundamental because what do you do with your data?
5:15You don't just, you know, stick it in a data lake house and just let it sit there collecting dust. You want to understand it. It's only useful. It's only worth having there if you can do something useful with it. So this is kind of this concept we like to refer to these days as data intelligence, the idea that, you know, you want to be able to talk to your data, understand it, figure out what's there, interact with it, learn what it's trying to tell you, kind of make effective use of it. So, you know, we offer ways to manage both unstructured data like documents and structured data. You can think of that as like, you know, SQL or kind of individual entries into a database.
5:50and a lot of the work these days is on the question of how do we use all the cool modern AI that's out there to give people new ways to understand it and to give new people ways of understanding it there are a lot of people who in yesteryear would have had to like write a notebook and write Python and do a SQL query or write in Scala or Spark to be able to get insights from their data and today I think AI is going to make it possible for pretty much everybody to do this I have no SQL, but I've been able to do this for my data at Databricks. Yeah, and the data lake, lake house concept, it really started during the supervised learning phase of AI when everyone was labeling data and training models with labeled data, was largely sort of text reading models or image reading models.
6:50And you came in after the generative AI revolution where suddenly really those silos are breaking down. I mean, at the time during the supervised learning phase, I used to talk to people about, you know, you train your data, but there are all these pools or lakes, some data's around, and sooner or later they're going to merge and we're going to have an ocean of data, a lot of it labeled. But the science kind of overtook that. You don't need to do that any longer. I wanted to talk to you about tau, which I've got to call it up.
7:44Let me see here. Tau, tau. So test time adoptive optimization, which I found fascinating. I read your paper. before I start asking about it, what really fascinates me is it looks like it's a way to create an ever-improving model, at least in narrow domains, which is something that I haven't seen before. Maybe it exists and I just haven't stumbled on it. But But first of all, can you talk about training and tuning? How you, is that what, I mean, I guess that's what Mosaic was doing, right? And that's what Databricks bought you for. But is, what's, talk a little bit about your role in that at Databricks.
8:43Yeah. So I think it's worth taking a step back and just asking, you know, yeah, why am I even here? Like, what is my purpose and what do I contribute? You know, I'm an AI researcher. every day yeah yeah we're trying to solve real problems for customers like why do you need a guy like me around databricks and you know my best answer is there are a lot of cool things happening in the research world that are clearly very valuable somehow and there are a lot of important problems our customers have that they need to get solved and they really wish they had the magic to do it and my job is to connect the dots between these and tau is one of those examples so you know customers want to get their data and their models to interact with each other.
9:26They want to be able to bring all that data to their models and have the models be able to speak fluently or answer questions about their model or reason in whatever domain they care about. And there are lots of ways of doing that. And those ways are things like you can prompt or do rag to kind of bring your data in. You could fine tune a model on your data. There are lots of ways to do this. You could even have a model write calls to a SQL database. And fine tuning is probably the most powerful of these techniques. It lets you literally change the model and teach it more about your data directly and intrinsically, as opposed to having the model kind of start fresh every time you interact with it.
10:04And fine tuning is a bit of a tricky thing to do for a lot of reasons. But one of the reasons we saw again and again was that fine tuning is really demanding in terms of data. You basically, to fine tune a model to do something, you need examples of doing that exact thing with your data. So if you want a model to serve as a customer service bot, you need to have examples of user asking question and what the customer service agent should say in return. And you need thousands of those examples. And what we found was that people have lots of incredible data. It's very rare that they magically have the perfect fine tuning set.
10:42That's just not the kind of data that gets created in nature. In nature, you have lots of documents that accumulate and you have, you know, some telemetry data and you have some user inputs and you have, you know, some manuals on how to respond or some guides or some decision trees or whatever else you have floating around, plus a SQL database of relevant information. But in nature, you don't have this perfect, pristine data showing what an interaction should look like with an LLM. It's just a very, it's a hard, it's a lot to ask of someone. And, you know, the correct answer when you come to a customer and they say, I really want an LLM that's fine-tuned for my task is not, well, go annotate 10 ,000 examples and come back to me when you're done.
11:22That's not an adequate answer. And I think to date, that's been the answer that the research community has had for all of our customers. Just, you know, it's your problem. You didn't give me the right data. And I think that's a bad answer. That's not an answer that's going to lead to people successfully deploying AI. Nobody has time for that. Nobody has resources for that. By the time they do that, the field will have moved on or their business will have moved on. And so the question we asked ourselves with Tao was kind of, on the one hand, what can we do to simplify this? Like, is there something we can take away or something we can, some constraint we can remove that will make this easier?
11:56And on the flip side, does the technology exist to actually do anything useful with that? If I have the customer give me just inputs, but no outputs for how the LLM should behave? Well, do I have science that can help me to actually tune a model on that? Or that it sounds kind of unreasonable, like you have no idea what someone actually wants the LLM to do. You know, there could be many good responses to one input, depending on what the user wants. And so my job, in a sense, is to move back and forth between these worlds. On the one hand, I'm always asking, like, what's making our customers' lives difficult, or what did they wish they could do?
12:32And if I had a magic wand and I could make their life a little easier, what would I do to make that easier? And on the other hand, okay, my magic wand is really not magic, it's science. What techniques exist? What techniques can I modify? What can I build that will bridge these gaps? So half my time, I feel like I'm a researcher. Half my time, I feel like I'm a product manager. I actually spend a huge amount of time with our customers, just talking to them about what they're trying to do and understanding their problems, not the ones in abstract that I think many of us think people have, but really where they're stuck.
13:03And if only I could remove this one thing. So with Tao, that's exactly what happened. We said, getting labeled data is really hard. How can we simplify this? And there are lots of ways you could simplify it. You could ask someone to write a detailed rubric of how something should behave, or write a bunch of judges or graders that tell you for any arbitrary output, how good is it? It turns out those things are also really hard. And it would be nice if we could just simplify it further. You don't even need any labels at all. That means that all you have to do is show up with a bunch of inputs reflecting what people might ask your LLM.
13:37And you just get an LLM that's good at those inputs. Now, the nice thing about that is it's really easy to get inputs. Just set up a really basic LLM. You can set up Llama, give it a really basic prompt, and ask your friends or your teammates or yourself to just interact with it and ask it a bunch of questions. It's probably going to do a bad job. But you don't care whether it gives you a good answer. You just care that you've asked it a good question. And so this is great. This means that it's really easy to get the training data because it's something that people can naturally provide the system.
14:06There's no real work involved. We're not asking someone to sit down and literally write 10 ,000 questions. Just put this in the wild. Beta test it. Prototype it. Ask your team to do an hour hackathon and just ask it all the questions that they might want to. Pretend to be the customer. Ask a customer to prototype it with you. and you can collect the data you need. Then you get to the really tricky part, which is, what the hell do you do with that data? Because it sounds, I mean, I hope I'm not like being overly dramatic, but it sounds kind of insane. Like all I know are the inputs, and you're going to tune a model that will give me the outputs I want.
14:38That's like, what's going on? And, you know, this is a place where we turn to techniques from reinforcement learning. And reinforcement learning is a set of techniques where rather than giving yourself inputs and good outputs, you give yourself inputs and just an environment where a model can explore. You can think of this like, you know, when, um, when, you know, people have built models for go, for example, deep mind built this model that was really good at go or game playing, let it explore the environment and figure out for itself. You need to have an environment it can explore and some kind of signal that tells you what outcomes are good and which outcomes aren't.
15:13This is why people love games so much because you just, you explore the environment, you know, when you're scoring points, you know, and you're winning and, you know, you get good signal. And here, what we did is the environment is in some sense, this special model we call DBRM or the Databricks reward model. This is a type of model that's very common in the reinforcement learning world. And all it's meant to do is tell you how good is a particular output. And, you know, it's, that sounds very generic, like just generically, how good is it? But the way you train this is you You show it lots of examples of, you know, pairs of outputs for a given input and which one a human liked more.
15:50And you just show it tons and tons of this data. And it turns out it actually generalizes pretty well. So, you know, you combine that with a bunch of inputs you have, do some synthetic data generation and kind of filter it with this reward model, use the right techniques to train the model after that. You know, a lot of, and this sounds really easy, a lot of very careful work on the part of the scientists who worked on this to just get the details right. And it turns out that it actually, there are very reasonable choices you can make a lot of the time. If I showed you two outputs and I showed you the input, you'll probably have a pretty good intuition as to which one is better.
16:25You won't know which one is right. But as long as you can kind of keep steering things in the right direction, the model will get better and better. And that's what we saw. Not only did this actually lead to improvements, even though you had no examples of how the model should behave, but it even beat training on labeled data, which is kind of crazy. And I can say a bit more about that at some point if you'd like, but it's, you know, it's not as crazy as it sounds, I think, but it was just very surprising to us. And this worked across a pretty wide range of tasks, like document question answering and SQL, which again was very surprising to me that a reward model could kind of generalize in that way.
17:00It's, it's really exciting for the future. Maybe enterprise tasks are more similar than we think. yeah i mean but but supervise fine-tuning you need these question answer pairs in in uh for uh a language model uh and what you're talking about you you still need question answer pairs right but you're getting them through reinforcement learning with human feedback i i wouldn't call them question answer pairs here because I think it's it ends up being more complicated than that I would call it questions and kind of steering on which answers are better or worse because it's not really none of them are there's nothing in tau that ever checks is the answer right or even is the answer good there's just asking is the answer better and which answers are better and trying to continually move the model toward better answers, which is very, I don't know, I found that to be a little bit mind bending when I started this project.
18:03I have a bunch of reinforcement learning kind of enthusiasts on my team and specialists on my team who have been, you know, spouting the gospel of reinforcement learning for as long as they've been here. And I've just been politely ignoring them and patting them on the head. And, you know, I've learned actually maybe there's something to what they do and I should listen to them a little bit more. Yeah. So in this process, you do use RLHF? That's correct. Yeah. And then you, once you have that, that data, then you use pure reinforcement learning. That's correct. So we're kind of using, I would think of this as techniques from the RLHF literature and just kind of combining them in the right way.
18:42I won't claim that there's anything we invented that's mind bending from like a scientific perspective. I, you know, I think there were, when we released this, there were some folks on Twitter who were like, this doesn't sound very novel. And my response was, it works. And getting it to work is the hard part. Like that's the novelty is easy. Like getting stuff to work for real customers is hard. So the details of like, okay, we're going to take some inputs. We're going to generate in the right way. We're going to select in the right way, build the right reward model. All that stuff took time. This reinforcement learning step after you've gathered the data is happening at test time.
19:18Is that right? It's very similar to DeepSeq and this test time compute training that has become kind of a paradigm for refining models. but it's different in that you're only doing it during the tuning phase. You don't do it during the inference phase. Can you talk about why that's an advantage? And what really fascinates me is unlike test time compute training is temporary, It improves the inference accuracy. But once you're done with that inference, the model reverts to its training state and you lose that improvement. But in your case, it continually improves the model. Can you talk about that?
20:29Yeah, yeah. So I think there was actually enormous debate internally about whether we should use the word test time. because it's not really it's like using test time compute for training and so is it really test time there was kind of you know there there was some debate over this word so the way i think about it is even ignoring the word test time the idea is that you can if you really want to be reductive about it this is a synthetic data generation and usage technique where you give us the prompts we generate synthetic data based on those prompts and then we integrate that synthetic data back into the model and then you can repeat this process many times with more prompts as you collect them over time and the nice thing about synthetic data generation at training time is like well let's take a step back if you were to just use let's say o1 or r1 which are both excellent models you're going to use kind of an uncertain amount of compute at inference time you don't know how much is going to get used.
21:28The model may explore for a while. It kind of, you know, there's some uncertainty there. And, you know, that can get expensive as well. We saw kind of in practice, 01 and 03, we're using about five times as many tokens somewhere around then to solve problems, which is not unexpected. They're designed to do that. That's the whole point. That's why they're so powerful. And one other way of looking at this is what if you use all that extra compute to generate the synthetic data. And then you create a standard model that when you actually go to use it at inference time, doesn't use any extra compute, just behaves normally.
22:04And the idea here is we expect that for our customers who are using this technique, if they have a bunch of inputs already, they've probably already got an application deployed. They at least have some amount of confidence that it's worth doing this training. And so So they're willing to invest a little more in training to make inference a lot cheaper. So the idea is, OK, let's spend some more compute at training time to generate better synthetic data. And then let's generate a normal model that when we use it at inference time has exactly the same inference cost as, you know, Lama had we not done Tau.
22:37So it's a bit of kind of a it's taking this test time compute paradigm and flipping it on its head a little bit. The test time compute is to help you build the training set. But at the end of the day, you're training a model that you use normally. And we think this is kind of the right mix for customers who are deploying applications at high volume. If you're doing something like Tau, you're trying to cut down on the cost of inference time anyway. Often you're trying to train a custom model to replace something else that you might have. So let's really lean into the fact that we can help you reduce your inference cost and a little more upfront investment, probably worth it to you.
23:11Yeah. And then once you're in inference mode, you're collecting more queries, more prompts that then you can use in tuning mode, just the tuning time compute. uh does do you periodically go back and retune the model on the data that's built up or is there a way of doing that in the background even if the model is being used to uh for inference or or do you split it into a uh two instances of the model and one you're improving through this tune time compute, the others doing inference out in the wild, and then you swap the models periodically. I mean, how does that work? It's a great question. It's something we're still kind of experimenting with customers on just to kind of get it in exactly the right shape.
24:17I usually make recommendations to customers to be kind of as conservative as possible when they're doing something with AI. They're already doing something with AI, and that's an aggressive choice to make. We're still learning how to do that. So, you know, don't push your luck in some sense. So what I typically tell people is don't deploy something that you haven't tested. And so my usual recommendation is deploy the model, you know, collect the data, and then redo tau on all the data you've collected so far, test that model, and then, you know, even consider A-B testing it to make sure that, you know, for whatever metrics you're looking at in production, this model isn't going to make things worse, and then deploy it.
24:53And the idea there is, you know, first of all, it's just playing it safe. And, you know, again, you're making a very aggressive choice to use AI at all right now. It's exciting. But, you know, be careful. And the other piece is that you don't necessarily want to just want to train a model that you've already done Tau on, on more recent data. You can end up in this interesting place where the model will kind of get extra good at the latest stuff and may forget things because you're emphasizing the latest data. Sometimes that makes a ton of sense. If you're in a chatbot application, let's say for retail, you may actually want to use the latest data because, you know, you don't want it to be talking about Valentine's Day when it's Thanksgiving.
25:31Sometimes you do want to mix all the data together just to, you know, make sure that you're fully representing the earlier stuff. So there are, this kind of gets into more application specific considerations where like, I'm guessing the big recommender systems powering like ad placement at Google and Facebook. I'm guessing those are actually trained mostly on most recent data because human behavior changes so quickly. But for a lot of our enterprise applications, taking things one step at a time and kind of combining everything together, again, it's the safest possible choice to make. And then you can experiment and try lots of different things.
26:01But I wouldn't recommend kind of trying to update in real time, especially without testing the models first, just to be extra safe. Yeah. And so you mentioned sort of overriding previous parameters, which has been the problem that restricts continuous learning, is catastrophic forgetting. So yours doesn't get around the catastrophic forgetting problem. No, I wouldn't make that claim. I would say very narrowly, this gets around the need for labeled data, or at least assists you by reducing the need for labeled data. When it comes to catastrophic forgetting, it still puts you in the same trade-off that you're always in of like, do you train on the latest data and hope that the model doesn't forget?
26:52Do you combine all the data together? Do you upweight more recent data? I don't think we have an amazing answer for catastrophic forgetting. I will say forgetting is a lot less catastrophic than it used to be. Back in my day, five whole years ago, catastrophic forgetting was a very severe problem in general when you did fine tuning. And these days, models are a little more robust to that just because they're so much bigger and they're able to store a lot more knowledge and also because we have better techniques for managing catastrophic forgetting. To give you one example, there was a paper we worked on last year as a lab.
Read the full transcript
27:25It's called Laura Learns Less and Forgets Less. It's using Laura, you know, this parameter efficient fine-tuning technique where you don't train all your parameters. People used to think of this as like a compromise. You're doing this for the sake of efficiency because you can't afford to do full fine-tuning, but you're going to get worse results. What we found a lot over the past year has been actually, Laura is a regularization technique as much as anything else. It learns less and forgets less. It's a way to reduce overfitting and a way to reduce forgetting. Now there's a tradeoff there. Sometimes it'll, you know, learn too little and, you know, forget or learn too much and forget too much.
28:02Sometimes it'll learn too little and forget too little. There's kind of a balance. but Laura has actually been a really interesting technique just to manage learning and forgetting, or at least to give you more opportunities to explore the trade-off between the two. Deeply surprising to us, but it actually, one of the most common scenarios I see Laura is for customers who come to me and they say, I'm overfitting and I don't know what to do about it. And I just say, use Laura. It sounds counterintuitive, but it works great. And just for listeners, so we understand the process, you take a model in the old days, or the the current paradigm you take let's say an open source model uh and you want to fine-tune it so you get a bunch of uh label data and you uh fine-tune the model on that label data and and then you release the model in the wild and and then the reasoning models are allowed at test time to run further on the data or explore the data.
29:09Or maybe you can help me with my vocabulary there. The way I would describe it to folks who are a little bit familiar, at least, is it's automated chain of thought. You're kind of letting the model figure out its own chain of thought rather than guiding it through it. But you're kind of saying to the model, it's totally okay for you to spend some time thinking and exploring and considering and doing chain of thought. get back to me when you're ready. Whereas in the more traditional paradigm, we basically ask the model for the answer right away and hope that just doing one pass through the model is enough for it to figure things out.
29:40Yeah. But in that case, the improvement that happens at inference time is temporary. And what strikes me about Tau is that during tuning time, you're doing doing something similar in that you're uh you're gathering uh data from from inference time and using it to tune the the the model uh uh fine-tune the model of using reinforcement learning uh but But those changes, like supervised reinforcement learning, they change the weights of the underlying model so you don't lose that improvement. Is that right? That's right. And I think there's, I'd add a little more nuance to this, which is that, you know, how did OpenAI train O1 in the first place?
30:43Or how did DeepSeq train R1? Well, what they did is they let the model think, kind of, they did the same thing of kind of like tune time compute, if you want to call it that, where the model was given the chance to think and spend a lot of time coming to the right answer. And what they did is they had a way to check whether the answer was correct or incorrect. They threw out all the cases where the model thought for a long time and the answer was incorrect. And that became their new training data. And so they trained a model that knew how to kind of reason and do its own automated chain of thought.
31:13I think one of the big differences here is for our customer tasks, we have no idea how to determine what answer is right versus what answer is wrong. For DeepSeek or for OpenAI, they're focused on domains like math and coding where you can't actually check the answer. There's a math equation. It leads to the right answer. There's a proof. You can check whether it makes sense. There's a program and you can check whether it executes and has the right behavior. Whereas for a document question answering task or a summarization task, there is no right and wrong. And so instead, we're looking for better and worse.
31:49That's kind of an if you want to look for a key difference here. The other part is we're not trying to train the model to actually spit out its whole thinking trace. We're just trying to train a model to get right and wrong answers. But we're trying to generalize this to basically any domain that a customer might want to address, which is, you know, they include a lot of hard, really fuzzy problems, like, you know, answering questions about health, where there's no one right answer necessarily, or no one right way to give an answer. And so we have to have something a little more general to help with that.
32:19The other piece is that we expect the user to bring their own prompts. And if they bring their own prompts, we can kind of automatically specialize the model for a domain. Whereas in the DeepSeek case, they kind of have to generate a lot of math problems and keep generating new math problems over and over and over again to get more examples. We're kind of relying on the user to bring their own prompts because, you know, who are we to say what their problem is? We should just let them describe it to us using the prompts and we found it's very easy for users to get lots of prompts it's just hard to get a lot of responses yeah in this process before i get on to what tal can do in terms of uh increasing the power of of open source models uh what is the um can you talk about the the reinforcement learning process uh and how that's done you mentioned dbr dbrm is that what it's called deep data bricks reward model remord model i'm sorry yeah how exactly does that work so we're using you know there's a bit we're probably not going to say about this because there was a lot of like very detailed algorithmic work and you know a little bit of innovation that went on there.
33:33But the general idea is that we're generating a lot of data. We're figuring out which examples are better and worse using, or which outputs are better and worse using DPRM. And then we're applying relatively standard RLHF techniques on this. So it's nothing that will blow your mind or nothing that will surprise you. A lot of the work is in just getting the details so this whole process works well and doesn't completely fall apart. And so honestly, the most interesting parts are the experiments where this went wildly off the rails and we had to kind of spend some time getting the hyper parameters exactly right getting the order of the steps exactly right figuring out how many outputs to generate all the ugly stuff that again you know for my fan club on twitter or my anti fan club on twitter um you know not novel enough um nah the novelty is overrated working systems are what matter Yeah.
34:22So the DBRM is Databricks flavor of a reward. Exactly. We collected a ton of data on tasks that we thought were going to be most relevant to our customers. A lot of stuff involving documents or structured data to unstructured data or vice versa. Writing tasks, kind of some math tasks. but we really kind of focused our data on the things that we thought would be reflective of what our customers were doing. And we feel pretty good that everybody's trying to understand the data they have in Databricks and that's documents for the most part, a little bit of SQL. Yeah. The, and that process you, you use, I mean, there's the RLHF part that, that provides data, but on the reinforcement learning of tuning, that is autonomous.
35:23That's correct. And using other LLMs to judge the outputs to mark whether it's to reward or not reward. Is that right? So that's where, yeah, that's where DBRM is most helpful is that DBRM kind of gives us our signal of what we should reward and what we shouldn't reward. If you imagine that game playing environment, in some sense dbrm tells you what your score is at the game it helps you understand which strategies led to higher scores and which strategies led to lower scores and so that's kind of if there's kind of one key piece to this whole thing is the work we put into making dbrm really good and really specific to enterprise tasks yeah and so with this you're able to bring like a LAMA model up close to the level of O1 or O3, O3 Mini or R1, without increasing the inference cost of LAMA that's not fine-tuned.
36:27Is that right? So I'll be really careful how I answer this. As a good scientist, I have to be a little bit pedantic here. I'm going to sound like a lawyer for a moment. For all my friends in the scientific community, I want to represent this properly. The thing that I can say with certainty is at least for the tasks that we tried and for the models that we tried with the data sets we tried, because everything in science is contingent on those things. If you have bad data, nothing good is going to happen. And we're never going to claim that something universally works on every task. But for what we tried, which we hope is representative of what our customers do, we saw really big improvements in performance by using tau.
37:06And not only were they big objectively, but they were actually big compared to doing supervised fine tuning. This was the part that got me most excited and was a bit of a surprise for us. For every one of our tasks, we had a set of several thousand inputs and several thousand labeled outputs. So we could do supervised fine tuning. That's the traditional way of solving these tasks. We tried doing tau where we just dropped the outputs and instead used tau on the exact same inputs. And it in fact led to better results across the board, which is crazy, right? We're getting rid of labels. We're throwing out information and it leads to better results.
37:42And, you know, again, I don't want to promise this is universally true. It's something we saw for these training sets on these tasks. My hypothesis here, and it's just a hypothesis, is even for really good labeled data, data that, you know, is used widely in the academic community that's quite popular that our team has looked at or our team has done our best work to try to create, it's still really hard to create good labels. Yeah. And so DBRM in some sense is it allows us to put all of our effort into building this one really, really good reward model. And that may be easier actually than trying to create several sets of several thousand good labels.
38:19You know, I don't think this is necessarily, we didn't invent magic. I don't think that in general, having no labels is better than having labels. Like obviously the more information you have, the better you should be able to do. I just think getting good labels is really hard. That was our takeaway. Like, wow, even for these world-class benchmarks, the labels may not be perfect. And so DBRM kind of shines in comparison. And so if it's hard for my team and for world-class academic researchers to create good labels, it's not a fair fight for our customers to have to do that. They're busy, they have other things going on.
38:53They're not sitting around getting PhDs in machine learning and spending their whole day trying to build the best labels. So let's just, you know, not only should we reduce the work for them, but it's probably leading to better results if they don't even try. And we instead kind of put a lot of work into DBRM to make their lives easier. So it's been this really, I don't know, it was really rewarding for me, no pun intended, to kind of see that maybe not only can we make customers' lives easier, but in fact, it might be better to do that. Usually you would think of this as a compromise. Now, the comparisons to O1 and R1 and GPT-40, I think that varies a lot by the task.
39:29It's also a bit of an unfair comparison in a sense because OpenAI has to make O1 good at everything. Tau is trying to make a model good at one thing. But that's kind of where the unfair advantage is for anyone who has some interesting inputs and domain knowledge. you just have to make the model good at your one thing um open ai is spending a huge amount of effort to try to make a model good at everything and that's where you can kind of compete with the closed models by building something yourself and where we see a lot of opportunity for kind of custom fine-tuning of open models like llama yeah and i guess the question is uh you can get that improvement through supervised fine-tuning if you have enough label data right it's it's just the what is the cost of doing it that way versus the cost of doing it the Tau way?
40:20And the Tau way, since you're generating prompts or queries through inference time, and if you have a big user base, you're going to get a ton of that. It's just a cheaper way of doing it. Is that right? Yeah, cheaper. but I, you know, the word cheaper is doing a lot of work there because it's cheaper in terms of cost, but it's cheaper in terms of time and effort and frustration. I don't know, Craig, have you ever tried to label data? No, I've watched it done on a small, you know, on images, but yeah. I'd love to, like, this is actually something I try to do with my team periodically, just so they remember, like whenever we say, oh, just have the customers label data.
41:07I want them to kind of feel the pain and understand how frustrating this is. Like whenever we say, oh, we're going to go and buy label data, just remember, try being the user of the rubric and the specification you created or the product you created. It's really frustrating. Like getting my team to do this for more than like 20 minutes before they start checking their email and checking Slack is really hard. You start seeing people drop off the Zoom call. Oh, gotta go. It's lunchtime, got a meeting, sorry. And now we're going to ask our wonderful Databricks customers to go do that. So it's really kind of, you know, it's cheaper, but I think that underplay is like, it's, is the problem going to get solved with AI or is the problem not going to get solved with AI?
41:49And I want to give our customers at least more at bats to try to, to try to solve problems with AI. And this is a key enabler for that, I hope. Yeah. And so the process is you start with a bunch of queries that if you've been using the model in some application, you're getting off your customers. But otherwise, you have people interact with the model to develop the queries. Then you use those queries. Then you go through a reinforcement learning with human feedback phase where you score the answers to those queries that's had right, whether it's good or bad, or which you prefer. and then you take that data and use reinforcement learning to to replicate those high scoring behaviors and the question becomes can you identify that and if you can you can steer the model toward doing more of that and less of the bad stuff on your task and that's really it's another intuition for fundamentally what we're doing here It's easier said than done.
43:15Identifying the good ones is, if you knew how to do that, you probably wouldn't be generating the bad ones in the first place. But DBRM has been really helpful for doing that on enterprise tasks. And so I've kind of, again, surprisingly well based on what I expected. But I don't know, is that intuition kind of a little bit helpful, kind of different way of looking at it? And I guess this is the question on supervised fine tuning, because in this case, so you've you've gone through this process you improve the model you you put it back out for inference uh people interact and it generates more queries or prompts and then you can pull that in go through the same process and tune the model further and it at least on that uh specific domain, it has sort of continual improvement.
44:11Is that right? That's correct. You're kind of seeing more of the set of possible inputs your users might provide. And that just gives Tau more data to work off of. So it's kind of, it's a bit magical that you just, you get improvement without intervention. That was what we were going for is how can I train and then improve without really needing to do anything and just kind of let the process work itself out. Now, obviously this will encounter limits. You know, you, there, there's only so much you can figure out without giving more clarity on what task you're trying to solve. But I think for many of our customers, it's certainly the result showed it's good enough to get you, you know, an LLM that will like, you know, that will do a lot better than the base llama model at your task and will start to rival the closed models.
44:56And for a lot of our customers, that's, that's winning. Like that's what winning looks like relatively easy improvements in your model. That's something you control, you own is relatively cheap to serve. And then if you're really getting a lot of uptake, that's where you can go in and do the more detailed kind of hardcore sort of stuff that, you know, you can, you can start to label data or get thumbs up, thumbs down data from users, or kind of do more advanced stuff to make this even more powerful. But just getting to the point where your biggest problem is making it even more powerful, that's success.
45:28That's a huge amount of progress from where I think a lot of folks are today with AI. Yeah. And on supervised fine tuning, if you had some magical way to generate the supervised, the label data, you could do the same thing. You could keep fine tuning the model over and over and get it stronger and stronger. It's just, it takes more compute, more man hours. Yeah. It's just effort. It's, you know, know, the person who's building this AI system at an enterprise is also trying to do their day job and also trying to, you know, they're either trying to do their day job or if they're lucky enough that this is their day job, they're trying to prove to their business leaders that AI has a place in their organization.
46:10Because I think that's still something we all have to prove with this new technology. And so just like, yeah, let's get out of people's way. Let's make sure their time is being spent carefully. Yeah. And what about, this is on narrow tasks, what about on multi-tasks, multi-task LLMs. I mean, you can apply it, given you have enough prompts in your enterprise, you can apply it to more than just one task, right? Yeah, and we saw really good results with this where we kind of put together a prompt set geared toward a bunch of different enterprise tasks, you know, some documents, some question answering, some reasoning, some chat sort of things.
47:00And we got general improvement across that family of tasks, which made us confident as people move toward agents, you're going to probably want one LLM that can handle many different kinds of tasks for you. You can use the same model in many parts of an agentic workflow. And this gives us some hope that you can actually train something that's not just specialized to one task, but can do many tasks. I'm still not gonna claim it can do every task or anywhere close to it. that would be, you know, I think that would be a little bit of hubris at this point. But, you know, I do feel good that at least we have results indicating someone can hopefully have an LLM that can both do agentic routing and can also, you know, generate SQL queries and can also, let's say, you know, solve customer service requests.
47:45And if you have good, you know, prompt sets for each of those, you can get an LLM that's going to do a pretty good job with those and you know that again these were surprises things worked a little better than we expected which is you know it's the happy part of science most of the time things work much worse than you expect yeah and where are you going with this i mean what's next on your plate are you are you working on developing tau uh to to to to broaden it or make it more autonomous or uh yeah Yeah, I want to have reasoning LLMs for everyone as well. The same thing I mentioned before about how right now, if you want to train a reasoning LLM, you kind of need a way to tell right from wrong and correct from incorrect something a lot.
48:28It's a much more rigorous thing than saying better and worse. So we're thinking a lot these days about how we help our customers come into their own domain and get a model that will not just answer questions well, but actually reason. And so reasoning to help the model answer better and to also give either you or your users insight into why the model made the decision it made. I think that we all want a little more insight into how AI is behaving. And so that's really important to me. The other piece is that the same idea is behind Tau. The idea that, you know, DBRM can help you figure out which data is better and worse and how you steer the model.
49:02That's also really useful for a bunch of other things you might want to do with AI. Like it's a way to do evaluation without needing any labeled data either. It's kind of the same idea. You have your LLM give multiple outputs and you look at which ones are better and which ones are worse. And if your LLM is getting better over time, and that's a great way to tell whether you're making progress. So I think there's just a lot of hope of kind of reducing the burden on our users when they want to build, measure, customize AI systems. And just getting them to do the thing that they're good at, kind of coming back to where we started.
49:31The idea that on the web, everybody has their own website and they can make of it what they will. It shouldn't be that hard to build a website because the hard part is solving a problem with a website or doing something interesting with that website. And I think the same should be true about AI. We should make it really easy to build AI because the hard part should be going and solving your problem with your expertise and making a difference in the world. Yeah. Yeah. Yeah, well, I'll leave it there, but I do have one question remaining. The judges, the LLMs that are acting as judges, are you using proprietary models, reasoning models to act as judges?
50:07So for judges in various contexts, judges show up a lot throughout Databricks, and it's actually a combination of a bunch of things. Some of them are doing prompt engineering on top of closed models, and some of them are actually variants of DBRM or kind of cousins of DBRM that were trained specifically to tell better and worse for a specific kind of judgment. Like, is the output actually factually grounded in the document? And so, you know, it's an interesting landscape right now where if you were to hit the Databricks judges in our APIs, you'd actually see a mix of all sorts of different things depending on what we found to work well.
50:39I'd love it to all be cousins of DBRM at the end of the day. But the most important part for me is, you know, solving customer issues. And, you know, If we can't offer something that's better than prompt engineering a great model that somebody else has built, I will always err on the side of giving the customers the best thing we have rather than making myself feel good. Yeah. Okay. Well, let's leave it there. That's fascinating. And I'll watch the progress of Tau and see how it becomes adopted or adapted. Okay. Jonathan, that was great. In business, they say you can have better, cheaper, or faster.
51:22You only get to pick two. What if you could have all three at the same time? That's exactly what Cohare, Thomson Reuters, and Specialized Bikes have, since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 % less for compute, 70 % less for storage, and 80 % less for networking.
52:13Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head Head to oracle.com slash IonAI. IonAI all run together, E-Y-E-O-N-A-I. That's oracle.com slash IonAI.
From the publisher
This episode is sponsored by Oracle. OCI is the next-generation cloud designed for every workload – where you can run any application, including any AI projects, faster and more securely for less. On average, OCI costs 50% less for compute, 70% less for storage, and 80% less for networking. Join Modal, Skydance Animation, and today’s innovative AI tech companies who upgraded to OCI…and saved.
Try OCI for free at http://oracle.com/eyeonai
What if you could fine-tune an AI model without any labeled data—and still outperform traditional training methods?
In this episode of Eye on AI, we sit down with Jonathan Frankle, Chief Scientist at Databricks and co-founder of MosaicML, to explore TAO (Test-time Adaptive Optimization)—Databricks’ breakthrough tuning method that’s transforming how enterprises build and scale large language models (LLMs).
Jonathan explains how TAO uses reinforcement learning and synthetic data to train models without the need for expensive, time-consuming annotation. We dive into how TAO compares to supervised fine-tuning, why Databricks built their own reward model (DBRM), and how this system allows for continual improvement, lower inference costs, and faster enterprise AI deployment.
Whether you're an AI researcher, enterprise leader, or someone curious about the future of model customization, this episode will change how you think about training and deploying AI.
Explore the latest breakthroughs in data and AI from Databricks: https://www.databricks.com/events/dataaisummit-2025-announcements
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI




