912: In Case You Missed It in July 2025

8 Aug 2025 · 33 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

A July 2025 “In Case You Missed It” roundup (episode 912) covering data-centric machine learning (DMLR), benchmark contamination and alternatives, generative-AI risk in financial services (especially for RAG), human decision predictability, and causal AI tooling.

Guests (and backgrounds)

Lilith Batlia (legal-tech ML researcher; ML Commons Data Perf/Data-Centric ML work). Sinan Ozdemar (discusses benchmark limitations). Dr. Sebastian German (published on mitigating genAI risks in financial services; knowledge-intensive/regulation-heavy domains). Dr. Zohar Bronfman (neuroscience of unconscious decision timing; human predetermination). Dr. Robert Ness (Microsoft Research; causal AI models; PyTorch/causal inference libraries).

Key claims + examples

DMLR shifts iteration from model to data (fix noisy labels; Data Perf benchmarks; low-resource language datasets via Common Crawl). Benchmarks leak via online answers; “Chatbot Arena” uses blind human preference, but raises judging/ownership concerns. Finance LLMs and even safety guardrails may be misaligned because models aren’t trained on finance-specific corpora; evaluate in-domain and red-team using NIST/ML Commons taxonomies. Neuroscience suggests decisions are detectable before conscious awareness (predict purchases/conversions). Causal AI requires explicit causal assumptions (e.g., guild membership confounding side-quest spending); use Pyro/PyTorch/DoWhy-style libraries to handle inference under specified graphs/models.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Data-Centric Machine Learning

0:45 to 4:49

A discussion on data-centric machine learning and its importance in legal tech.

“And so this is now a topic that is, you know, this isn't just like, oh, there's some analogies here that might be relevant to your industry.”

AI Benchmarks and Their Challenges

4:49 to 6:51

Exploration of the issues related to AI benchmarks and potential solutions.

“He's one of the biggest names in data science, period.”

Model Selection and Generative AI Risks

6:51 to 12:49

Insights into selecting models for specific domains and the risks of generative AI.

“In it, Sinan Ozdemar and I discuss approaches to circumventing the limitations of benchmarks.”

Risk Management in Financial Services

14:00 to 16:06

Explore the importance of risk assessment and management in financial AI applications.

“And the way that we wrote our paper very much should be seen as a case study.”

Selecting the Right LLM for Specific Use Cases

16:06 to 19:15

Learn best practices for choosing language models tailored to specific domains.

“If we're trying to select an LLM for a particular use case, what do you recommend we do?”

Human Behavior and Predictability in AI

19:21 to 21:51

Discusses the predictability of human behavior and its implications for AI.

“Zohar Bronfman thinks we may be, and he has the research to back him up.”

Consumer Behavior and Data Insights

21:51 to 25:00

Understand how consumer behavior can be predicted using historical data.

“whether as something reaches consciousness, it can override or kind of change some of these or veto some of these processes and so on and so forth.”

Causal AI and Its Applications

25:00 to 28:00

Explores the principles of causal AI and its practical implementations in data science.

“Having AI knowing your next move might sound far-fetched, but given how much of our activities take place online, AI tools may soon be able to detect behavioral patterns we might not even consciously recognize ourselves.”

Understanding Causal Inference with PyTorch

28:00 to 31:15

Learn how to implement causal inference using PyTorch and related libraries.

“Some extent you saw this, you mentioned you interviewed somebody who talked about Stan.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:02Jon Krohn:This is episode number 912, our In Case You Missed It in July episode.

0:19Jon Krohn:Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. This is an In Case You Missed It episode that highlights the best parts of conversations we had on the show over the past month. We'll start off with a conversation I had in episode 901, in which I asked Lilith Batlia why data-centric machine learning research, or DMLR, has become a byword for accuracy in the field of legal tech. The impetus for having an episode when we talked about it on the train already a year ago was this idea of data-centric machine learning. And so this is now a topic that is, you know, this isn't just like, oh, there's some analogies here that might be relevant to your industry.

0:58Jon Krohn:Data-centric ML is relevant to every listener. Anybody who's working with data, this is relevant. And so tell us about data-centric machine learning research, DMLR. And my understanding is that you fell into DMLR as a result of how messy the data are in the legal space. Yeah, that's right. So in my first R &D role, I was really focused on algorithms and on finding the best classification algorithms for these classification tasks that we've discussed. At a certain point, I realized that the label data I was working with was so noisy, just had so many mislabeled instances and all of that, that it really curtailed my ability to evaluate the performance of the algorithm, just because I couldn't necessarily trust my data.

2:09So that led me to be very interested in what Andrew Ng coined data-centric AI, And I ended up getting involved with a working group at ML Commons called Data Perf, where we were looking to benchmark data centric machine learning. That ended up leading to a few different workshops that we've organized at iClear and ICML. Data Perf also became a NeurIPS paper.

2:45and yeah, yeah, basically it turned into a whole community. So now there's a DMLR journal, there are the DMLR workshops at these conferences, and then Dataperf morphed into the data-centric machine learning research working group with ML Commons. So we have a lot of different things going on. We're working in partnership with Common Crawl, the foundation that curates the datasets that most LLMs have been trained on. We're partnering with them on a challenge that will result in a low resource language dataset that will be publicly available. So if you're interested in joining the working group, please do get involved.

3:28Again, it's with ML Commons. You can go to that site and sign up to the working group.

3:35Jon Krohn:We'll be sure to have a link to ML Commons in the show notes. When you say low-resource language, this is languages for which there are not many data available online. They could be rarely spoken languages, or for whatever reason, languages that, even if they're spoken relatively commonly, they aren't represented on the internet. Exactly. That sounds really cool. Those acronyms that you were saying earlier, where this DMLR initiative was getting traction. So conferences like iClear, ICML, NeurIPS, these are the biggest conferences that there are, academic conferences that there are. And so really cool that you get such an impact there.

4:17Jon Krohn:And it's also interesting to hear the connection to Andrew Ng there. Because he, so he, I have in my notes here somewhere, I'm kind of scrolling around in here. Yeah, so at the inaugural DMLR workshop, Andrew Ng was the keynote. Yes, yes, exactly. And he was involved with Data Perf as well. He's on that Data Perf paper. Okay, so I'm now very clear on the importance of DMLR, the traction it's getting, and bigwigs like Andrew Ng being involved. Probably most of our listeners know who Andrew Ng is. He's one of the biggest names in data science, period. and if you aren't already familiar with him, he was on our show in December, so episode 841 you can go back to.

5:02Jon Krohn:We'll have a link to that in the show notes as well. So now I have a clear understanding of data-centric machine learning being very important, gaining traction, but our listeners still might not have a great understanding of what it is. Yeah, so the best way I can explain it is that in traditional machine learning paradigms, you're iterating on the model. You're iterating on the model architecture, on the learning algorithm, all of those sorts of pieces. And that's where you're really focused on improving performance is by iterating on the model. With data-centric machine learning, you're iterating on the data.

5:49So you're holding the model fixed and you're improving the data, you're systematically engineering better data. And then there are all these different questions, right? So there's the question of whether to aggregate labels or not. There's a really interesting paper, Doremi, that looked at weighting different domains of the pile to get the best LLM pre-training performance. So there's, yeah, it can go lots of different ways. There's another paper I'm thinking of, I can't remember the name, but they looked at selecting the best data points for training a model a priori, so not even active learning, where you're starting with the results of the model to determine which additional data points you should have labeled, but just with a data set from scratch, using linear algebra to figure out which data points are worth labeling.

6:48Jon Krohn:From DMLR, we turn to AI benchmarks in episode 903. In it, Sinan Ozdemar and I discuss approaches to circumventing the limitations of benchmarks. You right at the beginning, near the beginning I was talking about benchmarks, you talked about contamination. And so what is the resolution there? This seems like a really tricky problem. How do we prevent leaks once a benchmark's been out and the answers are online? I mean, I guess one solution is to just not have answers online. Tell that to the internet. Well, because here's the thing. If a benchmark literally comes with the answers, that's the whole point of the benchmark is you're supposed to know the right answer.

7:32So the same place where you get the questions for the benchmark also has the answers to the benchmarks where you can validate that it's correct. So it's impossible to not have the answers not on the internet.

7:44Jon Krohn:Couldn't you have something like, it could be like Kaggle, exactly. You could, but then who owns it? Who owns the results? Someone has to own it. Well, someone has to, because if it's going to be hidden from everybody else, someone now is in charge of holding those answers. So who is it? The developer, I guess in kind of the same way that's, so, okay, here's an interesting idea. So what about a solution like Chatbot Arena, where in Chatbot Arena, there's no correct answer necessarily. So it's run by Berkeley, the LM Sys Lab, if I'm remembering correctly. I think it's Joey Gonzalez's lab. And so Joey Gonzalez has actually been on this show talking about it.

8:29Jon Krohn:If I can find that episode quickly. Yes, episode 707. You can hear from the Berkeley professor that was in his lab that this chatbot arena was devised. And so in the chatbot arena, it's different from benchmarks in the sense that you don't have a specific set of questions and answers. You pit two LLMs against each other and you as a human evaluator of the arena, you don't know which two you're seeing output from, but you pick one as better than the other. And so first of all, I'd love to hear thoughts on the arena, but the reason why I'm bringing the arena up is that in that situation, I mean, so you're talking about like ownership.

9:11Jon Krohn:You could have a similar kind of thing where for a benchmark where somebody creates a training set like Humanity's Last Exam, you could have a holdout answer set. And yeah, I mean, some like a university like Berkeley could be administering it. You know, people submit their responses and then they get a grade back. Yeah, a few things. I'm a fan of the arena in general. The idea of blind judging from a human, for me, is one of the best ways to really get a good sense of an LLM's usability. Now, a couple of things, caveats there. If I'm just a layperson talking to a chatbot, to your point, I'm not coming in with structured questions.

9:58I'm just going to pick the one I like the most. And that might come down to which one's talking the way I like it to talk, which kind of leads to the whole sycophancy thing, right? when OpenAI said, well, we rely too much on people's thumbs up and thumbs down, and that's what got us in trouble. The LL Marina is pretty much a thumbs up and a thumbs down. That's all we're really doing is saying, I like that better. I'm not telling you why. Just because it cursed once, I thought that was cool. We have no idea. And sure, at scale, when you aggregate these, you'll get a much more stable answer. But again, at this point, we're just judging preference as opposed to knowledge.

10:36And again, without that structured data set. Now, also, I think you mentioned this, there is no answer to any questions on the arena, right? You're just shown response. You are not coming in with a question. You are just kind of shown answers and it's up to the human to decide which one is correct. So whoever is judging it behind the scenes, how are they doing it? Are they paying a human being to read each one and actually comparing it to the right answer? Or are they going the LLM as a judge route where they're saying, well, we have yet another LLM who is given a reference answer and this answer, and it is asked to say, how closely does it compare?

11:15We don't know. And again, a lot of it just comes back to what actually is the right way to judge the system? Who has the right to judge whether or not the AI was correct or not? That's a big question. And again, that's why we have benchmarks is that is our current proxy to that question, which is, well, if we all agree that Pablo Picasso painted this thing, and that's one of the answers they can pick from, it's on the right track to knowing general world knowledge. But if it just comes down to which one do you like talking to better, like an arena would be, you're going to miss a lot of the actual important pieces of information you're trying to get out of that LLM.

12:00I'll say one more thing. It's funny you brought up the arena. That's actually one of the allegations from Lama 4. Again, total allegations. But one of the separate allegations from Lama 4 was they released a trained to test model specifically for the arena that was different than the Lama 4 we all got in the end. Again, total allegation. But those rumors start bubbling up when people notice discrepancies. And who's to say those discrepancies are correct? They're all just our own interpretations and our own expectations, maybe not being met by what we were shown. There's no way to prove this.

12:42Jon Krohn:Finding the right way to judge a system is a huge question, and it's one that I am sure will continue to perplex and challenge us over the years. In episode 905, my guest, Dr. Sebastian German, and I continued this conversation specifically regarding ways to select models for the domains we might be working in. There's a second paper that you also recently published. So your first author on a paper that was submitted to Archive in April of understanding and mitigating risks of generative AI in financial services. So mostly so far in this episode, we've been talking about generally how models fare under RAG.

13:22Jon Krohn:But in that paper, it's related to risk of Gen AI and finance. You emphasize that most foundation models are not trained on finance-specific corpora bodies of knowledge. So what are the limitations this creates for LLMs in general, but particularly for RAG? And I'm assuming that this same kind of sentiment, you know, you looked at it with finance specifically because Bloomberg is, you know, as a financial services company largely. But do you think that the same kind of limitation would apply in other sectors as well? Yeah, absolutely. So, yeah, I gave a little bit of a teaser of this paper earlier and an answer as well.

14:04And the way that we wrote our paper very much should be seen as a case study. Finance here or financial services, in particular capital markets and asset management, is the case study that we use to make the point that we really need to think about risk and risk taxonomies and risk management in our domain, in what we are trying to build. and as you say we made the point yeah models are not necessarily trained on on financial domains we see that both in the helpfulness and the harmlessness angle often you know complex financial tasks are not being able to be sufficiently handled by large language models by themselves but also in our paper we make the point that even safeguards that are dedicated models or systems to provide these kind of first paths like is this safe is this unsafe judgment they're also not trained on financial services.

14:58And if you use them out of the box and say, look, I use LamaGuard, I use ShieldGemma, I use Aegis, I'm safe now, right? You're protected against a particular view of safety that is very much grounded in categories that are relevant to broad populations, to things like chatbots that help you do productivity day-to-day tasks. The typical applications that you would see in those AI productivity tools, no matter which one you use, They all have similar mechanisms, but those are not necessarily the same risks that we are under in financial services. Those are not the same obligations that companies, organizations in healthcare are under or law or any other highly domain-specific knowledge-intensive domain that has a lot of specific regulation, jurisdiction-specific regulation, considerations about whether just refusing to answer or giving disclaimers is enough or whether questions should be blocked altogether.

15:55And there's just this difference of view that can be capsulated in a single model that a provider can give that very much is focused on a different use case. Nice.

16:06Jon Krohn:Yeah. So I don't know. Do you have guidance for us? If we're trying to select an LLM for a particular use case, what do you recommend we do? I mean, like practically. How can we move forward with all the information that you provided in this episode? in selecting an LLM for a particular use case for a particular domain, particularly if we want to be applying it in RAG situations? Yeah, so in our paper we also have a list of best practices and recommendations that we have for especially for knowledge intensive domains and regulation heavy domains. Not necessarily everything has to be followed if you're building something for for a much broader general population, but especially for these kinds of domains all I can do is spray my mantra, evaluate the system in the context that it's deployed in.

16:59If you are building something for healthcare, well, you better evaluate it in the context of healthcare. If you are building in the context of financial services, you better evaluate your subject matter experts in financial services. And specifically on the safety angle, our paper makes a couple of suggestions here. There are very good starting points. There are taxonomies such as the NIST, risk management framework for AI. there are other industry collaborations ongoing, there's ML Commons. Those all provide more general purpose taxonomies, but just taking them as a starting point and then from there, adjusting them to your domain can often save a lot of time.

17:38And especially if you're a large organization with a compliance or risk department, it will help them also understand how one can classify and then categorize these kinds of risks. And other recommendation we make is to organize red teaming events and or do any other kind of red teaming. Red teaming in this case is this practice that had to start in the Cold War where you have users trying to be malicious. So we get people in the same room and we say, look, for the next couple of hours, try and break the system, try and play evil. Here are some instructions on how to do this. And then afterwards, we can look, how often was this actually broken?

18:16How often did the system give financial advice? How often did it refuse? And from there, we can quantify the risk surface. Since we were talking earlier about this unknown risk surface, well, just measure it, and then you have it. So that's kind of the main takeaway that we have. We give pretty specific advice for how to go about this and how to set up risk management frameworks. and all this needs to go hand in hand also with, again, this evaluate in the context that's applied. Make sure you invest a lot in evaluation. Don't just take the word of the large negative model providers that their benchmark scores are going to translate into all the downstream applications.

18:53And if you follow that advice, you're going to have a system that is in the end much more trustworthy, reliable, robust, and you're going to have users that are going to keep using it rather than trying it twice, getting really bad answers both times and never touching it again.

19:09Jon Krohn:So the best way to know if you've got a system that's fit for purpose is to test, test, and test again. This is how we can measure AI behavior and ensure the model we choose does what we want it to do. Determining human behavior is also a critical topic for AI practitioners. Are we so predictable? Dr. Zohar Bronfman thinks we may be, and he has the research to back him up. He explains human predetermination and desire in this clip from episode 907. Something very interesting that Benjamin LeBay and countless others have shown is that you can have a neurological, a neuro, the neural basis of some conscious idea that you have happens hundreds of milliseconds before you have the conscious thought.

19:56Jon Krohn:And this is a very disturbing thing to think about because you kind Most of us go around through the day with this illusion that you have some kind of control over what thoughts come into your head or what action you take next. But in fact, what these experiments show is that you become aware of a decision after that decision has already been made subconsciously in your brain. And so, yeah, I don't know. There's a lot of planning to dig into there, but maybe talk to us about this a bit more and then maybe tie it into your belief in AI's ability to anticipate user behavior. So I think we were talking about, you know, once you get exposed to something in the realm of neuroscience and AI, you lose sleep.

20:46This was probably the biggest sleep deprivation I had because it's mind-blowing, right? If we think about it deep, we might end up in a rabbit hole of, am I just an agent carrying my neurons or something like that? Which I don't think is very easy to disprove, by the way, in all honesty. So it's mind-blowing. It's obviously a lot to digest. But yes, by the way, those experiments happened first sometimes during the 80s. since then, it's been replicated and reproduced and in different settings and in different environments, in different technologies, in different animals, so many times that I don't think it's any more even just an open question.

21:32It's a truism, it's given. Now, obviously there's room for interpretation, but the fact that there are brain processes that are directly causally related to decisions we make and that we don't have access, we don't have conscious access to those processes, I think is already completely agreed upon. Obviously, you can ask the questions of how elaborate these processes are, whether as something reaches consciousness, it can override or kind of change some of these or veto some of these processes and so on and so forth. But the fact that these happen is like hard fact. Now, it means that much of what we are doing, okay, as humans, is predetermined by things that have nothing to do with our immediate desires.

22:20Okay, so you can put someone in fMRI, like a functional MRI that basically shows the blood in your brain and you know which areas are active. And you can tell 10 minutes or 15 minutes and they drive in a car simulator. You can know 10 or 15 minutes in advance before they reach the junction, whether they're going to turn left or right. so we and and you know it it has many contributors to it maybe it's a question of your stronger side maybe it's something that happened in the morning maybe your neck was sore and you know it's harder for you to look to the right by the way it's my case at the moment so there are many different ways that can contribute to these unconscious processes that you end up affecting your decision or your action.

23:08But what it also means, and this is something, again, that is quite known for many years, it means that as a consumer, your behavior is also affected by many things that you're not aware of. And it means that as a business that sells to consumers, you can probably know much in advance of your specific customers' behaviors before the event takes place. So you can predict the purchases that customer is going to make, the conversions or lack of those, lifetime value, best products, churn, and so on and so forth. And that ability to make those predictions based on their historical behavior is, like I said earlier, the biggest level we know in the industry for transforming businesses.

24:08So I'm basically saying if you collect data about your consumers as a business, there's a good chance you can start making predictions about their behavior in the future. And you can optimize their experience. You can optimize your processes. You can basically just make the most out of those precious interactions the consumers have with your business. my you know my personal and this is what Piken is all about my personal mission I want to bring these capabilities to as many small and mid-sized businesses as possible because they also deserve quote-unquote that remarkable technology that basically you know tells you what people are going to do even before they know what they're going to do and that's why we've invested so much in connecting LLM's data and machine learning together in one nice package.

25:10Jon Krohn:Having AI knowing your next move might sound far-fetched, but given how much of our activities take place online, AI tools may soon be able to detect behavioral patterns we might not even consciously recognize ourselves. We have to note, though, that AI doesn't make intuitive decisions the way we do. Instead, it relies on patterns and probabilities. Is this set to change? My final guest in this episode of In Case You Missed It is Microsoft Research's Dr. Robert Ness. In episode 909, Dr. Ness investigates how we can start to build causal AI models that go beyond the standard correlation-based patterns and probabilities that we're used to machine learning models generally detecting.

25:53Jon Krohn:I'm a data scientist at a gaming company, and I want to figure out whether the users of this game, when they tend to engage in more side quests, does that cause them to spend more money on in-game assets that they can be buying? And so there are potentially confounding variables out there that like things like being a member of a guild that you mentioned there. And so if we weren't collecting that guild data, we'd have to have more assumptions. We'd have to basically make the assumption that being in a guild doesn't matter or not. And so it seems like, so this has made clear that there's a lot of assumptions, more thinking potentially about your problem that you need to do if you're engaged in causal AI.

26:39Jon Krohn:So that's a great thing to understand about this. But to kind of get into the nuts and bolts, your book does a great job of using PyTorch code using examples to make causal AI or causality in general, which is often a very theoretical, difficult to understand topic, because of all of your examples and use of code in the book, it makes understanding causal AI more intuitive. And so let's say that you were going to, you were the data scientist, Robert, at this gaming company, what Python tools would you use to then do causal AI? How would you model this in order to come up with a causal conclusion?

27:27So one of the things that I had mentioned that I was trying to do with my book was to separate out the abstractions that have to do with statistics and computing, right? Like scale it up, you know, algorithmic complexity. from the causality. And what's cool about the libraries that we have today is that they can actually help us. If we're able to separate those abstractions, then we get to focus on one thing while leaving the nuts and bolts to be handled essentially by the library. Some extent you saw this, you mentioned you interviewed somebody who talked about Stan. And what's cool about STAN is that that inference algorithm Hamiltonian Monte Carlo is, I mean, you can go in there and understand it's not, you know, well, it is physics, but it's not rocket science.

Read the full transcript

28:21I was like, well, it kind of is rocket science. But you can still kind of just specify your model, specify what are the parameters, what's the model. And as long as you satisfy a certain set of requirements, I think mainly that all these things have to be continuous. that the inference will kind of just work for you without you having to like go implement your own inference algorithm. It's the same thing here, right? Where like if you can, as long as you can kind of specify your causal assumptions in some cases in the form of a graph, for example, then you can rely on say graphical causal inference

29:10algorithms from, say, probabilistic graphical models to kind of handle the inference there for you. If you're implementing it in PyTorch, I have plenty of PyTorch examples in the book, as long as you can incorporate your causal assumptions in the structure of the model in your algorithm and writes a kind of basic inference algorithm that has a differentiable loss function, then PyTorch is going to handle all of the nuts and bolts of the inference for you, right? That's kind of why we invented PyTorch, to say, well, if I can differentiate it, then I can get a gradient, then I can just turn it into an inference problem there.

30:06And so that's why, you know, so I do have examples there in PyTorch that are saying, okay, well, let's just, let's not worry about whether or not, you know, we need to use linear regression here or propensity scores or double machine learning or instrumental. These are all different types of kind of statistical methods for doing the inference you want. And you can learn all these things. Great. And there's great books for that, right? I think I can name a couple off the top of my head. But you can also say, let's work with some libraries that are just going to handle that stuff for us under the hood and kind of treat it as either an objective function to be optimized or as a configuration parameter in some model specification.

30:58and then focus on our ability to think causally and write that thinking down in the form of a model and focus on the actual domain that we're modeling as opposed to all the inference stuff that we need to do to get that to work. And so to answer your question, I talk a lot in the book about using deep probabilistic models like modeling with libraries like Pyro as well as some more conventional tools like the DoY, DoY from the broader PyY suite, which is a big collection of causal inference libraries. And so even in DoY, right, there are different types of statistical techniques that you can use to estimate a causal effect.

31:49But at the end of the day, you're thinking more about, you know, what are your modeling assumptions? and can you answer the question given your assumptions and your data and then if you can, you want to get to an answer and then all of the various statistical approaches you can take to arrive at that answer given your assumptions and your data. You can kind of just toggle between them and see which is giving you more stable results, for example.

32:14Jon Krohn:All right, that's it for today's In Case You Missed It episode. To be sure not to miss any of our exciting upcoming episodes, subscribe to this podcast if you haven't already. But most importantly, I hope you'll just keep on listening until next time. Keep on rocking it out there. And I'm looking forward to enjoying another round of the super data science podcast with you very soon.

From the publisher

In this episode of In Case You Missed It, we look back on five great interview episodes from July. Hear from Lilith Bat-Leah (Episode 901), Sinan Ozdemir (Episode 903), Sebastian Gehrmann (Episode 905), Zohar Bronfman (Episode 907) and Robert Ness (Episode 909). They’ll tell you why data-centric machine learning is so important across disciplines, starting with law, and how we can use AI benchmarks and “red teaming” to refine our search for the best AI models. 

Additional materials: ⁠⁠⁠⁠www.superdatascience.com/912

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
912: In Case You Missed It in July 2025Super Data Science: ML & AI Podcast with Jon Krohn · 33 min
Listen in VO