In short
Jazmia Henry discusses end-to-end foundation models for the energy industry at Collide, including data curation, custom tokenization/OCR, continued pretraining, reinforcement learning with physics-style environments, and production inference. She also covers two research themes: (1) evaluation failures in agentic RL systems leading to reward hacking, and (2) her NeurIPS 2023 claim that scaling with bad data degrades model performance.
Guest backgrounds
Jazmia Henry has degrees from Tulane and Columbia, is/was partway through a PhD at Oxford (social data science; paused), and previously worked in data/AI at Morgan Stanley and Microsoft. She is now foundational model engineer at Collide (Texas), building specialized foundation models for oil and gas.
Key claims
Bad, non-deduplicated training data can make models “forgetful” and less intelligent; compute can’t fix data quality. Current evaluation frameworks for deployed agentic systems are structurally inadequate, causing predictable reward hacking. Her mitigation is Grounded Continuous Evaluation (GCE) using simulation-based fine-tuning with determinism.
Notable examples
Job Safety Analysis (JSA) documents where a model must not be sycophantic; “smoke marbles” map/trajectory instructions; reward hacking via keyword-based shortcuts (e.g., guessing common equations) instead of correct reasoning.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOConversation Kick-off and Icebreaker
0:54 to 2:22
Hosts engage in a light conversation about Charlotte and a past ad shoot.
“This episode of Super Data Science is made possible by Anthropic, Excel Data, and Cisco.”
Industry Context: Oil and Gas
2:22 to 3:21
Discussion about the importance of the oil and gas industry and AI's role in it.
“You are an exceptional researcher, practitioner of AI in the real world.”
Jazmia's Work at Collide
3:21 to 11:40
Exploration of Jazmia's role at Collide and the application of AI in oil and gas.
“So tell us about your company and what you do there.”
Importance of Data in AI
14:00 to 14:26
Learn about the significance of data management in AI companies.
“So that's the first big portion that constantly changes.”
Importance of Data in AI
14:30 to 15:13
Learn about the significance of data management in AI companies.
“Data and AI observability, agentic data management, AI data engineering, Kubernetes native compute, and more, all on one platform that runs where your data already live.”
Building Effective Data Pipelines
15:13 to 17:48
Understand how to structure data for effective AI model deployment.
“Something that we talk about on the show a lot.”
Model Training Methods
17:48 to 19:34
Explore the methods of training AI models, including reinforcement learning.
“And so being able to put physical validation around it by creating physics environments and saying, OK, go ahead and do these things.”
Challenges of Reinforcement Learning
19:34 to 22:30
Learn about the challenges and intricacies of using reinforcement learning in AI.
“One of the most fascinating roles I think I've ever heard of on this show.”
Academic and Industry Experience
22:30 to 26:56
Discuss the value of combining academic knowledge with industry experience.
“So I'm using it in a sense of how people describe it.”
PhD Journey and AI Evolution
26:56 to 28:00
Reflect on the evolution of AI during the speaker's academic journey.
“And, yeah, I also now see that I, if I clicked on the ellipses on the dot, dot, dot on your bio, it does say left to focus on industry.”
Show all 21 chapters
Jazmia's PhD Journey and Industry Work
28:00 to 29:38
Learn about Jazmia's experience balancing her PhD with full-time work in AI.
“Where's the US to get more time because of the timeline to do it?”
Foundation Models in the Energy Sector
29:38 to 32:02
Explore Jazmia's work on petroleum engineering foundation models and their challenges.
“Especially with that kind of European three-year program.”
Systematic Failures in AI Evaluation
32:02 to 37:29
Understand the four systematic failures in AI evaluation frameworks and their implications.
“So at Collide, that oil and gas AI company that you're at right now, you are building RIGS, not oil RIGS, but R-I-G-G, lowercase s.”
Reward Hacking in Reinforcement Learning
37:29 to 42:00
Delve into the complexities of reward functions in reinforcement learning and how models can exploit them.
“So when you have reinforcement learning, what ends up happening is models are very smart in that they are able to identify the least that it has to be able to do in order to get a good reward.”
Understanding Reinforcement Learning in Models
42:00 to 44:40
Learn about the challenges of reinforcement learning in model training and how to maintain accuracy.
“But in the end, I want the final answer that it has to be correct.”
Grounded Continuous Evaluation Explained
44:40 to 47:20
Discover the concept of grounded continuous evaluation and its importance in AI model training.
“You mentioned grounding in there a few times, but I don't think you mentioned specifically that in the paper, your framework for mitigating this reward hacking is called grounded continuous evaluation, GCE, right?”
Scaling Data Quality and Model Performance
47:20 to 51:08
Understand how scaling bad data can negatively impact model performance and the principles behind it.
“You demonstrated that scaling bad data makes models worse, not better.”
The Importance of Linguistic Equity in AI
51:08 to 56:01
Explore the significance of linguistic equity and the need for representation in AI language models.
“I'm embarrassed that I wasn't sure whether, well, because it did also have a poster presentation, right?”
Jazmia's Diverse Background and Influences
56:01 to 1:02:47
Explore Jazmia's unique upbringing and how it shaped her career path.
“That's a, yeah, it's a nuance that I didn't appreciate as I was asking the question.”
Following Jazmia's Work
1:02:48 to 1:03:02
Learn how to stay updated with Jazmia's projects and insights.
“So for people who are also excited about what you're up to, what is the, what's the best way to follow you in your work?”
Book Recommendation: Value Systems Design
1:03:03 to 1:04:34
Discover a key book that influences Jazmia's decisions in AI.
“Um, occasionally I post articles on medium or towards data science.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:Today's guest had a paper accepted at NeurIPS, the world's most prestigious AI conference in 2023, that made a claim that nearly everyone thought was ridiculous, that scaling bad data can make models worse, not better. Two years later, the entire industry is quietly admitting she was right. Welcome to episode number 995 of the Super Data Science Podcast. I'm your host, Jon Krohn. My deeply impressive guest today is Jazmia Henry. Jazmia holds degrees from Tulane and Columbia and is partway through a PhD at Oxford. She's held data and AI roles at Morgan Stanley and Microsoft, and now she's a member of the technical staff for AI and machine learning at Collide, a Texas-based startup that's building AI infrastructure, including all aspects of specialized foundation models for the energy industry.
0:46Jon Krohn:In this episode, you'll hear about her cutting-edge academic contributions and her exciting work at the energy industry frontier. Enjoy. This episode of Super Data Science is made possible by Anthropic, Excel Data, and Cisco. Jasmia, welcome to the Super Data Science Podcast. How are you doing today? Doing great. Lovely. I am so glad to hear that. I'm not sure if this will make it through the editing process. There might be some kinds of filters that get applied, but right now I can hear birds chirping in your background, like many of them. It sounds wonderful. Where are you calling in from? Charlotte, North Carolina, right in the suburbs.
1:26It's beautiful. Yeah.
1:29Jon Krohn:I've only been there once. I was there for a couple of nights. We were shooting an ad for NVIDIA and Dell that was for Bloomberg TV. And it was so hot. We had to do like, we had some outdoor shots. And so the shooting started at like 6 a.m. to try to get it done before it was like insanely hot. Because it was over 100 degrees Fahrenheit that day. Oh, yeah. Yeah, that must have been the summertime. Yeah. Oh my goodness. And I had to wear a suit for this shoot. I know. Oh man, did you pair it with a cowboy hat? You gotta do the whole vibe. I just paired it with sweating and the makeup artist had to keep coming and try to dab it away and make it look natural.
2:14Jon Krohn:But Charlotte was super nice. Great food, super friendly people. Anyway, let's dig into the technical stuff here. You are an exceptional researcher, practitioner of AI in the real world. I was so excited to get you on the show. To give a little bit of context, let's talk about where you work. So we actually, I don't think we've ever had an episode. We're coming up on a thousand episodes of this show. I've been hosting for the last 500, 600 of them. and certainly in the 500 600 i've been hosting we've never had an episode focused on oil and gas which is really important i think maybe for some listeners even for myself you know we think you know clean energy alternatives would be great in the future um but for the foreseeable future we are going to need petroleum products as a big part of our energy mix no question and with geopolitical events right now, at least at the time of recording, there's a lot of interest in oil and gas and shipping in particular.
3:19Jon Krohn:So anyway, I thought it would be particularly interesting right now to get you on the show. So tell us about your company and what you do there. Yeah, I work at Collide. I had the benefit of getting reached out by some amazing oil and gas professionals. And our leader, his name is Colin McClellan, and he's been working in the fracking industry for like over 20 years. When you count up all of the members on our team that work in oil and gas, they have over 100 years of experience of doing work, either from being a geochemist all the way to being a petroleum engineer. And so it was really cool because the idea is, you know, you got a whole bunch of people who came out of the boomer generation who were working in oil and gas and they had a particular way of doing it.
4:03And then that's when you have a big drop off from the 90s, early 2000s, and then things pick up again during the Shell Revolution in the 2010s. And in between that, you have just a lot of knowledge that's just getting lost as those older people are disappearing, younger people don't know where things are. And so the idea is like, what happens if we take AI and we put it through all these documents? 90 % of the documents in oil and gas is all unstructured. So it's not like, oh, okay, well, we can go through and just, put everything in Excel sheets. A lot of the work is like handwritten documentation.
4:38It lends itself into leasing and titling. It lends itself into safety and regulation. And so what happens if we use AI to go across all of that, make sense of it, synthesize it, but also make things safer for those people who are the reason why we have access to electricity and are able to enjoy data centers, which are the things that give us access to AI in the first place. So it's a lot of cool, you know, work and definitely something that if you would ask me 10 years ago, like the industry that would have really made me energized, I would have been like, oh yeah, finance, you know, because that's where I was before.
5:17But now it's like been amazing to be able to work with people who are able to say, hey, the AI that you guys are working on actually saves lives. And that at the end of the day, helps you sleep at night. That was amazing.
5:29Jon Krohn:That is really cool. There's, I had a brief two year stint where I was working in finance. And well, I was going to kind of say the opposite because for me, I couldn't, well, it wasn't so much that what I was going to say is like, for me, it was hard to stay motivated about making money for its own sake. You know we were trading financial derivatives and they were Actually, it is interesting, interestingly related to what you're doing, because I was mostly trading things in the crude energy complex. So oil and gas futures. Oh, commodities. Yeah. Yeah. And some options, mostly futures. And, you know, we were never taking delivery of oil barrels or pulling oil out of the ground, but it was a very liquid market.
6:16Jon Krohn:And so for applying algorithms to it and trying to get some kind of return, there's a lot of opportunity and more liquid markets. So that was what we picked to work on. But yeah, for me, it was tough to kind of stay motivated. Whereas actually, when you're pulling oil out of the ground and selling that, you're providing something useful to the world economy in a way that you're not when you're doing. And there's all kinds of things that you could be doing in finance. I mean, obviously, there are lots of finance that aspects of finance that are providing utility. I mean, there's, you know, financing for any kinds of projects, whether it's energy projects or otherwise, happens by financial firms for the most part.
6:58Jon Krohn:And that is that is something valuable. But the stuff that I was doing, it wasn't something that the world economy needed. um so anyway i can see how it would be motivating potentially um yeah to be to be in your industry so tell us a bit more about your role and what you're doing at glide yeah um so i am the foundational model engineer there um you know every day is something different because you know we are very early stage and so we're ramping up with much larger customers right now and so But the original idea is if you go to Claude right now and you start asking it a bunch of oil and gas terms, it's going to like maybe know the definition of some of them.
7:43It's not going to know how much it's related to each other. I think in general, there's a sense, especially in Silicon Valley, that like oil and gas are like, oh, man, if we assist people in oil and gas and there's like a moral failing there. But the thing is, is that, you know, the people who are drilling oil out of the ground, those are 18 year old kids, 20 year old folks, you know, people who could be people's grandparents in their 60s. Right. Like it's that bimodal people who are towards the end of their career has been doing 40 years and people who are very fresh out of high school. You want those people to be able to make it home, regardless of if you want eventually things to move to wind or nuclear.
8:20or these are the jobs that people have. This is, you know, the things that's holding up a lot of the American economy. And so when you have people say, you know, well, everybody should be adapting to AI. They want to adapt to AI, too, but they can't. There becomes that issue of if I go into, you know, Claude and I say, hey, read this JSA, which is a job safety report, and tell me about the wells in this area and if there's been some past instances and it's going through the document and doesn't understand the document, and says, everything looks good. And you're like, okay, I wrote everything right.
8:54They're like, yeah, you wrote everything right. And then somebody loses an arm or somebody falls into the well bed and they're not able to continue to walk again. Then it's a big problem. You wanna have a system that not only knows how to read those documents and synthesize those documents, but also an AI model that's made with that in mind that says like, okay, this is a job safety report. I'm gonna shoot it off to identify all of the rules in the municipality they live in. I'm going to give them suggestions. I'm going to push back. You can't have a sycophantic AI when you have people whose lives are on a line.
9:34So when a person says, oh, I checked this valve, everything is fine. Oh, no, actually looking at the data, this data is different than this other data. Here's some possible things that could have happened. You need to go back and look at these other forms before you move forward. So that's been, you know, pretty cool work because it requires me to not only just go through and look at documents, which, you know, oil and gas has no shortage of, but it also requires me to go and talk to people who are actually in the field. Cause you know, there's these really cool documents where they might be talking about how big a well is and they might have a point in a map and it might say something like, okay, the point in the map of how big the well is, is you're going to smoke three marbles and you're going to walk east and you're going to smoke two more,
10:20Jon Krohn:you're going to go west. And it's like, what does that even mean? Like, I don't smoke. Nobody around me smoke. How will I know? That's an old school way of kind of judging the length of something. That is the most old school method of measuring I've ever heard. That is so funny. And so you want to have AI that's able to say like, oh, here's where that is. Here's that trajectory of how far that is. And so, yeah, I mean, that's the work that I do. But, you know, we have people on our team that do some amazing work that even builds around that. Anything from identifying like, hey, you know, this well that is an inactive well that's close to a school.
11:08And so you'll want to go ahead and do X, Y, and Z things to make sure to stay secure so that you're not, you know, having oil drip and lean out into these different areas. And then you have other folks that do work like, you know, leasing and titling work and being able to make sure that those documents go to the United States government appropriately so that everybody's compliant. So there's a lot of really great stuff that comes out of building that underlying foundational model and then everybody taking that model and building on top of it.
11:40Jon Krohn:I love it. Really cool work. To get into some more detail on it, you describe yourself as a full stack foundation model builder, which makes a lot of sense to me. So like end to end, all aspects of it. And you're at the very cutting edge of machine learning research. So and language modeling research, we're going to get into NeurIPS papers that you have for listeners who don't know NeurIPS. It's the most prestigious academic AI conference that there is. Um, and so, yeah, tell us what you mean by full stack foundation model builder. What is involved other than finding out how many marbles it takes to get to a rig?
12:16Jon Krohn:What is it involved? Yeah. So, um, uh, really four steps. Uh, the first step is data curation. So this is the step that everybody hates. I actually yesterday sat down with my CTO and we had a very large whiteboard where we're just kind of going through and coming up with a multistage plan. We've been working on just data curation over the past six months. We have people who focus specifically on that. And what that means is I might get a textbook of, you know, a thousand pages. And within that textbook, each paragraph might have a highlighted word. And then around the highlighted word, it has a definition and then it has different type of topics.
13:03And a topic might start off in chapter one and not be talked about again until chapter five, until chapter 20. And so how do you get AI to be able to look at that document and be able to synthesize of like, okay, yes, this thing might have been brought up in paragraph one, but that doesn't mean that paragraph two talks about this thing. This is not necessarily important. You need to skip over that when it comes to this and find the next time it's important in chapter six. You can do it different ways. You can come up with different graph databasing strategies and stuff, which are strategies that we use.
13:37But to enrich that data, you can also have somebody who's a subject matter expert who goes through and goes, oh, yeah, yeah. So the title DV tool means this right here. And then somebody might say DV in this other chapter over here. And then somebody might just say tool over here. But you know what means that when it's talking about hydrostatic pressure. So you want to have somebody who can help make sense of all of these things so that when we're creating that graph, you know, database, everything's appropriately attached. So that's the first big portion that constantly changes. And we kind of make jokes sometimes that at the end of the day, our company is just as much of a data company, if not more than an AI company.
14:22A lot of it's just building data models. so
14:26Jon Krohn:this episode is brought to you by Excel Data, the leader in autonomous data and AI most enterprises run their data across a sprawl of systems Snowflake, Databricks, BigQuery, Hadoop, Iceberg, Lake Houses, On-Prem clusters, ugh, and then stitch together a different operational tool for every layer Excel Data's Xlake changes all of that, Xlake provides a single architecture for hybrid compute, control, and intelligence. Data and AI observability, agentic data management, AI data engineering, Kubernetes native compute, and more, all on one platform that runs where your data already live. Xlake, the architecture built for what data need to become.
15:10Jon Krohn:See it at exceldata.io. That's A-C-C-E-L data dot I-O. Something that we talk about on the show a lot. In fact, in a recent episode of the show, episode 993, we had two book authors on the show. They have a brand new book called Architected Intelligence. Their names are Jacob Miller and Jeremy Mumford. And one of the key takeaways from that episode, if not the main takeaway, is that successful AI deployments are about having the right data for your model. And so data engineering can be a bigger part of any successful AI deployment than the model itself. So agreeing with you 100%. No, absolutely.
15:51And then, you know, there's that next step, right? What shape do you want the data to go into when you're using these data pipelines? You know, do you, I have built, you know, custom tokenizers in our office so that when we have people who are pulling things from our, so first off, starting off with, you know, bespoke tokenizers, but building on top of the embeddings models. so that when we have people who are doing other processes, which might not necessarily have to do with AI specifically, but still they need that knowledge base from that graph, they're able to use that and be able to, oh, okay, well, I have a document, a drilling document, and that drilling document is able to better be chunked and OCR works better because it's able to identify the information that it needs to pull from that document versus it's actually going through and, you know, having to put bounding boxes across multiple things within a document, we can better avoid that on a text-based level versus having to always defer to a more expensive vision model.
17:01So that's another portion there. And then there's a fun part, I guess, the part that everybody gets excited about, which is, you know, building the foundational model. There's multiple different ways that we do it. We do some continued pre-training. So just taking a model that we do benchmarks constantly and experiments constantly on the different types of models that are out there, both open source and the big box ones that you get behind an API. We compare which ones we're doing better at certain tasks. And then for those ones that are open source models that we feel like we can move forward with, do some continued pre-training specifically on oil and gas to make sure that we can extend that ability forward.
17:43But then there's also doing some reinforcement learning. You know, a model doesn't understand the world around us. And so being able to put physical validation around it by creating physics environments and saying, OK, go ahead and do these things. I want you to use Python REPL to create some code that's doing the mathematics. Yes, but it has to be defined within the space of physics. I can take a very long time to try to teach a model physics, but then it's going to overfit to physics, especially if I have a smaller model. And then somebody put oil and gas on it and everything kind of falls apart.
18:21Or I can just have the very basic rules of physics that are unchanging inside of an environment. And so doing that type of stuff, figuring out what things can we strip out from, you know, model training and instead put it in reinforcement learning is something that we do a lot of stuff with. And then the last part is inference, which is the, I find most exciting, but a lot of people find less interesting. You know, the way that you train a model is going to affect the way you have to serve a model. And so if I do some really cool environment that's super amazing and makes the model do better than all the other models out there and I'm like, okay, hey, I love this reinforcement learning thing.
19:01Like, it's great. Well, reinforcement learning models are bursty. And so that's going to affect the compute level that we can use when we're doing inference. That's very pricey. So suddenly you have to have this inference thing on so that you can keep your uptime at 99.9 % because you have SLAs you got to stick to. But you have this huge cluster that's supposed to be able to handle that burstiness. So being smart with how am I building the model and how can I serve it to the public so that they can get the best results from the model?
19:33Jon Krohn:This does sound like a fascinating role. One of the most fascinating roles I think I've ever heard of on this show. You're doing, like when you say end-to-end, really end-to-end, custom tokenizers, embedding tricks, reinforcement learning. You've got optical character recognition, so machine vision models, tons of text models that you've been describing there. Wow, really exciting place to be working. And then, of course, the inference aspects. So I actually have a question for you. I feel embarrassed that I don't know what it means for a reinforcement learning model to be bursty. What does that mean, bursty?
20:11Yeah, so the idea of burstiness is that so reinforcement learning models, there's a lot of calculations that is kind of running at a given time. And a lot of those calculations can be done on CPU. They don't have to be done on GPU. So the GPU is just sitting there idle, waiting for the reward to come out. And then from that reward, know the trajectory. So the direction that it needs to go into at the next portion of training, which is great because it makes your model do very well. But it also means that you have just like, you know, maybe I might have an 8 billion parameter model and I might be able to get away with putting it on an A100.
20:50Just one, you know, and I can move forward. But if I have that environment that is running and crunching a bunch of numbers and then suddenly it spikes up by sending a whole bunch of information at once to the GPU for the model to move to the next portion of training, that one A100 is going to fail. Suddenly I need multiples. So that can end up being an issue.
21:14Jon Krohn:I see. Yeah. So that's why they call it bursty because it's bursting suddenly with all of this activity that wasn't happening for a long time. And so that's something that a lot of, you know, when I was at Microsoft, when I first started there, I was working in reinforcement learning. And so before I moved over to AI and I find a lot of people now are starting to starting in AI and then are moving to reinforcement learning. And so there's a lot of people starting to figure out like, oh, we can use reinforcement learning to make the model better. And then they're running into that bursty issue and then they go like, oh, my goodness, I'm getting all these segfault issues, which is just like.
21:51The GPU is just kind of overloaded. It's shutting off. And so trying to be smart with how I build things, but also being way more conservative with the way that I dedicate GPU space. A lot of times people might go use 90 % of the overhead on my model and on a KV cache. I have to scale that down so that I can have space for those reinforcement learning results to come up so the model can get better over time and do that self-adaptive stuff we want to have a model doing reusing reinforcement learning.
22:28Jon Krohn:I got you. That's really cool. So I would personally kind of consider reinforcement learning to be a part of AI, but it's interesting how you describe them as distinct things. And I think when you're doing that, you're probably thinking of, correct me if I'm wrong, but when you're saying you move from reinforcement learning to AI and a lot of people go from AI to reinforcement learning, the AI you're describing, there's probably like large language models, right? Yeah, yeah. So I'm using it in a sense of how people describe it. So it's funny because like when I first entered the field, and again, all these things are underneath the umbrella of AI, but I graduated from graduate school, went and was working doing some quantitative methods work over at a bank.
23:07That was where I was in finance. and we were using models that like Monte Carlo simulations, but a lot of people might not consider that AI. And then I went and started doing, you know, machine learning and building recommendation systems. And a lot of people still might not consider that AI was doing recommendation systems and attaching them to natural language processing tooling at the time. This is your, you know, glove embeddings work before we started moving into transformers as much. And then that reinforcement learning piece and then over to AI. So I think it's more of how people have defined AI in the past, even though we know that they're all underneath that same umbrella.
23:49Jon Krohn:Definitely. And something pretty exciting here that I can't believe I haven't mentioned about you, at least in this conversation. Maybe it'll be included in my intro when I record that after you and I have this conversation. But you and I have something in common. I don't know if you know this, but we've both done PhDs at Oxford. oh yeah yeah and uh so yeah so when you're talking about you know grad school you know and you kind of you're being very uh you know you're humble about the grad school that you've done because you did uh correct me if i'm wrong on any of this i'm basically getting it off of your linkedin profile but you did a master's in political science and economics at Columbia.
24:30Jon Krohn:And then you went on to do a PhD at Oxford in social data science. And, you know, in that whole time, focused on these kinds of AI things, generative AI, responsible AI, model explainability, MLOps. It's a very practical, for a PhD, quite practical stuff that must be useful now in your role. Yeah, absolutely. Though, I must say, I did drop out of the PhD. I paused it. I see. You have to do the pause. I paused it last year. I probably had one more year. I was still working on a dissertation. Who knows? Probably by the end of the year, I'll pick it back up. But yeah, a lot of the work at Oxford, and also too, in between, I also got a technical fellowship at Stanford.
25:22um and like a lot of my work at oxford definitely informs it both was informed by my time of industry in industry but also informs my time in industry it's kind of hard for me to figure out where one begins and one ends yeah that's good so yeah it's
Read the full transcript
25:39Jon Krohn:been pretty cool well i wish i could do another one now because i really i didn't know i'm not i'm not going to i'm just saying i theoretically wish i could i mean i i guess never say never i have no plans to do a second one. But my point is that when, when you've been in industry already, and then you go to do a PhD, you have so much appreciation for how these things will be used in the real world. Whereas like, I had just been living in a bubble of textbooks and academic papers until I went into industry and you, you don't appreciate how special that time is. And And I think if I had been in industry first, I probably, you know, not to say that I didn't make good use of the PhD, like published a dozen academic papers.
26:25Jon Krohn:And, you know, so from that perspective of like academic success, it was successful. But if I had industry experience, I think I genuinely 10x in terms of the impact and the synergies that I could have noticed, you know, using the time well on things that will matter in industry later on, maybe even, you know, making impactful strides in the real world while doing the PhD. Yeah, as opposed to kind of just being in the bubble. And yeah, so I think that you, you know, you had a great path there. And, yeah, I also now see that I, if I clicked on the ellipses on the dot, dot, dot on your bio, it does say left to focus on industry.
27:10Yeah. Oh, it's totally fine. I mean, I, I will say though, to, to your point, I think that it is a interesting time right now in AI. I, if I had went straight to my PhD after Columbia, like if I went straight in, I think I would have just finished it. because it's like you know also oxford has a the european phd and the us phd is his own other i'm sure podcast where we could have a discussion on the differences of it but you know three years versus six years is a very different oh totally when it comes to but also it requires you to have a different what's the word understanding of what you want to do immediately like you start Oxford and people already know this is my advisor.
28:02This is my focus. Where's the US to get more time because of the timeline to do it? And so I can definitely see on one end, if I would have done it immediately, I would have just checked it out. And then, you know, because AI wasn't moving as fast at the time, we were, you know, like most in actual English processing, even though the the attention is all you need paper came out in 2017. Most of the industry was still using natural language, like the NLTK toolkit in Python, pretty much up until XactoT came out. And so you could still kind of get away with, you know, doing that and taking more time.
28:41I think it's because I started the PhD in 2023. And then all of these things are taking off at the same time. I might have applied in 22 before, right before Chaggy BC came out, started it. And then like, that was like, you know, on the run and I was at Microsoft. So I was right in the middle of it. I think that that is a lot of why I was like, okay, if I want to do research, that's cutting edge right now, it's an industry, but I'm sure the tide will change and we'll be right back to academia and I'll go ahead and finish the PhD and make my professor father very happy.
29:24Jon Krohn:Yeah, he'll appreciate it if you come back for sure. But regardless, am I correct in understanding based on your LinkedIn profile that basically throughout the PhD, you were also working full time at Microsoft and then later ISO? That is wild to me. Especially with that kind of European three-year program. I mean, yeah, you've got to be really productive. That's wild, Jasmine. Yeah, you got to be productive. You have to have, and this is, you know, for anybody who might be listening to this and thinking about taking that route, you have to be open and honest with everybody. So I was on a Microsoft research team when I first started at Project Bonsai over at Microsoft.
30:12And then I moved over to Microsoft AI. With my original manager, I had to have a conversation with him about what this PhD was going to be, what types of things I would have to publish. I had to, even during my application process, I applied with them knowing that I would be working full time and got advisors who had other students who were doing the same thing. You know, one of my advisors, Chris, Professor Chris Russell, one of his other people he was advising was at Google DeepMind. So, you know, having that understanding that like we're going to work on this together. But then, you know, throughout the time, having open and honest conversations, both with he and my other advisor, Brent Middlestat, and being like, hey, here's what's going on in industry.
31:00Here's what I'm seeing. Here's what I'm able to do. And here's what I'm not able to do. And even with me deciding to take a step back and go back to industry, that was a group decision. That was me coming to them with a paper I had worked on and a patent that I had worked on and them being like, oh, great. Where is this patent taking you? And I was like, to a company. And they were like, oh, if you're going to do a startup, you should probably, you know, consider because, you know, Professor Chris Russell, he's, you know, had his fair share doing that type of work. So I think that you have to have the right type of team around you.
31:37If you don't, it doesn't matter how much you want to do it. It's going to be very hard to, you know, get something like that going.
31:45Jon Krohn:For sure. Well, congrats on, you know, getting a couple of years into it. And yeah, hopefully, I don't know. I mean, not hopefully, like whatever happens, happens. You're going to make the right decision. But yeah, one way or another, it's going to be interesting. And regardless of whether you formally finished the PhD or not, you're doing frontier AI research. Let's dig into some of that. So at Collide, that oil and gas AI company that you're at right now, you are building RIGS, not oil RIGS, but R-I-G-G, lowercase s. So RIGS, it's a model name, I guess. You can tell us if RIGS stands for something.
32:21Jon Krohn:But my understanding from our research is that it is a set of petroleum engineering foundation models for the oil and gas industries. And so I guess that ties into a lot of the stuff that you've been talking about already in this episode. But as part of RIGS, you're tackling one of the most intractable challenges with large language models, and that's evaluation. So a few weeks ago, you dropped a new paper that identifies four systematic failures that make current evaluation frameworks structurally inadequate for assessing deployed agentic systems. And as a consequence of those failures, reward hacking is a predictable consequence.
33:04Jon Krohn:So tell us about these systematic failures, what reward hacking is, why it's a problem, and then what your solution is. Yeah. So I kind of was touching on how we use reinforcement learning with our models. And so there is this very interesting thing that has driven me absolutely insane for a very long time in AI, which is regardless of what type of sub genre of AI that you're doing, which is that it takes an insane amount of time for you to figure out if your experiment actually worked. Like you spend time, OK, I'm going to tune these parameters. I'm going to take this pathway. I'm going to do all this.
33:48Great. It's wonderful. You begin moving forward with it. Like, let's say for I'm going to talk about specifically large language models here. But let's say you have your data, you decided your chunk sizes for retrieval. You've done all these things that's like super fancy and amazing. You're like, OK, we're going to do QLaura because that's the that's the greatest thing to do. Great. You begin running training. You even do some checkpointing and stuff like that until your checkpoint is finished. you don't actually know how well your model is performing. So you have to wait for a checkpoint for eval.
34:22That could be anywhere from two hours if you're doing something very quick, fast and dirty with a really small model. Or if you're using a very large model, it can be days before you're able to actually sit here and go, okay, the model is moving in the right direction. Huge waste of time. And again, this isn't just large language models. This was even back when I was doing Monte Carlo simulations. It takes way too long. So my idea was, okay, what happens if we're already kind of spinning up these environments? What happens if we use that Monte Carlo simulation idea? Okay, I'm going to have a model that's training within these certain specs, within this data.
35:04And based off of how it's moved, like the direction of trajectories that is moving in, like within this reward model, because I can quickly pull out the calculations for the reward as it's going before it moves to a checkpoint. I'm going to do Monte Carlo simulations on those and see how well with these types of conditions it's going to be by the time it gets to the checkpoint. So that means I want to get my first burst of my first pass of reward that's coming out of that reinforcement learning model. it's probably about an hour or some odd of time, I can very quickly go, okay, based off of these trajectories, it's going to go in this direction versus the other direction.
35:46By the time I've had that do a few times, I'll be able to say like, oh, it's moving in the wrong direction. We need to go ahead and chunk this experiment. And then I can run ablation tests as well. So let's say like I have an environment that's, you know, doing a calculation, so on a return, I can go ahead and say, what happens if I have multiple different small variants of this experiment and I change just one small thing? So I'm going to weight the data more heavily towards equations for this set, or I'll weight the data more heavily towards oil and gas. And I can do the same thing of like, I can run multiple experiments at the same time, do simulations on all of them, see which one is going to be the best, and then move forward with the winner.
36:31And so that basic idea was like, okay, this is great because it means that we don't have to wait for days. I think the worst part, especially when you had a startup. So like back when I was at Morgan Stanley, back when I was at Microsoft, you just ate that time. You had a company that was big enough to where you can eat that time. You can eat that compute cost, whatever, who cares? Microsoft was charging itself to pay itself for the compute. It didn't matter. But like when you're at a startup, you have to be scrappy. you have to identify what can actually work. And also, you know, I might, we might have somebody who sells to our customer on Monday, hey, Riggs, you know, performed well, that X, Y, and Z thing by Friday, they want to go ahead and demo it.
37:18I don't have time to, you know, wait, because if the demo doesn't work out, we might not hit our numbers, we might not hit our funding round. So that was where that came from. Speaking about the failures, though. So when you have reinforcement learning, what ends up happening is models are very smart in that they are able to identify the least that it has to be able to do in order to get a good reward. Because over time, it's seeing the same things over and over again. It starts to go like, okay, fine, I'll just do X, Y, and Z. And that's another reason why these types of things, these way of evaluating can be so powerful because in that first hour, it hasn't seen the same thing over and over again.
38:04And so when I do a simulation on how well it's going to do by the checkpoint, if it doesn't look like it's going to be good, but then it comes out on a checkpoint and it did amazing, I already know what reward had because it went ahead and identified of like, oh, she's asking me these questions that talk about math. And I don't have to know the definition of something if I just guess at the equation or, oh, we're doing retrieval and I'm going to be judged based off of my retrieval at K and she's going to do three and five and judge me based off of how much I'd be getting three and five. So I'm just going to guess the five most common equations that are used in this set.
38:44One of them will be right. And then I can move to the next bit. And then suddenly if I give it away more equations, it falls apart. So that's another issue. And then another one of the failures that comes with the word hacking is the data that you're using to train is different than the way people are actually going to use the model. So I'm going to ask questions at eval, like there's something called the petroleum engineering exam. And at eval, I might have the model test itself on how well it can do with the petroleum engineering exam. Okay, given that you have this well and you have this, these amount of mouths and you have this tool, what's the hydrostatic pressure?
39:25Okay, great. People are not going to ask a perfectly formulated, created by academics question with multiple choice question and answer about oil and gas. They're going to come in and go, hey, tell me about this well. And it needs to be able to tell them about that well. And then they're going to go, what's the decline curve? Might even say, what's the decline curve? Some people just say decline curve. Some people just say decline. So being able to make sure that the model as well as able to converge to that, that's again where those simulations can come in at. because now suddenly I can begin injecting questions into the environment and see how well it does on those questions at the moment.
40:13So the model might be doing very well to be able to answer those SP &E questions. And I'm like, okay, great. This is exciting. But it might not do well at being able to understand the underlying text and definitions and how things are related.
40:32So so that that way I can make the adjustments as it goes on.
40:36Jon Krohn:Nice, yeah, so to kind of sum up this reward hacking issue, it's the situation with reinforcement learning models in particular where they are, the way that you program the reinforcement learning model to learn is by having some kind of reward function where a simple example that people can easily visualize is if you're training it to be really good at Tetris, then you can use the point score in Tetris as the reward because the higher the point score in the game, the better they're performing. But that's easy in Tetris, which is just a video game in the real world when the reward is more complex.
41:12Jon Krohn:You gave lots of examples there, but like, you know, for a self-driving car, how do you, if you were using reinforcement learning to, so is it, you know, every meter that you travel without an accident, you know, you get some reward. And then if you hit a pedestrian, and you get like negative reward. And so, you know, coming up with what the reward function is, is really complicated in the real world. And you can end up with, you might think that you've crafted a really good reward function, but the reward hacking happens when, as you described, the reinforcement learning model figures out some way of getting to like that Tetris high score or that self-driving car high score, getting a really high score in a way that you didn't imagine it would and that is different than what you how you wanted it to behave yes and that happens a lot with a reinforcement learning algorithm called grpo um which is just like a group policy i don't want my model to just know one thing which is like maybe math right i want to know uh be able to choose the right equation i want it to be able to exist within the physical constraints i wanted to answer or a question about oil and gas well, wanting to retrieve well.
42:27But in the end, I want the final answer that it has to be correct. And maybe the model might learn that, oh, I actually don't have to, like, I don't get as big of a reward if I know the math equation correct, for example. So I'll just not learn the math equation correct. I'll just make sure that I get the final answer correct by looking at some keyword that's in a sentence and know that every time there's keywords in a sentence, the final answer is likely going to be one of these four things. And so when you have that issue, you want to make sure that your model isn't over optimizing just to give a CBO award.
43:09And then by the time it gets in someone's hands, it's completely useless because it's just trying to get the final right answer. And we kind of see this a little bit with different models on different varying levels that use reinforcement learning. Like, you know, people have issues with a lot of models where they're asking a question and it will just like, you know, hey, you're right. This is actually the correct answer. And you're like, no, actually, that's not the right answer. It's like, you're right. My bad. And you're like, okay, like, I need you to be grounded. And essentially, that's what, you know, that paper is for it's try to have a level of groundedness that's deterministic within it.
43:53And also, too, another big portion is, you know, handling how much it takes to train these models with reinforcement learning. Many times you have to have a reference model, which is a whole other conversation that we can have. And so by having these deterministic environments, you don't have to have a reference model. You can simply take the weights of the model and how it was performing at a given time, use that as a simulation and then be able to compare how well your model is doing, which means lower compute, don't have that big old issue with burst, as big of an issue with burstiness and, you know, able to actually create a model that works well on a MacBook as opposed to having to always go for the much larger GPU cluster.
44:38Jon Krohn:Yeah, so great explanation there, Jasmia, thank you. You mentioned grounding in there a few times, but I don't think you mentioned specifically that in the paper, your framework for mitigating this reward hacking is called grounded continuous evaluation, GCE, right? Yeah. Yes. Yeah. Yeah. It sounds pretty cool. So yeah, using simulation-based fine tuning to make everything happen, get around these reward hacking issues and get all those kinds of intermediate steps that you were describing as well, not just the final answer. Yes, exactly. And, you know, injecting some determinism. It's so interesting that I think everyone is just kind of assumed that the non-deterministic nature of large language models is just a foregone conclusion.
45:23It's like, oh, non-deterministic, who knows why it decided that. And I think that it's cool that there is a level of non-determinism. I mean, like, even in this conversation we as humans um we are existing both on a non-deterministic and a deterministic frame like the non-deterministic frame is that i might not know exactly like i might not know exactly what words you might say next i don't even know what words i'm gonna say next yeah but i do know that in the context of this conversation that there's going to be certain guardrails there's going to be certain things that you aren't going to do like you're not going to jump up and down and start doing much of jumping jacks.
46:03Jon Krohn:Because that then would be outside of context. But, you know, like identifying that there are certain contextual cues that kind of inform us on different rules of behaviors that we kind of exist within. And we've all decided to trust each other with these things. Like, and anyone who violates that, we're like, whoa, that's a violation of a norm. And so, and here's me pulling in my undergrad philosophy major here, but, you know, we need to begin to inject those levels of norms and determinism in our models. And by having that GCE, you know, I try my best to explain things without using so much jargon, because not everybody comes to know the same backgrounds, but that's really where that idea comes from.
46:56I want the model to have a sense of this is a space where you have to exist and anything that deviates from that is not going to receive a reward.
47:06Jon Krohn:Really great analogy there, bringing into human deterministic or non-deterministic norms as well. I love that. Let's move on to another paper that, or actually I'm not 100 % sure it was a paper. Definitely it was a NeurIPS presentation. You can fill me in here. So it was a 2023 NeurIPS. You demonstrated that scaling bad data makes models worse, not better. Tell us about that. Okay, great. Yeah, that was a paper that I worked on. So this is where I might sound a little bit cheeky and arrogant here. When I made that paper, everyone thought that I was absolutely ridiculous. So there's this calculation that came out of OpenAI.
47:49It's a paper written by Kaplan et al. Essentially, it's this idea that is the underlying rule that people still kind of exist by, but less so, which is that if you have data, compute, and parameters, you can linearly extend the capabilities of a model as much as you just keep on adding on to it. There's some people who push back and say, well, data is most important. Those people quickly lost. I think there's some people who said, well, no, you know, parameter is more important. Those people also have been to loss over time. Compute has still remained that important thing where we're seeing it now.
48:30We have a compute shortage happening right now where people go like, OK, just throw more compute at it. Like if you just add more data and more compute, then the model is going to be even more amazing. And back in 2023, I was kind of looking at it from a perspective that comes from economics. So there's like a law of diminishing returns in economics. At a certain point, like money is great. For example, we live in a country that is, you know, very much fueled by money. But if our government is going through something like a recession or something like that, and they print like a bajillion dollars, then suddenly the value of money goes down.
49:07There's a level of printing that has to happen, but it comes at a cap, which suddenly the laws of supply and demand end up becoming an issue. And so that was the idea of the paper. The idea was data is great. Bad data is not great for a couple of reasons. One of them, if you have model data that hasn't been deduplicated, which a lot of these models in order to have just that much shared data, not only are they not deduplicated, but it's nearly impossible to deduplicate them because they're huge, you know, huge amounts of data. It begins to make the model get more forgetful over time, less intelligent over time, begin to overfit over time.
49:51It's so interesting because in the large-ingles models, we stopped talking about overfitting, even though that was one of the basic issues of machine learning. And then the same thing from just – I spoke a little bit about compute, but focus more on data. But also from a compute perspective, right? You can't compute your way out of bad data. So it doesn't matter how many GPU clusters I can have. I can have a machine with, you know, 10 ,000, you know, GPU clusters and all of it has really bad data and the model is going to perform worse versus if I have less compute and I have less data, but it's better data.
50:30It's more quality data. So, yeah, those were, you know, really great paper at the time. Again, a lot of people were skeptical. I got accepted into a workshop. I, you know, top 5 % paper, which is pretty exciting because I just assumed everyone was going to boo. You know, it ended up kind of, we're kind of starting to see that happening and at play now where you consistently see a 70 billion or 120 billion parameter model be able to outperform in a much larger, you know, 1 trillion plus parameter models.
51:07Jon Krohn:Sure. Yeah. And I did find that paper while you were speaking. I'm embarrassed that I wasn't sure whether, well, because it did also have a poster presentation, right? Yes. And so I think that's where I got confused. But so the paper title is Scaling Laws or the Laws of Diminishing Returns, How Scaling Law Quality Data Degrades Model Performance. And so I'll have that for people in the show notes, as well as your earlier, well, sorry, the paper we were discussing earlier, that came out only a couple of weeks ago called Beyond Static Snapshots. So I'll have the archive link to that one available as well.
51:45Jon Krohn:Jasmia, this has been a fascinating technical episode, but I do also want to cover a couple of other topics quickly before we wrap up. So you are a champion for linguistic equity. Yes. So dedicated to ensuring that the future of AI includes the voices of the marginalized through your pioneering work in creating the African-American Vernacular English data set. So AAVE. Tell us about that data set and why it's so important. Yeah. So we talked a little bit about, you know, my, my, my major. So my background when I was in college was a philosophy, English, and poli sci, and a focus in poli sci was on economics.
52:25I'm an economics nerd, but the idea came from kind of the way I was raised. My, My dad is an Afro-Latino Caribbean man, and my mom is a West African Black American woman. My mom, you know, Low County, South Carolina, Gullah accent. She just says all types of stuff that I'm like, that's not, I don't think that's fully English.
52:54And my dad, you know, spoke many different languages when he was growing up and has an accent. And I grew up in Atlantic, Georgia, and they have their own vernacular way of saying things, way of doing things. And one of the things that I kept on picking up was, and it started off with me doing my thesis in undergrad on this, which is that African-American vernacular English is like this tie that binds between the African languages and between English. and if languages can have a tie that binds to each other to where I don't necessarily have to understand exactly, like I don't have to understand tweet to kind of get a sense of what somebody is saying because the way that they might say something.
53:44And so I was like, this is very interesting. This is something that can help unlock a lot of richness for language models, for how we communicate with each other, for how we spread information. And so, yeah, I went and was working on that at Stanford because I was like, I think that there can be a way for us to create a corpus that is able to tie these two things together. I ended up not moving forward with it, though, because I began realizing that linguistic equity, yes, is about representation, which is very important. But there's also a beauty, and actually I was told this by one of my peers that works in a lot of African language work that actually worked on a model called AfroBert, which has recently been taken down.
54:37But there's a beauty in people being able to own and have ownership of their language. So yes, we want to have languages and systems and linguistic equity from a sense of language models being able to communicate across different people. But also having an understanding that there are certain individuals who feel unsafe by the powers that be. And so giving those powers access to be able to mimic their language, which gives them a sense of feeling safe. But then being able to use it for things that might be nefarious isn't the best thing to do. And so giving people ownership of their language, acknowledgement, but also ownership.
55:18and being able to say, I am a guest in an AAVE home, even though I identify, obviously I'm African-American, but identifying that with my parents being from different places, there is a certain level of guestness that I have to acknowledge and allow people to have ownership over their own thing. If they decide they want to come out with an AAVE language bot, somebody who's 80, somebody who comes from being an American descendant of a slave, then that's a tooling that they can use. And I've had some people reach out to me to do that. But I can provide the bridge. That doesn't mean that I need to be the person who crosses it, if that makes sense.
56:00Jon Krohn:I see. Yeah, that does make sense. That's a, yeah, it's a nuance that I didn't appreciate as I was asking the question. And yeah, I'm now enlightened on it. I didn't appreciate it as I did the research. I was like, So in that most recent response, you did mention, you know, your family background. And so I just want to have, I've got one last question for you, uh, before, well, technically I have three questions left for you, but I asked my final two questions. I ask every guest those same two, uh, and they're short, quick ones. So this is the last one that's like for you specifically. And it kind of, it goes back even further into your background.
56:36So, you know, we've talked about, you know, you did your undergrad at Tulane.
56:41Jon Krohn:We didn't specifically mention Tulane, but great university. You did your undergrad at Masters of Columbia. You did the technical research at Stanford HAI Lab, which is a super famous lab, you know, that's world leading. The Oxford PhD work. You've worked at Microsoft. You've been a data strategist at Morgan Stanley, a head of machine learning at the Motley Fool. And now you're at Collide at the front here, building end-to-end models for the oil and gas industry and publishing in NeurIPS. I mean, you're doing such amazing things all over the place. And so, I don't know. I don't know if you have kind of insights into how you became like this, that you're just able to achieve so much.
57:32Jon Krohn:You know, like what drives you? And, you know, do you have any tips for our listeners who would like to be, you know, competing at the frontier in so many ways like you have over your career? Sure. I laugh because I'm like, oh, I don't know if it's always a great thing to be like this. I think I've just always been a kid. I already talked about my parents a bit. I, you know, Atlanta is a what they call a chocolate city. It's very black city where people come from all over the world there. But, you know, my specific type of mixture is a more rare combination. And so my sisters and I, we just kind of existed within our own kind of pod.
58:22And then I have cousins who like all of my uncle is married interracially. And so all of my cousins, we all have like different racial identities. And so I kind of grew up knowing what it kind of felt like to like look like everybody, but not feel like everybody and kind of be excluded from some stuff and accepted in other stuff and try to like navigate that world. But also navigate that world with the understanding that we have a country that has multiple layers of privilege. And in some places I had an insane amount of privilege that were different than other people who looked like me and in other places that had less privilege than people who looked like me.
59:00And so I've always wanted to answer why that is. What does that mean? And how can I use that experience for positivity as opposed to walking around and being one of those people who's just like trying to create these strange, like, you know, intra-racial, I guess that would be called, race wars, like, well, Caribbean versus like all that silliness. or, you know, without being somebody who felt like I had to try to over-index by being somebody that I wasn't. Just saying like, okay, well, I'll just, you know, completely abdicate and I'll just only be around kids that were like, you know, white kids or whatever.
59:43Like having a beauty of the fact that I could exist in all spaces. And what would that look like? Like, so that's why I did my, you know, undergraduate the way that I did. But that's also why I took the jobs that I ended up taking, because all of them in some way were kind of helping me answer questions that I had for myself. When I first left college, I was working for the Clinton campaign and it was like, obviously I'm working in politics. So that's answering some questions that I have about what I can do to be able to help people. But then I went to graduate school and graduated and the thing that paid student loans was going to be in finance.
1:00:28But then learned very quickly that finances is a lot of the reasons why people have privilege versus not in this country and learning that a lot of the questions that a lot of things that are separating people might appear racial in some aspects because they are, but also are class interplays. and Atlanta definitely gave me a masterclass in that. And so me just trying to learn how to navigate what that means to be a petite bourgeois versus being somebody who might be, you know, a proletariat or in some other space. And then, you know, Motley Fool, same thing, was the distinction of that, right?
1:01:11This idea of a company that's like, we want to give the average person the same access that everybody gets when they have a financial advisor. and me being like, great, I'll do that work. And then we go into Microsoft and working at first for a reinforcement learning team that was actually in industrial manufacturing. We were building reinforcement learning agents that were powering robotic arms for companies like Shell and PepsiCo and Abbott Nutrition when they had a Similac shortage. And doing that work and going, wow, this is really cool. Like I can actually help single mothers or just any mothers who are having a hard time accessing Simulac during COVID because I can create a robotic arm that can do that work so that people can stay at home and be able to avoid, you know, being too close to each other during those COVID things.
1:02:04So I think my career has always been guided by that passion of trying to figure out where I exist and how I can exist in a way that can give of whatever privilege that I have while also protecting other people who don't have it, while also acknowledging that I don't have all of the privilege in the world, but trying to find ways to express to people who do ways that we can all work together.
1:02:30Jon Krohn:You brought together in that final answer, you know, economics, social aspects, philosophy. In a few minutes, you managed to capture a lot of your influences and a lot of the most important factors driving so much that's happening in our world today or at any given point. And, uh, yeah, so I'm looking forward to see where, seeing where your journey takes you next and how you continue to, to answer questions about the world for yourself and for us through your publications. Really fascinating. So for people who are also excited about what you're up to, what is the, what's the best way to follow you in your work?
1:03:08Yeah. Uh, best way is LinkedIn. Um, occasionally I post articles on medium or towards data science. but I'm most consistent at LinkedIn. So that's the best place.
1:03:21Jon Krohn:And you probably also crossbook when you do something on those, on Medium or towards data science, you probably post better on LinkedIn as well. I do. Great. So one stop shop. Fantastic. And then I apologize that I was supposed to warn you about this last question before we started recording. I've forgotten a couple of times with guests recently. I always ask my guests for a book recommendation. And yeah, it doesn't need to be something in our space. It can be in our space. It can be fiction, nonfiction, whatever you want. Do you have anything for us? Yeah. Uh, values, value systems design by Batya Friedman.
1:03:59She wasn't the only one who wrote it, but she's the head author of it. That book has guided all of my decisions in AI, regardless of what function of AI I was working in. Essentially, the idea of the book is how do you make sure that you're building systems that are rooted in the values that are important to us as other humans? Because the technology that we create is simply an extinction of human values. And so it's absolutely amazing work. I think everybody should read it.
1:04:32Jon Krohn:Nice. Thank you for that recommendation, Jasmia. And thank you for this whole episode, taking time out of your busy day to, yeah, enlighten my listeners on so many topics. Really enjoyed this conversation, Jasmia, and hope to have you on the show again in the future. Yeah, it'd be awesome. Very interesting episode indeed with the brilliant Jasmia Henry in it. She covered her work at Collide where she builds foundation models for the petroleum industry in which 90 % of documents are unstructured and AI's accuracy can be the difference between a routine shift and someone losing a limb. She described her full stack foundation model building as having four stages, first curating and structuring data, then building bespoke tokenizers and embeddings, training the model through continued pre-training and reinforcement learning is the third step, and then finally optimizing for inference at scale.
1:05:25Jon Krohn:She also talked about how reinforcement learning models are bursty because they idle the GPU during reward calculation and then dump enormous loads on it all at once. And finally, that reward hacking happens when a model discovers the laziest path to a high reward rather than actually learning the task we wanted it to. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Jasmia's social media profiles, as well as my own at superdatascience.com slash 995. Thanks, of course, to everyone on the Super Data Science Podcast team, our podcast manager, Sonia Breivich, media editor, Mario Pombo, partnerships manager, Natalie Zajski, researcher, Serge Massis, and our founder, Kirill Aramengo.
1:06:07Jon Krohn:Thanks to all of them for producing another super episode for us today for enabling that super team to create this free podcast for you. we are deeply grateful to our sponsors you can support the show by checking out our sponsors links which are in the show notes and if you'd ever like to sponsor the show yourself you can find out how at johnkrone.com slash podcast otherwise you can help us out by sharing this episode with folks who would like to learn about end-to-end foundation model training reinforcement learning and so on review the episode on your favorite podcasting platform or on youtube subscribe if you're not already a subscriber.
1:06:43Jon Krohn:But most importantly, I just hope you'll keep on tuning in. I'm so grateful to have you listening and I hope I can continue to make episodes you'd love for years and years to come. Till next time, keep on rocking it out there and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
Jazmia Henry joins Jon Krohn to break down what it actually takes to build end-to-end foundation models for the energy industry. From wrangling decades of handwritten oil-and-gas documents into usable training data, to bespoke tokenizers, reinforcement learning, and inference at scale, Jazmia walks through every stage of the stack. Along the way she explains why reinforcement learning models are "bursty," what reward hacking is and how her Grounded Continuous Evaluation framework fixes it, and revisits the 2023 NeurIPS paper that argued, to widespread skepticism at the time, that scaling bad data degrades model performance.
Additional materials: https://www.superdatascience.com/995
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(10:06) The User Agnosticism Tenet
(20:02) The Zillow Offers parable
(23:25) Why workflows should come before agents
(29:57) Why data engineering is the bedrock of AI
(52:41) Why velocity is the only durable moat




