In short
Episode 998 is a “best of May 2026” roundup from Super Data Science Podcast. It highlights: (1) AI agents accelerating cybersecurity risk. Jon discusses Anthropic’s Claude Mythos Preview (announced April 7; not public) and Rubrik’s view that AI agents can find and exploit software vulnerabilities at machine speed, making prevention/detection insufficient and requiring rapid breach recovery.
Guests
Anika Gupta and Cal Adubebe (Rubrik; cybersecurity/data security leaders).
Key claims
Mythos-like capabilities will proliferate (Rubrik frames Mythos as the visible point on a curve); AI increases attack surface via faster, chained actions; organizations lack “trust infrastructure.” Notable example: a misconfigured boolean flag in Anthropic code reportedly exposed the Cloud codebase; also internal data exposure risks like accidental sharing of employee salaries. (2) Notebooks as agent “working memory” (Dr. Trevor Mance; Marimo pair skill with Claude Code/Codex). (3) Full-stack foundation model building (Jasmia Henry; data curation, tokenizers/embeddings, continued pretraining + RL, inference).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI and Cybersecurity Concerns
0:20 to 2:26
Discussion on Anthropic's Claude Mythos and its implications for cybersecurity.
“Anthropic's recent Claude Mythos preview has reset expectations of what AI can do in the wrong hands, finding and exploiting software vulnerabilities at machine speed.”
Evolving Cybersecurity Challenges
2:26 to 4:25
Exploring how AI changes the landscape of cybersecurity and response strategies.
“But what the reality that now with AI agents, you can find and exploit these vulnerabilities much, much faster and at machine speed that requires machine speed response.”
AI Agents and Coding Workflows
4:25 to 6:42
The integration of coding agents in workflows, particularly in data analysis.
“We don't yet have the right trust infrastructure in place in most cases to prevent models from causing harm on their own.”
Utilizing Notebooks with AI
6:42 to 8:52
Discussion on the use of notebooks in AI coding and data exploration.
“And like, you know, we had access to files and permissions, but we didn't broadly use these permissions to take large scale actions.”
Building Foundation Models
8:52 to 13:20
Insight into the stages of building foundation models for AI.
“when you're coding, but when you're exploring data, when you're doing data visualization, it could be interesting for our audience to hear what a power user of Merimo, how they use that that kind of tool.”
Introduction to NeurIPS and Full Stack Models
14:01 to 16:46
Learn about the NeurIPS conference and the concept of full stack foundation model building.
“For listeners who don't know NeurIPS, it's the most prestigious academic AI conference that there is.”
Data Curation and Model Building Steps
16:46 to 20:44
Discover the critical steps in data curation and how they impact AI model creation.
“is that successful AI deployments are about having the right data for your model.”
Failing Fast in AI Projects
21:46 to 26:13
Understand the importance of failing quickly in AI development and organizational behavior changes needed.
“link, you're supporting our show, notion.com slash superdata.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:This is episode number 998, our In Case You Missed It in May episode. Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. This is an In Case You Missed It episode that highlights the best parts of conversations we had on the show over the past month. My first clip is from episode 989. Anthropic's recent Claude Mythos preview has reset expectations of what AI can do in the wrong hands, finding and exploiting software vulnerabilities at machine speed. I speak to Anika Gupta and Cal Adubebe of Rubrik about why the old cybersecurity playbook of prevention and detection is no longer enough and how AI agents themselves are becoming a new source of data exposure inside organizations.
0:50Jon Krohn:And then now, more recently, just a couple of weeks ago at the time of recording, on April 7th, Anthropic announced Claude Mythos Preview. And so this is a model that famously they haven't released to the public. You know, there's only people testing it privately. Anthropic's own estimate, however, is that Mythos class capabilities will proliferate to other labs within six to 18 months with open-weight versions to follow. So, yeah, the reason supposedly why they didn't release it to the public is because of its prolific ability to identify cybersecurity issues and exploit them if, you know, it's not being used by a friendly person.
1:31Now, I also do think it's brilliant marketing to say that we can't give it to the public.
1:39Jon Krohn:But yeah, from inside rubric, is mythos the inflection or just the most visible point on a curve you'd already been planning against? You know, I think when you look at mythos, it's definitely the most visible point on the curve that we've already been planning for. Because even if you look at models like Opus, for instance, they also are able to find vulnerabilities. They just require a little bit more handholding to get there, whereas Mithos, you can point it at an open source repository and tell it to find me all the vulnerabilities and it will find you the vulnerabilities. And this is just the nature of where AI is going.
2:17The challenge is that the cybersecurity industry as a whole has been super focused on attack prevention and attack detection because it used to be the case that attackers would enter into your system. They would sit there for weeks, if not months, sifting through your data, finding the opportune time to actually exploit this vulnerability and really perpetrate the attack and then actually do that in a visible way such that the company then has to react to it. But what the reality that now with AI agents, you can find and exploit these vulnerabilities much, much faster and at machine speed that requires machine speed response.
3:04And no longer can you hope that if you detect an attacker, detect a breach, okay, maybe you can cut that off before the attacker has done damage. Instead, you have to assume that you've already been breached and you need a plan for how are you actually going to recover and ensure that in that recovery, you're minimizing any impact to your overall business. And that's the business that Rubric has been in. So I feel like there's never been a more important time for what we've been doing because what we're ensuring is that you can actually restore your applications, your data, your identity, your infrastructure back up and running extremely quickly.
3:46And in the case that an attacker gets into your system and destroys your data, your identities, or your infrastructure. What I find really interesting is it's like this twofold problem. And we're kind of feeding into it from the enterprise perspective. There's still massive FOMO. Everyone has to use AI. Everyone has to be more productive and it's coming from leadership down. And so we're very quickly opening up actually the surface area that models that are like mythos can actually chain together vulnerability. So it's one, you've got this problem of adoption as increasing surface area. And then two, these models themselves, and I know we're going to nerd out a little bit more about it, but I get so passionate about this.
4:28We don't yet have the right trust infrastructure in place in most cases to prevent models from causing harm on their own. So you don't even need the bad guys.
4:37Jon Krohn:Right. Yeah, it's pretty wild. I mean, it was Anthropic themselves that actually had, there was some flag, like a Boolean flag, that they got wrong in some, like somehow it was possible for them. You guys might be able to explain this better than me because you actually are cybersecurity experts. Cal's been one for three weeks. But there was like a flag that was set the wrong way in some code that was pushed to GitHub. And so that allowed people to see the entirety of the Cloud code base. How does something like that happen? Well, I think that's an interesting case. There's a lot of different scenarios where things have been exposed publicly that wasn't supposed to be.
5:19There are cases where code has been exposed externally because the AI agents posted it into the wrong repository. That's not what happened in the Anthropic case. In an anthropic case, there was, again, a Boolean flag that was misconfigured and sent out. But the reality is, is that the challenge that we have today is that AI agents, they're very outcome focused. They're going and trying to say, how do I get this job done in the least steps possible, the fastest way possible? And they're not taking into account the same rules that us as humans, as employees working in an organization know to follow.
5:55You know, they haven't gone through all the security training that we've gone through. They haven't gone through the developer best practices training. Now, obviously, they have that information in their models broadly, but they're not optimizing for the same thing that we're optimizing for. And what that means is that inevitably, AI agents are going to make mistakes. Inevitably, they're going to do things that you didn't want them to do in their pursuit of this outcome of a task that you've given them. I find it really interesting that these models, because we've been investing in the capabilities for them to iterate autonomously, think of it as an exhaustive search that's now happening across all of the data assets and tools and configurations of that tool.
6:40Humans in the past, we actually had a lot of security through obscurity. We might have had permissions. I've been studying my cyber stuff. You have. I've impressed. We had security through obscurity. And like, you know, we had access to files and permissions, but we didn't broadly use these permissions to take large scale actions. And now we have these very hyperproductive agentic systems that can do an exhaustive search of everything that's available to it to accomplish the task. And so a lot of cyber controls really were never designed with this reality in mind. And it's not just about the mistakes that AI can make or what it can publicly expose.
7:19there are other kinds of data exposure risks that you have even internally, right? You could say, hey, like, you know, salaries of an employee could be accidentally shared with another employee because you've hooked up your AI agents to your employee and compensation systems. And that is really scary, too. So there's a lot of implications, both internally and externally, to exactly what Cal is talking about. And as the models themselves are able to take just more and more and more and more steps independently, it's really hard to diagnose like where in that 20 step process did this happen and what caused it.
8:01And you're certainly not going to be able to fix that after the fact. You have to figure out how you're actually going to fix this during, like while this agent is actually running.
8:10Jon Krohn:While agents are creating new cybersecurity risks, They also are reshaping how we work productively with data. In episode 991, I speak with Dr. Trevor Mance about why code notebooks have become the natural working memory for AI coding agents. Trevor walks me through the Marimo pair skill, which lets you drive a notebook from your agent collaborating with Claude Code or Codex in real time as you load, explore, and visualize your data. Beyond the coding, the development experience, you also obviously have a lot of experience doing data analytics, data visualization, data science in kind of a notebook environment.
8:49Jon Krohn:So maybe even walk through for us, you've now talked us through what your workflow is like when you're coding, but when you're exploring data, when you're doing data visualization, it could be interesting for our audience to hear what a power user of Merimo, how they use that that kind of tool. Yeah. I've always thought that notebooks hold this really special place in the ecosystem because they are this environment that marries your code and data together. And so although I've been sort of this bare bones text editor for my traditional software development, I've always used some kind of notebook when I wanted to do data work or some sort of live REPL.
9:27I think my first thing I was in was R Studio and then I switched over to Python and started using Jupyter and now I'm a Marimo user. Yeah, so I usually would spin up a notebook from the command line and just jump in and start loading my data. And now I find myself with the Marimo pair skill, like I'm often now like the way that I, anytime I get an issue or often the place that I start to work on a problem is now in my agent encoding tool. So for me, that's often cloud code. If that looks like some task that I want to like load some data, explore that data, I'm going to start asking to start up a marimapare session and then start describing a very high level, sort of, these are the data sets that I want to look at, these are libraries I think that I want to do.
10:12And sort of being a lot, like I was saying before, a lot more declarative about the approach that I want to take. I know exactly what I want to do with my data.
10:19Jon Krohn:And this is maybe a really dumb question, but when you start doing that marimapare session, that's in the command line or that's in a notebook environment? That's in your agent. So it would be like open code, cloud code, codex. You can just say, like, hey, I want to start working on my notebook. The Merimo pair skill teaches the agent how to start Merimo or connect to Merimo for you. And then you're driving Merimo from your agent. Whoa. So it would be kind of like you'd have it open in a browser window. And so you'd be like, you'd be working and, you know, say you could have a cloud code running a terminal.
10:53Jon Krohn:And so So you're typing into that terminal, but then the agent is changing things in your Marimo notebook. Exactly. It's open separately. Yes. And you can also collaborate on it. You have the Marimo notebook interface. So if the code that it generates or there's some tweak that you want to make to a cell, then you can go edit that by hand in the Marimo notebook. But then you can ask Claude, you say, hey, I circled something that's interesting in this notebook. Can you tell me about it? And Claude has access to that state now because of the way that they share this context. So it's like all that rich information that historically has been like trapped in that environment now is more context that you can offer to the agent beyond your file system.
11:32Jon Krohn:That's so cool. I love that. Does it do things like if you, so if it had kind of been developing a notebook, had been getting going on some data analysis for you, and then you go over to your Chrome browser or whatever, and you make a change to a cell, could you end up in a situation where either in the notebook experience or in the terminal window, it's watching what you're doing and it's just kind of like, hey, I like what you did there. Yeah. Right now, what we're trying to get our better feedback loops, if you've made edits, for us to stream those back, there's a feature inside of MCP that's called channels I think we might explore in that direction.
12:14At the moment, it's a lot. Claude has, or your agent has like access to running code inside the kernel. So if you've made a change, it can just ask, it can grab any state that it wants inside the notebook. It can see your cells. It can take screenshots of cells now and look at those. So if you say like, hey, there's something interesting here. I'm not quite sure what it is. Go off and, you know, use the other tools at your expense to help me understand what this is. At your expense. Exactly. Then that's something that's now context that you can feed back into getting somewhere with your data. So I like to think of it as when I started really adopting agent encoding tools for traditional software development, software ideas started to feel a lot cheaper in the sense that, okay, that's not going to take me an afternoon.
12:58I could do that really, maybe just do this proof of concept. And I've had so many of those times inside of notebooks where there's something I know that I should look at or explore, but I don't want to figure out the API or I can't remember exactly this thing. And instead I could just, I can just say, maybe we should try five, like these different validations first, and then, you know, let it work on that problem and then make some plots and bring me back when, when I can make my decision.
13:19Jon Krohn:So far, we've looked at using AI agents and securing against them, but what does it actually take to build the foundation models behind these agents from scratch? In episode 995, Jasmia Henry walks me through her work as a full stack foundation model builder. We cover all four stages of the process, the often unglamorous slog of data curation, building bespoke tokenizers and embeddings, model training and reinforcement learning, and then finally, the inference layer that serves it all to end users. You describe yourself as a full stack foundation model builder, which makes a lot of sense to me.
13:54Jon Krohn:So like end to end, all aspects of it. And you're at the very cutting edge of machine learning research and language modeling research. We're going to get into NeurIPS papers that you have For listeners who don't know NeurIPS, it's the most prestigious academic AI conference that there is. And so, yeah, tell us what you mean by full stack foundation model builder. What is involved other than finding out how many marbles it takes to get to a rig? What is it involved? Yeah, so really four steps. The first step is data curation. So this is the step that everybody hates. I actually yesterday sat down with my CTO and we had a very large whiteboard where we were just kind of going through and coming up with a multistage plan.
14:46We've been working on just data curation over the past six months. We have people who focus specifically on that. And what that means is I might get a textbook of, you know, a thousand pages. and within that textbook, each paragraph might have a highlighted word and then around the highlighted word, it has a definition and then it has different types of topics. And a topic might start off in chapter one and not be talked about again until chapter five, until chapter 20. And so how do you get AI to be able to look at that document and be able to synthesize of like, okay, yes, this thing might have been brought up in paragraph one, but that doesn't mean that paragraph two talks about this thing.
15:28This is not necessarily important. You need to skip over that when it comes to this and find the next time it's important in chapter six. You can do it different ways. You can come up with different graph databasing strategies and stuff, which are strategies that we use. But to enrich that data, you can also have somebody who's a subject matter expert who goes through and goes, oh, yeah, yeah. So the title DB tool means this right here. And then somebody might say DV in this other chapter over here. And then somebody might just say tool over here. But you know, it means that when it's talking about hydrostatic pressure.
16:03So you want to have somebody who can help make sense of all of these things so that when we're creating that graph, you know, database, everything's appropriately attached. So that's the first big portion that constantly changes. And we kind of make jokes sometimes that at the end of the day, our company is just as much of a data company, if not more than an AI company. A lot of it's just building data models.
16:30Jon Krohn:Something that we talk about on the show a lot. In fact, in a recent episode of the show, episode 993, we had two book authors on the show. They have a brand new book called Architected Intelligence. Their names are Jacob Miller and Jeremy Mumford. And one of the key takeaways from that episode, if not the main takeaway, is that successful AI deployments are about having the right data for your model. And so data engineering can be a bigger part of any successful AI deployment than the model itself. So agreeing with you 100%. No, absolutely. And then there's that next step, right? What shape do you want the data to go into when you're using these data pipelines?
17:11Do you, I have built custom tokenizers in our office so that when we have people who are pulling things from our... So first off, starting off with bespoke tokenizers, but building on top of the embeddings models so that when we have people who are doing other processes, which might not necessarily have to do with AI specifically, but still they need that knowledge base from that graph, they're able to use that and be able to, oh, okay, well, I have a document, a drilling document. And that drilling document is able to better be chunked and OCR works better because it's able to identify the information that it needs to pull from that document versus it's actually going through and having to put bounding boxes across multiple things within a document.
18:05We can better avoid that on a text-based level versus having to always defer to a more expensive vision models. So that's another portion there. And then there's a fun part, I guess, the part that everybody gets excited about, which is, you know, building the foundational model. There's multiple different ways that we do it. We do some continued pre-training. So just taking a model that we do benchmarks constantly and experiments constantly on the different types of models that are out there, both open source and the big box ones that you get behind the API. We compare which ones we're doing better at certain tasks.
18:44And then for those ones that are open source models that we feel like we can move forward with, do some continued pre-training specifically on oil and gas to make sure that we can extend that ability forward. But then there's also doing some reinforcement learning. You know, a model doesn't understand the world around us. And so being able to put physical validation around it by creating physics environments and saying, OK, go ahead and do these things. I want you to use Python REPL to create some code that's doing the mathematics. Yes. But it has to be defined within the space of physics. I can take a very long time to try to teach a model physics, but then it's going to overfit to physics.
19:29So if I have a smaller model and then somebody put oil and gas on it and everything kind of falls apart. Or I can just have the very basic rules of physics that are unchanging inside of an environment. And so doing that type of stuff, figuring out what things can we strip out from, you know, model training and instead put it in reinforcement learning is something that we do a lot of stuff with. And then the last part is inference, which is the I find most exciting, but a lot of people find less interesting. you know, the way that you train a model is going to affect the way you have to serve a model.
20:04And so if I do some really cool environment that's super amazing and makes the model do better than all the other models out there and I'm like, okay, I love this reinforcement learning thing. Like it's great. Well, reinforcement learning models are bursty. And so that's going to affect the compute level that we can use when we're doing inference. That's very pricey. So suddenly you have to have this inference thing on so that you can keep your uptime at 99.9 % because you have SLAs you got to stick to. But you have this huge cluster that's supposed to be able to handle that burstiness. So being smart with how am I building the model and how can I serve it to the public so that they can get the best results from the model.
20:48Jon Krohn:Agents are getting smarter every day, but even the The smartest agents get stuck without the right context and the right tools. That's where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side. And now, their new developer platform is turning that workspace into infrastructure developers can build on. What sets Notion apart is that the collaborative workspace and the platform you build on are the same thing, with permissions, context, and governance baked in from day one. Workers for instance are Notion-hosted sandboxes where I can run database syncs without standing up my own infrastructure.
21:22Jon Krohn:That means I can pull guest research, episode analytics, and my consulting client data from all the scattered systems where they live, then keep them synced in Notion databases automatically. My agents and my human team work from one single source of truth. Learn more about Notion's developer platform today at notion.com slash superdata. That's all lowercase letters, notion.com slash superdata to try Notion's developer platform today. And when you use our link, you're supporting our show, notion.com slash superdata. We're rounding up a great month with episode 993, in which I sit down with Jeremy Mumford and Jacob Miller, the co-authors of the brand new book, Architected Intelligence.
22:01Jon Krohn:Jeremy and Jacob argue that the most expensive AI mistake an organization can make is failing slowly, sticking with prototypes long past their sell-by date because the traditional software mindset says you have to. We discuss how the define-build-feedback loop changes when going from zero to prototype is nearly instantaneous, and why narrowly defined lanes of ownership are one of the best antidotes to the organizational antibodies that fight against failure. another part that i like about the first chapters of your book is how you write that failing slowly is costly and failing quickly can become your superpower so if you're in an organization and we're going to get to speed again well at least a few more times actually in this episode speed is kind of a key thing today it seems but being able to fail quickly uh is something that's so important.
22:57Jon Krohn:However, organizational antibodies like the health system, something that was designed to keep an organization healthy, these structures in an organization, let's call them antibodies to give it kind of like a biological analogy, they actively fight against failure. They're always trying to make sure that everything succeeds. And in this day and age where things are moving so quickly in terms of technologies, capabilities, what you could be doing with your platform, that historically useful antibody is now, it's like an autoimmune disease. So what are your tips on forcing failure fast on AI projects?
23:43It's an interesting problem because I think enterprises, they have these organizational antibodies, as you describe, and they start to really want to stick with projects and ideas maybe longer than they should. And I think also they are kind of accustomed to the traditional way of software engineering and building products, which is that the initial product will maybe not have full functionality, but if you stick with it long enough, you can build out to the point where it's useful. But now building with AI and building AI infused products and features, going from zero to prototype is nearly instantaneous.
24:27And so it requires a mindset shift and an organizational behavior shift that can be hard to execute at the enterprise level. Just to give an example, I remember early on at Pattern, we wanted to build a text-to-SQL chatbot. And having interviewed hundreds of people for AI engineer roles and talking to people across the industry. I swear every single company tried to do this. We all tried to build some sort of text-to-SQL chatbot that would just solve the problem of being able to retrieve data from our endlessly complex databases. And the thing about all these demos is that within like a day, you could have an amazing demo that seemed to kind of have something amazing on the surface.
Read the full transcript
25:14and so a lot of companies kind of committed to that and they were like well if we just stick with it we'll be able to kind of keep going with it i think the key thing there is that now that it's so easy and quick to build prototypes you really need to prove value on day one if you can't prove value on day one then you need to toss it out and start all over again and that's a really hard thing to tell people to just toss out their ideas but i think we're seeing that shift I look at AWS just this last week. They've kind of announced that they're doing this massive shift from their old AWS bedrock system to this new project mantle.
25:51And they talked about how they had six engineers over like two or three months just completely rebuild their entire AI infrastructure stack from the ground up. That's the kind of commitment and resolve you need to have in this new age to just say, we got to start over from scratch and build something with immediate value from the beginning. And so that's really what I think it comes down to. With one approach, when we're creating something of value, the product cycle is that you define what we want to build. We then build it and then we get feedback. And it used to be that building was the bottleneck.
26:26So you could be really bad. So if it's define, build and feedback, and if build was always the bottleneck, you could always just be bad at definition and feedback because that was never the bottleneck anyway. And so when building no longer becomes the bottleneck, the critical, the crucial step is definition and feedback. And so one of the organizational changes we've seen help is actually narrowly defining these lanes and say, hey, you two, three people, you can make whatever decisions you want. This is what you're chasing after. This is what success looks like. And you guys can try and fail over and over and over within that lane.
27:02And by involving fewer people, it ends up that those who are working on the project feel much more comfortable trying things out and failing because you're not failing in front of such a large audience.
27:14Jon Krohn:All right, that's it for today's In Case You Missed It episode. Be sure not to miss any of our exciting upcoming episodes, including upcoming episode 1000. Subscribe to this podcast if you haven't already. But most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
In this month’s episode of ICYMI, Jon Krohn explores how AI agents are simultaneously creating new risks and unlocking powerful new ways of working with data. Hear from Anneka Gupta, Cal Al-Dhubaib, Trevor Manz, Jazmia Henry, Jeremy Mumford, and Jacob Miller, discussing why the old cybersecurity playbook breaks down in the age of Claude Mythos, how the notebook became an AI agent’s working memory, what it really takes to build a foundation model from scratch, and why failing slowly is the most expensive mistake an AI team can make.
Additional materials: www.superdatascience.com/998
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(00:40) Why Claude Mythos Changes Everything About Cybersecurity
(08:11) Why Your Notebook Should Be Your Agent’s Working Memory
(13:19) What It Actually Takes to Build a Foundation Model From Scratch
(20:46) Failing Slowly Is the Most Expensive AI Mistake




