In short
Marimo, an open-source “reactive” Python notebook replacement for data/AI that targets Jupyter’s reproducibility and maintainability failures.
Key claims
Jupyter notebooks are hard to version-control (JSON diffs), easy to break via out-of-order execution (“hidden state”), and often miss package/version provenance. Marimo keeps code and outputs in sync like a spreadsheet (reactivity), supports “lazy” execution to avoid expensive auto-runs, and stores notebooks as executable Python files for CLI/pipelines and reuse/import. It can hide code to publish a performant data app (including WebAssembly in-browser serving). It also offers an AI “generate with AI” editor that can generate code using in-memory dataframes’ schema.
Notable examples
changing A in cell 1 auto-updates dependent outputs; selecting points on a scatterplot returns a dataframe; deleting a cell invalidates dependent in-memory objects; sandbox mode records installed package versions and recreates isolated virtualenvs.
Guests
Dr. Akshay Agrawal, co-founder and CEO of Marimo; Stanford ML PhD (2018–2021) focused on vector embeddings/dimensionality reduction; previously worked on open-source ML/optimization research (e.g., PyMDE, CVXPY).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction of Marimo and Its Background
0:25 to 1:41
Discussion about Marimo, an open-source replacement for Jupyter Notebooks.
“Akshay, welcome to the Super Data Science Podcast.”
Meaning Behind the Name Marimo
1:41 to 3:48
Dr. Agrawal shares the significance of the name Marimo and its cultural context.
“And so the reason why you're there once a week is because AIX has invested in your startup, which you are co-founder and CEO of, and it's Marimo.”
Reproducibility Challenges with Jupyter Notebooks
3:48 to 4:52
Dr. Agrawal discusses the reproducibility issues faced using Jupyter Notebooks during his PhD.
“And they were like, okay, like this is the thing.”
Common Reproducibility Pain Points
4:52 to 6:08
Elaboration on specific pain points related to reproducibility in data science.
“So anybody who has used Jupyter Notebooks has experienced difficulties with reproducibility.”
Key Features of Marimo
6:08 to 9:39
Overview of Marimo's unique features like being open-source, Git-friendly, and reactive.
“And so you can't really reuse code across them easily without jumping through some hoops.”
Reactivity in Marimo: Enhancing User Experience
9:39 to 12:30
Discussion on how Marimo's reactivity improves data exploration and reproducibility.
“So I've got some of the key adjectives here around Merrimo from our research.”
Interactivity Features: Sliders and Plots
12:30 to 14:00
Exploration of how interactive UI elements like sliders enhance data analysis in Marimo.
“And like, I don't want that to start automatically.”
Enhancing Python Notebooks with Interactivity
14:00 to 15:14
Learn about the interactivity features in Marimo that enhance data exploration.
“And even if you do, like the things aren't going to automatically run with the updated value of your slider.”
Marimo's Transition from Notebook to Data App
16:01 to 18:14
Explore how Marimo allows users to easily convert notebooks into data apps.
“I I appreciate the tangible example, something that's always useful for our listeners, for sure.”
Commercial Strategy Behind Open Source Marimo
18:14 to 23:26
Understand the commercial strategy for Marimo as an open-source tool emphasizing user empowerment.
“growth, PLG, where you have users that are volunteering to try this product out.”
Show all 23 chapters
AI Assistance in Marimo for Enhanced Coding
23:26 to 28:01
Discover how Marimo integrates AI to assist users in writing code more effectively.
“lot less difficult, a lot more streamlined.”
Building Merimo: A Design Journey
28:01 to 31:28
Learn about the process and inspiration behind creating Merimo from scratch.
“And that actually is a great segue to my next question, which is around, you've previously said that, that there was a joy to rethinking a system from first principles, as you said about creating Merimo.”
Reproducibility Challenges in Jupyter Notebooks
31:29 to 34:24
Explore the reproducibility issues faced by Jupyter Notebooks and how Merimo addresses them.
“for any of our listeners out there that are thinking about building a product.”
Innovative Features of Merimo
34:25 to 36:45
Discover how Merimo enhances data exploration and analysis workflows for users.
“So basically, if you start Marimo in a mode that we call sandbox from the command line or however, every time you install a package, Marimo has a very nice slick package installation UI.”
Prioritizing Features in Product Development
36:46 to 38:54
Understand how Merimo's team prioritizes features based on user feedback and needs.
“So it really enables like really, really tangible like data exploration, data analysis workflows that were like really, really hard, if not impossible to do before.”
Building a Strong Open Source Community
38:55 to 40:54
Learn about fostering community engagement and contributions in open source projects.
“so open source projects often promise community but not every project earns it so what do you think separates tools that spark tons of devotion like the pluto project that you were describing there in Julia.”
Getting Started with Open Source Contributions
40:55 to 42:00
Find out how newcomers can begin contributing to open source projects effectively.
“and it depends on how you want to contribute.”
Getting Started with Open Source Development
42:00 to 43:13
Learn how to begin contributing to open source projects effectively.
“Would you recommend, yeah, where would you recommend they get started?”
Understanding Convex Optimization
43:13 to 45:22
Discover the principles of convex optimization and its applications.
“learning framework TensorFlow that probably most of our listeners are familiar with to the vector embeddings computation library.”
Real-world Applications of Convex Optimization
45:22 to 47:24
Explore examples of how convex optimization is applied in various fields.
“Convex optimization is just the subset of those problems we know we can solve super efficiently, super reliably, provably.”
Making Machine Learning Actionable
48:24 to 51:01
Learn about the importance of making mathematical concepts actionable in ML.
“As I've actually all of your explanations today, you do a great job of taking complex concepts and making them seem really approachable.”
Book Recommendations
51:01 to 52:51
Get insights on two impactful books from the guest.
“And yeah, I'm so glad that Sean introduced us.”
Connecting with Akshay Agrawal
52:51 to 56:00
Find out how to follow and connect with Akshay's work.
“Two great and very different recommendations for us.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:Welcome to episode number 911. I'm your host, Jon Krohn. Starting with today's episode, we're trying out much shorter intros. So I'll quickly tell you that you're in for a treat with today's guest, Dr. Akshay Agrawal. Akshay has built a clever open source replacement for Jupyter Notebooks based on the pain points he experienced with notebooks while completing his PhD at Stanford. Akshay is smart, well spoken, humble. I think you'll really enjoy this one. This episode of Super Data Science is made possible by Dell, NVIDIA, and AWS. Akshay, welcome to the Super Data Science Podcast. Where are you calling in from today?
0:36I'm in the Bay Area in Redwood City. Thanks for having me, John.
0:39Jon Krohn:Classic AI tech founder in the Bay Area. What are the odds? You were introduced to me by Sean Johnson, who is an investor at AIX Ventures in San Francisco. If people are interested in hearing an episode involving, you know, what investors are looking for in AI startups and where the opportunities are for AI startups or anybody building AI solutions for enterprises. That's episode 895. It's a great episode. Sean is amazing. I guess you spend time kind of with him in person in San Francisco sometimes. Yeah, I do. At the AIX office, I do spend time with him. And he is amazing. He definitely is. I've known him for a bit over a year.
1:25Yeah. And I'll see him usually almost every week.
1:29Jon Krohn:Oh, wow. Really? From what I could see in the background, and people can see it in the video version of 895, it's a beautiful looking office. Yeah. The AIX office is very, very nice. Yeah. Up in San Francisco. Nice. And so the reason why you're there once a week is because AIX has invested in your startup, which you are co-founder and CEO of, and it's Marimo. And so I want to start off with that name because our research research has dug up that Marimo means bouncy play ball that grows on water. And so it's a species of ball-shaped green algae native to Japan. Does that relate in any way to your company name?
2:12I mean, yes. I mean, sort of. I can give you two stories, okay? Like the outer story and like the actual inner story. So the outer story is that Marimo is a next generation computational notebook for Python, data, and AI. and well what are the previous generation ones or what are related ones so the one that most people know about are is the jupiter notebook now jupiter it's a very large sphere a very very large sphere there's another notebook for a different programming language called uh pluto jail but it was a smaller sphere it's kind of cuter as some would say and marimo is the next generation one for python even smaller sphere very adorable very beloved uh by uh especially in japanese culture but also like plant enthusiasts and like marimo moss balls like one thing that's actually cool about them is that they're actually assembled from like these like tiny strands of moss like like that all get clumped together and they like kind of like live at the bottom of lake beds and they like kind of like intermix and like mingle and like um like together they're like the greater they're greater than the sum of their parts and like marima notebooks in ways that we we might discuss are sort of similar like the collections of cells that like when taken together are greater than the sum of the parts so that's sort of like the story conceptually of what the name is
3:40Jon Krohn:nice there's a lot to that yeah nice give me the also there's there's also an also which is um during like the pandemic uh my partner and i we had a marima moss ball in our apartment and we like really loved it. And they were like, okay, like this is the thing. Like we love this thing. Let's just, let's just call. Yeah. It just makes sense. It also abbreviates well import Marimo as MO, which is you need that two character abbreviation for Python, Python libraries. I love that. That's, that's a great story. And they are super cute. I hadn't, I wasn't aware of them before, but I, I looked, I looked at some and so I'm going to, I'm going to make a special request to our video editor, Mario.
4:23Jon Krohn:to overlay some cute photos from the internet of marimo balls uh yeah to so that at least the people watching on youtube can see some of those as you were describing it and people who are in the audio only format you're just gonna have to look it up if you're driving don't do it while you're driving but when you get to a stoplight check them out they're super cute nice so um Your work at Mary Mo was inspired in part by reproducibility pain. So anybody who has used Jupyter Notebooks has experienced difficulties with reproducibility. And I have lots to go into around that later in the episode. But just as a starting point, that kind of reproducibility pain that you experienced as a PhD student at Stanford, especially with things like co-authors handing you broken notebooks and model training, often beginning without checkpoints, version control, or recorded data providence.
5:20Jon Krohn:Tell us about those pain points that you experienced as a PhD student and how you came up with your solution. Yeah, definitely. So I guess for a bit of background, so I did my PhD in machine learning. It was like 2018 to 2021. So I did it at Stanford. And like I specialized in like vector embeddings, high dimensionality vector, dimensionality reduction. So like finding structure in these like high dimensional spaces. And so when you when you work on things like that, like it's really important to be able to see your data while you work on it. So like notebooks and like Jupyter notebooks are like super, super useful.
5:58Like you just like without that iterative experience, like that kind of research is just like would be way harder. so like very grateful for working in jupiter notebook for having jupiter notebooks um but like i guess as i did that research like there was there were a few like shortcomings some of them having to do with like reproducibility that sort of like came to the forefront as i like worked on them like worked with jupiter notebooks on an almost daily basis um and so when it comes to like reproducibility and maintainability i guess there's a few different aspects so uh for one, the file format for Jupyter Notebooks, they are stored as these JSON blobs.
6:42And so you can't really reuse code across them easily without jumping through some hoops. So you end up just duplicating your notebooks for slightly different experiments. Whereas if you were writing software, you would just like import functions from one module to another. But also this JSON blob like made like version control pretty hard because like, like you'll run a notebook and you might not even change any code, but like maybe the image changes or like some metadata changes and you get like a big diff. So you're, it's like hard to know like how your experiment code also like actually changed over time.
7:22Another issue that would come up. And so like for, for some uses of jupiter like this is like a feature but i'm about to describe like for some uses it is a feature for the kinds of stuff where like you you're like no you're doing science or you're doing data engineering and you need things to be deterministic it is not a feature in my opinion at least uh so in a jupiter notebook you can like run cells like out of order so you can like say you know you can run cell number one then cell number three then cell number two then cell number one again and then maybe you forgot to run cell number three cell number three maybe depended on cell number one and you just totally forgot you never ran it and the point is like now your program memory is in some like strange sort of quantum state that you don't understand uh and then like you'll figure out like four hours later eight hours later like oh crap like none of my experiment results are valid um and so like that that problem is like kind of like broadly referred to as hidden state and it just comes down to the fact that in a jupiter notebook there is no um like the notebook runtime doesn't know how your cells are related to each other um and so that that was like a big pain point like if you contrast that to like an excel for example like if you update a for like a value in a cell everything recalculates and you know the world makes sense and that doesn't happen by default in a jupiter notebook and i'm talking a lot but like the last thing and we can go into any of these is like package requirements like so like when you're working with a notebook like you might forget to like track what packages like they're like you if you use pytorch like this specific version of pytorch and this specific version of transformers you might just like forget to like put that down in like a requirements.txt or pi project toml like researchers often did but then that means when you pass your notebook to someone else, they can't get the same results that are serialized in it.
9:18And so that's like, that's sort of just another angle in which like this sort of reproducibility problem sort of becomes sort of front and center. All right.
9:28Jon Krohn:So yeah, we're going to get into more detail on reproducibility, all of the issues over the course of this episode, all the pain points that probably many of our listeners have experienced and the solutions that you have for them. So So I've got some of the key adjectives here around Merrimo from our research. So it's open source, it's reproducible, it's Git-friendly, it's AI native, and it's a reactive notebook. So a few things to dig into here. The open source thing, for example, you guys are doing very well. Four million downloads, I understand, already at the time of recording, which is very cool.
10:00It must be amazing to see people just voluntarily being like, this is a solution for me.
10:06Jon Krohn:I need this tool. you know, something that you've spent all this time writing and developing and leading the product on. So that must be cool. But what I want to talk about is the word reactive here. And so the idea is that it's a better tool, Marimo is a better tool because it nudges users towards better data science practice by being reactive in this way. Tell us more about that. Definitely. Nudges is a good word. One of our users described it as gentle parenting. Okay, so yeah, what reactive means is it's just a, what it means is that Marimo keeps the code on the page in sync with the outputs you see.
10:49So basically, say you have three cells. The first one says A equals one. The second says B equals A plus one. And the third one is C equals A plus B. And then you output those. You output the variable in each, right? So you see in the first cell, A equals 1, B equals A plus 1 equals 2, and then C equals A plus B, 1 plus 2 equals 3. Now what reactive means is if you go back to the first cell and change the value of A from 1 to 2, the next two cells will, and then you run that cell, so you change the value of A equals 1 to 2, and then you run it, the next two cells will automatically run with the updated value of A.
11:33And so their outputs will update to three and four or something. So they'll recalculate like a spreadsheet, sort of like you're used to. And why does this matter? Why is this a good thing? I guess two things. So one is actually, it just leads to way faster data exploration. Instead of having to manually change the value of a variable and then hunt for all the other cells that use that variable, Maremo will just automatically run them. And you can just see really quickly in real time, like, you know, what do the new outputs look like? So that's one. But the other one, which is maybe more subtle, is the reproducibility aspect, which is like, because if you forget to run a cell that depends on a variable you just changed, now you can't trust any of the outputs in your notebook.
12:23And, you know, your experiment is kind abort. So that's the main, that's what reactivity gives you. And like one thing I'll just pause, because if there's machine learning practitioners or AI folks in here, the next question that always comes up, which I'll anticipate is, oh, but like some of my cells are going to start like kicking off like a GPU training job. And like, I don't want that to start automatically. Like, you know, it's expensive. I can't interrupt it. And so for those situations, Marima does have a button that says make execution we call it lazy and so if you do that and if you run a cell marimo will just mark the other ones that depend on that cell as stale but it won't auto run them and i'll give you a button to run them so like you still get guarantees on like the reproducibility in the state um but you can like you know rest easy that you're not accidentally spending your open ai credit or whatever you're doing that's expensive nice and so i guess that ties into how
13:20Jon Krohn:Now, you've previously mentioned that Merimo's reactivity makes UI elements like sliders and plots substantially more useful than than Jupyter widgets. And I guess it's because of what you just described, where in the Jupyter notebook, you don't know necessarily like like those, any kind of UI elements, any plots that you're outputting, any sliders you might have in the notebook, you don't know in the Jupyter notebook whether those are being updated properly based on other changes that you've made. Is that kind of where that reactivity becomes useful there, or is it something else? Yeah, no, that's essentially it.
13:57So in a traditional style notebook, if you had a slider, first of all, getting that hooked up to even propagate its value back to Python is kind of challenging. You have to have callbacks and stuff. And even if you do, like the things aren't going to automatically run with the updated value of your slider. So So it's like, okay, what was the point? Whereas in Marimo, you import Marimo as Mo into your Marimo notebook. And now you have access to a bunch of UI elements, such as sliders. So you could replace the variable A, I said, instead of it just being an integer, it can be a slider. And now as you like scrub the slider with your mouse, like the other cells are automatically running as well.
14:39And so that's like the interactivity feeds into reactivity to give you like a truly dynamic experience that can really, really speed up your data exploration. One concrete example, like I guess from the PhD, that sort of like really helped me, what I really wanted to achieve and I think we achieved pretty early was like, you can even output plots in Marimo where you can select like a cluster of like a scatterplot like with your mouse. And then those points get sent automatically back to Python as a data frame and then for downstream analysis. So like it really opens up like new sort of kinds of interactive experiences.
15:14Jon Krohn:This episode of Super Data Science is brought to you by the Dell AI Factory with NVIDIA, two trusted technology leaders united to deliver a comprehensive and secure AI solution. Dell Technologies and NVIDIA can help you leverage AI to drive innovation and achieve your business goals. The Dell AI Factory with NVIDIA is the industry's first and only end-to-end enterprise AI solution designed to speed AI adoption by delivering integrated Dell and NVIDIA capabilities to accelerate your AI-powered use cases, integrate your data and workflows, and enable you to design your own AI journey for repeatable, scalable outcomes.
15:53Jon Krohn:Learn more at www.dell.com slash superdatascience. That's dell.com slash superdatascience. Nice. That's really cool. I I appreciate the tangible example, something that's always useful for our listeners, for sure. And another feature that seems really cool to me is that when you have the notebook maybe into a state that you're happy with, you could actually hide the code cells, and then all of a sudden, you have a data app just ready to go. Is that right? Yeah, that's exactly right. So a Marima notebook collapses the space between notebooks and data apps. Now, all your notebooks can be just notebooks.
16:35They don't have to be anything more. But if you want to, yeah, just push a button and then boom, you don't have just a notebook. You have a full-fledged data app, which, by the way, is actually pretty performant. Unlike Streamlit, in Marima, you slide a slider, it's only going to run the things that depend on the slider. It's not going to rerun the whole notebook. And yeah, that data app, you can serve it on the server. You can even run it entirely in the browser with a technology called WebAssembly. so really easy to share these these artifacts nice i love that and then so it just occurred to me this actually wasn't in in our research but if this is an open source project marimo
17:15Jon Krohn:um but you know you have a venture data like sean's invested in your company aix and other investors are expecting a return what's the commercial angle for you so marimo as the open source uh what we're doing is we're providing like the best programming environment for individuals to like work with their own data and you can still use it sort of successfully at your own company um or at a company we have many people using marimo in like companies very large as well as small cutting-edge startups as far as like commercialization goes whereas marimo is the best place for individuals to work with their own data like there's still like a lot of gaps in order to like work with data in your enterprise at scale especially doing like rapid experimentation with like very large data sets um and so there's a lot of unsolved problems there and like that's where we uh plan to commercialize yeah that sounds like um a really great strategy a great commercialization strategy so people can get the full use of your tool as individuals and so i think this is something that would often be called product-led growth, PLG, where you have users that are volunteering to try this product out.
18:26Jon Krohn:And in a situation like yours, where you're getting millions of downloads, it's obviously working well. It's about as well as PLG goes. But then there's kinds of features that enterprises are looking for around security, for example, maybe collaboration, maybe file storage. There's all kinds of ways that it would be impossible to offer those things for free to people. I mean, there would be substantial costs to the creator of that kind of functionality. And so it makes perfect sense to then be offering those as a commercial offering. And it's something that we've seen going back a couple of decades.
19:02Tools like RStudio, they followed a similar kind of
19:05Jon Krohn:trajectory. They become the default IDE at that time using R, which was very popular for data science, data analytics some years ago. And it still is, I mean, it still gets used today, but Python is now the lingua franca in data science. And it's great to see Marimo developing a great environment. And I love how you can get so quickly from a notebook to a working data app. That's something cool that I haven't seen. I haven't seen it done exactly like that before. I think that could be great. So on to my next question. You've talked about a vision where you go from, and as we've been talking about this whole episode, where you go from these kind of error-prone JSON-based scratch paths that Jupyter Notebooks are into a full-stack developer platform where exploration and deployment converge just as in having a data app there ready to go instantly.
20:10So when tools collapse, the boundary between notebook and application, that seems to change not just workflows, but also how roles work.
20:21Jon Krohn:So as this line blurs, what kinds of new responsibilities and skills should our listeners adapt? Or does this just mean that they don't have to kind of develop engineering skills? Um, yeah. So how does it change things for our listeners who are, you know, maybe, you know, developing data applications, developing, uh, models for deployments. Um, and then after you've kind of thought about the, how, how changes things for an individual, how does it change things for an organization? You know, that, you know, in terms of research, engineering operations, this something like Merrimo seems to break down a lot of silos.
20:58Yeah, that's a really good question. I think it's an astute observation. So So in terms of the individual, like the practitioner, data scientist, or even ML engineer, AI engineer, data engineer, I think the way that I think about it is that Marimo gives them new capabilities to make their work just far more useful in the organization. And that gets into your second question, too. How does that change things within an organization? So the data app is one good example. like previously like you may have had to you may have done your experiment in a notebook and then you're like okay where's the front-end engineer who can help me like actually like build an application around this to make my work actionable you know at best you might like try to reach for something like streamlit but like you would hit performance bottlenecks there rather quickly it's like with marimo like there's no migration phase like your notebook should you want it to just change a couple of variables to like sliders or drop downs or tables or whatever you need and then all of a sudden you have a data app that you can then share with your team but also like to your ceo so like actually in a number of companies that are using us ceos are are using marimo notebooks that their data scientists made for them sometimes the ceos are actually making their own marimo notebooks to like run their own operations uh just because like the barrier entry is like so low and it's even lower like once you had considered like llm integrations that we have like you know just like vibe code your way through like a you know a pretty simple sort of tool that you can make so so that is one way that i think um marima like gives the data scientists new capabilities and also brings new people sort of into the fold there there is another way um and it it has to do with what you mentioned of like marima notebooks being stored as python files because actually every Marimo notebook is an executable script.
22:53It's just a Python file. Like you can go to the command line and say Python my notebook.py, even pass a command line arguments. So like from your notebook, now you actually have a workflow that you can like run as a pipeline or as a cron. And so that's like yet another way by like reducing that friction, we've now made you hopefully like a lot more productive. And also there's a lot of also's, But because it's a Python file, you can actually say like, from my notebook, import my function, or like import my class. So and when you do that, that means that like, you know, there's famously that for like many years now, people have talked about like the notebook to production handoff that researcher to engineer handoff, like that makes that handoff like a lot less difficult, a lot more streamlined.
23:43because as Marimo nudges you to write better code and it gives you these reusable functions, you can actually give your notebook to someone and they can import it and just use the logic. You could even write tests for your Marimo notebook. Marimo works with PyTest. So that's another way that it makes notebooks actually more useful in the engineering context as well.
24:05Jon Krohn:Nice. Probably during that part of the video version of this, we would have had the camera all on you, Akshay, but I was nodding my head aggressively throughout that whole time because that kind of functionality that you just described there is so useful. Like that would be an amazing functionality in terms of being able to call my notebooks from the command line or be able to pull in classes into other programs that I've running or other scripts that I'm developing. That is super cool. And it's pretty amazing because, well, maybe it's just because it flows naturally from topic to topic. But our next question that I had for you was on exactly that.
24:49Jon Krohn:So let's get to what I had right after that, which was around AI assistance. So this is something that, you know, a lot of our listeners, including myself, are huge fans of Cursor as an IDE. And part of that is because LLMs make it very easy to write code, all of a sudden whole functions are just appearing before our eyes. So it's kind of like an autocomplete on steroids. If people have the experience, for people who haven't used cursor, it's kind of similar to what something like Gmail does today or even a Google search where you kind of get the next few words suggested. But it ends up sometimes it's like whole function.
25:36Jon Krohn:It appears in front of you in cursor. And so it sounds like Miramal also has a kind of AI assistant involved. Tell us about it. Yeah, definitely. So in the Marima editor, so Marima as a notebook works like most notebooks, you see something in your browser, Python cells, where you can type code in, you see outputs. And there's also, you know, there's a few different points for, I think, at which like AI is sort of integrated into our editor. so the most prominent one you'll see is a button that says generate with ai right and so that's what it kind of sounds like so you can generate python code or sql code i should mention so marima has native support for sql um and you can use essentially like whatever ai back end you like so you can connect to open ai anthropic gemini um you can also use olama to run local models but the point is what makes a notebook a unique experience it makes our notebook a unique experience for using an llm is that um well you actually have data in memory right and that's really different from like traditional software engineering where it's like just text files here you're like you have a running notebook session you have data frames in memory you have tables in memory so what you can do is like you can say something like hey you know using my data frame and you tag your data frame using at df, do a group by the country column and aggregate over all numerical columns.
27:08And then what Marima will do behind the scenes, it'll pull the data frame from memory, look at the schema, get some sample values, and pass that in as part of your prompt so that when you get the generated code, it actually has the column names in it and it'll actually work. and that's like something subtle but like you you won't be able to get that from like just using cursor on like the text files and that that alone is like a huge productivity boost i think nice
27:34Jon Krohn:i'm so glad that you were able to you know i feel kind of bad almost when you know uh like someone like you startup founder comes on the show and is talking about their product and i'm like how about this other product no no so yeah so i'm but you know chris is something that a lot of people have experience using. And so thank you for that explanation of how it works in Merimo and also the kinds of things you can do. It sounds like you've been thinking about ways that data scientists in particular or AI ML engineers in particular, data analysts, people who are kind of used to the Jupyter environment, how they would benefit most from an AI assistant, which adds kinds of features like you just described there around data frames being populated like you'd expect.
28:20Um, so super cool.
Read the full transcript
28:23Jon Krohn:And that actually is a great segue to my next question, which is around, you've previously said that, that there was a joy to rethinking a system from first principles, as you said about creating Merimo. So, you know, you get to have an opinion on the two, on how the tool should be designed. And so it sounds like if our research is correct, you began Merimo with a 2500 word design document so yeah tell us about that experience of you know coming up with a with this opinionated tool from scratch what that was like oh yeah it was a ton of fun um so this was like right after I finished my PhD so uh 2022 at that point like we talked about like I used notebooks a ton during my PhD there were still some gaps and so when I got the opportunity to do is that you know i had some free space and like i got to just like study the landscape you know got to you know really dig into like what do people like about notebooks interactivity seeing your seeing your outputs live and also like what are people doing like in other ecosystems like what kind of innovations are you know what kind of ideas are people playing with but what are they innovating on and so that that process took me through um uh i did research basically right and like i that's when i discovered the pluto notebook for the julia programming link which would felt like something out of like the future to me because like it has these concepts so pluto predates marima and it's a big source of inspiration it has reactive execution it has ui elements it's stored as a julia file um it doesn't have the web app thing but i looked at that and i'm like well wow that's like really close to a data app and you know we can build that functionality in.
30:12So that was a huge inspiration for me. Python has its own challenges of like making something like that work. So then the next step was, yeah, like writing like a design doc, like, well, how do you, it's like, you're threading this needle, like, how do you design a batteries included experience where you have a notebook, it doubles as an app, and it triples as a script or as a module right and like what i didn't want was like a bunch of like kludges to like you know it had to feel seamless and it had to feel frictionless and that's why there was like this design iteration and i was lucky though because i wasn't designing in a vacuum so i guess part of what we didn't chat about so like before we got funding from aix ventures um marima was built sort of with like close input from researchers at stanford slack national laboratory department of energy laboratory as well as with some like uh industry research researchers at a couple of startups um and so like i would have ideas and like throw them at them and they would give feedback and like you could see some of these early design documents and like some of the ideas were just totally horrible but like i got to work with them to like refine them and make them a lot better i think that was that was really special time nice that does sound like a great experience
31:29Jon Krohn:and yeah, hopefully something that's kind of inspiring for any of our listeners out there that are thinking about building a product. It sounds like you had a great way of getting going on that, thinking about the best parts, the worst parts of kind of the development experience for data scientists. You've put a great product together. It's really cool. Something that we talked about earlier, right at the onset of the episode, and that would have come up, I'm sure, as you were developing the product in that kind of Word document is this reproducibility issue. So the Jupyter notebook has been a very standard tool for a decade now for data scientists, but you've cited a study previously showing that only 4 % of Jupyter notebooks reproduce their original results due to issues that you called a hidden state trap, among other kinds of problems.
32:23Jon Krohn:Do you want to dig into these kinds of problems with reproducibility that Jupyter Notebooks have and how Merimo resolves them? Definitely. So just to caveat this discussion, I think every notebook does have its place and with Jupyter Notebooks, I think they are well suited to REPL workflows where maybe you're not shooting for reproducibility, but you're quickly really hacking things. The issue is though then, I think, because it's such a great environment for interactive coding, people started using them and really relying on them for issues where first use cases where reproducibility is paramount like data engineering like science like machine learning and so the question was dig into some of those reproducibility issues right and how marimo solves them um so two main classes so one is the issue of hidden state where in a jupyter notebook you run a cell jupyter is not going to run the cells that depend on it like the cells that use its variables and it's up to you to remember all the complicated dependencies that might flow through your notebook and you're going to forget some i i often did right and so now your notebook is in some weird inconsistent state so that that's one issue i to make that like really like pronounced i guess is like so in a traditional notebook like jupiter you can like say like you have a cell that like it's got a bunch of code in it um and maybe one thing that that cell does is like create some like PyTorch model class, you know, like instantiate some class.
33:51And then say, so you delete that cell. And then you continue coding elsewhere in the notebook. You deleted that cell, but say you didn't realize that that was the cell that defined the model class. And you really wanted that model class around. But in Jupyter, if you delete that cell, that model is still in memory for the time being. And so your rest of your notebook will work as you want it to work. But you come back the next day and you run the notebook from scratch nothing works and you're like what happened and it's like oh crap i deleted the cell that defined the model and like i really needed that so in remod if you do that if you delete the cell that defines your model variable it'll tell you it'll remove the model from memory and like it'll invalidate the other cells you'll be like yo that that variable is no longer around like you can't do this anymore and so it'll catch bugs like immediately when when you introduce them uh so that that's one big way and the second big way is package management and So Marimo has a special sort of, I guess, opt-in package manager, which I think is really neat.
34:54So basically, if you start Marimo in a mode that we call sandbox from the command line or however, every time you install a package, Marimo has a very nice slick package installation UI. we will save the package that you installed and the version as like a comment or an entry in the notebook file itself. And now when you come back to the notebook a second time and try to run it, and you do run it, Marimo will create an isolated virtual environment for you that has just those packages. So that means that like you can just send a single notebook file around and like people can just run it without even thinking about what packages they need to install, which makes them more reproducible, but also they're just like a lot more portable.
35:39It's like really easy to just create these single standalone tools and share them.
35:42Jon Krohn:I had more aggressive head nodding there while you were describing all that. I'm going to have to start using Marimo myself. Very, very cool. So something else that I think is really cool about functionality that we've learned to that I want to, it's something that's so visual and so easy to understand. It's something that would feel to me like magic as somebody who has really only used Jupyter Notebooks before for this kind of script development. So it sounds like it's possible, correct me if I'm wrong, that because of the way that you've thought about data analysts, data scientists, data people's experience from the ground up with developing this tool, it allows you to do things like highlight, select the data with your mouse in a scatterplot and then get it back as a data frame.
36:35Jon Krohn:Is that right? Yeah, that's right. Yeah, that's an example. That's crazy. Yeah, it feels really nice. It's like, you know, and I really needed this during my PhD and I didn't have it, right? So it really enables like really, really tangible like data exploration, data analysis workflows that were like really, really hard, if not impossible to do before. And so how do you decide when you're doing product development development how do you decide i guess to put something on your roadmap at all or to you know then prioritize highly a feature like this where yeah like how do you how do you decide you know which of these kinds of force multipliers to include in your product there's two angles i guess especially earlier on right like when we didn't have that much feedback and it was me and my founder just developing and jamming like with our built-up experience of like seeing like okay like here are some issues that we've hit with working with traditional notebooks and we know a lot of others have like we just we have strong opinions and we think this will land and we're just going to trust our gut so there was a lot of gut trusting in the beginning and there still is um now with the product being you know more mature um there's still gut trusting for like big new you know, features that we're working on, but we have a big community too, right?
37:57And so we can chat with them. And actually they come to us. They're very vocal about what features they want and they don't want. And I think like the way we think about it is like, is this something that like enables a broad class of users? So like data scientist, data engineer, ML engineer, AI engineer to be far more productive than they would have been otherwise in their sort of previous tool of choice. And I think like that, that is like something we call them like big rocks. Like what are the things that really moves the needle for folks? And it needs to move the needle a lot. And it needs to move the needle for a lot of people.
38:38Of course, we care a lot about craft and design and visual design and usability. And so those things are just like always top of mind and they're like ongoing. but for like the big new features that they have to be big rocks sort of how we think about it
38:50Jon Krohn:nice i like that big rocks and so you talked about community and how you can leverage them so open source projects often promise community but not every project earns it so what do you think separates tools that spark tons of devotion like the pluto project that you were describing there in Julia. What do you think separates those from open source communities that fade? And what are you going to do to ensure that you're in the first category? Yeah, it's a good question. I think, honestly, it sounds basic, but a lot of it comes from the maintainers just being like really kind and open like you gotta you gotta be open like if you if you want community you have to you know you can't just say i want community but like there are people out there who are excited you need to encourage them to like contribute and like be really really vocal about like how much you appreciate your community and also like help them make prs if they want to be you know super responsive on like your discord Like we try to respond really quickly to issues and especially like bug, bug reports.
40:07Like one of the feedback we usually get from our community who like file issues is that they're like shocked by how quickly we fix their bugs. Like if someone files a bug, you triage it and like fix it and ship a release the same day. Like you've kind of won a supporter for life. Oh, wow. Yeah. That's the thing that we've noticed something we've also noticed like from other popular Python projects, like from Charlie Marshall, Astral's UV project. And I think my community is many levels, right? Not everyone has to ship code to your project, although many people may want to, but just encourage all kinds of engagement.
40:45Yeah, and just make it a fun place for people to hang out. Nice.
40:49Jon Krohn:If we have listeners who would like to contribute to the project, how should they get started? There's a few ways, and it depends on how you want to contribute. So if you want to contribute code, you can check out our GitHub issues. And some of the issues are tagged as like good first issue. And like these are like great places for new contributors to kind of get their feet wet and learn what the code is like. And some of them are even like improved documentation. Which is, by the way, just generically a really good way to start contributing code to a project. Improve the documentation. other ways you can contribute is you can just like file a feature request or a bug a bug report we have a little check box saying are you willing to submit a pr for this and if you are and if it aligns with our roadmap then we'll like work with you um and then more generally we have a discord so you can get the discord link if you go to marimo.io slash discord and there like we have lots of free-flowing conversation.
41:50And I think that there's a lot of good touch points to get involved.
41:54Jon Krohn:Fantastic. And what if this is a listener's very first, you know, what if they've never contributed to open source before? Would you recommend, yeah, where would you recommend they get started? How should they get their feet wet with open source development? Yeah, so actually we do have, I think, a number of contributors who have made their first ever contribution with us. For that, I'd say like, you see a typo in our docs, something like that. I think that's a really good way to contribute. The docs ones are...
42:28I think I could be wrong, but I think if you're just making a simple change, you might be able to click edit in the repo itself, like online on github.com, and it'll create your fork for you and kind of simplify that process. But yeah, make a docs change. We have a contributing.md guide in the GitHub repo that tells you about, oh, you'll need to make a pull request. That's a pull request. And it'll walk you through that workflow.
42:55Jon Krohn:Nice. All right. And then that brings me to maybe one of my last technical questions for you, which is, this is completely beyond Merimo. And so this is, going back to the research that you've done, you focused on machine learning and optimization. and as an engineer, you've contributed to several open source projects from the deep learning framework TensorFlow that probably most of our listeners are familiar with to the vector embeddings computation library. You're going to have to tell me how people pronounce this, but it's PyMDE. Perfect. Oh yeah. Okay. All right. So it's, yeah, it's P-Y and then the letters M-D-E.
43:36Jon Krohn:And I'll have a link to that in the show notes. And so, yeah, so that vector embeddings computation library, PyMDE, as well as there's a convex optimization parser, which is also just letters, CVXPY. CVXPY, yeah. CVXPY, that makes sense. And so, you know, half of your published research relates to convex optimization. For our listeners unfamiliar with the topic, could you expand on how convex optimization differs from machine learning and the kinds of questions that they can answer? Precision, interpretability, scalability. Yeah, tell us about convex optimization. So mathematical optimization in general is like you have some variables that you're trying to make an assignment of values to.
44:29So, for example, say we're trying to choose a good stock portfolio. The variables is how much to invest in each stock. And then you need something that tells you whether or not your assignment to variables is good. So that's a mathematical function. It's an objective function that says, hey, what do I predict my return will be if I make this investment? And maybe it trades off. It factors in risk into it as well. And then you have some constraints like, well, I can't invest more money than I have. and maybe I'm not allowed to short stocks. So you have some constraints. And mathematical optimization is then the process of finding an optimal assignment of values to the variables to minimize the cost or maximize the reward while satisfying constraints.
45:22That's mathematical optimization. Convex optimization is just the subset of those problems we know we can solve super efficiently, super reliably, provably. And the use cases are somewhat different. So in some sense, like machine learning, a lot of especially classical machine learning, logistic regression, SVMs, all these are actually under the hood. They're using convex optimization techniques to fit the models. In terms of like use cases out in the wild, like what are people using CVXPy for today? It's well, well, one huge one is financial portfolio construction. So like many billions of dollars daily are like allocated through portfolio optimization problems, which are solved with convex optimization problems.
46:11Energy management, like a lot, there's a lot of usage there as well. Like real-time control, like controlling a vehicle, landing a rocket, like SpaceX uses convex optimization to land, I think it's like the Falcon or something, using software developed by our lab. And so these are cases where you can model the world and you have some understandable constraints and you can really exploit the structure. Machine learning, sometimes you, there's typically not many constraints that are kind of implicit. You don't really have as good of an understanding of like, well, what is the model doing? And also, you're just trying to find a solution.
46:51that's kind of good enough that you're going to test it out in the wild on unseen examples. And so the use case of finding the optimal assignment, it doesn't even necessarily always make as much sense. That said, there's a lot of overlap. So the vector embeddings library that you mentioned, for that to fit these embeddings using a GPU, even though the problem was in the machine learning domain, we use techniques from convex optimization to solve it really, really efficiently.
47:23Jon Krohn:On this podcast, I'm always going on about how Claude Code is mind-blowing, but now Claude Cowork is making my jaw drop as well. For example, I recently wanted to quantify how healthy my sales pipeline is for my AI consulting business. I simply asked Claude to estimate my sales for the coming quarter, and it brought info from relevant Google Sheets and my Gmail to create a professional spreadsheet of clients with estimated revenue for each one. Whoa, this might have taken me a day. Instead, it was done flawlessly with Claude Cowork in minutes. Claude is the AI for minds that don't stop at good enough.
47:54Jon Krohn:It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. Ah, and you'll appreciate that I can ask Cowork to show me data, such as my sales spreadsheet, and it provides an interactive chart right in the conversation. For problems worth solving, get started with Claude at Claude.ai slash That's claude.ai slash superdata. And check out Claude Pro, which includes access to all of the features mentioned in today's episode.
48:24Jon Krohn:Claude.ai slash superdata. Nice. That was a beautiful explanation. As I've actually all of your explanations today, you do a great job of taking complex concepts and making them seem really approachable. And it isn't just the language that you use. You also have this really accessible tone. You make everything just kind of seem light and relaxed. I really like that. It's been a joy interviewing you. Actually, on the note of you making things feel so accessible, on your personal webpage, so akshayagrawal.com, which I'll have a link to in the show notes, it says that your goal is to make machine learning and math accessible and actionable.
49:09Jon Krohn:What does actionable mean in this context? What's the biggest gap today between mathematical tools and being able to take action with those tools? oh i love that question like when you you're in school or you're taking a course especially like a traditional course i.e most of them like you're like you know they're teaching you like okay this is logistic regression and then like you write down like the optimality conditions by hand you're computing some gradients and you're like okay great like i did a bunch of I don't know homework on my paper and I submitted a piece that that's fine but and that's good like I've done a lot of that in my life but like to make something actionable and to make like I think like what's really cool about our field is that like all that math you can do real things with it and like that's why like PyTorch and TensorFlow and Jax and like PyMDE and CVXPy all exist it's about like you know using concepts from math and and in order to like affect real change in the world um and so like that's sort of i guess been the theme of like the kind of projects that i've chosen to work on from tensorflow to cvxpy pymde and marimo like even though there is no math necessarily in the marimo code base like what it does is give you a really really tangible interactive environment to work with your data.
50:42And so it's making your data actionable. Like it's actually useful and like you can actually run your notebook as a script or share it as an app or reuse the code in it. And so like, I guess that theme is sort of what resonates to me.
50:56Jon Krohn:For sure. You're building a platform to make the change that you want to see in the world. I've loved everything you've said about the Marimo product today. And yeah, I'm so glad that Sean introduced us. It's been a great episode. Before I let you go, Akshay, I always ask my guests for a book recommendation. And actually, I can see and our YouTube viewers can see a huge bookshelf of books behind you. So what have you got for us? OK, I can recommend a couple. One that I most recently read that really stuck with me, as well as one that I'm currently reading. So I most recently read American Prometheus, which is the biography of Robert Oppenheimer.
51:38And so like most people I watched, like many people I watched Oppenheimer, like I want to learn more and I want to learn a lot more. So I read the book and it was really fascinating because it like actually like, you know, talks about. It talks about not just like, you know, the events that happened and not just like the witch hunt of Oppenheimer that happened after after the war, but also like what is like the social process of like doing science and like what are the social factors that influence? someone you know who is as much of a you know genius as Oppenheimer was to like make certain decisions that like led him to like develop the atomic bomb and I thought just like having that broader context was like super interesting so that's one another one that I'm currently reading that I would recommend because it's already making an impact on my life is the design of everyday things so a classic and just like product design it's taking me a long time to read it because I'll go like two pages and like, oh man, like I need to go and like improve something in Marimo.
52:40Um, and so like, if you want something, if you, if you, if you work in design or just are like observing the world, um, I think that that's a good way. It's already sort of made our product better.
52:51Jon Krohn:Fantastic. Two great and very different recommendations for us. Love it. And then for people, our listeners like me who have really enjoyed this conversation with you today, what's the best way to follow your work or connect with you going forward so that we can continue to get your thoughts after the episode? Definitely. So I am on the major social media platform. So me personally on X, my handle is Akshay K. Agrawal. I sometimes post under the official Marimo channel, which is Marimo underscore IO. I'm also on Blue Sky as well. Those are the best ways to follow me personally. And then Marimo also has various social channels.
53:40So we're on YouTube, we're on Blue Sky, we're on Discord, and we have a newsletter, marimo.io slash newsletter, if you want to subscribe, which I personally write once a month. Nice.
53:52Jon Krohn:Love that. Thanks so much for taking the time out of your busy founder schedule with us. Running an early stage tech startup like this must be exhilarating, but also very time consuming. So it means a lot to me and to our listeners that you took that time out. And yeah, thanks for providing us with such a great episode. Thanks, John. Really appreciate it. It was a blast.
54:16Jon Krohn:I hope you enjoyed that episode and I hope you liked the shorter intro to today's episode. Reach out to me on LinkedIn with a DM or a comment if you weren't happy or if you have any other ideas how we can improve the intro or any other part of the show. Really, I always love to hear from you. We assume that if you listen to the entire episode like you have today, that you'd probably still like the full outro that we usually do. So here you go. In today's episode, Akshay Agarwal covered the core reproducibility problems with Jupyter notebooks, such as hidden state traps where cells can run out of order, JSON file formats that break version control, and missing package dependency tracking that makes notebooks impossible to share reliably.
54:57Jon Krohn:He talked about how Merimo's reactive execution model that automatically runs dependent cells when variables change ensures your notebook state always matches what's displayed on screen. He talked about advanced interactivity features in Merimo, including UI elements like sliders that trigger real-time updates, the ability to select data points directly from plots and receive them as data frames, and instantly, an instant conversion from notebooks to deployable data apps. That's cool. Akshay also talked about the technical innovations that make Merimo notebooks stored as pure Python files, enable command line execution, function imports, and seamless integration with existing development workflows.
55:33Jon Krohn:and finally we talked about how you can get involved with open source projects like these Maribou notebooks if you'd like to as well you can get all the show notes including the transcript for this episode the video recording any materials mentioned on the show the urls for Akshay's social media profiles as well as my own at superdatascience.com slash 9-1-1 thanks to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, our media editor, Mario Pombo, our partnerships team, which is Nathan Daly and Natalie Zajski, our researcher, Serge Massis, writer, Dr. Zahra Karche, and our founder, Kirill Aramengo.
56:12Jon Krohn:Thanks to all of them for producing another excellent episode for us today. For enabling that super team to create this free podcast for you, we are oh so grateful to our sponsors. You can support the show by checking out our sponsors' links, which are in the show notes. And if you're interested in sponsoring an episode yourself, You can get the details on how by making your way to johnkrone.com slash podcast. Otherwise, please help us out by sharing the podcast, sharing this episode with someone who would enjoy this episode. Review the episode on your favorite podcasting app or on YouTube or wherever you watch it or listen to it.
56:48Jon Krohn:Subscribe, obviously, if you're not already a subscriber. But most importantly, I hope you'll just keep on tuning in. And I'm so grateful to have you listening and hope I can continue to make episodes you love for years and years to come. Till next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
Reproducibility, Python notebooks, and data science communities: Software developer Akshay Agrawal speaks to Jon Krohn about Marimo, the next-generation computational notebook for Python, how he built and fostered a thriving community around the product, and what makes this notebook so versatile and accessible for users.
Additional materials: www.superdatascience.com/911
This episode is brought to you by Trainium2, the latest AI chip from AWS and by the Dell AI Factory with NVIDIA.
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.




