Reinventing the Python Notebook with Akshay Agrawal

10 Mar 2026 · 46 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Reinventing the Python Notebook with Akshay Agrawal

Podcast Overview

  • Title: Software Engineering Daily
  • Description: Technical interviews about software topics.
  • Episode Title: Reinventing the Python Notebook with Akshay Agrawal
  • Episode Description: Discussion on the evolution of interactive notebooks, their limitations, and the introduction of Marimo, a next-generation Python notebook.

---

Key Participants

  • Akshay Agrawal: Creator of Marimo, former Google Brain researcher, and PhD in machine learning.
  • Kevin Ball: Vice President of Engineering at Mento, independent coach, and co-founder of multiple companies.

---

Introduction to Notebooks

  • Notebooks have become essential for:
  • Data science
  • Research
  • Data exploration
  • Challenges with Traditional Notebooks:
  • Hidden state: Complexity increases with project size.
  • Non-reproducible execution.
  • Poor version control ergonomics.
  • Difficulty in reusing code in software systems.

---

Marimo

A Next-Generation Notebook

  • Purpose: To address issues in traditional, imperative notebooks such as Jupyter.
  • Core Features:
  • Reproducibility: Functions stored as pure Python, enabling version control with Git.
  • Reactive Execution: Automatically updates dependent cells when a variable changes.
  • Interactive Sharing: Convert notebooks into interactive web apps without needing collaborators to recreate environments.

---

Key Issues with Traditional Notebooks

  1. Hidden State:
  2. Users may forget to run cells, leading to inconsistent states.
  3. Can lead to major issues during complex projects.
  1. Ergonomics:
  2. Difficult to maintain code across different notebooks due to JSON storage format.
  3. Encourages code duplication and messiness.
  1. Shareability:
  2. Requires collaborators to set up environments, hindering collaboration.

---

Marimo's Design and Execution Model

  • Inspired by Reactive Notebooks:
  • Incorporates ideas from Pluto (Julia) and Observable (JavaScript).
  • Execution Model:
  • Cells reactively execute based on variable dependencies.
  • Uses static analysis to track variables, ensuring consistency.
  • Seamless Integration:
  • Supports modern software engineering standards and practices.
  • Enables easy conversions to CLI scripts for job automation.

---

Comparison with Traditional Notebooks

  • Traditional Notebooks:
  • Operate in an imperative style, leading to hidden state issues.
  • Marimo:
  • Promotes a reactive model ensuring coherence and clarity.
  • Facilitates the transition from data exploration to production.

---

Use Cases and Adoption

  • Diverse Applications:
  • Used not only in data science but also by backend engineers for data pipelines.
  • Allows easy creation of interactive web apps for internal use.
  • Interactivity:
  • Users can create sliders, text inputs, and charts that reactively update based on changes.

---

Integration with AI and Future Directions

  • Potential for AI Integration:
  • Exploring how to utilize agents with runtime information to enhance data workflows.
  • Educational Engagement:
  • Plans to collaborate with universities for educational use, leveraging the interactive nature of Marimo to teach concepts effectively.

---

Conclusion

  • Invitation to Engage:
  • Encouragement for listeners interested in exploring Marimo and its educational applications to reach out for collaboration.
  • Emphasis on Innovation:
  • Marimo represents a significant advancement in how data exploration and software development can be integrated within a single tool for better productivity and collaboration.

---

Closing Remarks

  • Call to Action: Explore and provide feedback on Marimo's functionality and its role in modern data science and software engineering.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Limitations of Traditional Notebooks

0:45 to 1:40

Exploring the challenges posed by traditional notebooks in complex projects.

“Kevin Ball, or Kate Ball, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders.”

Introduction to Marimo

1:40 to 2:36

A discussion on Marimo, a next-generation Python notebook created by Akshay Agrawal.

“Yeah, I'm excited to get to talk to you.”

Akshay's Background and Vision

2:36 to 4:00

Akshay shares his background and the vision behind creating Marimo.

“And it is a notebook, but as we'll talk about in this episode, it's different than traditional notebooks in many ways.”

Issues with Traditional Notebooks

4:00 to 5:14

Delving into specific issues encountered with traditional notebooks like hidden state and version control.

“So one, like the one that sort of really trips up many people, myself included, is this idea of hidden state.”

Enhancing Shareability and Interactivity

5:14 to 6:50

How Marimo aims to improve shareability and interactivity in notebook use.

“ergonomics of using notebooks as part of modern software projects.”

The Shift Towards Reactive Notebooks

6:50 to 7:42

Discussing the advantages of a reactive model in notebooks over traditional methods.

“with the document and see how results would change.”

Dependency Graph and Performance

7:42 to 9:16

Understanding how Marimo utilizes a dependency graph for efficient execution.

“It's not me hacking away in a REPL somewhere.”

Ensuring Reproducibility

9:16 to 11:15

Exploring Marimo’s approach to reproducibility and managing dependencies effectively.

“And instead, Marimo basically just statically reads the code of every cell that you have and determines what are the variables it defines.”

Marimo's Reproducibility Features

14:00 to 17:49

Learn how Marimo ensures reproducibility through package management and file formatting.

“Like Maremo will just run it for you or market a still and like really be really loud if you've done something sort of, I guess, yeah, it just won't let you step out of its reactive execution model.”

Notebooks in Software Development Lifecycle

20:01 to 22:35

Understand how notebooks integrate into the software development lifecycle and their usage in data pipelines.

“Okay, so I want to come back to the UI widgets because that was the thing you talked about as well.”
Show all 20 chapters

Interactivity and Usability in Marimo

22:36 to 28:00

Explore the interactive features of Marimo and how they enhance data exploration and reporting.

“nice to just have a report alongside it to see like, what was the shape of the data that day, et cetera.”

Using Marimo for Sports Analytics

28:00 to 29:20

Learn how a well-known sports team utilizes Marimo for analytics applications.

“will just use the ui elements to speed up exploration but let's well just like you know i was just on the on the phone with like a i guess i i have to speak about them anonymously right now They're a big Marimo user.”

The Value of Notebooks in Data Exploration

29:20 to 30:40

Discover the advantages of using notebooks for data exploration and sharing insights.

“And like, honestly, that's not something I anticipated as a use case when I first started making Maremo right after my PhD, but it was really cool to hear people being empowered by that.”

Transformation in Data Programming with AI

30:40 to 32:30

Explore the transformation in data programming with the rise of AI and new tools.

“in the data, so I don't even know what to make.”

Frontiers of Notebook Technology

32:30 to 34:20

Understand the advancements and future directions in notebook technology.

“It finds like with the syntax of how the notebook is stored.”

Integrating Agents with Notebooks

34:20 to 36:30

Learn about integrating agents with notebooks for enhanced data workflows.

“in terms of use of notebooks, data access, data exploration, that sort of world?”

Ensuring Correctness in Data Workflows

36:30 to 39:20

Discover how Marimo ensures correctness and validity in data workflows.

“We're not exactly sure, but it seems like that could just be like a really valuable primitive, like a sandbox in some sense for agents to have.”

Extending Marimo Functionality

39:20 to 42:00

Learn about extending Marimo's functionality through APIs and plugins.

“So those are the two main semantic things that we check for.”

Exploring AnyWidget: The Future of Interactive Notebooks

42:00 to 44:19

Learn about the development and significance of AnyWidget in notebook environments.

“And so people can target that with code generation tools, but that's about the extent.”

Marimo's Origins and Future Directions

44:20 to 47:19

Discover the founding story of Marimo and its goals for educational engagement.

“And in terms of hooking into a reactivity model, it's actually quite nice.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Akshay Agrawal:Interactive notebooks were popularized by the Jupyter Project and have since become a core tool for data science, research, and data exploration. However, traditional imperative notebooks often break down as projects grow more complex. Hidden state, non-reproducible execution, poor version control ergonomics, and difficulty reusing notebook code in real software systems make it hard to move from exploration to production. At the same time, sharing results often requires collaborators to recreate entire environments, limiting interactivity and slowing feedback. Marimo is an open-source, next-generation Python notebook designed to address these problems directly.

0:43Akshay Agrawal:Akshay Agrawal is the creator of Marimo, and he previously worked at Google Brain. He joins the show with Kevin Ball to discuss the limitations of traditional notebooks, the design of reactive notebooks in Python, how Marimo bridges research and production, and where notebooks fit in an increasingly agentic AI-assisted development world. Kevin Ball, or Kate Ball, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through latent space.

1:21Akshay Agrawal:Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.

1:39Akshay Agrawal:Akshay, welcome to the show. Thanks, Kevin. It's great to be here. Yeah, I'm excited to get to talk to you. Let's start out with a little bit about you. So can you give a quick background of who you are and how you got to our topic today, where you got to with Marimo?

1:54Software Engineering Daily Host:Sure. Happy to. So my name is Akshay. I've got a background in computer systems, but also machine learning research. I spent a little bit of time at Google Brain. This was a while ago, back when Google Brain existed before it was taken over by DeepMind. So this was 2017 to 18, I worked on the TensorFlow team. And then after that, I went back to Stanford to do a PhD in machine learning research. And through that experience and my time at Google, I realized what I really enjoy doing is building open source developer tools for people who work with data. So after my PhD in 2022, I started working on Marimo.

2:31Software Engineering Daily Host:And Marimo is, it feels like a next generation open source Python notebook. And it is a notebook, but as we'll talk about in this episode, it's different than traditional notebooks in many ways. It's reproducible, It's stored as pure Python, so you can version it with a Git, execute it as a Python script, and share it as an interactive web app too, if you want to.

2:52Akshay Agrawal:So let's maybe look into that, right? I feel like notebooks have been a part of the Python ecosystem for a while. And when they came out, there was this big breakthrough of like, oh my gosh, I can do this kind of interactive exploration of data and then share it and do all these different pieces. What's wrong with the state of notebooks? Why Marima?

3:11Software Engineering Daily Host:Yeah, it's a great question. and traditional notebooks like Jupyter notebooks and things like them, I think have been extremely useful in like research and education. And I use Jupyter a lot during my own PhD. So my thesis was on vector embeddings. Now I would make a lot of low dimensional plots of high dimensional data after doing some dimensionality reduction, right? So lots of scatter charts and stuff like that. And it was really, really useful to do that in like a Jupyter notebook because I needed to run code and then see what the results of my algorithm were. And like, there was like a back and forth between you as the algorithm developer and your data, right?

3:49Software Engineering Daily Host:And so that's great. There's nothing wrong with that. The issues that I ran into and that others have run into with like these traditional notebooks are a few, there's a few issues. And I think it comes down to like maybe two or three. So one, like the one that sort of really trips up many people, myself included, is this idea of hidden state. So in a Jupyter notebook, and I'm using that as shorthand, but just like the default experience of a Jupyter notebook using the IPython kernel, it's an imperative paradigm, right? Like you run a cell, it mutates memory, and then you run another cell and then it mutates memory.

4:24Software Engineering Daily Host:And what you're really doing though is oftentimes once you get past the just exploratory phase, you're kind of writing a program, like even in a notebook, right? But because it's imperative and because Jupyter doesn't know how your cells are related, you may one run cell, then forget to run other cells that depended on that cell. And then all of a sudden the code on your page doesn't match the variables in memory. And this can lead to like a lot, at least in myself, like, you know, I would do things like I would delete a cell and then it would delete some variable that I forgot was defined in that cell.

4:56Software Engineering Daily Host:And it was still in memory and like other code referred to that variable. And then four hours later, I would realize I just had a bunch of inconsistent state and I had to restart my notebook, restart my analysis. And so this broad idea of hidden state is one thing that I think is really challenging that I wanted to solve with Marimo. And then other things include the ergonomics of using notebooks as part of modern software projects. The default file format for Jupyter notebooks is great for communication because it stores not just the code, but also or your plots and things like that in a single JSON file.

5:29Software Engineering Daily Host:But it makes it difficult for any kind of software engineering-like task, right? Because you don't want to version base64 data with Git. Like, it's just not going to work, right? And then also, you might want to reuse some code you wrote in a notebook in like another notebook or a Python module, right? But you can't really. Those aren't Python files. So then you have the issue where like, I've been guilty of this. you duplicate your Jupyter notebook like 40 times and then it's just a total mess. So that was another thing I wanted to solve. And then finally, the third thing has to do with like shareability and like sort of interactivity.

6:04Software Engineering Daily Host:So on the one hand, Jupyter notebooks are really great because you have plots, you have code, you have visuals, this document you can share out. Like if you share out some analysis or some like research investigation with a collaborator, they need to also have Python installed and Jupyter installed in order to like interrogate your results right like they might want to say like if I change some parameters how might this plot change and that feels like kind of unfortunate I mean as a matter of practice like my PhD advisor couldn't do that right and he would have questions and then I would have to go back and change things and so with Marimo we wanted to make it so that any notebook could also really easily double as like an interactive web app where like you can promote any variable or anything to like this UI component, like a slider or something, so that people could for themselves just interact with the document and see how results would change.

6:53Akshay Agrawal:So if I'm hearing you properly, I'm going to play back. So two of these things sounded like essentially saying traditional notebooks or Jupyter notebooks or whatever were in a lot of ways just kind of a fancy REPL. They're a REPL environment that has a UI baked in and is able to be a little bit more approachable in interweave plots and interweave descriptive text and things like that. But they have all the drawbacks of a REPL in the sense that it's not really meant to be something that's packaged and duplicated. It's your exploring and, as you said, having a dialogue with your data. And what I'm hearing is you want to keep that vibe of dialogue with your data, but kind of bring this up into first world software development.

7:36Akshay Agrawal:Like, hey, this is reproducible. It's data-driven. and it's state, it like keeps track of all the things. It's not me hacking away in a REPL somewhere.

7:44Software Engineering Daily Host:Yeah, that's exactly right. And like, I think one of my teammates on remote, Trevor, we were just talking about this the other day. And the model of just a REPL works well, as long as it's just you sort of working on that network, like no one else. But as soon as you bring in a second person and like, or even the second person can be you a day from now, right? Like, yeah, like you kind of want some guarantees, some reproducibility, some, yeah, Yeah, that's exactly right. All right.

8:10Akshay Agrawal:So let's maybe talk then about how you solve some of these problems. And I'm particularly interested in kind of understanding what is the execution model that you're adopting that is not just this sort of imperative REPL?

8:22Software Engineering Daily Host:Definitely. So in designing Marima, we took inspiration from other projects as well. And the two biggest ones are these other really cool notebook systems for other languages. So Pluto JL for Julia, which in turn is inspired by Observable for the JavaScript ecosystem. And what those two notebooks are, and what Marimo is, at their core, they're what are called reactive notebooks. So reactivity is sort of the alternative to this imperative style of notebooks. And so what this means is, say in a Marimo notebook, if you run a cell, say that cell defines some variable X. When you run that cell, Marimo then by default will automatically run all other cells that read the variable X.

9:06Software Engineering Daily Host:And so it reacts to your code execution of one cell to keep the outputs of the rest of the notebook in sync with the action you just took. The way that this works is that there's no runtime tracing involved. And instead, Marimo basically just statically reads the code of every cell that you have and determines what are the variables it defines. and what are the variables it references. And from there, it just builds this dependency graph. And it's kind of like Excel in some sense, right? It's not magic. And we have like configuration. So like automatic execution may not be even desirable for all notebooks, especially if like some downstream cells are going to take a long time to run.

9:47Software Engineering Daily Host:So you can make the executor, like we call it lazy so that you run a cell and then Marine Mobile will just mark the effect cells as stale, but it won't run them automatically, but it'll give you one button that you can push to bring all your code and outputs back into sync. So that's the core thing. And so that can minimize hidden state. And also, as you play with it, you also find that it actually enables you to do data exploration a lot faster because you just change a variable value, hit enter, and then you see everything change, and etc.

10:17Akshay Agrawal:Well, and this is a paradigm that I think web user interface frameworks have very much gone towards. It's this kind of reactive model. I think Vue did it first and you see Svelte and React and all these folks like taking this very data-driven reactive model. And it does allow very fast and easy keeping things in state. So looking at that dependency graph, then how much overhead does that end up creating? Or like, are there any things in Python that are hard to statically analyze? I'm not super deep on Python in particular. I know some languages, it can actually be hard to trace the dependencies.

10:49Software Engineering Daily Host:So on the question of overhead, it's like totally negligible. Python has a built-in AST module that you can use to do static and semantic analysis of code. And like most heavily used libraries in Python, that's implemented in C, so it's really fast. So the overhead is negligible, especially because it's static analysis. So it's like we parse and analyze once, and then every other execution, there's no overhead. So yeah, you don't notice that. And in terms of what is hard, I guess the scope, is the static analysis itself, is the semantic analysis that difficult? And I think what we chose to do is to like, there's like two approaches you could take to making at least two approaches to making a reactive notebook in Python, right?

11:32Software Engineering Daily Host:What we did is our data flow graph is based only on variable definitions and references, right? So like I mentioned, a cell defines a variable X, you run that cell, we run all other cells that read the variable X. That's easy to implement faithfully for the Python language. You could take another approach, which would not be based on definition references, but just on memory access. And for example, you could say, if my cell mutates or touches some very, if this cell appends to some list, I'm going to run all their cells that read that list. That's extremely hard to do reliably. And so we explicitly don't try to do that at all.

12:11Software Engineering Daily Host:Because you can't, it's impossible to do just static analysis in Python. So then you're going to have to start tracing user code. and you are invariably going to miss, like it's impossible to implement that 100 % correctly. And then you get into an uncanny valley when a user runs a cell and they won't know what else is going to run. So we're just really upfront. We're like, hey, we only track variables, definitions and references. So if you mutate things, go for it, but be aware that we are not there. Yeah, that's an escape hatch, right?

12:40Akshay Agrawal:Exactly. If you're mutating things, make sure you do an assignment afterwards. Exactly.

12:44Software Engineering Daily Host:That's exactly right. And you know, the benefit of this is that like the rule sets exceedingly clear to the user, right? Like they can understand it. It's like, like a sentence long. And then also it encourages the users to write like functional code, right? Which especially for like data-driven stuff, machine learning, it's kind of what you want to do anyway. So it's good. And like, I think like one of our users described it as gentle parenting or something like we nudge data, data scientists, machine learning engineers, et cetera, to write just generally good code.

13:15Akshay Agrawal:There's a lot of value in that. I actually, I was, this is a slight aside, but I was hearing somebody describe, they said, if LLMs generate no other value, they have dramatically upgraded the quality of code coming out of graduate schools.

13:28Software Engineering Daily Host:That's fair, actually. Yeah. I like it.

13:32Akshay Agrawal:Okay. So let's keep going down this road. So now you have, instead of a mutable REPL, you have a well-defined dependency graph. And you mentioned the next thing was around sort of reproducibility and not checking in these massive binaries or things like that. So how does Maremo handle treating this stuff as code?

13:51Software Engineering Daily Host:Yeah, so that's a great question. So the data flow graph gives you some amount of reproducibility insofar as that you can't like run a cell and then forget to run some other cell. Like Maremo will just run it for you or market a still and like really be really loud if you've done something sort of, I guess, yeah, it just won't let you step out of its reactive execution model. So that handles like reproducibility and execution in some sense. And then I'll get to the file format. I guess it is related actually, but Marimo does have sort of an optional built-in package management system that's like powered by the UV package manager.

14:26Software Engineering Daily Host:So if you opt into it, when you import a package, Marimo will detect that import, a module, resolve it to a package, prompt you if you want to install it. And if you do, it'll add it in a comment block at the top of the Python file. Python has a standard called PEP 723 for this. So basically we'll document all the packages your notebook has used. And then the next time you run that notebook, we'll use the UV package manager to create an isolated virtual environment, install just those packages, and then you're off to the races in this sort of reproducible package environment. So we handle that as well.

15:00Software Engineering Daily Host:But the reason that was an easy feature for us to add is that we actually decided to store our notebooks as Python files instead of as these JSON files that sort of Jupiter has historically used. And the way that this works is each cell is represented as like a function. There's like some decorator to like demarcate, like this is like a cell that's going to be going into the notebook. And you can think of each cell as a function mapping the variable references it uses to the definitions it creates. And at the bottom of the notebook, there's a Python. You can do an if name equals main guard, which is like when you run the script, that's what's going to run.

15:38Software Engineering Daily Host:And so there's an if name equals main guard that will then say app dot, like it will run the Marima notebook and the cells in a topologically sorted order. So basically all that to say, you can go to the command line, say python my notebook dot py, and it'll run it as a script. You can even parameterize with CLI args.

15:56Akshay Agrawal:You're already though, getting to a place where now, now this stuff is pluggable because it's just Python code. You can import the functions to wherever you can do what So before we dive down that road, which I am interested to go down the implications there, what, if anything, is lost by storing it as Python rather than this sort of proprietary format?

16:15Software Engineering Daily Host:Yeah, so there's definitely something lost. So the main thing that's lost is by default, when you're using Jupyter, not only is your code, but like your plots, for example, like are stored as base64 encoded data in the file so that you can just like put it on GitHub and like you can immediately see a record of your analysis. So I actually think the IPython notebook file is a really valuable artifact. I just don't think it should be the artifact that your development is centered around. And so what we do to sort of bridge the gap is that there's a configuration setting that you can turn on, which will basically automatically snapshot your notebook as an IPython notebook alongside the Python file.

16:54Software Engineering Daily Host:So there's a little underscore underscore marima directory, and you'll save my notebook that IPyNB in there alongside your Python file. so that you can try and get the best of both worlds. There's other things that we ended up having to sort of implement our own versions of because for example, one thing that's nice about a Jupyter notebook file format or like just something that stores the outputs is that when you load up the notebook, like you can see the previous runs execution without running the whole thing, if that makes sense, right? Whereas we start from just the Python file and that file doesn't have those outputs.

17:29Software Engineering Daily Host:So we implemented our own sort of like session cache, which is stored in some sort of directory that is hidden from the user. And so to replicate some of these nice features that flat file format did provide.

17:42Akshay Agrawal:That makes sense. Well, and since you have the dependency, you have all the variables already like labeled. You know what you need to save. Exactly. Yeah. In mobile application security, good enough is a risk. GuardSquare uses advanced, multi-layered code hardening techniques and automated runtime application self-protection and mobile application security testing, combined with real-time threat monitoring to deliver the highest level of mobile app security. Discover how GuardSquare brings all these together to provide mobile app security for your Android and iOS apps without compromise at www.guardsquare.com.

18:23Akshay Agrawal:Why is there always a meeting bot in your Zoom call? Blame Recall.ai. Recall.ai powers the meeting bots and desktop recording apps behind products like Cluely, HubSpot, and ClickUp. They handle the hard infrastructure work, capturing clean recordings, transcripts, and metadata across Zoom, Google Meet, Microsoft Teams, in-person meetings, and more. So developers don't have to build it themselves. If you're building a meeting note taker or anything involving conversation data, Recall.ai is the API for meeting recording. Get started today with$100 in free credits at Recall.ai slash software. You know Fidelity is a financial services leader, but did you know that inside Fidelity is a community of technologists working together to shape the future of finance and tech?

19:13Akshay Agrawal:Fidelity is always investing in tomorrow. From emerging tech to cutting-edge tools that will transform what comes next. Their technologists are encouraged to keep learning so they can expand their skill sets, explore new ground, and stay ahead of this rapidly evolving industry. And right now, Fidelity is hiring technologists to join their team. Fidelity technologists get the best of both worlds. Startup energy that's grounded in the stability of a financial institution. That means support, resources, and amazing benefits. Bring your skills to a culture where you're empowered to dream big and build the tech that drives an organization and makes a real impact on people's lives.

Read the full transcript

19:53Akshay Agrawal:Find out more at tech.fidelitycareers.com. That's tech.fidelitycareers.com. Fidelity is an equal opportunity employer. Okay, so I want to come back to the UI widgets because that was the thing you talked about as well. And I think that's interesting. But this has brought me into this question or discussion topic around how notebooks fit into the broader software development lifecycle. Because I think one of the things I have seen in places where IPython notebooks were tended to be used before is they were heavily used by, for example, a data scientist or data science team. They were used for data exploration.

20:31Akshay Agrawal:and then if you wanted to then take something that was there and package it for reuse or embed it in a product or whatever, it was like a whole effort, porting, new code, all these different pieces. But to me, it sounds like this Marimo file is literally just Python. Once you have something that works, you could use it.

20:50Software Engineering Daily Host:Yeah, yeah, that's correct. So you can, and we do, and many of our own sort of internal utilities that we write for our team, just like internal tools, they happen to be in Marimo notebooks that are reusable. as Python files. You can even say like, from my notebook, import my function from my notebook, import my class, like that syntax kind of just works. There's like some details, the function needs to be pure, so that it serializes correctly, which by the way, ends up you end up writing better code, right? So this is, and so yeah, you totally can. And so it really does blur the boundaries of what you can use a notebook for, which I think is really exciting, because like, we see like all kinds of use cases from the traditional data science and research use cases to like backend engineers, like emailing us, telling us like, yeah, we're doing our data pipelines with Marima Notebooks just because we can.

21:37Software Engineering Daily Host:So I want to hear more about that

21:38Akshay Agrawal:because yeah, my experience has all been notebooks sort of often research ML data science communities and then that's its own thing. So how are you seeing the integration happening? When do you choose, if you have one of these backend engineers, when are they using a notebook? Why would they choose to do that over something else? And like, what is the process there?

21:59Software Engineering Daily Host:Yeah. So I think there's at least two reasons. One that we'll touch on with the interactive components that you alluded to earlier. But even without that, so I think often when making simple data pipelines, not super complicated ones, but simple ones, it can be helpful to prototype a data pipeline in a notebook because similar to what I was mentioning of having a back and forth with your data, right? Like you write some code, you see if the data is put into the shape you want it to be in. Sometimes it's easier to do that with visual inspection. So yeah, because these are due with visual inspection, it can be nice to do it in a notebook.

22:35Software Engineering Daily Host:A notebook's also a good choice because with data pipelines, the job runs and it can be nice to just have a report alongside it to see like, what was the shape of the data that day, et cetera. Right. So it's nice to prototype as a notebook, but now with Marimo, not only is it nice to prototype as a notebook, it's also really easy to just run it as a cron job or as a script. You don't need to reach for other tools that orchestrate IPyNB files. You just say Python my job dot py dash whatever. It's just a Python script. So that's one area where we do see natural usage.

23:12Akshay Agrawal:So this already is getting me to a place that I'm curious. So often, if I'm doing a big data analysis job, I will want to do something interactive on a subset of my data. and then I'm going to want to run something async when I do the full data because it's going to be big, slow, expensive, what have you. But maybe I want that same visualization, right? I want that report right in there. So like, is there an easy way within Marimo? And maybe I'm just missing something obvious to be like, plug, okay, right now we're using this local subset of data. For this one, you're going to do a remote call, fetch it from here.

23:43Akshay Agrawal:You're going to do what have you. Can you plug into those? I want to run my big data off in the cloud somewhere async fast. or slow, but get it back into my notebook?

23:54Software Engineering Daily Host:Yeah, you totally can. So it's not necessarily productized, but we have a number of primitives. And so one primitive you can use. So Marimo is a notebook, but it's also a library that you typically only use in the Marimo notebook. So you can import Marimo as Mo into your Marimo notebook and you get some primitives. One of those primitives is, am I running inside in an interactive session or am I running as a script? And so you can just use Mo.runninginnotebook to parameterize where the data is being fetched from. And I think that would do what you're asking, essentially. Or you could use that to do what you're asking.

24:29Akshay Agrawal:I'm also wondering about, yeah, can I, well, maybe this comes back to the publishing side and publishing things as web. So maybe we'll come back to this. But I'm thinking like, okay, I've run this on my temporary thing. I still want my interactive view, even though I'm running this asynchronously. Can I run it async and still publish the web version of Marimo so I can see the report at the end or something like that? Yes.

24:49Software Engineering Daily Host:Yeah, you can do that too. that is from the CLI, we have like a Marimo export and then choose your file format of choice, such as HTML. So Marimo export HTML that will run it as a script, but also generating an HTML report at the end.

25:06Akshay Agrawal:Got it. Oh, that's super cool, right? So what I'm hearing then, just thinking about life cycles here, right? It's like, okay, I'm tinkering with it. I'm exploring it. I'm running it locally as a notebook subset of data. I have this flag. I say, okay, I think this is ready. put in my flag saying, when you run as a script, run against this source instead of that source. And then I run it, generating an HTML report that looks the same as my notebook, lets me go and look at it, do what have you.

25:31Software Engineering Daily Host:Yeah, yeah, yeah, that's correct.

25:33Akshay Agrawal:That's really cool. Digging into that shareable web side of it. So what is interactive? What does that web generation look like? Can I still go and tinker with cells or change things? Like, What is the output starting to look like there?

25:48Software Engineering Daily Host:Yeah. So interactivity in Marimo starts, let's say, in an interactive edit session. So you're working on your notebook, you're in browser or VS Code, wherever you're using your notebook. We do have a VS Code extension as well. So I guess just taking one step back, repls are like, everyone thinks of a repl as interactive, right? Because it is, right? Like you run a cell and then you see what happens, you run something else. in like notebooks traditionally like when you want to see what happens when you change a value of some variable you have x equals five then you hit backspace then you change the value of x to six then you hit like shift enter and then you hit shift enter a bunch more times right then you see what the new thing looks like and then you go back to the road then you hit backspace x equals seven and then you do it again and again and it's like very tedious and like that's the kind of thing like any normal person would be like that there should be some ui element right to control the value of that variable.

26:38So in Marimo, you can import Marimo as Mo into your notebook.

26:42Software Engineering Daily Host:And then the Mo.UI module gives you access to a bunch of different UI elements, ranging from the very simple to like sliders and text inputs and drop downs to sort of more complicated ones like interactive scatter charts, like selectable scatter charts and things like this. And the way that it works in Marimo is that you can assign a UI element to a variable. And so like x equals mo.ui.slider. Then when you output that variable in the notebook, so if you make x the last expression of the cell, Marimo will display the slider. Then if you scrub the slider, then what Marimo will then do is then automatically run all other cells that refer to the variable x.

27:25Software Engineering Daily Host:So it hooks into the reactive execution system. And then every UI element has a value attribute that gives you the value that was assigned or like associated with it on the front end, gives it back to you in Python. And so like, just like that with no callbacks required, now you have like user interface interactivity in your notebook. And you can use that to say speed up data exploration. but as you can imagine you can also use that to make like really simple interactive data apps or web apps whatever you want to call them right like any kind of internal tool and so that's another big use case of marimo for different folks different sort of i guess pathways some people will just use the ui elements to speed up exploration but let's well just like you know i was just on the on the phone with like a i guess i i have to speak about them anonymously right now They're a big Marimo user.

28:15Software Engineering Daily Host:They're like a very well-known sports team. And they're using like Marimo for a bunch of like analytics and stuff. And they've made like a bunch of Marimo apps that like you type in a player's name and you see like a big table of like a bunch of stats about that player, et cetera. And it's just like fancy interactive web apps that they deploy on an internal site for the rest of their team to use. And the way that works in Marimo is that any notebook from our CLI. You can type at the CLI, marimorunmynotebook.py. And it'll serve it as a read-only web app. Code cells are hidden. And then now some non-technical stakeholder can interact with your data.

28:55Software Engineering Daily Host:And so you were mentioning backend engineers. And surprisingly to me, we talked to this one company, we have a case study published in our blog about them. Their name is TaxWire. They were like, yeah, all our backend engineers, we all use Marimo on a weekly basis. It's because we'll make all these internal web apps about this tech software we're building and its use cases. And then we embedded inside our internal Next.js app and it just makes our lives easier. I'm like, okay, cool. That's awesome. And like, honestly, that's not something I anticipated as a use case when I first started making Maremo right after my PhD, but it was really cool to hear people being empowered by that.

29:29Akshay Agrawal:That's awesome. So I think what I'm hearing too with this is like you could decide which variables you're exposing via ui versus which are loaded from somewhere versus how you're managing all of that so if you had this example of loading up a whole bunch of data i mean i guess your your sports example is that it's got a whole bunch of data in the back end but lets you configure which player you're looking at maybe which slice of data you want to analyze based on exactly yeah that's exactly it and i guess the

29:59Software Engineering Daily Host:value of the notebook here in this case because it may not necessarily be obvious it's just that like it's it's like this like progression from like i just have some data here and i don't really know what it is and i'm just like kind of playing with it and then like oh actually there's something useful here and like okay now i want to expose this to other people and it's just like you can just like stay in the same tool and just incrementally like abstract it a little bit and promote things to ui elements and then add some explanatory text it's just like a really seamless path. Whereas like, if you had never opened a notebook to look at your data in the first place, it might be hard to then, you know, think about, okay, what is a React app I'm going to write to like, I don't know, empower the rest of my team.

30:39Software Engineering Daily Host:Cause like, I don't even know what's in the data, so I don't even know what to make. And so I think that's what the value in the notebook is just making it really easy for you to first get a feel for your data. And then from there, making it really easy to make like a tool that's good enough for you and your colleagues in yourself.

30:55Akshay Agrawal:Yeah, that is super valuable. I mean, I have an example where I've just whipped up some internal scripts to analyze things for me, but it outputs in text because I think in text, if I had thought ahead and used a notebook, suddenly I could share it much more easily with different folks. Another question that I have in this is I think one of the things that's going on in the software development world right now is things are changing incredibly rapidly, We've got LLMs, copilots, agentic tools, all these different things. Is there a similar transformation going on in the sort of data and data programming worlds?

31:31Akshay Agrawal:Or how are you seeing Notebooks fitting in with this kind of new industrialized coding era we're getting into?

31:38Software Engineering Daily Host:Yeah, I think there is a similar transformation. It's a few different ways. So I'll give one example. So one of our users is like a very large sort of public company. And they told us they have hundreds of like Marimo apps deployed internally. And they said what made it really easy for them to adopt Marimo and the reason they adopted it so quickly was because it turns out Claude is like really good at writing them. Because it is a pure Python file format. It's like all these things you can get around, but like, well, I guess Jupyter Notebooks aren't interactive by default anyway. So you can't make web apps with them.

32:18Software Engineering Daily Host:But with, yeah, Marimo being a pure Python notebook, Cloud can write it easily. It can also run the notebook as a script to check if it's doing the right thing. We also have like a Linter CLI tool called Marimo Check, which will like report any errors. It finds like with the syntax of how the notebook is stored. And honestly, like I use Cloud also to like make really quick internal, like things where I would have made a Marimo notebook by hand. Now I have Cloud help me make a Marimo notebook of it. And it does a good job. Like, like, especially like compared to like earlier in the project's life cycle, Marimo is now, I think popular enough that it's like in distribution.

32:54Software Engineering Daily Host:So it works pretty well. I guess that is more similar to like how Claude is like just speeding up software development in general. Right. In terms of for data specifically or for ML research specifically, I think that people are still figuring out. I was just talking to someone on my team and he told me that he's seen a bunch of VC things out early this year saying, 2026 is going to be the year that AI and agents revolutionize how we work with data. So I'm like, I told him to send that to me because I haven't seen one of those yet. But I guess people are talking about it. I just, I'm not too familiar.

33:30Software Engineering Daily Host:I think like Databricks says that 80 % of something, something databases are created with agents, but I think that's like neon or something.

33:36Akshay Agrawal:and so i mean it's it's sexy right now to say okay we're gonna use this is gonna be revolutionized with ai i do feel like software development is the one place i am actually seeing that play out and so yeah kind of interesting to see how that happens um where i think there's oh go ahead i

33:54Software Engineering Daily Host:was just gonna say i think that the basic things that i think many people have been trying like maybe just being quietly integrated in a bunch of products but like text to sequel type of things like schema discovery, like it seems natural that some form of that will benefit from code generation and sort of agentic workflows. But I guess we'll see. Yeah, we'll see if these predictions end up being true.

34:15Akshay Agrawal:On this sort of forward looking prediction side, what do you see as sort of the frontier in terms of use of notebooks, data access, data exploration, that sort of world? And what kinds of stuff are you working on internally to address it?

34:30Software Engineering Daily Host:The frontier. And that's a good question. frontier is always hard to opine on because well it's the frontier so i can talk about what we're working on and see if i'll back out and see if any of that is going to be getting us to the frontier so let's see so some amount of work we have like a good amount of work is like sort of we're close to parity with the jupiter and collab ecosystems but we're not at full parity so we do actually honestly have like a good amount of parity work that we're doing like getting Marima working seamlessly inside of JupyterHub, which is like a multi-user hosted sort of Jupyter deployment that many universities use.

35:09Software Engineering Daily Host:We also have a free hosted notebook that's similar to Google Colab called, tongue in cheek, it's called Molab, Mo for Marima. Nice. And so we're working on that quite a bit this year. I guess one thing in terms of frontier, so we have a speculative project, which I don't exactly even know what exactly it is. And part of the project is to figure out what it is, but it's driving Marimo headlessly, like potentially with agents. There was this paper that came out recently that someone on my team sent me that I've so far only skimmed the headline and intro, but it's something to do with like, instead of using, I think the claim of the paper was that Python repls or repls in general can be very valuable ways to like dynamically create context for LLMs or agents.

35:57Software Engineering Daily Host:And in that paper, I think they use a Jupyter kernel as like the thing that the LLM has access to. And so we actually did some, one of our engineers did a lot of work to like sort of modularize Marimo towards the end of last year. So we're getting closer to a place where like you can use the kernel headlessly without our UI. And one project that he's really interested in exploring is like, well, what if you gave an agent access to that kernel, could that somehow like speed up whether it's like data exploration workflows or research workflows? We're not exactly sure, but it seems like that could just be like a really valuable primitive, like a sandbox in some sense for agents to have.

36:39Software Engineering Daily Host:Like, so I'll give you like guys one example, and this is not related to like necessarily headless, but like enabling agents to work more effectively with data. So like in Marimo, we have a built-in like AI assistant sort of system. And one thing that you can do is when you write a prompt, like for generating some code, you can tag a variable, say like a data frame. And when you do that, we inspect the data frame, see its schema, get like sample values, et cetera, and like dynamically generate context to give the LLM. And now you have like code that's like specialized to the data at hand, which is like more useful than like, say if you're like, I don't know, using cloud or cursor and it does, Yeah, so they won't know what's in the data frame unless you explicitly tell them.

37:21Software Engineering Daily Host:That's just one way that you can make, I guess, empower agents with runtime information. And I guess that's one thing that we want to explore more this quarter.

37:29Akshay Agrawal:Yeah, no, it's super interesting to kind of think about that because there's a couple different pieces that stand out to me. So one is the fact that you have kind of the dependency graph already mapped becomes quite interesting in terms of showing just the relevant context to the agent, right? So you might say, hey, you're changing something down in this one cell, but you don't actually care about that. Instead of forcing the agent to absorb the context of all the different steps, you could say like, register which step in the graph you're interested in. We'll do all the computation. Here you go.

37:58Akshay Agrawal:Here's the output. I do wonder, I feel like there is some sort of interesting opportunity here in terms of just like showing it the right things at the right time. Do you have a concept of lints or correctness checking on cells? So like you've got this dependency graph you're going through and maybe a step three or four ways down. It's like, oh, this is outside the bounds of what could be valid. So something must be broken upstream.

38:23Software Engineering Daily Host:Yeah. Yeah, we do. We do. We have a checker or a linter that will check your entire program for like semantic correctness as well as like syntactic correctness. So there's a couple of rules that Marema enforces to make sure that your graph remains like a DAG basically. one of which is you can't have cycles across cells, which I think is sensible. Although I recently learned that Excel has a feature that you can turn on in settings to enable cyclic calculations. And then you choose like the number of iterations you wanted to go until convergence. I got a spreadsheet that was all ref errors. I'm like, why did you send me the spreadsheet with all ref errors?

38:59Software Engineering Daily Host:No, no, no. You have to enable a fixed point iteration. Anyway, sorry. This is a digression. We don't allow that.

39:06Akshay Agrawal:There's a lot of value in keeping it to a tag.

39:08Software Engineering Daily Host:I'll say. Yeah. Yeah. So Marima has to be a DAG and we enforce that. That's one rule that we check. And the other is actually you can't redefine the same variable across multiple cells. And the reason is like we actually allow you to reorder cells for like presentation purposes. There's like a column view and stuff. So those are the two main semantic things that we check for.

39:27Akshay Agrawal:I was also wondering though, in terms of like data range validations, right? So for example, you have a dependency of a set of different computations and you might know something about the shape of the data and say like, okay, this data needs to actually be in this shape. If not, flag something and don't keep going or what have you.

39:44Software Engineering Daily Host:Oh, that's super interesting. Yeah, we haven't explored that.

39:47Akshay Agrawal:The reason I think about that is like with agentic coding, which I've been diving down in, like the more you can programmatically, deterministically limit the sort of scope of possibility and then give that feedback to the agent, the more it's able to independently iterate. Yeah, that's very interesting. There's a lot of things we could play with. But that does sound interesting. So you mentioned in Marimo, you have these things as functions with decorators around it. And you're already building in this kind of some amount of static analysis, some amount of linting and things like that. What hooks do you expose to your end users in order to kind of plug into that?

40:22Software Engineering Daily Host:So right now, to be honest, we don't have the biggest extension API surface area. In terms of extension points, not in the file format, but we have standardized on this protocol called AnyWidget for building third-party interactive widgets. The developer of AnyWidget, Trevor Manns, actually works at Marimo now. And so that's one way that you can plug into Marimo. You can also hook into our display protocol for objects. We support the IPython display protocol, but we also have some additional hooks. and then in terms of the file format itself like you can actually write marima notebooks by hand in vim or whatever your text choice how did you know i use vim well if you're holding software engineering daily that is my anyway it's my guess but so you you can do that that's not an extension point but it is designed well so that like you actually still the file format will guarantee you still get like code completion intellisense and all these things like but we haven't opened up an actual, like put another way, Jupyter is famously, I think, well designed in terms of like the internal protocol.

41:29Software Engineering Daily Host:Like there's like the kernel protocol. They have a wire protocol. They've got a bunch of things that like third party developers can hook into. We haven't done that yet just because it was too early for us to. I think things are just moving really quickly. It's starting to change. We now have our own still internal, but semi-public APIs for ourselves because we're now consuming Marimo in many different ways. And I think eventually over time, this will evolve into like some of these will be opened up to the public. But right now it's, I guess in terms of the DAG, it's like our file format specification is public and fixed.

42:03Software Engineering Daily Host:And so people can target that with code generation tools, but that's about the extent. That makes sense. Cool.

42:09Akshay Agrawal:Let's look a little bit actually at any widget. You mentioned that as one of the places that you plug in. It's an open source. you hired the developer or whatever the sequence. I'm a big fan of that type of thing. What is it? How does it interact with the reactive data model that you have? And are there any constraints or things to know going into it?

42:29Software Engineering Daily Host:Yeah. So I'm going to caveat this by saying that Trevor is a far better spokesman for AnyWidget than I am, but I will try to channel. And he was just on, I think, TalkPython with me and had a great hour-long conversation about AnyWidget. But it is both a spec and a tool set. for making like reusable widgets, I think really focused on for using them in like interactive like notebook environments. And so what any widget was born from, my understanding is from talking to Trevor is, so Trevor also has a PhD and sort of in like the biocomputation space was originally his focus. And he was found himself having to make widgets for like all kinds of sort of domain specific tasks in the Jupyter ecosystem.

43:16Software Engineering Daily Host:And they were kind of really difficult to build and maintain and test. Like web programming has advanced a lot in recent years and like IPython widgets had not kept up. And so I think sort of out of some of those difficulties and frustration, like Trevor sort of built AnyWidget to like make it a lot easier to implement and maintain these widgets and also to make it a lot easier for like different frontends to consume them. Like so before any widget, if you made some kind of domain specific widget that worked in JupyterLab, then you would have to go and then customize it to work in CoLab and then also make sure it worked in the VS Code extension.

43:53Software Engineering Daily Host:And now by like making the spec, like you can make an any widget, just make it once. And because people have agreed to support it, it'll kind of work anywhere. And so that's been really valuable. And it was really valuable for us. So like my co-founder, Miles, discovered any widget pretty early in its lifecycle. and he was like, oh, we should use this. I'm like, I never heard of this. It was my original version. He was like, no, no, no, it was really good. It was a really good bet because I think it has emerged as the standard for interactive notebooks. And in terms of hooking into a reactivity model, it's actually quite nice.

44:26Software Engineering Daily Host:So basically any any widget, you can wrap it in a mode.ui.anywidget wrapper and then it basically binds it to our reactivity model and makes it into just like any other UI element that's first part in marimo yeah so it hooks into the data flow graph in the same way and i think that the value of it is i mean people make all kinds of really really cool widgets and like we just like there's no way that like us as like a team of seven would be able to satisfy everyone but like widgets that were originally developed for jupiter is a scatter plot widget called jupiter scatter which you know lets you see like 10 million points on the scatter plot really efficiently and zoom in, zoom out, etc.

45:07Software Engineering Daily Host:That now works in Marimo today. And it's also reactive, which sort of gives it superpowers that you might not have had in traditional networking environment.

45:16Akshay Agrawal:Got it. So yeah, it looks to me like it's essentially a vanilla JavaScript spec that if you meet that, then you can wrap it up and it'll just plug in. You have a wrapper, Jupyter has a wrapper, other folks who support this have a wrapper, and it'll just kind of work anywhere.

45:31Software Engineering Daily Host:Yeah, yeah. And we're going to be focusing a lot on AnyWidget this quarter as well. One of our employees, Vincent, he runs our YouTube channel and does a bunch of things. He never ceases to amaze me how far he can get with vibe coding these really cool AnyWidgets. I don't know. He vibe coded this AnyWidget for robotic simulations with this humanoid person. I don't know. It's really cool. So it really allows you to expand your imagination and get creative. Nice.

46:01Akshay Agrawal:Well, we're getting kind of close to the end of our time here. Is there anything we haven't talked about yet that we should talk about before we wrap?

46:09Software Engineering Daily Host:I think we've covered all the basics. So I guess we didn't really talk about Marimo's origins. And then I can talk a little bit about his origins and a little bit about where we're going, like at least in the next few months. So I started Marimo after my PhD, like I mentioned, out of both appreciation for notebooks and frustration with them. And I actually originally got funding from a national lab at Stanford. It's a lab called Slack, which is a particle accelerator lab. And there are a bunch of scientists who use Python and like had basically the same gripes as I did. And so they were really excited to sort of partner with us to bring like a new open source programming environment into the world.

46:52Software Engineering Daily Host:So we have our roots in academia in that sense. and this quarter one thing that we're really interested in doing is like engaging a lot more with universities to help them try out marimo for education help support them like incorporate it into their classes because i really do feel like that like combination of reactivity and interactivity can really just make concepts just like a lot more intuitive like it's just somehow like if you learn like some new numerical algorithm it's just so much easier to just change your parameter and see what happens as opposed to just like abstractly think through it.

47:26Software Engineering Daily Host:And in fact, Marima's main inspiration, Pluto, for the Julia language, it was originally designed exclusively for education and still actually is advertised in that way. It was at MIT where it was designed for a computational thinking class. And I don't know, it's just something I care a lot about and something that we really want to support. So to the extent anyone in your audience is in the intersection of software engineering and education and you find Marimo interesting, please try it out or better yet, reach out to me and my team. We would be happy to chat with you guys and support you. Alright, let's call that

48:02Akshay Agrawal:a wrap.

From the publisher

Interactive notebooks were popularized by the Jupyter project and have since become a core tool for data science, research, and data exploration. However, traditional, imperative notebooks often break down as projects grow more complex. Hidden state, non-reproducible execution, poor version control ergonomics, and difficulty reusing notebook code in real software systems make it hard to

The post Reinventing the Python Notebook with Akshay Agrawal appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Reinventing the Python Notebook with Akshay AgrawalSoftware Engineering Daily · 46 min
Listen in VO