In short
Podcast Notes: Mamba and Software Package Security with Sylvain Corlay
Overview
- Podcast Title: Software Engineering Daily
- Episode Title: Mamba and Software Package Security with Sylvain Corlay
- Description: Discussion with Sylvain Corlay, the CEO of QuantStack, about the company, its projects like Mamba, software supply chain security, and recent developments.
Key Participants
- Sylvain Corlay: CEO of QuantStack, involved in open-source projects in scientific computing.
- Gregor Vand: Security technologist and founder of MailPass, co-hosting the discussion.
QuantStack and Mamba What is QuantStack?
- Open-source technology company focused on tools for data science and scientific computing.
- Key projects include:
- Jupyter: Active contributions to collaborative editing features and visual debugging.
- Conda Forge: A community-driven package management repository.
- Mamba: A fast alternative to Conda, developed to improve package management speed and efficiency.
Relationship Between QuantStack, Conda, and Mamba
- Conda: A general-purpose package manager for multiple platforms, not limited to Python; supports creation of multiple environments.
- Mamba: Built as a faster alternative to Conda, initially started as a community project due to performance issues with Conda's dependency resolution.
- Mamba is a drop-in replacement for Conda, written in C++ for speed and efficiency.
Mamba 2.0 Release Key Features of Mamba 2.0
- Significant refactoring for better software engineering practices.
- New features include:
- Support for package mirrors.
- More protocols for downloading packages.
- Enhanced dependency resolution capabilities.
Discussion on Vendor Neutrality and Open Source Importance of Open Source in Scientific Computing
- The need for transparency in scientific research necessitates open-source tools.
- Historical context: Scientists have long shared code to prevent dependency on costly proprietary systems.
Vendor Neutrality Challenges
- Conda’s ties with Anaconda Inc. raise concerns about potential vendor lock-in.
- Hard-coded channels and cryptographic keys in Conda restrict its flexibility and community governance.
- Mamba offers an alternative route to ensure project direction is not dominated by a single entity.
Software Supply Chain Security Mamba’s Approach
- Mamba's implementation of the Conda Content Trust protocol allows for community channels like Conda Forge to operate securely.
- Mamba aims to provide a framework addressing regulatory constraints for software security.
Future Directions for Supply Chain Security
- Focus on ensuring package authenticity and reproducible builds.
- Collaborations with organizations like Chain Guard to enhance reliability and security in package management.
The Role of WebAssembly Integration with WebAssembly
- Mamba is exploring how to leverage WebAssembly for educational tools like JupyterLite, enabling better scalability and accessibility.
- JupyterLite: A browser-based system allowing users to run Python notebooks without server dependencies, demonstrated in educational environments.
Community Engagement and Future Vision How to Get Involved
- Open invitation for developers to contribute to both Mamba and project-related repositories.
- Public meetings and platforms for collaboration on open-source initiatives.
Vision for Package Management
- Aspiration to create a community-driven ecosystem for managing scientific software.
- Ambition to educate future generations on programming without barriers related to infrastructure costs.
Conclusion
- Sylvain Corlay emphasizes the ongoing evolution of Mamba and its role in improving software package management.
- Encouragement for listeners to engage with the Mamba community and contribute to the open-source movement.
Additional Notes
- Acknowledgment of Wolf Volprecht, original creator of Mamba, and ongoing efforts in the broader community.
- Call for collaborative and inclusive approaches in the future of software development and educational tools.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Quantstack is an open source technology software company specializing in tools for data science, scientific computing, and visualization. They're known for maintaining vital projects such as Jupyter, the Conda Forge package channel, and the Mamba package manager. Sylvain Corle is the CEO of Quantstack. He joins the podcast to talk about his company, Conda, Mamba, the new Mamba 2.0 release, software supply chain security, and more. Gregor Vand is a security-focused technologist and is the founder and CTO of MailPass. Previously, Gregor was a CTO across cybersecurity, cyber insurance, and general software engineering companies.
0:39He has been based in Asia-Pacific for almost a decade and can be found via his profile at vand.hk.
0:59Hi Sylvain, welcome to Software Engineering Daily. Hi Grigor, thanks for having me here. Yeah, very exciting to have you here today, Sylvain. I think a lot of the listeners will know quite a bit about you and the projects you work on, and you might not need any introduction to them. but equally we'll have a lot of listeners today who know nothing which is also exciting and you get to talk to some completely new people in terms of what we're here to talk about today generally speaking it's Mamba and through the the company Quantstack so I think main thing is just to sort of set the scene here actually just to ask like what is you know you're the CEO of Quantstack and based in Paris what is Quantstack and sort of what's the relationship to Mamba how did it come to be, I think, a key maintainer of Mamba?
1:45Maybe just start there. Sure. So, yeah, so QuantStack, it's a team, mostly. It's a team of open source maintainers of key projects of the scientific computing ecosystem. So some of the main projects that we are active in are Jupyter. We're very active in the Jupyter project. The team comprises over 10 people working full-time on the project, and we've been some of the main drivers in the recent innovations in Jupyter, such as collaborative editing, the visual debugger for JupyterLab, the new version of the new flavor of the Jupyter notebook that came out recently. So Jupyter has been one of the main projects that we've been active in in the past years.
2:27We're also very active in the open source package management ecosystem for science, mostly with the Mamba project that we're going to talk about today and Counterforge. And finally, there is a new chapter that we've started actually a few weeks ago with the Apache Arrow project in the wake of the recent layoffs at Voltron Data. A bunch of maintainers of Apache Arrow showed up at Pellet Tappers and they told us that they were looking for a new home. And so we were studying this new chapter, which is very exciting. And so Quantstack, more than a team of open source developers, we're not a startup.
3:03So we operate under a service consultancy model where people and companies that depend on these tools in their operations or in their products contract us out to do bug fixes, maintenance, and sometimes add new features. So it really started as this form of self-employment for myself and a couple of others. So at the very beginning, we did not have any sort of strategy for growth. And nearly all of the business was inbound. And this was like this for a few years. And eventually it became a bit more deliberate. And now we are a team of about 30 people. The team is not just in France. So we have like the biggest group is in France, but about for half of it.
3:52And then we have a significant team in Germany and also folks in Austria, in the UK and Spain now. Awesome. And just to clarify, has it always been quite scientific leaning in terms of the focus or did that come out of a project? No, it's always been indeed focused on sciences since the very beginning, just from our professional backgrounds and the projects that we were focused on. I think that's a good sort of just distinction to make in terms of where any of this has come from. So again, for anyone not familiar with any of these sort of companies or packages that we'll be talking about today.
4:31So we're here to talk about Mamba and it has a sort of intertwined relationship with some other things like Conda Forge, Anaconda. Do you want to just maybe speak at a high level to kind of what the interrelationship is between them? Yeah, so assuming that people don't necessarily know what Kanda is, maybe I should define it. So Kanda is a general purpose package manager that works on multiple platforms like Windows, Linux, OSX. And it's very popular in the scientific computing ecosystem. So maybe I should better define what it's not because there is a lot of confusion about it. So Kanda is not a Python package manager and that it's not a package manager for the Python programming language.
5:14It's more similar to YUM or DPKG, like, you know, apt-get and RPM, you know, the classical Linux package managers. It's in that what is installed is binary packages and already pre-built assets. It's different from Linux package managers in that we can create multiple software environments in multiple locations in the file system, and it's cross-platform. Yeah. And again, in the context of a lot of science-based projects, is it fair to say that they're often working with multiple environments and actually that could be using quite different packages or languages? And that's part of the inspiration behind all of this?
5:59Yeah, so it came out of the Python scientific computing community, even though it's not a Python package manager. And it started from the observation that a lot of the popular Python packages were actually built upon Fortune or C++ code bases for efficiency and had a thin layer of Python for usage in a Python interpreter. and these posed significant distribution challenges. This whole thing started before Python had a wheel format for distributing packages that would embed binaries in the Python packages. So this was the raison d 'etre of the project to really enable a better story for distributing binary packages.
6:47The other thing is, I think the notion of environment is really key. I mean, people use Docker as some kind of way to bundle a bunch of stuff together that you can easily distribute. But it's still more of a way to distribute something where there could be a lot of mess. Package managers are really key in distributing software in a way that's reproducible, in my opinion. And reproducibility is also another key problem in scientific computing and in science in general. So being able to switch back and forth, like if you have, let's say, a Jupyter notebook that you've written maybe 10 years ago, which does some number crunching and produces a few plots that you use for a scientific paper, how do you run it today?
7:40How do you reproduce the set of packages and data sets that you were using to produce these plots? And this is one of the reasons why having a strong environment story for creating such bundles of packages is really important in scientific computing. Yeah, makes a lot of sense. So, yeah, I feel we've got a lot to cover today. So we're going to sort of dive straight in. Mamba 2.0, I believe, has not that long been released. What would you describe as kind of the big changes there? And, I mean, I guess Mamba generally is written in C++. So is that correct? So Mamba is really meant to be a drop-in replacement for Kanda initially.
8:23And it started as, so basically when the Kanda community grew, there is this community channel of packages called Kanda Forge that is starting to overgrow the rest of the ecosystem and had tens of thousands of packages. And the way the Kanda solver was devised was actually crumbling under the weight of KandaForge. And Mambas was created by Wolf Volpresch, who was an employee of Konstak at the time, as initially just a hack. Basically, I'm going to delegate the solving of environment and the dependency resolution to another library written in C++ and use Kanda for everything else. which was still fragile but proved very promising.
9:19So our first approach was to reach out to the Kanda project, Anaconda Inc., and tell them about it and ask if they would be willing to fund Quantstack to make this a reality. And at the time, they were not so interested yet in Mamba, so we decided to make it its own thing. So what Mamba is today is really an alternative to Kanda to install the packages of the same ecosystem that's fully compatible with Kanda, even supports the same common line options. But unlike Kanda, it's written in C++, which really helps with speed in some areas. There is another flavor of Mamba that is also very popular called Macro Mamba.
10:03And Macro Mamba is a single statically linked bundle of Mamba. It's essentially the same code base, but with different linkage. And if you want to install Micromaba, it's just four megabytes. It's a four megabytes download while an installer for like a basic installation of Kanda requires a Python interpreter and a bunch of dependencies in the space environment. And download sizes are in the tens of megabytes or maybe 80 or 90 megabytes over 100 and some platforms. So Micromaba has also proven very useful in CI workflows where we want to bootstrap an environment very quickly and fully self-contained and can be used to create kind of environments from scratch.
10:48Yeah, exactly. So Mamba 2 released in just the last couple of months, is that right? Yes, that's right. Okay, so talk to us about that. Yes, so Mamba 2 is, first of all, it's the result of a very boring refactor of Mamba. Really, Mamba was, I would say that the Mamba first major release was sort of the result of a rush. Basically, there was such demand for it in the community that we were rushing to cover the features of Kanda. And it was entirely built to be used as a command line utility. But then some people started using it as a toolkit. They wanted to use the internal components of Mamba.
11:30and these were not necessarily written in a way that was very solid and because if we assume that people use it as a common line utility, we can make all kinds of assumptions that are not necessarily satisfied if you are making a toolkit or a library that ought to be used in web services and be thread safe and whatnot. So basically, Marba 2 is in many ways almost a rewrite of Marba in a more deliberate software engineering with a more deliberate software engineering bottom-up approach, right? There are a bunch of new features, though. First, as of Mamba2, we can specify mirrors for package channels, which is really important for things I'm going to talk about later.
12:14And this will soon be utilized in popular installers such as those of KandaForge and others. And also, we support more protocols for downloading packages. And one key protocol that we wanted to support and is enabled in Mamba2 behind an experimental flag is getting packages from OCR registries and other cloud storage solutions. Gotcha. So let's sort of dive into one of the kind of key topics I think would be great to cover today. And, you know, I'd be curious, I think a thread throughout this would just always be what has, or maybe there's been no change, but what has MAMBA 2 maybe brought to this in any different way over version one.
12:54But let's talk about vendor neutrality, effectively. and what can you say on the basis of like for example what is let's take mamba specifically for now and then you know if you would like to branch out into any other areas like what's kind of preventing a single organization and i guess in this case it might be quant stack um you know from for example dominating the project's direction like what would you say to that before we get to mamba specifically i think just wanted to take a step back and talk a bit about but more like why open source at all? And I think there is a specific reason in the case of scientific computing that is not necessarily as present for other areas of computing.
13:39And that, in my opinion, there would be a very deep contradiction for physicists, for example, to try to understand nature with a tool that they don't have the right to understand. And so this contradiction is the key reason why scientists, long before the open source movement was a thing, were sharing code and sending descriptions and in-depth documentation of their code to each other. This is really the origin of the World Wide Web, which was started at CERN and started by a physicist. And then if we look at more recent events and the reason why the Python Open Source Scientific Computing Committee is so big is that it was also built as a reaction to the scientific computing world being held by a few corporations that were imposing very high costs on licenses for computing tools.
14:39And these costs were preventing people from certain countries from engaging because they couldn't afford the licenses. They also prevented, in some cases, students from using the tools. And so they were really building a world garden around scientific computing. So everyone in this field has this atavik sort of need for openness and making sure that we are not giving ourselves up to one corporation, right? And so distribution of packages and package management is really like one of the key areas that in which having one actor dominating you know how you distribute and install packages can be very harmful and so so obviously here in the case of conda we're talking about anaconda inc which in many ways is a company that i really look up to and has done great achievements in this area making scientific software more usable and accessible But Kanda has challenges with respect to vendor neutrality.
15:45And so obviously, if you check out the Kanda documentation, it's pointing to the Anaconda distribution and channels. In some areas, Anaconda channels are hard-coded in the Kanda codebase. One more, I would say, problematic thing is that there are some root cryptographic keys that are also hard-coded in the codebase that are used in package signing protocol, use the update framework, which is some kind of extension of, it uses asymmetric key pairs for signing packages and ensuring their authenticity. But at the moment, only Anaconda Inc. can sign packages that could be verified with a Kanda client.
16:25So first, Anaconda Inc. and the Kanda community, they have taken steps to resolve this. So some, for example, some of the documentation has been fixed recently. There are open issues, I don't know, maybe PRs open as well, removing the hard coding of Kanda channels. But I think in the end, like the elephant in the room is the branding and identity of the project, which is that Anaconda and Kanda, you know, really share an untruthier part of the name. Like the visual identity and the logos of the company and the project are very similar. And even if they really wanted to solve this, they are really facing a difficult situation.
17:06So, yeah, this is one of the main reasons why just the existence of Mamba is important because it allows, you know, projects like KandaForge to potentially rely on another implementation of the Kanda Content Trust Protocol. And, you know, in case of takeover or failure of the company, you know, we have fallbacks that we can also use. So without even getting into the technical benefits of Mamba that people could disagree with. Yeah, so I think just sort of then talking about, I guess, you've laid out a very clear reason why there are aspects of Condor Forge plus Anaconda that their approach, it can work, but there's just various obvious drawbacks if you just take it sort of at face value.
17:55in terms of then how Mamba has approached this. And, you know, if we think about things like community engagement and, you know, I'm very curious about what policies, like how have you set out policies in terms of, you know, ensuring contributors, regardless of affiliation, for example, have equal opportunity, which I think is kind of what you're getting at there. And maybe then just a follow-on from that is things like the signing of packages. What policies are then around that if it's not this one entity who can effectively design their own policy around that and you can use it or not? Where has Mamba gone with that?
18:34So it's not as much as who can put code into the codebase as much as what's in the codebase in terms of hard coding of channels and hard coding of keys. I think these two things are really important if we want people to be able to, for example, start a new community channel and have a key signing ceremony for starting a well-devised software supply chain security policy for the channel and then use Mamba. There is no place where we would prevent this from happening at the moment. And so, yeah, I think that's the main thing. For example, if you download the main micro Mamba binary, everything can be overridden and it's not hard coding any channel by default any package source by default so you have to configure it and choose where you're going to get your packages from and you know i think an obvious kind of place to go next is really supply chain security we've covered a couple of products and frameworks in the past on this topic but mamba i believe like this is kind of a core tenet as well of what Mamba is supposed to be helping with.
19:48So maybe could you just speak a bit to sort of what does it bring in in that way? And again, if there's any, I guess, comparisons, you know, effectively with Conda Forge or otherwise as to what is different and why, I think that's very interesting to understand. So it's a broad subject, so Conda, supply chain security. So one thing that is currently in the Marba codebase is an implementation of the Kanda Content Trust protocol, but in a way that would allow anyone to install the public keys on the system so that we could check packages. It's going to be really important for a community channel like KandaForge to be able to sign their packages in the future as it's increasingly becoming almost a regulatory constraint.
20:39So with the recent laws that were enacted in the US and are also coming to the EU, state agencies won't be able to use packages that are not implementing these kinds of good practices. As a consequence, it's going to percolate in the entire industry. So preventing CandaForge, which is the de facto main source of packages for scientific computing, from signing their packages, I think can be really harmful. That's why I think having an independent implementation of the current protocol, without even getting into any kind of innovation, is really important. And it's really bound to vendor neutrality.
21:26And getting back to this question of vendor neutrality, actually, I think one thing I didn't do is I think we have a path forward that should probably allow everyone to continue operating in a way that's satisfactory, including Anaconda, including Kanda Forge and everyone. And would also resolve this kind of branding issue around the Kanda project. Because what I really think is that Anaconda is really trying to resolve the situation. They acknowledge that there is a problem at the moment. And so they have opened up Kanda for a more community-led governance. They have transferred over the Kanda trademark to the NumFocus Foundation.
22:09And a proposal for the future would be to actually, rather than trying to have the Kanda project become more open, but still be bound by this kind of weird situation with the name is to create a broader organization that would encompass Mamba, Kanda, but other clients like Pixie and that would not be named after either of these projects. And Kanda would just be one of the members of this broader community. And this community should be built upon common standards and have an open governance for deciding where the future should be. I mean, I believe, at least at the moment, you have like biweekly dev meetings that you actually host that, you know, for the project that people can come and be a part of.
22:56Is that correct? Does that sort of play into what you've just been talking about? Like, is that almost like a grassroots thing there where how can the people that have the same vision can come together and not just talk about that? But is that a sort of forum for talking about this as well? So Kanda has their own sort of governance and regular meetings. And they also now have their really nice system for proposing changes in Kanda with the CEP process. We also host public meetings for the Mamba project. Kanda Forge is its own thing as well and is really focused on the tools for building packages.
23:39But obviously, they have a vested interest in the tooling. And so, yeah, this is why my take is we need to actually acknowledge all of that and have, you know, sort of a broader community movement. And to not continue with the snake-based names, we could, you know, should call it something else completely. Like not Rattler, not Kanda, not Mamba. Maybe call it, I don't know, scikit packaging or whatever. like something that really conveys the idea that we are about package management in a generic way and we come from these we have these scientific roots yeah i mean this is this is a bit of a sidebar but i was kind of curious here why mamba but just based on i mean i think obviously at the beginning of the episode you you said you had approached anaconda to suggest this this concept But at the same time, what was then the thinking to continue the snake naming if the idea was to have something that kind of didn't sit completely opposed to, but is meant to represent a different direction?
24:46So first, in the very beginning, we weren't really thinking too much about it. Like Mamba was probably the result of a name search for fast snake. and speed was the main reason of why it was started and we were just continuing the series of snakes in that ecosystem around Python and Canda and whatnot and since it was just a demo in the very beginning we went to Anaconda as a natural outlet for this demo and as a potential client that could fund this work then we continued working on it but outside of B-label cycles and as a side project for WorldFort Project at Quantstack and a bunch of others.
25:34And eventually, the problems that Kanda was facing became key for some of our clients. And we managed to put some Kanda and Mamba-related deliverables in some client contracts and it became a thing. So it's only recently that we realized that we needed to be more thoughtful about it. and make it have this standing in the community that, oh, maybe this should be a thing and it should be part of a bigger thing that includes Kanda and Pixie and whatnot. So just hopping back in from that, I'd like to just touch on the supply chain security bit a little bit more, and then we're going to switch gears a bit to WebAssembly because I think that's an interesting place to go and leave this stuff behind.
26:20But yeah, I mean, just in terms of the software supply chain side of things yeah again just what can you kind of speak to in terms of i mean i'm for example i'm more familiar with for example node package manager and nvm and all of that kind of ecosystem and sort of understanding what policies and what sort of decisions have been made around that to by no means fix a lot of problems but at least enhance certain issues that have come up in the last couple of years so again what is sort of mamba doing to that end yeah so one of the features that actually came with Mamba2 was the support of package mirrors which is great from a vendor neutrality standpoint but actually brings more challenges with software supply chain security in that if you have a network of mirrors, could there be a bad actor?
27:11Could there be a mirror that is compromised? And there is a number of approaches to be a bad actor in a network of mirrors for package channel. And this is actually sort of by chance a good reason why the CADA content trust protocol, which is based on the update framework, is a really good fit because it really addresses some of the key challenges that could happen in software distribution. For example, could someone managing a mirror freeze their packages at a certain date before a security update? All of the packages that they host would still be legitimate. Everything would be real. But how can we prevent that?
27:57So the signature should presumably work, right? So Tuff addresses this by requiring some kind of cryptographic heartbeat in that when you try to get a package from this channel, or it will download a cryptographic key that actually has expired. Unless the root origin of the package sources has re-signed it for this content, it's going to not consider this content as valid. So there is a number of attacks that could be done on a network of mirrors that Kanda Content Trust really addresses as well. So that's why it was important for us to implement it. Now, I really wish that we could fix the upstream codebase in Kanda so that we can install alternative keys for CounterForge and other channels.
28:43And then we will be able to enable package mirrors for these key channels used by the community. But really, distribution is just one part of supply chain security. There is another entire field that Mamba is not addressing at all, which is reproducible builds and guaranteeing that the content of the package that is shipped is legitimate. which is also a concern for many organizations. So we've worked with companies that build their own Conda-based distribution from source and don't get any binary from the internet, but use effectively CondaForge as some kind of Wikipedia of how to build stuff because it's actually really hard to build the entire scientific computing stack from scratch.
29:33And so CondaForge is a really good source of information on this. Yeah, I think we had an episode on not that long ago with a company, Chain Guard, and that's sort of their goal is, you know, reproducible builds, you know, containers that they build from source. And I don't know what, if any, interaction they have with Anaconda-based packages at the moment. But yeah, just if any listeners are interested just in that pure topic, generally, that's maybe an episode to go to. So, I mean, what would you say in terms of are there plans to kind of deliberately work with more, whether it's open source projects or companies that are aiming to go this direction on reproducible builds?
30:14Or like, is it a sort of major concern at the moment? Not specifically reproducible builds because we don't have funding for this. It's a hard enough problem that we can't just hack our way around it. we're talking about making sure that the packages that are uploaded on channels have been produced by the people we think they're produced by right so sort of really signing at the build time and uploading them to kind of package servers let's switch gears to web assembly kind of exciting that there's quite a large amount of support now for web assembly from the mamba ecosystem could you maybe just sort of speak to what inspired i guess the decision sort of add add this support and like what has been that sort of journey of evolution for this support and what's kind of being produced as a as a result of this yeah so this is probably the thing that i'm the most enthusiastic about these days in my work overall like and it all started from a grant proposal.
Read the full transcript
31:15So we were actually writing a grant for the French government, because, you know, the French company, to develop a Jupiter-based platform for, you know, secondary education, so high school education in, you know, get kids to learn Python. Around the same time, a group of high school teachers here in Paris worked on an interesting project called Baston, which is a pun in French which kind of sounds like Python but it's a slang word for fistfight and so Baston was a sort of fork of the classic Jupyter notebook but instead of coding out to a server for executing the code that you would type it would use this Python distribution in the browser called Biodite and so this whole thing started in 2019 and they got Baston to be in a workable state basically and they deployed it for the Paris school district.
32:11And it worked really well. And so other districts in the country started showing interest. And so they signed agreements with other districts and progressively they expanded this project to the entire country. So the thing has matured a bit, became more solid over the years. And now, so this project, like the deployment is called Capital. Capital has half a million registered users, and they have over 200 ,000 user sessions per week, and all of it is entirely served from one machine. So how come? The main reason is that by using WebAssembly, by running user code in the browser, we can become free of having to run, have a Docker image running in the cloud for each user session, right?
33:08So this scalability is crazy. Just to give a comparison, UC Berkeley runs a data science class, of course, called Data 8. It's a really big one. I think they have over 10 ,000 registered students, but there is a team of DevOps engineers that operates the Kubernetes-based deployment of Jupyter, underlaying this, and it costs over $100 ,000 per year to run in cloud compute. So there is a team of DevOps engineers and then there is significant hosting costs for allowing this, which is essentially having one Docker image per user session. Now, if you don't need this, all there is on this server hosted literally in the basement of high school here in Paris is a content management system for the user notebooks.
33:59And that's it. So now if we start making multiplications, you realize, oh wait, France only has so many high school students. It's not a very big country, right? And not all of them learn Python anyways. But if you consider a bigger country like Nigeria, they have over 200 million people at the moment. But the forecast is that the population is probably going to grow by at least 100 million mediums in the next 25 years. And most of these kids are going to go to high school. And presumably in the 21st century, they would want to learn programming. And at this scale, is it even feasible to have a Kubernetes-based deployment of Jupyter to learn Python?
34:41I'm not so sure, right? And if Nigeria wanted to do this, they are not the home of Microsoft or AWS or Alibaba. they would probably need to rent that space on someone else's cloud, right? While if they use a system based on what I just described in WebAssembly, they will be able to host the platform in a sovereign fashion. So to me, this is really enormous because this model can be used to teach programming to a billion kids. It just works, right? Because basically you run everything on the browsers of the end user, right? Just to put this in context, you know, what is, just to take super basic points here, what is like a very base level of machine needed to run that just in a nice way?
35:38On the client side, that's what I mean, yeah. Well, depends on what you want to run, obviously. The example you gave with, for example, the high school kids learning Python. Okay, so what kind of Python is a high school kid going to write? they are going to learn how to compute the greatest common denominator. The complexity of that code is probably lesser than the rendering of the UI. So for very basic things, we don't need much. And you can already expose this kind of interactive computing environment for them. So the problem of Capital and this deployment that I talked about earlier is that they forked the original classic notebook from 2016.
36:19And this code base, obviously, is not maintained anymore and poses many challenges in terms of accessibility, in terms of security. And so this was the reason why we proposed to build a JupyterLab-based solution that would also use this new JupyterLab-based notebook. And we called it JupyterLite. And as soon as we started the JupyterLite project, it also grew in adoption very quickly. Now if you go to numpy.org and you scroll a bit you'll find a Jupyter Lite console. You can try NumPy in the browser and there is nothing running in the cloud. If you go to the scikit-learn documentation you will find runnable code snippets and now they are going to run their MOOC on Jupyter Lite as well.
37:00If you visit the SimPy project documentation you will also find a console there to try out SimPy in the browser. So after we developed Jupyter Lite we realized that we wanted to start expanding a bit what you can do in the browser. And Pyodide is really meant to be about Python and the Python packaging ecosystem, and we wanted to do a lot more. And we started this project called Emscutting Forge, which is a distribution of Kanda packages for the browser that goes way beyond the Python ecosystem. One thing, for example, is that we recently released a JupyterDite terminal with an emulator of bash that runs in the browser.
37:43And then you can start typing bash commands like grep, sed, touch, cat, less, whatever, and see it reflected in your file system and address these files in notebooks running in your browser. And there is another ongoing project to build all packages for WebAssembly and all of this using the same Conda-based and Mamba-based package manager. And even beyond education, I think this is going to be really important for scientific publishing and long-term reproducibility. For example, today, any binary that runs on your computer, be it ARM or x86, is probably going to require some kind of emulator to run on a machine in 20 years.
38:35while WebAssembly is a web standard. So presumably WebAssembly binaries should be runnable by web browsers in 20 years. So what I think is that a bundle of a Jupyter notebook and a bunch of WebAssembly packages and small data set all serve statically at a given URL is like a time capsule. It's like a website from the 90s that we can still see today. So this time capsule has a research paper doing some number crunching and data analysis and some discovery could still be runnable in 20 years. And this, to me, is a real revolution. So we're trying to push as much as possible the boundaries of what's possible in the browser.
39:17Yeah, there's a really fascinating way of looking at it. I mean, at the end of the day, the browser has kind of become almost the OS of the average user these days, generally. You know, I don't, most, I would say, unpowered users are basically just working through their browser for most things. But that said, there's still a way to go for what could, in theory, be run in the browser. And obviously, you've given this great example of Jupyter Lite and made possible by Emscription Forge. And I guess my question is, Mamba as a whole, is there going to be a fork in the road where if the WebAssembly side of things is, I don't want to say takes off, but you know what I mean, is there going to be a point where you have to actually make a choice between what the focus of mamba is potentially so this refactor that i spoke about earlier where we try to make mamba more of a toolkit is really paying off here because we are writing javascript bindings to some of the components of that toolkit and so that we can actually do the dependency resolution in the browser and download things from CDNs from the browser directly and create a counter environment in browser.
40:32So, yeah. Although in our team at the moment, there are probably more people working on WebAssembly focused than on the core C++ code base of Mamba. But for obvious reasons, the initial left is really significant. And there is a lot of code to be written for pieces that are just simply missing. Gotcha. So, I mean, we are slightly cruising towards the end of the episode today. I think just in general, it would be, I just want you to have a sort of a platform to be able to speak to developers here as much as being able to speak on behalf of Mamba and Quantstack generally. I mean, what would you say, like, what's the most important thing that you think people should know about sort of your vision for the future of package management in terms of picking up Mamba versus anything else or just generally.
41:29For me, this WebAssembly story is the crossover between Jupyter and our effort on package management. And just having this idea of providing a platform that can be used to teach programming to billion kids, potentially, and that will just work at this scale. If I have the opportunity to do this, I think it's probably the greatest opportunity of things that I could do professionally in my life. And so I'm taking it. I want to try it, right? Maybe this is going to be another platform that's going to be that. But I think we have a shot. So I want to do this upholding the principles that we laid out in the very beginning about the open source movement and the fact that the tools should be opened and openly governed.
42:16So join us. That's the message to the people. I think there is something happening now in the package management ecosystem as well as in the Jupyter ecosystem. That's really important. And I've read somewhere that there was a survey. Is it IBM? I'm not sure. They were trying to estimate the number of users of Jupyter. And the answer was probably in the order of magnitude of 10 million people in the world, which I think is probably fair, like certainly more than a million. And it's probably not 100, right? So like 10 million is probably not a crazy number. If we ever get to 100 million, it's going to be with a tool like JupyterNight.
42:53That's what I think. And yeah, we need help. Yeah. Well, just off the back of that, where's the best place to go? What's the best sort of inroad for someone to come, you know, when you talk about joining you, like where should they start? We have a number of easy fix issues in the relevant repositories and GitHub. Both Jupyter and Mamba have public meetings that you can find references to in our website. so yeah it's easy to engage but just a pull request fixing you know the tiny usability thing that you find annoying when using it is already super welcome awesome i mean i think that's just great advice for anyone looking to get interested in open source full stop but i think you know you've heard it from the the source today which is you know it's a very welcoming community it sounds like and exactly just lend a small hand and then who knows where that can go so sylvan it's been fantastic to have you here today i think you know again some quite meaty topics and obviously mamba is doing some pretty huge things and obviously there's always going to be some you know different groups in every area of tech and everyone has you know their different approaches and i think it's always great for our listener base to get to kind of hear from those they might have maybe seen some online discussions or or even just they've really read a pull request or GitHub issue kind of thread and navigating to hear from the voice.
44:20So I think it's been a really valuable discussion. Thanks. I wanted, before we leave this, to make a shout out to Wolf Volprecht. So Wolf is a former employee of Quantstack, and he's the person who started the Mamba project while he was at Quantstack and did a lot of the initial work. Now Mamba is maintained by a team of four or five people working on the code base. And Wolf actually is still a big driver behind this broader community. He founded a company called PrefixDev, which is the company behind Pixie. It's a set of tools built in Rust for addressing other package management issues and also seeks compatibility with the Kanda ecosystem in terms of package format.
45:07And I would say that Wolf in the past few years, not just with Mamba, and also what's going on with Pixie at the moment has been one of the main drivers for change and innovation in that space. And maybe that's another message for people. They should really follow what's going on there. Fantastic. Yeah, well, thanks for calling that one out. And I hope we get to catch up again, you know, in who knows, maybe a year's time or something like that. And we'd love to be seeing where things are going, especially on the WebAssembly side. That just sounds a very exciting place to kind of be epicenter, so yeah.
45:40Thank you, Gregor. Thanks so much.
From the publisher
QuantStack is an open-source technology software company specializing in tools for data science, scientific computing, and visualization. They are known for maintaining vital projects such as Jupyter, the conda-forge package channel, and the Mamba package manager. Sylvain Corlay is the CEO of QuantStack. He joins the podcast to talk about his company, Conda, Mamba, the
The post Mamba and Software Package Security with Sylvain Corlay appeared first on Software Engineering Daily.
