scikit-learn & data science you own

19 Nov 2024 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast Episode Notes

Episode Title

scikit-learn & Data Science You Own

Episode Summary In this episode of Practical AI, hosts Daniel Whitenack and Chris Benson discuss scikit-learn, a popular machine learning library, with guests Yann Lechelle and Guillaume Lemaitre from the company Probable. They explore the significance of scikit-learn in the data science field, its applications, and the future direction of open-source projects in AI.

Key Contributors

  • Yann Lechelle: CEO of Probable
  • Guillaume Lemaitre: Open Source Engineer at Probable
  • Chris Benson: Principal AI Research Engineer at Lockheed Martin
  • Daniel Whitenack: CEO at Prediction Guard

Topics Discussed

  1. Introduction to scikit-learn
  2. Definition: A foundational library for machine learning in Python, heavily used for building classifiers, regression models, and data preprocessing.
  3. Use Cases: Applications span various industries including healthcare (disease detection), finance (fraud detection), and predictive maintenance.
  1. Probable and scikit-learn
  2. Company Background: Probable is a spinoff from the French research center INRIA, focusing on maintaining and developing scikit-learn and other open-source data science projects.
  3. Vision: To create a suite of technologies that enhances data science practices while remaining open-source.
  1. Importance of Open Source
  2. Open Source Mission: Yann emphasizes the aim of creating more open-source tools for the data science community, counteracting the trend of proprietary software dominance by large tech companies.
  3. Community Engagement: The governance of scikit-learn is designed to ensure it remains community-driven, with contributions from developers around the world.
  1. The Future of scikit-learn
  2. Sustainability and Governance: Discussion on the balance between profit and the mission to keep scikit-learn open-source, emphasizing a structure that supports long-term community benefits.
  3. Technology Evolution: Exploration of how scikit-learn may adapt to changes in technology, such as the rise of generative AI and large language models, while maintaining its relevance.
  1. Practical Applications of scikit-learn
  2. Callback Feature: Introduction of new features that enhance model introspection and transparency, essential for industries dealing with compliance and regulatory requirements.
  3. Integration with Other Tools: Discussion on libraries that complement scikit-learn, such as Scrub for data preprocessing and Scope for bridging machine learning with database management.

Key Takeaways

  • Community Contribution: The power of scikit-learn lies not only in its features but also in the active community that continually contributes to its improvement.
  • Market Relevance: Despite the emergence of advanced AI models, scikit-learn remains crucial for a significant percentage of practical machine learning applications.
  • Future Preparedness: The importance of being adaptable to technological changes and user needs to ensure continued relevance in the fast-evolving AI landscape.

Conclusion The episode highlights the critical role of scikit-learn in the data science ecosystem and the concerted efforts of organizations like Probable to maintain and innovate within the open-source community. As data science evolves, so too does the need to create tools that empower users while remaining accessible and sustainable.

---

Additional Resources

  • [Probable Website](https://probabl.ai/)
  • [scikit-learn Documentation](https://scikit-learn.org/stable/)
  • [TechCrunch Article on Probable](https://techcrunch.com/2024/02/01/probabl-is-a-new-ai-company-built-around-popular-library-scikit-learn/)

Sponsors

  • Timescale: Purpose-built performance for AI applications.
  • WorkOS: A platform for adding enterprise-ready features to applications.
  • Shopify: Start your e-commerce journey with a trial at Shopify.

---

This episode serves as a valuable resource for data scientists, tech enthusiasts, and anyone interested in the practical applications of AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:03Welcome to Practical AI, the podcast that makes artificial intelligence practical, productive, productive, and accessible to all. If you like this show, you will love The Change Log. It's news on Mondays, deep technical interviews on Wednesdays, and on Fridays, an awesome talk show for your weekend enjoyment. Find us by searching for The Change Log wherever you get your podcasts. Thanks to our partners at Fly.io. Launch your AI apps in five minutes or less. Learn how at Fly.io. Okay, friends. I'm here with a new friend of ours over at Timescale Avthar. Suothen. So Afthar, help me understand what exactly is Timescale.

0:45So Timescale is a Postgres company. We build tools in the cloud and in the open source ecosystem that allow developers to do more with Postgres. So using it for things like time series, analytics, and more recently, AI applications like RAG and search and agents. Okay, if our listeners were trying to get started with Postgres, Timescale, AI application development, what would you tell them? What's a good roadmap? If you're a developer out there, you're either getting tasked with building an AI application, or you're interested in seeing all the innovation going on in the space and want to get involved yourself.

1:17And the good news is that any developer today can become an AI engineer using tools that they already know and love. And so the work that we've been doing at Timescale with the PGAI project is allowing developers to build AI applications with the tools and with the database that they already know, and that being Postgres. Yes. What this means is that you can actually level up your career. You can build new interesting projects. You can add more skills without learning a whole new set of technologies. And the best part is it's all open source. Both PGAI and PG Vector Scale are open source. You can go and spin it up on your local machine via Docker, follow one of the tutorials on the Timescale blog, build these cutting edge applications like RAG and search without having to learn 10 different new technologies and just using Postgres and the SQL query language that you will probably already know and are familiar with.

2:08So yeah, that's it. Get started today. It's a PGAI project. And just go to any of the timescale GitHub repos, either the PGAI one or the PG vector scale one, and follow one of the tutorials to get started with becoming an AI engineer just using Postgres. Okay, just use Postgres and just use Postgres to get started with AI development, build RAG, search, AI agents, and it's all open source. Go to timescale.com slash AI. Play with PGAI. Play with PG Vector Scale. All locally on your desktop. It's open source. Once again, timescale.com slash AI.

3:04Welcome to another episode of Practical AI. This is Daniel Whitenack. I am the CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Doing very well today, Daniel. How's it going? It's going great. I was saying that I'm really pumped to be talking about something that's near and dear to my heart over many, many years. Because today we have with us Jan, who's the CEO at Probable, and Guillaume, who's an open source engineer at Probable. Welcome. Thanks for having us. Well, Jan and Guillaume are working on data science that you own, including projects like Scikit-Learn, which is, of course, very near and dear to me, along with other data scientists all around the world.

3:59So Jan, if you could, since you're coming from the CEO perspective, help us understand a little bit, maybe for those that have heard of Scikit-Learn or some of the other projects that you're involved with, but they haven't heard of Probable. if you could give us a sense of what is Probable. As you mentioned kind of in the lead up to this conversation, it's a slightly different kind of company that came about in different sorts of ways than other types of startups. So yeah, if you could give us a little bit of context, that would be great. Well, very glad to be on the show with you to date. And Probable is a company that is typically known as a spinoff from a research center in France called Inria.

4:45And INRIA is the place where this technology, PsychicLearn, has been developed over the past 10, 15 years. Not many people know that. And the project has been somewhat protected and sort of incubated within that research center. And after all that time, as you know, PsychicLearn has been adopted or even probably participated in creating the field of data science. because it is applied math and essentially has created a sort of paradigm for how data scientists approach data science, typically through two functions, fit and predict. And the French government has a national strategy for AI, like many, many countries.

5:31And the government decided to double down on psychic learn. And they came up with a budget. They entrusted the research center with that budget. But then they also asked for the project to be break-even at some point. And the team said, okay, break-even is fine, but we don't do that in the research center. We don't break-even. So why don't we call an entrepreneur to try and help us figure it out? And they called me. So I have a track record as a software engineer and an entrepreneur in tech for the past 25 plus years. but I'm not a data scientist. So I did my due diligence and I sort of dug deep to find out what this project was all about under the hood.

6:16Is it any good? Is the community any good? And of course, Psychic Current is this quite amazing jewel of a technology that every data scientist on the planet uses. I discovered that it was downloaded 1.5 billion times cumulatively 80 million times a month, 22 % in the US, only 3 % in France. So this is a project that is used all over the world. And Probable is essentially the spinoff that takes all of the team, including Guillaume here, from the research center and turns it into an open source company that inherited the mission that was initially given to the research center. And the mission is to build a suite of open source technologies, including Psychic Learn, but above and beyond Psychic Learn as well, for data science.

7:12So the scope is large. The mission is noble. And this is what we're building, essentially. So Probable is a one-year-old company that has already started doing many, many things. And Guillaume is the representative here for Psychic Learn, this technology that is used, again, by every data scientist on the planet. Well, this has brought up a lot of interesting questions on my end. And I really love the part of your pitch and at least how you framed it on your website and in your materials online about data science that you own and the open source side of this. which I know from experience, there can be some interesting challenges around finding business models that really work with open source technologies.

7:59And we've seen technologies where companies start with the posture towards open source and then gradually become more closed over time. So I'm wondering from the leadership perspective, it sounds even in the way that this company was formed that there is a posture towards stewarding the, you know, PsychitLearn and these types of projects. But from your perspective, what is your posture towards stewarding these projects in the open source side? And how do you view the business element of this to make it sustainable in the longer term? So that is the hard question, but it is the one that is important here.

8:41Psychic learn is a technology that is, again, applied math. It's not rocket science, but it's applied math, and it's quite intricate. The thing is, the scientific community uses it day in, day out, and everyone depends on this. So typically, when I discovered the scope of the project and the mission that was entrusted to the research center, I realized that this project is bigger than me, number one. Number two, the mission is to actually create more open source. In other words, in 2024, it's even more acute. Typically, big tech keeps on amassing so much power, concentration. And we could argue that they do not distribute as much as they should.

9:28So that's not a judgment, but it is a fact. And psychic learn is precisely the contrary. It actually enables so many companies to do data science. So with that in mind, before creating the company, we decided to craft a sort of architecture for the company that would respect that. And so before we created the company, before Guillaume joined as a co-founder, before he even incorporated the company, we had a template that actually created the governance, the shareholding structure, and also leveraging a new law in France that allows us to do a sort of B Corp. So a company with a mission, where the mission is clearly stated in the bylaws, and that mission is to create open source for data science.

10:17So in a way, we've created a sort of constrained environment that is unlike many companies because it's by design. This company by design has created guardrails so that the governance cannot take this company too far on the right, let's say, proprietary technology or even changing the license. That's not in the cards. And we've created a sort of mechanism where, you know, if we do not uphold to the mission, then we can actually lose some of the assets, such as the brand. We are the official brand operator, but the brand belongs to the research institute, right? Still. So there are many mechanisms, trigger mechanisms that force us, including shareholders that we would bring in, to actually bind with the mission long term.

11:08Gotcha. You've raised so many questions for me that I want to ask. I actually want to take just a moment and kind of go back because it occurred to me as we're talking about this, for some folks listening who may have never even used Scikit-Learn, they might have heard the name and stuff, and you talked about it being applied math. Could you guys expand on that a little bit for somebody who hasn't had a chance to ever actually utilize it themselves in terms of what it's doing and kind of catch them up to us in the conversation a little bit? And then I'm going to pepper you with a few more questions because you've got me really interested.

11:42You hit so many topics on that last. Yeah, so maybe I can give a bit of background. So, basically, the tagline is machine learning in Python. So, it goes back to the, let's say, statistical route. So, the simple answer is we try to make predictive modeling. So try to use mathematics to form like data doable in the future to like give answer to a specific question, to a specific paradigm. The big difference with generative AI or deep learning is just that all the statistics that you have there are simple stuff. So they are fundamentals and deep learning build on those, but are just like much more, let's say, costly to train, costly to in the inference states or not in the same scope as well.

12:35And is like the de facto choice when you want to have like tabular data. So Excel spreadsheets data structuring this way. So that's a de facto way of training to be able to hit those spreadsheets and give back some labels or some regression, let's say. And whatever is like image or NLP or like this is like, let's say, deep learning and Transformers is more in that area. So we are more like back to what was machine learning like a few days ago, but that have many, many, many applications. Well, considering how incredibly popular and foundational in the data science world, could you kind of give me a little bit of a landscape view?

13:21And I'm not sure which of you would be the right one to answer, so you guys pick between yourselves. But a little bit about kind of how that fits into the data science landscape with AI coming in, just so that with people listening, they can kind of go, Ah, I see how it fits into the many organizations and tools that are out there. How do you think about that for that? And then I'll get back a little bit more to the organizational stuff that was talking about a few minutes ago. So maybe I can answer like partly, which is by giving use cases and to see that, for instance, with a partner as well that we worked over the years to give like, where do you find machine learning?

14:02And for instance, machine learning can be found in healthcare, where you want to know if the drug works or not. Then if you want to find diseases as well in some type of data, it could be as well like fraud detections in banks, in insurance, predictive maintenance, and those type of like all applications that you have since like many years. so let's say the use cases are very very large and what brings like cyclone on that is that this is not I mean it's not for one of the use case I mean it was fought from the beginning to be general enough such that you can apply to any of those use cases and to come back to let's say classification and regression programs let's say or unsupervised learning as well but like that you can apply the tool anywhere in that field So maybe, Jan, you have something more to add?

14:55Perhaps also at the macro level is to say, you know, Psychic Learn does a lot of things, including deep learning. But to be frank, when you want to do deep learning, typically you'd go to PyTorch or TensorFlow. But for everything else, you know, Psychic Learn. In other words, in the great AI family of algorithms, there is machine learning. And within machine learning, you have deep learning. Within deep learning, you have other categories of algorithms such as transformer-based models that lead to LLMs. So it's basically Russian dolls of sorts. And scikit-learn is the biggest provider of algorithms in the machine learning space.

15:40And in fact, if you look at the downloads, typically, Scikit-Learn is downloaded as many times as PyTorch and TensorFlow combined, which is crazy because now everyone is talking about LLMs, of course, but also deep learning because deep learning is currently in a spring state, not quite a winter yet. So, of course, deep learning and Gen.AI is a wonderful breakthrough. That being said, I like to simplify sometimes the 80-20 Pareto distribution. So I had the intuition that 80 % of the use cases out there use scikit-learn when it comes to machine learning. And people actually tell me, no, Jan, you're wrong.

16:26It's more like 90, 95%, right? Because in terms of, you know, technology that is robust, that is tried and true, that is used to actually, you know, turn a profit or return on investment. Banks and insurance companies, right? Guillermo was mentioning fraud detection. Fraud detection typically uses scikit-learn. And that actually saves money. That actually, you know, banks would be losing money without that. So it is actually quite essential. But again, it's applied math, right? So scikit-learn is only a facilitator to this category of problems.

17:16what's up friends i'm here with a friend of mine a good friend of mine michael greenwich ceo and founder of work os work os is the all-in-one enterprise sso and a whole lot more solution for everyone from a brand new startup to a enterprise and all the ai apps in between. So Michael, when is too early or too late to begin to think about being enterprise ready? It's not just a single point in time where people make this transition. It occurs at many steps of the business. Enterprise single sign on like SAML auth, you usually don't need that until you have users. You're not going to need that when you're getting started.

17:54And we call it an enterprise feature. But I think what you'll find is there's companies when you sell to like a 50 person company, they might want this. They actually, especially if they care about security, they might want that capability in it. So it's more of like SMB features even if they're tech forward. At WorkOS, we provide a ton of other stuff that we give away for free for people earlier in their lifecycle. We just don't charge you for it. So that AuthKit stuff I mentioned, that identity service, we give that away for free up to a million users, one million users. And this competes with Auth0 and other platforms that have much, much lower free plans.

18:28I'm talking like 10 ,000, 50 ,000, like we give you a million free because we really want to give developers the best tools and capabilities to build their products faster, you know, and to go to market much, much faster. And where we charge people money for the service is on these enterprise things. If you end up being successful and grow and scale up market, that's where we monetize. And that's also when you're making money as a business. So we really like to align, you know, our incentives across that. So we have people using AuthKit that are brand new apps, just getting started companies in Y Combinator, side projects, hackathon things, you know, things that are not necessarily commercial focus, but could be someday.

19:03They're kind of future-proofing their tech stack by using WorkOS. On the other side, we have companies much, much later that are really big who typically don't like us talking about them. They're logos, you know, because they're big, big customers. But they say, hey, we tried to build this stuff or we have some existing technology, but we're sort of unhappy with it. The developer that built it maybe has left. I was talking last week with a company that does over a billion in revenue each year. And their skim connection, the user provisioning, was written last summer by an intern who's no longer obviously at the company and the thing doesn't really work.

19:35And so they're looking for a solution for that. So there's a really wide spectrum. We'll serve companies that are in a, you know, their office is in a coffee shop or their living room all the way through. They have a, you know, their own building in downtown San Francisco or New York or something. And it's the same platform, same technology, same tools on both sides. The volume is obviously different. And sometimes the way we support them from a kind of customer support perspective is a little bit different. Their needs are different, but same technology, same platform, just like AWS, right? You can use AWS and pay them$10 a month.

20:03You can also pay them$10 million a month. Same product. Or more, for sure. Or more. Well, no matter where you're at on your enterprise ready journey, WorkOS has a solution for you. They're trusted by Perplexity, Copy.ai, Loom, Vercel, Indeed, and so many more. You can learn more and check them out at WorkOS.com. That's W-O-R-K-O-S.com. Again, WorkOS.com.

20:50so Jan you were kind of already going there and I I love the direction that you're going with this but I think maybe I could tee up a softball for you here because I'm personally passionate about the answer to this question and you probably have a better view on it. But there might be people out there maybe listening to this podcast who are thinking, well, now that we have Gen AI, we have large language models, I could put in a prompt to one of these models to do fraud detection or to find entities in text or to make some prediction of a classification. And sometimes that works. And so maybe there's people thinking, well, there's these general purpose large models out there.

21:38How does that change the way that something like Scikit-learn plays in industry? And I personally would argue and think that this actually makes Scikit-learn more valuable, if anything, rather than less valuable in terms of the ways that it can be combined even as a tool that's orchestrated with Gen.AI models. But I'm curious your perspective on this from the business side, and maybe Guillaume has some ideas on the technical side. Yes. So, Psychiclaren typically is this one technology that is patrimonial. In other words, it belongs to everybody. In fact, there's another stat when you look at the figures that are public, by the way, the number of dependencies.

22:24So Psychiclern is actually used by nearly 900 ,000 projects on GitHub. So there's nearly a million projects that depend on Psychiclern. And there's a new law that I discovered recently, someone mentioned that, Lindy's effect, which means that something that's been used long enough will remain important for long enough. So not saying that PsychicLearn will go the way of COBOL, but PsychicLearn is here to stay and we are with the community, the guardians of that. So we're going to make sure that PsychicLearn remains there forever for companies that actually need it in a stable version. And of course, Guillaume and the team are building up new features as we go.

23:10So there's a dedicated effort, and I should say that we have carved out nearly 10 people in a team are doing only that, contributing to Psychiclern and the other associated libraries. Now, your question, Daniel, is whether Psychiclern will be obsolete in, say, a number of years because general purpose technology has made it irrelevant in some ways. Number one, Psychiclern is extremely frugal. It actually works on CPUs. And it is well-controlled, well-understood. It's actually quite predictable in some ways, whereas deep learning is usually known as a black box where it's really, really hard to introspect.

23:56And so Psychiclern does produce for certain categories of problems things that are actually working quite well, More so than large language models, for sure, today, and more so than any sort of deep learning-based technology that we understand today. Now, it is possible that with additional data, additional training and techniques, and even evolutions on the transformer-based model, we could improve and probably render obsolete psychic learning. But to us, and then Guillaume and I, we talk about that with the team. We also experiment with LLMs. And we are also trying to figure out how we can use these new technologies to actually help our first persona.

24:45And that is the data scientist. So we are a technology provider to help data scientists and increasingly so the data scientists in enterprises, because we will be creating value-adding services and solutions so that we can generate revenue to sustain our mission. So the goal for us is to actually project ourselves while contributing to open source, but also create a sort of business value proposition, not dissimilar to Red Hat, because that is the closest type of company that we identify with in terms of spirit. To that point that you're making right there, I'd like to get back to something that you said earlier that feels like you're kind of tying back to it anyway there.

25:31And that's that you talked about the mission to create more open source and the mission that you're trying to create this environment that you're describing by design, you said. And that with Scikit-Learn here to stay for the long haul, it's going to be something that is not going away soon. It's solving such a high percentage of the problems. Could you describe a little bit about kind of what you're thinking around that in terms of further developing this particular set of software and the ecosystem around it so that we have the benefit for many years to come? How are you approaching that? So the company is built with multiple business units, if you wish.

26:12That's a big word for startup, right? But we have multiple revenue lines and multiple activities, even within the open source team, which is dedicated. So Guillaume perhaps can elaborate on some of the other libraries that we support that complement Scikit-Line. So that's one way to answer the question. But also we are building a new product, which I call reversible SaaS. So we are building a product that will provide additional value to data scientists. and the goal is to create a sort of, I don't want to use the term copilot because that is too close to LLMs, but it is the spirit. We are building a companion to augment the work of data scientists all the way to teams.

26:58So that is an additional product on top of PsychiCon because PsychiCon just works. And so we don't want to change that. And contrary to a company that would build a SaaS solution with a proprietary approach, we want to say, okay, whatever you guys use is fine. We need to find a way to add new value. And some of it will be open source, fairly modular. But for those companies that have more money than time, that need more service than be on their own, we'll have a solution for you and we'll make your life easier. And data scientists are a new breed. It's a new type of job. It's not been around for very long.

27:41And in a way, when I talk to people, so I've been in code forever, and you know this, right? The developers, when they get hired, they are turnkey in some ways, right? They have their Git environment, and they know how to peer code, and that's all pretty standard. But when you talk about data scientists, it's actually quite artisanal. It's an art and a science at the same time. And you're manipulating two objects, actual code, but data scientists are not coders. And you're manipulating actual data. It's not code. It's patterns. And so data scientists have a difficult task, which is to combine these two things and create value for the enterprise.

28:21And then they talk to business units and they're like, what do I do with this model? How do I put it in production, right? So there is a huge conundrum to solve. And that's what we're going to do, additionally to building open source that are modules that people can use. Maybe, Guillaume, you can elaborate on some of the other libraries that are key to actually help. Yeah, so within Probable, we have the open source team. And so we worked for many years on PsychicLine already. But we see the importance, as a community, we see the importance of putting models into productions and as well getting closer to the data sources.

28:57So we are just like working on libraries that should like make those come together. So for instance, we have a library that is more on the MLB side that is called SCOPS. We work a bit to make like the persisting more secure in some way. But we look as well on how to bring databases like SQL words into like closer to the machine learning models. so like how can you transform data with states with different tables and how you can be in your python words we were caring so much about like sql for instance and how you can bring this into psychic learn and within psychic learn as well we want to improve like whatever is visualization evaluations inspection of models which is on the top of just like training an algorithm because So we want to augment all those aspects beyond those.

Read the full transcript

29:52And either it's in Cyclearn or either this is a library connected to Cyclearn, let's say. So the one before is called Scrub, by the way. So it's like scrubbing data. So Scrub and Scope are two libraries that we look at. So as we're kind of talking about the libraries now, you have this robust open source contributor community built up around Cyclet Learn and the various projects within it. how does Probable work with those? How have you guys set up that relationship? What does the governance look like on that? Because you have both your core team that you alluded to earlier that's working at Probable on this, but you also have that larger open source community.

30:36How does that all work? Can you kind of tell us how that's evolved? I imagine it's quite mature by now. And that's the point. The maturity means that by design, we decided to not affect the license of scikit-learn. We're not branching it out. We're not care for it. And so the governance of scikit-learn being so sane already means you don't touch it. If it ain't broke, don't fix it. So the governance is unchanged. So the center of gravity was at INRIA, the research center, but also involving people all over the world. I don't know, Guillaume, how many... contributors, maybe 200? Or even more. I think in a year, you have more than that.

31:18You have maybe 300, 400. And the core team is, let's say, half of it might be around France, around Paris, around Probables. But then there's another half of 10-ish persons around the world that contributes almost every day, let's say, by communicating with the community. And as Ian mentioned, we didn't want to change that. nothing changed in that regard so the only thing that we actually did more we did more to bring transparency so to explain to people so now that we're improbable we feel that because we are a private entity we need to communicate what are we doing and on what are our roadmap and which community items are we going to work on just to like bring more trust such that i mean we don't go in the dark and that nobody knows now what we're doing.

32:12So we try really to pay attention to every six months to mention which of the items that are defined by the community, they are not defined by probable, but from the items, which one we have the capacity to work with the human resources that we have, let's say, at hand. So we really want to show that. And by design, the open source team that is full-time on Scikit-Learn and other open source libraries means it's a cost center to the company. So that cost center is by design, and we know that's a cost we have to cover. So we will cover it through different types of activities. So for instance, and this was something that was done in the past where brands were sponsors.

32:57So either they hire someone that becomes a core developer and they're naturally sponsoring someone to build up this technology, or they were giving money as a donation to the research center. But now that the team is with us, we are translating this into a contractual sponsorship framework. And so brands who want to contribute to Psychic Learn and help us compensate for salaries will get something in return, exposure. And if they actually put more money into it, then we'll have conversations around the roadmap. Find a way to make it converge in a win-win kind of way. Because Guillaume, for instance, can say, this brand wants us to do something, but it makes no sense for the community.

33:46Then we won't take their money for the sponsorship type of business. However, if companies want to pay us to do a certain type of paid for software, we'll look at it. But that's a different branch of the company. So we've really clearly separated. And by design, we know there's a cost to it. And that cost is actually, if we are doing well, it's compensated by the fact that we have done good by the brand. In other words, hopefully the community will actually resonate with what we're doing. And so they'll pay us back by actually appreciating what we're doing, which will carry the message further.

34:24So we think that there is a self-fulfilling prophecy if we actually keep adding value to the whole scheme as opposed to removing value. And I will not name certain projects that have chosen a different way. But on the other hand, going back to the governance of the company, when a company flips and becomes VC funded or only VC funded, VCs require a sort of return on investment that is too radical. And so that sort of forces a change of posture vis-a-vis the community and the licensing scheme. In our case, we've actually created a structure that is balanced in terms of shareholding groups. And so we will ultimately have, that's the goal of the structure, the architecture is to have as much money from public support than from private support.

35:14So it's again, sort of balanced.

35:37You know, when we started podcasting back in 2009, an online store was just the furthest thing from our minds. Now we have merch.changelog.com. And you can go there right now and order some t-shirts. And that's all powered by Shopify. What did we do before Shopify? I'll tell you, we did nothing. We couldn't sell. There were other ways, of course, but they were very hard, very difficult. Shopify let us build out an entire front end, obviously branded like ChangeLog is. It's amazing. Merch.ChangeLog.com. And our favorite feature is we use their API to generate a new coupon code, a personalized coupon code for every guest that comes on our podcast and they get a free t-shirt from our merch store.

36:20And that's so cool. They choose the shirt they want. They use the coupon code. It arrives free of charge to them and life is amazing. But also you can go there right now to merch.changelog.com and buy some threads yourself. And that's awesome as well. So upgrade your business and get the same checkout we use with Shopify. Sign up for your$1 per month trial period at shopify.com slash practical ai all lowercase go to shopify.com slash practical ai to upgrade you're selling today again shopify.com slash practical ai

37:19So as we come back out of break here, I'd like to, I want to turn to kind of a fun question for you. And I'd like, I'd like each of you to take a swing at it because it's not specific to being the CEO or being doing the technology itself. If each of you could describe kind of a a cool use case, something fun or interesting, or that's really captured your imagination with scikit-learn and kind of share that with the listeners in terms of something that, that just kind of really took you as your thing. I I'd love to hear, I'm expecting it to be a bit different coming from each of you and your different roles, but I'd love to hear kind of how, how you see that and what's the thing that sticks out in your mind.

38:02Jim, you start because I have to think about it now. so it's a very technical one let's say but so during my phd i was doing classification so which is something that i was trying to find people that has a specific type of concept so prostate cancer versus people that didn't have it and inside that space you had one very specific problem which is called imbalanced data and is what introduced me basically to scikit-learn because I had that prem I was using scikit-learn for the specific issues and how to tackle down those those type of issues and what is really funny is that so it's how I got introduced to scikit-learn and speak for instance with the developers and I developed one library which is called imbalance learn that is merging as well with scikit-learn like it's compatible in some ways and for many years I maintained that package even when I was maintaining as well scikit-learn and over years years after years we did everything by the book basically in that library we implemented the arguments that were inside the literatures and everything was fine until that as part of inria and now probable we we have as well time to educate ourselves and to try to as well then bring through the documentations of psychic learn to explain some concept of people and by doing this we find out that most of the research there didn't look at the problem properly.

39:26And by communicating with other core dev, we just found out that a huge part of this thing was just wrong and that you should look at it in another way. And then it's pretty funny because with this, we found some useless stuff that was, for instance, inside InBalance Learn. But then now we have better content. We went to conferences to explain these programs and people start to tell us, oh, yes, actually, that's right. And it's fun that you come and say that whatever you were doing like five years ago or 10 years ago is actually like obsolete or not good or, I mean, we wouldn't expect from there.

40:02And it's something that I find very like fun when you do open source because you can, like, you are just here to contribute to something and just to bring like the best of what you do to everyone. And everybody will be like thankful for that. even I mean and you are not defending your own like let's say scientific paper or like that's all what is true and and for me that's like one experience that come from my PhD from like now eight or nine years ago to up where I am now and then like I see like an evolution where I was with very good people and you could like correct errors that you do in the past and actually that will benefit everyone afterwards because that's landing inside the documentation of second or even inside the library.

40:47And then like everybody will just like, it says a million of users will be affected and say, oh, actually that's good. And this is something that's, I would have stayed in academia, for instance, three that wouldn't have happened because you wouldn't have like time or be critic enough because you would have been like in the books, like pursue books and go like this. So that's one anecdote, let's say. That's good. I might be the CEO, but I do have the imposter syndrome. Because scikit-learn is so impressive. It's day in, day out. I mean, that team and Guillaume is very humble and very discreet.

41:25But the amount of knowledge and the amount of technicality that is trapped inside this library is mind-blowing. And you haven't met the other members of the team. It's pretty much very, very hard to compete in terms of the amount of CPU cycles that go in there. So PsychitLearn is the gift that keeps on giving in some ways. And the team is just out of this world and nice. And it's just a pleasure to work with that team all the time. Now, the more I discover PsychitLearn and the more I find it amazing because of what the brand means to people. And so last week and today, actually, we just released, and if you allow us to actually put the link in the notes.

42:10Of course. Absolutely. We released the very first official Psychic Learn certification program. And what's amazing is that we, so this is the first time, so we're doing it step by step and the system works. People can register, they can pass or fail the test. But without advertising, we had like within a couple of days, 600 registrations all over the world. A lot from India, actually, because people in India, they do work also remotely for other clients across the world. So they do need a stamp of approval to showcase their ability to provide a service. So very interesting that this brand almost instantly can promote a sort of service that is value-adding.

42:55So that's the one thing. But then on the more technical level, I fell in love with one new feature that came out with 1.5 of Psychic Learn, developed by another co-founder and core developer, Jeremy. And that is the callback feature. Why? Because Psychic Learn, in fact, is a platform. It is a platform. And the callback feature allows us to provide extensions, if you wish, where people can hook into the inner workings of Psychic Learn as they are building new models. And in fact, I find that to be essential because we are entering an age of liability with regards to AI. Companies need to be able to introspect.

43:40They need to actually find out why the model is producing such and such results. And so introspection is critical. And as I said earlier, deep learning is sort of a black box type of approach, which I love, by the way. Again, in 1992, I was building deep learning models in the middle of the winter of AI at the time. But psychic learning is actually quite introspective, quite transparent, frugal, as I said. And so callbacks are yet another feature that provides actual introspection into how we build models. Because talk about insurance companies, fraud detections, you've got human beings at the end of the spectrum being handled by algorithms.

44:24And so that is critical. And I think we fulfill a very important need with these features. So again, PsychicLarine, the gift that keeps on giving, and I'm impressed every day with a bit of an imposter syndrome because that team is just so powerful with this tool. And speaking of this team, Guillaume, I'm going to throw a question at you. People out there have been listening to this. They're kind of going, okay, I want to dig into this. So you're going to get some new developers that are going to come. How should they engage? How should they find and get started in the projects to develop? What's a good onboarding path for those developers?

45:05Probably the best onboarding path is like if you have a chance that's inside your local community, there's some people that do what we call first-time contribution to open source or like coding sprints, go speak to those people because, I mean, they will help you to get onboard. But then if you are behind your computer and then you don't know where to start, it's where we have documentations that describe what do we call contribution. Because contribution is not only coding. It could be speaking, debugging, documenting, organizing sprints, and those type of things. And we have what we consider as contribution and how you can help, basically, and where you can help.

45:47So, of course, the natural thing is to come and code. And then we explain you how to start with that. So this is on the documentation channel, the documentation webpage. And afterwards, everything is online and public. So there's nothing private. So we have different channel of communications. The main one is GitHub, and it's going through the issue tracker or the pull request, depending on which side you are. And the core developer will be, I would say 24 hours over 24, because we are around the world. So that's like if I'm sitting there's somebody else in Australia or in the US that probably can just answer to you.

46:26And then we'll just give you feedback. And then is where your journey starts. You should not be shy and you should not be scared of making a mistake because we are not judgmental. We all started by that stage of saying, I don't know what I'm doing and I need to ask people, what should I do? And that's a normal step. And afterwards, you just grow with the community and then the community bring you over. I mean, but the most difficult thing is, yeah, is the first step, like engaging and saying like, so I'm the imposter syndrome as well. But that people say, I don't want, I mean, like this, like those very skilled people, they will never want to speak to me.

47:04And that's not the case. So just come and just try your best. And then people will just communicate with you for sure. Great guidance there. As we wind up, I'd like to get for each of y 'all, for both Probable and for Psych-It-Learn, kind of what you think about for the future. And, you know, and I'll let you define what time span the future is, you know, whether it's, you know, a few months or years out. But I'd really like to wind up, paint us a picture of when the duties of the day have finished and you're just relaxing and you're thinking about what's possible going forward. What do you think about?

47:43I'll go with the mission. The mission is bigger than me, bigger than us. And so that's why the governance creates a self-sustaining model. So, of course, it's not trivial. So there's a lot of work to achieve the mission long term, but that mission ends up with an IPO. In other words, this company is not meant to be sold or wrapped up. The goal is to do an IPO so that this company can carry on with the mission and allowing people to invest and be part of that story. And that's why earlier Daniel asked a question about investors and all that. So we do have 70 individual investors, including people who were contributors or are contributors to Psychic Learn, who don't have the chance yet to be employees full-time of the company.

48:33So the goal is to create this sort of dynamic vehicle. And if we look at the North Star, there is no such company today that is the provider of open source machine learning technology. That company does not exist. And we aim to be that because we need that in an age where there's too much concentration within just a handful of players. That's not okay. It's not okay for the global South. It's not okay for Europe, which is lagging behind. But it's not even okay for the US. The US may have big tech, but that's not okay as a single model. We need people to own their data science. That's why that is our tagline.

49:15That was good. Guillaume, what are your thoughts? Yeah, so maybe more on Probable. I'm really thinking that we have missions, let's say, to help more data scientists. But I will speak more about Psychitlone and the ecosystem. So for me, the mission is we should stay focused on what's happening out there and make sure that Psychitlone is still relevant. So we have the foundational model. That's fine. But we need as well to understand where this is deployed and how this is used, because we can make such progress that bring, for instance, make it easier to bring databases to Cyclone or to bring Cyclone models into productions and to reduce friction and everything.

50:02And as well bring values on understanding the model. I mean, we are speaking about AI acts as well in Europe now. So I'm sure there's plenty of, let's say, areas where we can have really an impact. And then there as well, technology that's moved very fast. So for instance, before we knew Pandas, now this is Polar. So we need to move like in a fraction of seconds saying, how do we like deliver value to the user that just makes a switch? And still can you scikit-learn? Like can we do like accept those things? And then so we have to make this audit of what's happening. So this is difficult to say where we will be in five years, because in five years we have all those things that can, let's say, We have the full chain of machine learning that probably will be here, so we should be aware, but we should be aware of whatever moves very fast around us to stay relevant.

50:53That was well said, too. Gentlemen, you guys have done a fantastic job of teaching the rest of us about this, and thank you very much for coming on the show today. You're welcome. Always a pleasure.

51:14All right, that is our show for this week. If you haven't checked out our ChangeLog newsletter, head to changelog.com slash news. There you'll find 29 reasons, yes, 29 reasons why you should subscribe. I'll tell you reason number 17, you might actually start looking forward to Mondays. Sounds like somebody's got a case of the Mondays. 28 more reasons are waiting for you at changelog.com slash news. Thanks again to our partners at fly.io to Breakmaster Cylinder for the beats and to you for listening. That is all for now, but we'll talk to you again next time.

From the publisher

We are at GenAI saturation, so let’s talk about scikit-learn, a long time favorite for data scientists building classifiers, time series analyzers, dimensionality reducers, and more! Scikit-learn is deployed across industry and driving a significant portion of the “AI” that is actually in production. :probabl is a new kind of company that is stewarding this project along with a variety of other open source projects. Yann Lechelle and Guillaume Lemaitre share some of the vision behind the company and talk about the future of scikit-learn!

Join the discussion

Changelog++ members save 9 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Timescale – Purpose-built performance for AI Build RAG, search, and AI agents on the cloud and with PostgreSQL and purpose-built extensions for AI: pgvector, pgvectorscale, and pgai. 
  • WorkOS – A platform that gives developers a set of building blocks for quickly adding enterprise-ready features to their application. Add Single Sign-On (Okta, Azure, Google, Microsoft OAuth), sync users from any SCIM directory, HRIS integration, audit trails (SIEM), free magic link sign-in. WorkOS is designed for developers and offers a single, elegant interface that abstracts dozens of enterprise integrations. Learn more and get started at WorkOS.com
  • Shopify – Sign up for a $1/month trial period at shopify.com/practicalai

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
scikit-learn & data science you ownPractical AI · 52 min
Listen in VO