Pausing to think about scikit-learn & OpenAI o1

17 Sep 2024 · 50 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Practical AI - Pausing to Think About Scikit-Learn & OpenAI O1

Episode Overview

  • Podcast Title: Practical AI
  • Episode Title: Pausing to think about scikit-learn & OpenAI O1
  • Hosts: Chris Benson & Daniel Whitenack
  • Focus: This episode discusses recent developments in the AI ecosystem, particularly focusing on scikit-learn's seed funding and OpenAI's new model, O1.

Key Discussion Points

  1. Scikit-Learn Seed Funding
  2. Announcement: The organization behind scikit-learn, Probabl, announced a seed funding round.
  3. Mission: Their slogan, "own your data science," emphasizes the importance of controlling one's own data and analyses.
  4. Future Plans: Probabl aims to support the data science community while maintaining an open-source ethos.
  5. Launching an official scikit-learn certification program in Q4 2024.
  6. Focusing on the data scientist's role in the pre-ML ops phase, such as model selection and data munging.
  1. OpenAI O1 Model
  2. Overview: OpenAI's new model, termed O1, focuses on advanced logical reasoning and introduces a "thinking" feature that pauses before generating responses.
  3. Comparison to Previous Models:
  4. O1 is slower than version 4.0, which affects user interaction.
  5. It utilizes chain-of-thought processing, generating in-depth reasoning before providing an answer.
  6. Application Scope:
  7. Best suited for complex tasks, particularly in coding and mathematical reasoning.
  8. Emphasizes accuracy in reasoning tasks over general conversational capabilities.
  1. Implications of the New Model
  2. User Experience: The latency in response time may lead users to adapt their expectations and utilize O1 for different tasks than those suited for faster models.
  3. Limitations:
  4. Knowledge cutoff date is set to October 2023.
  5. Lacks the ability to browse the internet and does not support file uploads, which limits its applicability in real-time scenarios.
  6. Marketing Perspective: Concerns were raised about how OpenAI presents O1's features, suggesting it may oversell the novelty of its "thinking" capabilities.

Key Takeaways

  • Open Source vs. Proprietary Models:
  • The hosts advocate for open-source models like scikit-learn, which provide more control and transparency compared to proprietary systems from companies like OpenAI.
  • There's a recognition of the essential role traditional data science practices will continue to play even as generative AI becomes more prominent.
  • Community Engagement:
  • The discussion encourages participation in communities (like MLOps and Slack groups) for sharing knowledge and resources.
  • Learning Resources:
  • DataCamp and Codecademy offer courses on Scikit-Learn.
  • Purdue University is hosting a "Data for Good" competition, encouraging students to engage in meaningful projects.

Conclusion The episode highlights the evolving landscape of AI technologies, focusing on the balance between proprietary advancements and the enduring relevance of open-source tools. It encourages listeners to explore both realms and engage with the broader AI community for continual learning and development.

Links and References

  • [Probabl Seed Funding Announcement](https://papers.probabl.ai/announcing-major-milestone-empowering-the-future-of-data-science)
  • [OpenAI O1 Announcement](https://openai.com/index/introducing-openai-o1-preview/)
  • [Purdue Data4Good Competition](https://business.purdue.edu/events/data4good/)
  • [MLOps Community Homepage](https://home.mlops.community/)
  • [Latent Space Homepage](https://latent.space/)

Sponsors

  • Assembly AI: Leading Speech AI models for voice data.
  • Fly.io: Application hosting platform with unique networking and scaling capabilities.
  • Speakeasy: SDK generation platform for API developers.

---

Feel free to explore these resources to deepen your understanding of the topics discussed!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:15Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Learn more at fly.io.

0:50what's up friends i'm here with a new friend of ours over at assembly ai founder and ceo dylan fox assembly ai is where you can turn voice data into insights chapters transcripts summaries and so much more with their leading speech ai models so dylan give me a glimpse into what you're doing with speech ai models at assembly ai so at assembly we're building industry leading speech ai models for various tasks like speech-to-text, streaming speech-to-text, speech understanding to help developers easily convert voice data, whether it's live or pre-reported, into super accurate text. And then to help developers extract a ton of information and metadata around voice data or even around the text that they just were able to convert from that audio data.

1:37So these are things like picking out entities or PII that was spoken in voice files or summarizing voice and audio data down into custom summaries. It's things like being able to detect how many speakers spoke and who said what and what the names of different speakers were. So we bundle all those things into a super simple API with really great docs that developers can just sign up to for free to start, use the API, build into their apps, and then build these really cool AI apps and products and workflows and automations on top of voice data with. I dig it. Okay. Can you take me a little deeper into the opportunity for developers?

2:16Because it seems like there's a lot of voice data out there and there's a lot of trapped value in that voice data. There's so much voice data being created on the internet now. Podcasts, videos, phone calls, voice messages, audio books, virtual meetings. It's crazy. And you can now transform and understand all this voice and audio data in ways that were not even possible a year, 18 months ago. So what we're seeing with the help of these new AI models that we're creating at Assembly, developers and organizations are just racing to build all these new applications, workflows, automations that leverage the voice data they have either within their organization or within their product to build really cool new products and services workflows that are just like taking off in the market.

3:03So at Assembly, we're building the industry leading models for all those different apps and workflows, whether it's speech to text or speaker diarization or speech understanding capabilities to summarize voice data or extract entities from voice data or mask PII from phone calls for various types of automations that might be built. And we're exposing that through a super simple, super scalable API that's just constantly being updated and constantly getting better. And so we're seeing a crazy amount of developers and companies just build really cool apps and services on top of our API every day.

3:36It's really only just getting started, especially with the model updates that we have planned over the second half of the year that are coming out. They're really excited to launch to the developers on our API. Okay. Constantly updated speech AI models at your fingertips. Well, at your API fingertips, that is. A good next step is to go to their playground. You can test out their models for free right there in the browser. Or you can get started with a$50 credit at assemblyai.com slash practical AI. Again, that's assemblyai.com slash practical AI.

4:20Welcome to another fully connected episode of the Practical AI podcast. This is Daniel Whitenack. I am the CEO and founder at Prediction Guard, and I'm joined by my co-host Chris Benson, who is a principal AI research engineer at Lockheed Martin. In these fully connected episodes, we try to keep you fully connected with everything that's happening in the machine learning and data science and AI world, and hopefully share some things with you that'll help you level up your machine learning and AI game. How are you doing, Chris? It'll be fun to catch up on a few things that have been happening over the past couple weeks today.

5:02There is so much going on. Oh my gosh. Always. We'll have to pick and choose what we have time to hit here. How are you hearing about AI things these days, Chris? Maybe that's even something that people might be interested in knowing. maybe we've talked about this a few times on the show, but there may be new listeners who kind of got into the show after Gen AI stuff and they're trying to figure out where to keep up with news and learn things. The question's the inverse. It's like, where are you not hearing about it? Because we're getting it from every angle. Or where do you, how do you filter through the noise?

5:39That's the issue there is there's so much noise now. There's so many, like, you know, I, uh, you I know both of us have our kind of workflow on how we're consuming new things going on out there and have for many years as we've been doing the show. And I think there's so much more that is coming in. And the quality varies hugely in terms of how you might do it. I know we're going to talk about a few topics today. And I know that on at least one of those to be seen in a few minutes, some of the info is kind of straight down the line, gives you the facts. and others are very hypey. We've talked a lot about how hypey things are.

6:19But I think that's one of the challenges with this field. It's moving so fast and it's just one of the dominant topics in mainstream media now. And you have a lot of folks out there reporting on it, some of which know something about it, some of which don't know as much. So yeah, the filtering is the big challenge. It's been slightly different for me, I think, in the sense that, well, one, I founded a company So my attentions are slightly maybe distracted with one thing or another, or maybe exposed to things in different ways. But also, I think my habits online have changed even over the past year or so, whereas I was maybe more at least not always posting, but looking kind of for things on Twitter or X as it is now and seeing, you know, a pretty active AI community there, of course, where I had, well, you as well.

7:22You know, we both kind of had different journeys into this, but definitely I had this sort of more background on the data science thing and a lot of kind of data science discussion was at least in my world happening on Twitter or X. then kind of all the whole AI world took over and also, you know, Twitter changed to X and things are still happening there to one degree or another. But I sort of started ignoring that a little bit. And so at least one of the things that we'll talk about today, I kind of looked at and learned through LinkedIn, which to be honest, if I'm completely honest, I always, you know hated kind of scrolling through linkedin um throughout the years well it's very corporate yes but yeah maybe it's just the the people that i've learned to follow or look for things from or yeah i i don't know it's it's hard to find the proper channel through which you're getting a good signal to noise ratio but i i found a little bit better there recently i don't know if you've seen a similar thing no i think you're very right about that is linkedin's gotten better about producing good information in that way.

8:39And it used to be once upon a time, I had my workflow sources and then you'd see it kind of showing up on LinkedIn in the hours and days afterwards. And I also would use Twitter, but I also have moved away from X as well. It's just not giving me what I'm looking for most of the time. And so a bunch of new sources and aggregators put together with some filtering on that. But yeah, LinkedIn is definitely on the upswing versus where it was maybe two years ago. Yeah. And I don't know, I guess I'm lucky enough to be part of a few Slack channels and discords now where, you know, either my coworkers or, you know, other collaborators or different discord channels whether that be you know latent space or the ml ops community oh yes like these things pop up things as well and people find them around and so yeah i would definitely encourage people to of course to think about you know we we have a slack channel with the the podcast you can find at practicalai.com slash community but there's also really great ones out there with latent space ml ops community there's of course ones uh you know collaborate with your co-workers figure out where they're finding good info but yeah it is i think a little bit more fragmented now which makes it hard to to pick apart some of those news stories now i'm gonna say people need to join our slack community because um we've had some really good conversations and questions and suggestions arise there.

10:18And the folks who are participating really know their stuff. And they've pointed me to a few things in recent months that have been really useful to have. So I'm spending more and more time taking pointers from that. Yeah. Awesome. Yeah. Well, the one of the first things that I wanted to mention on on the show here was something that did you know i did run across on linkedin which i i just had missed till it was posted recently as related to their recent seed round of of funding but there's uh you know this company probable which had an announcement of a round of seed funding right yeah i guess they would classify to seed funding this sort of as i've learned this phases of funding are somewhat strange and all you know have different definitions depending on where you're at but yeah this is a company probable which is i'll use their exact words because i you know um they've they've chosen them well it's the official operator of the scikit learn brand And they talk about this funding representing a powerful step forward in our mission to help professionals truly adopt our slogan, own your data science.

11:39So a few interesting things there. I am really happy. So we, of course, people are probably saying, well, why aren't you talking about OpenAI 01 as the first thing you're talking about? We'll talk about it later. Don't worry. But this one I thought was really interesting. in that, you know, hey, I'm excited to talk about something maybe that intersects something other than OpenAI and the next-gen AI model. But Scikit-Learn, of course, has a huge place in my heart throughout the years of operating in data science. So to see a brand or a company that's really putting an effort behind that brand of Scikit-Learn, advancing that and also advancing a slogan forward to the future of, you know, own your data science.

12:32One, thinking about kind of owning as in tools that you can operate internally and privately and with your own data, but also tools that are around data science. So, you know, maybe there's a future for data science after all. I think there is. And I love, you know, the reason we put this first is because as huge supporters of open source and people being able to kind of control their own destiny with data science and machine learning and AI technologies, I love having these companies that are out there supporting open source and it gives us options. And so if you just want to do the open source yourself and maybe you're just a data scientist working on your own side project or something, you can do it.

13:22Or if you're a corporation and you're looking to have dependable partners to work with around open source, you've got that too. So it makes it, in its own funny way, it makes it more accessible to a wider audience and gives us those choices. And I just love it when companies are doing that. You know, just as a two second aside, it's one of the things that surprised me, you know, with Facebook and their models is they're open sourcing it, you know, which the other big companies we've historically talked about aren't. But back to probable, what are some of the things you see in this in this announcement that really gets you going, Daniel?

14:01Yeah, well, I think, you know, for people that aren't aware, and maybe just coming into kind of the AI world with, you know, ChatGPT and all of that, Scikit-learn has been around for a long time and is kind of a primary open source set of tooling for the Python community and the data science community that allows you to build a wide variety of models. So everything from neural networks to decision trees to random forest models to support vector machines and clustering algorithms, just a huge number of sort of algorithms and metrics and evaluations and models all within this very kind of consistent API, well-supported API, widely used library that is scikit-learn.

14:56And so I think I'm excited to see that one could think, well, maybe everyone that was interested in supporting and contributing to those things has kind of jump ship to what's fancy and shiny with LLMs and all this stuff. But I think I'm encouraged that there is a strong backing behind this. And I think it's needed moving forward. We've talked about this on the show where there is going to be a need for kind of hybrid systems between traditional machine learning and statistical learning and generative AI. and there will still be a need for kind of smaller performant models in a variety of contexts.

15:38And maybe those are better, faster, more secure for a whole variety of things as they've been useful for decades now. So I think that I'm encouraged for what's there. Just so people know, and we'll link to this announcement, but they talk about what's next for Probable. So one of the things they talk about is an acquihire, securing talent, but also the launch of an official scikit-learn certification program coming in Q4 2024. And then the release of a product. So obviously this is a commercial company, so they will have something productized. and they say that they're aiming at augmenting the work of data scientists in the pre-ML ops phase.

16:29So this is very interesting to me. I think it's an interesting niche that they're focusing on there, not trying to cover or recreate ML ops tooling that's already out there, but focus on the data scientist role in that kind of pre-ML ops period, which in my mind involves, of course, data munging and model selection, feature development, all of these sorts of things. You know, I think that, and it's funny that you bring that up because I think that gets lost, you know, with all the hype on the AI and particularly Gen AI these days, you know, there's still so much data science going on. And I would suggest that just core everyday data data science is still by far bigger in terms of being present in the number of organizations.

17:20It may not be a big fancy glitzy effort, but it's pretty core to most organizations. And yet, the AI hype tends to get all the press. So it's really good to see them kind of acknowledging that that section is still there and that it does need support and being willing to do that a little boldly and stepping a little bit away from where everybody else is going to. So yeah, I hope they do really well in that capacity.

18:01Okay, friends, I'm here in the breaks with Annie Sexton over at Fly. Annie, you know we use Fly here at ChangeLaw. We love Fly. It is such an awesome platform and we love building on it. But for those who don't know much about Fly, what's special about building on Fly? Fly gives you a lot of flexibility, like a lot of flexibility on multiple fronts. And on top of that, you get so I've talked a lot about the networking and that's obviously one thing. But there's various data stores that we partner with that are really easy to use. Actually, one of my favorite partners is Tigress. I can't say enough good things about them when it comes to object storage.

18:42I've never in my life thought I would have so many opinions about object storage, but I do now. Tigris is a partner of Fly, and it's S3 compatible object storage that basically seems like it's a CDN, but is not. It's basically object storage that's globally distributed without needing to actually set up a CDN at all. It's like automatically distributed around the world. And it's also incredibly easy to use and set up. Like creating a bucket is literally one command. So it's partners like that that I think are this sort of extra icing on top of Fly that really makes it sort of the platform that has everything that you need.

19:17So we use Tigris here at Changelog. Are they built on top of Fly? Is this one of those examples of being able to build on Fly? Yeah, so Tigris is built on top of Fly's infrastructure and that's what allows it to be globally distributed. did. I do have a video on this, but basically the way it works is whenever, like, let's say a user uploads an asset to a particular bucket. Well, that gets uploaded directly to the region closest to the user. Whereas with a CDN, there's sort of like a centralized place where assets need to get copied to. And then eventually they get sort of trickled out to all of the different global locations.

19:51Whereas with Tigris, the moment you upload something, it's available in that region instantly. And then it's eventually cached in all the other regions as well as it's requested. In fact, with Tigris, you don't even have to select which regions things are stored in. You just get these regions for free. And then on top of that, it is so much easier to work with. I feel like the way they manage permissions, the way they handle bucket creation, making things public or private is just so much simpler than other solutions. And the good news is that you don't actually need to change your code if you're already using S3.

20:24It's S3 compatible. So like whatever SDK you're using is probably just fine. it all you got to do is update the credentials so it's super easy very cool thanks annie so fly has everything you need over three million applications including ours here at changelog multiple applications have launched on fly boosted by global anycast load balancing zero configuration private networking hardware isolation instant wire guard vpn connections push button deployments that scale to thousands of instances, it's all there for you right now. To pull your app in five minutes, go to fly.io. Again, fly.io.

21:29Well, Chris, one of the things, you know, we started out talking about Probable and this funding that they have, of course, this is related to Scikit-Learn and owning the data science process with Scikit-Learn. They're obviously very committed to the open source way of going about things. They have a long-term vision for people still to be able to maintain their own data science process with their own data, with models that they own internally. that's of course in in contrast to some things out there which um you know there there's a mix of this i think there is a validity to people trying to create their own proprietary models and that's their way of owning certain things but certainly the models that people would think of right now with ai models or maybe those from open ai and you know we got a another one of those over the past, whenever it was, I don't know, all the days blurred together for me.

22:36But this rumored, what was it, sort of strawberry codename slash O1 model, which people have been talking about. And certainly we want to cover that on this news show. So yeah, O1, Chris, how has O1 struck you in its first days on the AI scene? Yep, it's an interesting animal. And, you know, it is entirely proprietary. And on this show, as we've noted, we usually give those the second spot, not the first spot. But it's interesting. I've been using it some over the past few days. some of the new features that it talks about are advanced logical reasoning, and it has a capability where it slows things down in terms of processing.

23:33And so that can be, it's not the same experience as the 4.0 that we've been used to, where we've been, you know, I know on my iPhone, on 4.0, when I'm using that, I'm talking at it these days, and it talks back, and it's fairly conversational due to the speed. You can't really do that with this O1 preview as it's currently released. It's a little bit too slow for that. When I've tried, I've had to wait a while for my responses, but it's taking a different approach. Yeah, I guess it brings up an interesting question. So for people that are not familiar with the O1 model, it's a model that really operates on this principle of thinking through a series of stages of reasoning before giving a kind of final answer.

24:21Some of what this might have been called in terms of how a completion would happen or how you would structure a prompt or a completion in the past might have been chain of thought processing. And so that takes a bit more generation, more text is generated, you know, there's a pause in the result. And this brings up a question, you know, you were talking about this latency element, Chris, and here they've intentionally slowed things down. And I was wondering for a while, you know, with Grok and very quick, you know, completions out of these models, does what role latency would really have, you know, at a certain point, as text is generated, our human minds can still only process, you know, a certain amount of natural text at a certain speed.

25:07And then here, OpenAI has intentionally slowed down the generation process, which I don't know if you'd call it, maybe it's still generating at the same or a similar token per second rate, but it's generating more, right? It's generating this kind of chain of reasoning or chain of thought kind of steps and then producing something on the output, which I'm assuming is just like a generation in. And there's then some special token that they have in their prompts that they train the model on, which is like now give the answer to the user, right? Whatever that special token is, which, you know, we don't have a ton of information about.

25:48But yeah, I don't know. What are your thoughts on this kind of latency versus reasoning versus also the human interaction element of this? Before I dive into that, I want to note that having read quite a bit about the model over the last week or so, a lot is a bit speculative in terms of, you know, you'll read articles where people are saying clearly it's doing X. Yeah. But they don't really know because OpenAI has not specified, you know, exactly how they're doing it. So I want to note. Yeah, so anything we say here may or may not be accurate. This is our best chain of thought on the topic. And I think it's important to say that.

26:29We're not speaking in perfect factual. We have no direct line to the OpenAI technical team and revealing their secrets. Nor do many of the authors that we've been talking about. So I guess my impression has been, it's a little bit startling. you know, the latency thing kind of kicks in after you've been using the 4.0 model for a while. And it forces you to start realizing that there are definitely different use cases for using the 4.0 and the 4.01 preview. And I think that that is notable because it's really the first time that OpenAI has offered a new model that everyone didn't just instantly switch to that as the thing to use.

27:13It was kind of like they said when we went from 4 to 4.0, well, there may be cases where you go to either one, but in practice, I saw people just going to 4.0 pretty much nonstop at that point. Whereas this one, the latency kind of forces you to change your ways. And they also have given guidance on the prompt engineering that the way you prompt the O1 preview is not the same anymore because what they're doing behind the scenes and, you know, with the notion of possibly multiple concurrent addressing of your prompt in different ways and they verify them against each other, all things that I've read that are unverified at this point in time, that because they've changed the way they're processing on the back and the latency is now there, that there might be different types of things.

28:03They tend to highlight coding. They tend to highlight math and other critical reasoning skills where you'd want to go to the 01 preview rather than back to the 40. And they're claiming that it's more accurate. I've seen numbers like 15 % increase in complex reasoning tasks to support that. And there's some drawbacks, which we can talk about in a moment. But I think I am still trying to figure out the tasks that I'm going to assign to each model in my own mind. And I'll find myself stumbling a little bit on that. I probably at this point over the next week, trying this stuff out when I'm coding, I don't have a reason to do a lot of math in my day to day stuff just on a constant basis.

28:47But for coding, I'll probably be spending more time on the O1 preview than I, whereas I was using the O itself before that. So it'll be interesting to see. And so with that difference, as you look back to the two between the 01 preview and just the 4.0 we've been using, I just want to throw out into the mix at some point in the not too distant future, we've been led to believe that GPT-5, which is a much larger model, according to OpenAI, will be coming out. And so it kind of leaves you wondering a little bit, you know, is that which model that we're on now, is that going to be closer to? Are there two different sets of kind of prompt engineering approaches that we have going forward at this point there's a lot of unanswered questions yeah and it was you know kind of curious to me and maybe slightly revealing i don't know um it's hard to read into everything that's going on behind the open ai curtain but the fact that the sort of big advance here was apparently some sort of, you know, RLHF preference tuning around this sort of chain of thought generation versus kind of, you know, you could think of any number of things that could change.

30:10You know, we saw this in the past. We saw a wave of mixture of experts models, right? Which, you know, that was a big change in the model. We saw, you know, of course, the model size. We saw changes at a certain point in terms of how models were trained and aligned. But here, this is just sort of a different prompt set that is used in this RLHF process. And I think it's intriguing and interesting that they're applying it a different way in the UI and the chat GPT UI. You might want to define the acronym as you're going there. just uh yeah so like this is all inference on my own with my own chain of reasoning but when they say sort of oh we're quote pausing the model in the generation like we're allowing it to think and then all i think that's somewhat confusing because obviously the model is not you know well there's a little marketing thrown in there and this is my own personal opinion but the model is not thinking, it's just generating text, right?

31:18And the difference is that it's generating text that is more explanatory or exploratory, representing a series of decisions that can be made leading to accomplish the goal or the task that you prompted it to accomplish. So So when it's pausing like that, what I'm assuming is that they have a pre-trained model. Maybe it's the GPT for something internally. I don't know what they call it internally. The parent model, right? And then they've curated a set of prompts that in the training set, complete with the complete chain of thought, whatever, whether they synthesize that data or use humans to created or, you know, whatever, probably some combination of that, right?

32:09And then they fine tune or preference tune or align the pre-trained model to that prompt set that they've curated, which includes all of this chain of thought stuff using this process. I mean, they refer to RLHF, reinforcement learning from human feedback. So this sort of reinforcement learning driven loop to align the model and then you get the sort of 01 model that is the process i'm assuming happen on the back end and when you're in chat gpt you know it's not like there's some button like you know they have it thinking thinking they say that really it's just i'm assuming it's generating text and then at a certain point similar to like people ask like these models just generate text how do they know when to stop well they don't know they generate a special token which is an end of statement token and then the program the computer program stops it from generating more text after that token is is generated in the same way here i'm assuming after a certain level of text is generated in the chain of thought it generates a special token that's what i was referring to before of like now show the answer so you know now generate the answer token.

33:28And then that's how they control the UI. So all of that is inference on my part. Again, I could totally be wrong, but that really is a similar process to what we've been doing now for a couple of years with these models. There's not anything fundamentally different about that process. And so I found it interesting that the big reveal here was sort of a different model created with RLHF, which is basically what everybody else is doing. Maybe they're using this interesting methodology in the UI. So I don't know what that reveals. It could mean this is a cool thing to hold us over until GPT-5, which will be this fundamentally world-changing different process, methodology, architecture, model that's going to be crazy.

Read the full transcript

34:19Or it could just mean there's very much a diminishing returns here in terms of the methodologies that are available to improve this wave of models. I think that that's as good of an educated guess, not privy to their internals as I've heard. So I think if you haven't nailed it completely, you've probably nailed parts of it. This model being out there in the consumer market, so on my iPhone, I use it a fair amount. I find it interesting that they're trying to create a user experience that's a little bit mystic. You know, as you pointed out, it says thinking while the delay is going on, which kind of there's an implication there, especially if you're not like us and in the industry where we're talking about AI every day and that's what we do.

35:09If you're somebody out there who's just kind of just your typical average person consuming the technology, there's an implication there. And especially when you talk about they're using the word advanced logical reasoning, things like that, that I think it's a little bit of kind of marketing hocus pocus that's being applied to a fairly mundane set of processes, as you pointed out, using the technologies that have been. They may be constructing how their models are interacting, as you pointed out, in a way that's unique on their side, but it's probably not revolutionary. It's an evolutionary decision that they've done to try to make their model more accurate.

35:50I find that a little bit questionable in terms of kind of how they're marketing that to the general public. A little bit worrisome. There's some of the voices out there that have been out, you know, talking about this coming out or expressing concern. Again, the lack of visibility into what's actually happening makes it really hard to verify or not, you know, how they're approaching it.

36:32Well, our friends over at Speakeasy have the complete platform for API developer experience. They can generate SDKs, Terraform providers, API testing, docs, and more. And they just released a new version of their Python SDK generation that's optimized for anyone building an AI API. Every Python SDK comes with Pydantic models for requests and response objects and HTTPX client for async and synchronous method calls and support for server-sent events as well. Speakeasy is everything you need to give your Python users an amazing experience integrating with your API. Learn more at speakeasy.com slash Python.

37:17Again, speakeasy.com slash python.

37:29Well, Chris, just to give people a sense of this one thing, and then I think I could get maybe a last impression from your end, and then maybe a summary, but I was trying one. So I think the idea at least, or part of the idea with the way that they're providing access to this model and promoting its usage is for more kind of researchy, kind of deep reasoning type of things. So the simple prompt that I gave was determine a new problem in physics related to density functional theory for my PhD focus, write a concise summary for me. So that was my prompt. The only reason I did that prompt is because that was the subject of my PhD research.

38:21You know something about it there. Hey, you know, this would have maybe been nice back in the day. And so the UI experience, for those that have not tried it yet, and you're just listening, it kind of paused, said thinking for about 10 seconds. Then it gave the executive summary or the summary that I asked for. But there's a little kind of drop down that I can click. and it said thought for 10 seconds and the steps that it says that it kind of talked through or formulating a research problem breaking down dft applications considering quantum embeddings changing gears and pioneering new dft modeling and it gives a little summary of those and seems to be you know more text generated there and and that sort of thing so the problem statement is um relevant.

39:19The text is relevant to some of the things that I would know about in that case, but also pretty much what I'm aware of from what people have already been exploring in the past. It's not like all of a sudden this model knew, like it shocked me with a, wow, that would be a really interesting and profitable area of research in this topic that I'm aware of. But it was generally in an area that's interesting. So I guess, you know, kind of good-ish, but not mind-blowing. And so, yeah, that kind of brings up part of what I'm wondering here around, where is the proper place for this in my day-to-day workflow?

40:03Maybe we just haven't figured that out yet, because I can definitely see a lot of places where I'd be like, well, I don't want to pay up for this, because it is, you know, going to be way more expensive than like the 4.0 or 4.0 Mini, right? I could see a lot of places I could apply for 4.0 Mini or any number of LLMs that would operate in a similar way. Where am I going to apply this in kind of in the software that I'm building? I don't know. I haven't really pinpointed that yet. You know, and I agree with you. I want to go back to something you mentioned, though, as you were talking your way through the example.

40:40and that's that it gives you these kind of intermediary steps that it says it's following along the way and from my standpoint i think that's more once again on the marketing side tried to reinforce that that reasoning marketing message that they're that they're driving on that you know as we talk about where things what's an appropriate set of tasks to use with each of these models and especially with this 01 preview that's out right now it warrants knowing that there are a few drawbacks to the model, which we have not called out yet. One of them is that it still has a cutoff date on knowledge, which is tied back to October of 2023.

41:23So, you know, it's almost a year since it has access to that knowledge. And unlike some of the other models before it, which had internet access to go out and kind of make up for that and get some more current information to throw into the model that was trained up until the cutoff date. This one does not have the ability to browse the internet. So that alone may change kind of how you use it because so if you were, you know, going back to what I was suggesting, if I'm coding, well, it may be that if I'm worrying about whether it's Python code or Rust code or Go code, not a huge amount has changed in the coding world in that time in terms of what's available library-wise unless it's just the latest, greatest thing to come out.

42:07And so I can probably use it really beneficially in a coding context there. But as we record this, for instance, on this model, just to talk about current events, a potential second, it sounds like a second assassination attempt on Donald Trump occurred today in the news. And if I wanted to ask a model to get information about that, this would not be the model. And I'm just pulling that one out simply because it's a big news event that happened on the day that we're recording. And so in that case, if I'm curious to explore the story or whatever using a set of models, I might have to go to some of the other models that have the internet access and might be able to frame things that I'd be curious about for my consumption and my knowledge, whereas that current event would not be applicable here.

42:53And then finally, I wanted to point out that we're kind of used to being able to upload into the 4.0 series, the external documents from a RAG perspective, retrieval augmented generation, that's not available in the preview here. So that's yet another limitation. And when you add up the cutoff date, you add the lack of file upload and you add the lack of internet access, those are some fairly substantial limitations on this model at the current time. So there truly are certainly a set of use cases that you might go to different models to see, you know, current events, it's not going to be this one.

43:33Yeah. To kind of wrap up this section here, Chris, on 01, bringing it back to scikit-learn and the probable funding, I found it interesting that probable on their website has a bit of a manifesto as far as their values. And maybe I can share these and I'll leave it to you and the audience to see how these compare in this sort of scikit-learn ecosystem, open data science, own your own data science in the open AI world of AI. So their values, they talk about supporting the whole long-term ecosystem than individual stakeholder gain. Openness rather than proprietary lock-in. I think that's definitely an interesting one, especially, you know, I was talking to someone the other day even about this idea of model lock-in in the AI world in this proprietary sense.

44:39The other ones are interoperability rather than fragmentation, cross-platform rather than platform-specific, collaboration rather than competition, accessibility rather than elitism, and transparency rather than stealth. So if you're interested in any of those types of values, definitely check out the data science community around Scikit-Learn and other projects and give a shout out to what they're doing over there. So I think it's an important balance and cool effort to highlight in light of all the other crazy things happening in our AI world. I think those are fantastic values just in general.

45:29And they're very consistent with other open source and kind of commercial support of open source that we've seen in other companies that we've liked, whether they're in the AI space or the software development space. And I find myself certainly feeling very comfortable and gravitating, but I would actually conclude by saying, in my own experience working at some large companies, large corporations, and having a lot of friends and colleagues that I talk to at other big companies, they may use proprietary models to some degree, but nobody's betting their business. At least in the conversations I have, they're not betting their business on an entirely opaque proprietary approach.

46:13So you see it there in use cases, but not for the big things. For the big things they're looking at, at open source models that they can rely on and such. And I just wanted to call that out because that's been really notable to me over the last year or so is how strong that sentiment is. So, and I would say more power to it. Yeah, and we do like sharing some learning resources and experiences here on the show, and in particular as related to Scikit-Learn and in that community. If you're interested in those things, you can take a look. There's some learning resources around DataCamp. So DataCamp has a supervised learning with Scikit-Learn course, which I think you can try out for free.

47:03I think Codecademy as well has a course there. There's probably innumerable blog posts around the internet in terms of scikit-learn and what you can do with it. So maybe if you're coming from the Gen AI world and have tried a bunch of things with OpenAI, you could also dip your toes a little bit into the data science world. we're happy to welcome you in and try a few things with scikit-learn they of course have great documentation and examples on their on their site as well i'll also mention if you're hearing this podcast right after it airs i think there's still time for you but uh purdue university which i'm i'm close here but they're working with actually a bunch of partners including commercial partners, but they're running a data for good competition, which is close to Chris and I's heart.

47:58So they're supporting the TAPS organization, which provides grief counseling and care for those that have lost family members that served in the military. And so if you're interested in that, check it out. That's for graduate students and undergraduate students. And there's a huge, I think,$45 ,000 in cash prizes, but also a bunch of free training and AI that you'll get along with the effort. So I really encourage you, if you're a student, that's a great thing that's happening this fall that you could be a part of. So check it out. If you just search for Data for Good Purdue, the link will be there.

48:39And I think you could register. If you hear this right after it airs, go check it out and sign up right away. fantastic awesome chris well good to talk things through and um enjoy the rest of your week uh until we hear about o2 or gpt5 we'll we'll see you next time there's always another one coming see you next time thank you for listening to practical ai you know what's cool free stickers during the month of september we're mailing out changelog sticker packs to everyone who leaves us a thoughtful five-star review or blog post about our pods. Simply email proof of your review to stickers at changelog.com alongside your address and we'll mail out the goods anywhere in the world.

49:31Once again, that's stickers at changelog.com. Picks or it didn't happen only in the month of September. Let's do this. Thanks again to our partners at fly.io, to our Beat Freak in residence, the one and only Breakmaster Cylinder, and to our longtime sponsors at Sentry. Use code CHANGELOG when signing up for a new Sentry team plan and save$100. That's all for now. We'll talk to you again next time.

From the publisher

Recently the company stewarding the open source library scikit-learn announced their seed funding. Also, OpenAI released “o1” with new behavior in which it pauses to “think” about complex tasks. Chris and Daniel take some time to do their own thinking about o1 and the contrast to the scikit-learn ecosystem, which has the goal to promote “data science that you own.”

Join the discussion

Changelog++ members save 9 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Assembly AI – Turn voice data into summaries with AssemblyAI’s leading Speech AI models. Built by AI experts, their Speech AI models include accurate speech-to-text for voice data (such as calls, virtual meetings, and podcasts), speaker detection, sentiment analysis, chapter detection, PII redaction, and more. 
  • Fly.io – The home of Changelog.com — Deploy your apps close to your users — global Anycast load-balancing, zero-configuration private networking, hardware isolation, and instant WireGuard VPN connections. Push-button deployments that scale to thousands of instances. Check out the speedrun to get started in minutes. 
  • Speakeasy – Production-ready, enterprise-resilient, best-in-class SDKs crafted in minutes. Speakeasy takes care of the entire SDK workflow to save you significant time, delivering SDKs to your customers in minutes with just a few clicks! Create your first SDK for free!

Featuring:

Show Notes:

Join the Practical AI community — it’s free! Connect with us in the #practicalai Slack channel. The community is here for you as a place to below, to bounce ideas around, or to get feedback on a conference to attend or talk you’d like to give.

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Pausing to think about scikit-learn & OpenAI o1Practical AI · 50 min
Listen in VO