#240 Manos Koukoumidis: Why The Future of AI is Open-Source

4 Mar 2025 · 1 h 6 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Eye On A.I. - Episode #240

Host: Craig S. Smith Guest: Manos Koukoumidis, CEO of OUMI Episode Title: Why The Future of AI is Open-Source Date: [Insert Date] Sponsor: Sonar

Overview In this episode, Craig S. Smith converses with Manos Koukoumidis about the importance of open-source AI for the future of technology and innovation. Koukoumidis argues that the current closed AI models hinder progress, advocating for a collaborative, transparent approach in AI development through open-source practices.

---

Key Topics Discussed

  1. Open vs. Closed AI
  2. Definitions:
  3. Open Source: Includes open data, open code, and open models/weights. The aim is for reproducibility and community collaboration.
  4. Closed AI: Dominated by big tech firms limiting access and programming flexibility.
  5. Challenges of Closed AI:
  6. Innovation is stifled; closed models may hide vulnerabilities and biases.
  1. Benefits of Open-Source AI
  2. Innovation: Open-source AI can lead to faster technological advancements by harnessing wide-ranging community input.
  3. Safety: More eyes on the model fosters identification and rectification of potential risks and biases.
  4. Real-World Applications: Open-source models can significantly impact fields like healthcare, where AI can enhance diagnostics and patient outcomes.
  1. OUMI's Role in AI Development
  2. Platform Offerings:
  3. Fully open-source AI development platform.
  4. Tools for pre-training and post-training across different models.
  5. Accessibility: Simplifies the training process; researchers can easily utilize resources regardless of their compute infrastructure (cloud or local).
  6. Community Collaboration: Encourages collaboration with universities and other entities to develop better AI models.
  1. Concerns About Data and Bias
  2. Open Data: Critical for transparency; understanding the data used is essential for model reliability and fairness.
  3. Bias Issues: Acknowledges that biases may persist in pre-trained models and emphasizes the importance of addressing these during fine-tuning.

---

Key Takeaways

  • Open-source is the Future: Koukoumidis strongly believes that the future of AI lies in open-source methods, drawing parallels with the evolution of Linux as an open alternative to Unix.
  • Community Empowerment: By democratizing AI development and access, OUMI aims to empower a wider community of developers and researchers.
  • Call to Action: Koukoumidis encourages listeners and developers to contribute to open-source initiatives, stressing that collective effort is necessary for the advancement of AI technologies.

---

Conclusion This episode provides valuable insights into the open-source movement in AI, making a compelling argument for its necessity in fostering innovation and collaboration in the technology landscape. The discussion emphasizes the potential of community-driven efforts to reshape the future of AI.

---

References

  • [OUMI website](https://www.oumi.ai/)
  • [SonarQube](http://sonarsource.com/eyeonai)
  • [Thuma](http://thuma.co/eyeonai)

---

For more insightful discussions on AI and its implications, tune in to the Eye On A.I. podcast.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00For something to be open source, it needs to be open data, open code, and open models or open weights. This means that it should be possible for somebody to have the ingredients and the recipe to be able to reproduce the model. Yeah, it's good if you're giving them the model, the weights, but also they should be able to reproduce it. I think this is a very, very good definition. However, what we say is that we want to have open code, open data, open code, open models. And what we say open collaboration, what this means by this, which may be a little bit less clear, is that something that is open source by the OSI definition, it's not sufficient.

0:33If, yeah, the code, everything gets there, but it's too hard for anybody else to reproduce and build upon. Because then it is the whole purpose of open source being a community effort where it's easy for people to build upon and extend somebody else's work. AI coding assistants like GitHub Copilot, Google Gemini Code Assist, and Amazon Q Developer have quickly become essential tools for developers. They generate code with remarkable efficiency, significantly boosting developer productivity. But the widespread use of AI-generated code brings its own set of challenges. Bugs, vulnerabilities, and suboptimal code can inadvertently enter production, leading to issues with maintainability, stability, even costly outages.

1:25How can organizations minimize disruption and risk while maximizing productivity and innovation with these AI tools? With Sonar, the code quality and code security leader, organizations can amplify developer productivity in concert with AI assistants. They can improve developer experience with streamlined workflows and prevent quality and security issues in all code from reaching production. Sonar's AI Code Assurance feature found in their Sonar Cube solution offers a thorough validation process for AI-generated code, including the first in its industry to automatically detect and review AI-generated code from GitHub Copilot.

2:17The AI code assurance workflow helps to build trust in AI-generated code, assuring companies that proper due diligence has been performed and code is ready for production. Their AI code fix feature speeds up productivity and issue resolution with instant AI-generated fixes for issues discovered by their code analysis that developers can review and apply directly within their workflow. Sonar is helping redefine the software development lifecycle by use of AI and AI-agentic systems with tools that streamline the verification of AI code and provide guided remediation of issues, supercharging developers to build better applications faster.

3:11Join over 7 million developers from organizations like the DoD, Microsoft, NASA, and MasterCard who use Sonar. Visit sonarsource.com slash ionai to request a free demo and see how SonarCube can help your team with machine-generated software development. That's sonarsource, S-O-N-A-R-S-O-U-R-C-E dot com slash ionai. That's E-Y-E-O-N-A-I, all run together. and request a free demo. sonarsource.com slash IonAI. Very nice to be here, Craig. I'm Manos Kukumiris. I'm the CEO of UMI. Until about nine months ago, I was at Google Cloud. I was supporting the science and engineering for all the natural language AI services.

4:11I built also some of the multimodal ones at Cloud AI at Google. And about three months before ZGPT was coming out, I was bootstrapping the efforts to productionize CloudPalm and drove this effort almost until the general availability launch in May of 23, before it was called DeepMine and sorry called Gemini and moved to DeepMine. Before that, I spent some time at a startup. I was at Mera working on commercial AI. Also before that at Microsoft, where I was building something like ZadGBT in 2016, a multimodal open-ended chatbot, even built what now we call embedding retrieval and RAG back then. And before that, I was doing a PhD.

4:49I spent my first three years at Princeton and the last two at MIT. And before that, I was doing my undergrad in Greece. Okay. And UMI, what does UMI mean? UMI stands for Open Universal Machine Intelligence. We wanted to pick a name for the effort that, for one, clearly denotes that it's open. We want to make sure that this is very clear. And also we use the term appropriately, unlike perhaps what's happening in other cases. And we didn't want to say AI because we think it's a little bit overused term, because where everything started was machine learning, machine intelligence. So that's why we call this open universal machine intelligence.

5:36So universal means is we wanted something that open machine intelligence that is democratized, that people can use everywhere and anywhere. And that's why we end up with the name. Yeah. And UMI is a platform, right? So, I mean, I was interested in talking to you because open source, there's been a long running debate in the artificial intelligence world, certainly in the generative AI space about open source versus closed source. There were a lot of hand-waving in the beginning about, my God, what if this stuff is open source? Then all the bad actors in the world are going to have the most powerful AIs.

6:26People like yourself and Jan LeCun argued that open source is the way to go, and that if anything, having more eyes on the model creates a safer model. and since certainly I mean there's been a growth in open source you know starting with Llama by Meta and then Mistral with its models and then DeepSeek sort of came along and woke people up if they weren't paying attention to China that open source is a big part of China's AI strategy so I wanted to hear from your perspective just first generally how you feel about the open source closed source debate and then my second question is open source is an overused and under defined term and what exactly does it mean?

7:31There's open weight, there's open code, there's open training data and no one seems to agree on what open source really means. So let's start with just the open source closed model debate. Where do you stand on that and how do you see it developing yeah i agree craig if that's okay with you how about we start from the definition just make sure we're on the same page about when i say about open source what do i mean about open source great once it's clearer you know when i argue that that's the way forward you know it's clear what exactly i mean by this so um as you mentioned before there's a lot of efforts quite often many of them uh being called open source when they're not really um i think the osi the open source initiative definition uh for open source ai it's pretty on point it says for something to be open source it needs to be open data open code and open models or open weights this means that it should be possible for somebody to have the ingredients and the recipe to be able to reproduce the model yeah it's good if you're giving them the model the weights but also they should be able to reproduce it i think this is a very very uh good definition however However, what we say is that we want to have open code, open data, open code, open models.

8:53And what we say open collaboration, what this means by this, which may be a little bit less clear, is that something that is open source by the OSI definition, it's not sufficient. If, yeah, the code, everything gets there, but it's too hard for anybody else to reproduce and build upon. because then it is the whole purpose of open source being a community effort where it's easy for people to build upon and extend somebody else's work. So that's what we mean by open collaboration that for one, we have the principle that any of these AI technologies that we also are building that are easy for anybody else to use in a fully record, sorry, very easy for somebody to reuse because what somebody else did is fully recordable and the platform has developer ease of use as first principle.

9:40and also to extend a bit also the collaboration aspect is that we already have and we plan to have even broader open collaboration efforts where we'd say you know what here's the goal the problem we plan to solve anybody in the community is invited to help us make it better so this can be more inclusively developed it's almost similar to what was happening in some ways with linux compared to unix where it was collaboratively advanced but again to do this it needs to be open code open data open models but also in a way that's easy for everybody else to reuse build upon and contribute So that's our definition of open source.

10:12We think it should be nothing less because anything less is doing a disservice to the community and not helping us as humanity advance this technology in the best way possible. Let's now move to the debate. You know, why is this the best way possible? Actually, before you do, open data is the one that's a little confusing to me. Open model means the source code is published. Open weights mean the weights are published. I at the models I've looked at people describe the source material for the data but no one that I've seen really uh has uh you know a database that you can look into and see exactly what data is in there and I presume because there's a lot of concern about uh copyright liability and that sort of thing.

11:07But can you talk about what open data means? And are there any models, quote-unquote, open source models today that have truly open data? Yeah, yeah, very good question. So, you know, for an open data means that again, if you use the analogy of, you know, baking a cake, you need to have the ingredients, not just the recipe or the oven, but you also need to have the ingredients themselves to build it. Now, in terms of, as you mentioned before, quite a few companies, quite often they may say we have an open source model, we can open source the model, but usually they only mean the weights, not the data, not even the code.

11:48There is very few truly open source efforts right now, and it's mostly by universities or non-profits, like for example, AI2. Everything else, even though they may say they're open source, they're not really open source, which means it's very hard for anybody else to reproduce and build upon it. Okay, and somebody like DeepSeq who claimed when they came out that they're fully open source, including the data, they described the data. Is that helpful at all? So I'm looking forward to all the announcements that DeepSeq is going to be making this week. They said it's going to be announcing one more thing every day.

12:31But as it stands right now, the models they have are like R1 and others. They are open weights. They have released some code. Like yesterday, they released some inference code that helps do inference more efficiently. But still, nobody can fully produce their work. They had a report that detailed some of the steps. But if anything, that left a lot of questions across the industry about how exactly did they do it and what exactly data did they use and how exactly did they produce the data. So yeah, nobody can still reproduce, for example, DIPC-Card 1 because they're liking the actual data. And some of the details of the algorithms that it's not clear if everything is in the report.

13:12Yeah. But you guys are promoting full open data. And what does that mean? Exactly. So one good example is an agentic model that we released just about 10 days ago on a Friday that was called CALM, C-O-A-L-M. And actually it's an agentic conversational agentic model that even built GPT-4-O in some of the Berkeley benchmarks in the leaderboard. and to give an idea for this any of the data that we curated in collaboration actually that's say that it was UIUC Emre and Gökhan and the lecturer from UIUC that was leading this effort any of the data they generated and any of the training that we also helped them perform and scale up to 70b or even 405b which I'm not sure if any other university has trained the 405b model so far all of that is open which means right now somebody can go to the UMI repository and see where we have the the pointer to hiding face where we have uploaded the data sets and it takes right now two commands for somebody to produce they can do UMI install to install the solution and then UMI train with the name of the configuration file if they want the the smaller 8b models or the 7db or the 4 of 5b that's what takes two commands and as I tell people it shouldn't be any harder than that for anybody to reproduce somebody else's work end to end with the data, the code and everything.

14:37Of course, the model itself is based on LAMA, but at least all the post-training and everything else is fully open source. Yeah. And then presumably there are, even if it's fully open source, there are licenses attached. And is there a concern that if you fully open source in that regard, open data, that people will just take the model and not license it? Yeah, so there is that concern. But at the same time, this would be almost, you know, also if I'm saying the question correctly, you know, by design, the desired benefits of open source, which is you're saying, you know what, I'm going to help contribute to the common efforts as research community, as humanity, to make this technology better.

15:29And I'm going to tell somebody else all the critical ingredients so that they can also help me make it better. I think one good example here is I was seeing this just a couple, maybe weeks ago. It was the Human Genome Project. That it was, I don't remember how many, over 100 institutions, thousands of researchers that participated. it and somebody could have said like the same thing now with frontier foundation models hey we're going to give it a black box just so that we can monetize freely uh or we're going to make it open and share the knowledge with everybody and i mean i don't know where people stand in terms of politics you know with i think it was bill clinton but regardless of politics you know he made a statement that resonated with me that said that he's i think he said it would be virtual criminal if they were to withhold all these critical discoveries that are almost about life and death just so that as opposed to advancing humanity and helping everybody advance it instead they close in a box so that some companies can monetize out of it i think the same thing for foundation models the link may be less intuitive to um to life and death but actually it's there we're using foundation models not just for technology and the industry but for all the sciences from material science to climate science to healthcare i was actually maybe not to disclose specific names here uh in um in a in a discussion with uh quite a few healthcare leaders from some of the biggest institutions we have in the us and what he was telling that they use right now these models they adapt them to detect i think uh uh polyps i think colonoscopy if i'm correct or some of these procedures.

17:10And what they said for these procedures is that soon enough it's going to be malpractice if somebody doesn't use AI, not the other way around, because AI was like five times, if I can recall, better than humans, not 20 % or 30%, but five times or something better, which means, again, and then you get almost the question about it's almost like a matter of life and death. I mean, these are critical technologies. It's a disservice to humanity if they're not open and they're not advanced. And even if you keep those open, there's still enough opportunities for many enterprises, including ourselves, to monetize.

17:44So that's what I think things should be. And we're going to also be touching on the earlier debate. Yeah. Okay. And then more generally, how do you think this is going to develop? Do you think, and then we'll talk about UMI, but do you think that open data will increase? what will happen to the closed source operators? I mean, already DeepSeek has put pressure on them with its pricing. Yeah, yeah. I think eventually things should be more open because if anybody, even closed providers, are using data that they shouldn't be using, there should be a little bit more transparency over that or make sure that everybody's being a good player.

18:30So if anything, eventually things should be more open. But yeah, for pre-training, massive amounts of data, it's the whole web, sometimes even accounting for everything may be practically hard in many ways to be very thorough. I think for post-training, a lot of the research, for example, that we recently did or we're doing with university, it's a lot easier to do. And yeah, I think now your second part of your question was, okay, what will happen in the future, right, in terms of open versus close and the pressure that open models are putting? that was actually our strong conviction about a year ago or more than that when we were starting this effort so you know what it's clear that the future of enterprise ai and oh and beyond just enterprises is going to be open the same reasons why linux became better than unix and i could elaborate a little bit on the history and the parallels that are happening with ai and And once this starts happening, which, if anything, maybe with Dipsy could accelerate a little bit more, even perhaps a little bit more than what I was expecting.

19:37And if anything, then all the pressure is going to be a lot of pressure to the closed source companies. And they're going to start feeling a lot of pressure because they, I mean, arguably you won't be able to justify the greatly inflated valuations anymore. or for them to be able to continue raising at this high valuation so they can keep executing on the same, in my opinion, bad strategy because it will become clearer and clearer to more people and their investors that, hey, this just doesn't make sense. There is a better way to do this. If anything, history has taught us that there's a better way usually for these complex technologies to be developed and they may stop funding them or be less interested.

20:19Yeah. Okay, well, let's look at UMI. So tell us about the genesis and exactly what UMI's platform does and how it's different from last time we spoke from Hugging Face or Lightning Studio is one I'm familiar with, the other platforms out there. Yeah. So the reason why we started UMI Actually, sometimes people ask me, you know, what was the one thing that made you stardew me? I say, no, actually, it was not one thing. I not just want to do the right thing by humanity, but also I'm a pragmatic person. I want to build the conviction that something makes sense before I start doing it. But in terms of the right thing, that goes a little bit back to the debate, right, that we're discussing before.

21:09It was very clear in my mind that for this, generic technology should be like electricity, should be a common utility. It's going to be powering everything. for science's industry to advance, there needs to be an impeded access, and also to advance it faster, safer, and most efficiently. This can only be done if you have all hands on deck, all eyeballs on this technology, helping to make it safer and advancing. So philosophically, it was very clear that, you know what, this is the right thing for humanity. But then increasingly, I kept thinking, okay, but this may be the right technology, but will it happen in this way, or will unfortunately go in a different path just because, you know, it's not practical for open source AI to succeed.

21:52And then it was a sequence of things that made me realize, actually, you know what, actually this could work. And not only it could work, it's actually the most likely scenario, the most likely outcome. The first one was increasingly I was realizing how, especially for post-training, and there's so many tasks, so many capabilities that these foundation models have across so many modalities that we are already working on and many more that we haven't worked on yet that is not necessarily bottlenecked by the amount of compute you have but the amount of people creativity and people that you have to go and innovate in all these different directions and across all the different tasks and modalities to advance them which means that we an open community that is not just a thousand two thousand three thousand people but orders of magnitude larger can do a much better job if they have the right tools, which is what we're helping them at 2Me to give them.

22:47At the same time, besides the fact that, again, an open community could do better than a closed community, I was increasingly realizing that the current status quo is only serving a few companies. Nobody else wants it. So as I tell people, unless you are an aspiring AI oligarch, and the handful of ones that they are, they know who they are, nobody else wants this because imagine for example the outcome where uh let's say you know a specific closers model like gemini or open ai's models dominate everybody's going to go use those specific models on the very specific cloud provider on the preferred accelerators that this cloud provider has which means other cloud providers they won't have a competitive model to serve other accelerator providers including nvidia amd all the major ones they may be out of the loop because they know that the closed model providers that may have their own preferred accelerators where because the model is closed they can more freely optimize across software and hardware which means they can offer a better experience end-to-end so anyway and besides the cloud providers and accelerator providers then i started realizing that actually it's not just them even consumer companies don't want uh ai close the ai close source ai to dominate mark zuckerberg has said it publicly recently quite a few times that even for a company of the size of meta it would be problematic if ai one day is controlled by a single company let's say open ai or google and then they are at the mercy of getting permission from them to use the ai they need for their consumer products and the ai is going to be everywhere right so this is extremely critical so given all that i was realizing that there's so many tailwinds again most of the players in the ecosystem they don't want to succeed and if you help channel the efforts across the community and across them in a more meaningful way at the end what is the most likely outcome that open source would succeed or close source that is developed in a silo by its company individually to shoulder all the economic and human costs to make it happen so for me was becoming clear again and as history has told us with Linux versus Unix, that the open source will be actually the most viable path.

24:57And I think this becomes increasingly to more people, and that's why we started UMI about a year ago. It was a combination of, you know what, this is the right thing for enterprise to succeed. This is the right thing for humanity, and at the same time, it's possible. So it just makes sense to go build it. Yeah. Well, tell us how UMI works. I mean, how does somebody use the platform? Create an oasis with Thuma, a modern design company that specializes in furniture and home goods. By stripping away everything but the essential, Thuma makes elevated beds with premium materials and intentional details.

25:37I'm in the process of reorganizing my house, and I'm giving Thuma a serious look for help in renovating and redesigning. Thuma combines the perfect balance of form, craftsmanship, and functionality. With over 17 ,000 five-star reviews, the Thuma Bed Collection is proof that simplicity is the truest form of sophistication. Using the technique of Japanese joinery, pieces are crafted from solid wood and precision cut for a silent, stable foundation. With clean lines, subtle curves, and minimalist style, the Thuma bed collection is available in four signature finishes to match any design aesthetic.

26:31Headboard upgrades are available for customization as desired. To get$100 toward your first bed purchase, go to Thuma.co.i on AI. Ion AI all run together, E-Y-E-O-N-A-I. So for$100 off your first purchase, go to Thuma.co.i on AI. That's T-H-U-M-A dot C-O slash IonAI to receive$100 off your first bed purchase. Yeah. So UMI, as we launched it about three and a half weeks ago, so it's very, very recent, it's a fully open source platform on GitHub. It provides all the capabilities that somebody needs in enterprise or academia to advanced foundation models, anything from pre-training, all the different post-training techniques, full fine-tuning, parameters and fine-tuning, RL techniques, evaluation with all the most common benchmarks and a growing list of them, data curation, auto-evaluation with LLM judges, and many more.

27:46And the idea is that all the tools that somebody needs to do research in academia, let's say you want to develop the next Dipsick RL, the next Dipsick R1, you should be able to have all the steps in a single platform so that your whole recipe can be recorded fully. It's not you combine this with that, and there's dark knowledge and scripts, and so anything should be recorded in the same platform. First of all, for convenience, you have all the tools you need. The whole recipe can be fully recorded, which means it's easy for somebody else to go to produce it, cite your paper, and continue advancing it.

28:17It shouldn't be any harder. So that was part of the goal. Now, besides providing all these capabilities, we support both text and multimodal models, over 100 open models, pretty much all the ones that matter right now that are relevant for enterprises or researchers. They are supported. From hundreds of million parameters to all the way, a 405b LAMA that somebody can train. And also as a solution, we wanted to make it fully accessible, which means that, and seamlessly accessible, which means we want to make sure that somebody, whether they're getting their computer, Let's say they have just a MacBook with a couple GPUs.

28:56Or they're getting their compute from some cloud provider, WA, GCP, Azure, Lambda, Rampo, Together, or anywhere else. Or they have a big HPC that they can still run the same experiments, of course, assuming they have the necessary compute. But all it changes is that some simple deployment configuration says, how many GPUs do you have and how many do you want to use? but it is that the same recipe should be reproducible easy by somebody regardless of where they get the compute. So that's kind of what the platform provides. Now again, all the capabilities that somebody needs to build foundation models or adapt them to specific tasks and domain, whether it's pre-training or post-training, even evaluating data curation, all text and multimodal models and also running across any platform and across any scale.

29:46which as I mentioned before I'm not sure how many universities have been able to produce models at this scale and that's why those researchers from EIUC, we had a long relation with them, they said you know what, we really need your help to scale them up because we have been using other technologies but it's just impossible for us to do this and yeah, the good thing was very easy and it's also very easy for anybody else now to produce so this is the fully open source solution, we're starting to develop also enterprise solutions on top of that happy to discuss about it as well Yeah, on building models, how many, you know, on Hugging Face, there was, I don't know how many models are now, transformer-based models.

30:30how many people are really building models from scratch i mean they they write their first line of code and then you know code for however long it takes and then train that on on data that they've collected versus how many people are taking LAMA or some other open source and tweaking it or forking it or whatever they're doing, fine-tuning. Yeah, so the vast majority of the people are getting an existing open model because to get to that point takes significant cost. It doesn't make sense for them to produce it. They get an existing open model and they continue post-training it with different techniques to continue making it better for specific research or if they're enterprise for a specific problem they're trying to solve.

31:29Very few are working on pre-training or training from scratch, actually to be more precise, to training from scratch those foundational models. But the good news is that there is huge innovation that's still required, not just on the pre-training, but across all the other stages. And as I was mentioning before, an open community, especially when it comes to post-training or finding creative ways to get existing open models after the summary like meta has incurred all these massive costs to train them and make them better across all these different domains and modalities and capabilities and with creative approaches like deep six grpo and any new ones that they're gonna invent that's something that you don't need massive compute and there's huge opportunities us as a research community as humanity to continue to advance them sure it would be great if it was easy for more people to do pre-training but even on the latter stage where it's more of a how many people how many creative minds do you have to advance this and less about the compute constraints uh so there's huge opportunities again to advance even uh on getting an existing model and curating or finding creative ways to make it better uh and actually just something to add to my previous uh the previous question i forgot to mention is that besides providing the platform itself we already have also open collaboration with many universities and we plan to launch even broader collaborations where we say you know what here's the problem that we as a community make sense to solve let's all open and collaborate and make it better and uh does the platform require uh open code open um data as well as open weights yes exactly so for all any of these collaborations that we currently have to advance this foundation models we say you know what everything needs to be open and if it's built on the only platform for one it's very easy for you to do your experiments end to end and then the record the full recipe so that anybody else can easily reuse your data your code everything and reproduce this work and then make it better yeah yeah it's it's it's um it's an absolute necessary requirement.

33:43If you want open source to thrive, that's how we need to function. Anything less will just impede the success in open source. Yeah. And you think that people will start building truly open source models if they have this platform or, yeah. Yeah. i think for pre-training it may become my hope is eventually we'll be able to do more and more open source even though now the truly open source on the pre-training states are very very limited my hope is that in the future we may do more and more there as well but for post-training and or at least using existing pre-trained models and different continue to pre-train them or post train them in different ways that is already happening there's huge appetite from researchers in academia to do this and by the way to your earlier question that's how even we started with umi this started with a sunday discussion with uh ruslan salagudino from cmu where he told me you know what uh we are i mean he didn't say that but they are the top researchers among the top researchers in academia and and then he said that uh that's the top research it's it's my assessment um they said you know even even for us it's very hard to do this type of research it takes a lot effort for my students to put together all the pieces that they need to do this type of research and then none of them has scaled and done multi-node distributed training to scale up and train bigger models but it's too hard for any student to to do this so if you build a platform we'll support you and then he introduced me also to amit al-warkar who is also from cmu and actually in his words he said uh we are an untapped resource we want to contribute but those labs don't make it easy for us or possible to be precise.

35:36And also, you know, if you had the platform we described, that would be really helpful. So yeah, please build it. We're going to help you. It was the same response for many other academics. So that's why we started to build this. So it's easy for people to do the research, build upon it, and then give the, as we say, open source, the kind of Linux moment that it was missing. Yeah. One question I've had for a long time. If you take an open model, well, let's take LLAMA as an example. It's been pre-trained, right? So are there models out there, the code out there, the source code is out there that you can pre-train from scratch?

36:20Yes. There is a few, very few examples. One of the most notable is the efforts from AI2, where, for example, for the old paper, they said, okay, here's the data that we used, and here's the code, and then the resulting model. So we need more of that. And my take with this is, you know, but such models should not be developed by almost a single organization or small collaboration should be a much more open collaboration across the community. And that's what we're trying to help do. So it's amazing to have such efforts from Mera or AI2, but there was a missing tool, a missing capability, the Linux of AI, as we call it, that was missing for the community.

37:04And that's what we hope to help provide. And as a platform, it's something that can work with the models developed by AI2, by Mera, or any new models that the community develops. Any of the best tools. So that's why we're very grateful and excited for all these efforts that all these other organizations are doing, including Mera, even though it's not fully open source, open weights, it's still a huge asset for the community. Yeah. So if someone, because I talk to people all the time who say they're building on an open source model and they're fine-tuning it for one particular use case or another, but if you're doing that, if there are biases or frankly bad data in the original training data, it's baked into the model, right?

Read the full transcript

38:00It's encoded in the weights. So it seems like that's a disadvantage that you're not able to fully control how the model is going to react. or is it that taking source code, as you said, some of these models, and then pre-training it yourself is just too expensive? Yeah, yeah. So that's a very good question, Craig. So what is definitely clear is that if one had access to all the data in the procedure in which a model was pre-trained, it would be easier for them to reason about why the model is behaving in a certain way or what could I change as a researcher to make it better. This is something actually that some very notable researchers have mentioned in the presentations that I attended in the past.

39:00So, you know, we want to be able to do this research, but, you know, if we had access, if we knew these details, it would be a lot easier for us. So, yeah, that's definitely unfortunate. it that being said um if a model like for example what meta does with the llama models they're open and say and says you know what here as a community you can get it deployed whatever you want test however you want i'm not gonna control or uh limit you if you issue requests to my api that limit you in certain ways uh or because maybe you don't want to reverse engineer or do whatever you may want to you're trying to do with the model then it still leaves some good amount of flexibility for the community to identify the issues and say you know what i think it's it's it's demonstrating this issue and here and here's how by continuing to pre-train it or to be usually with some alignment technique mostly as a post-training technique i can align the model to do less of that bias and behave more in the expected way.

40:00So you could still rectify and fix some of this in the post-training. But yeah, to do fundamental research that helps us understand better how these big models behave, I mean, yeah, it would be really helpful to have more access to the whole recipe that would help to do more foundational research in how they behave. Yeah. I was talking to somebody recently about a project that they're working on. It's a fascinating idea is to take a model, an open source model, and train it on all the original religious texts in the world. but if you do that with a pre-trained model if you do that in the fine-tuning or with rag you've the the pre-training model already has a lot of that text in it and depending I mean, this is something that bothers me about these public models that have been trained on the internet.

41:16As the internet fills up with marketing material, for example, if you ask about a company, the answer you get back is not objective. it's skewed by the marketing content that's in the, and I would guess that companies with enough budget realize this and are pumping out marketing content onto the internet so it gets vacuumed up and biases the model in their favor. And so if you're working with a pre-trained model on something like this religious text, presumably there's going to be some bias in the underlying model depending on the distribution and the training data. So how do you adjust for that?

42:11Or is something like all the religious texts in the world, is that enough of a corpus to pre-train a model from scratch? Or is it, yeah. yeah so i would uh conjecture here that likely the old religious content maybe may not be enough or maybe a lot of i guess knowledge like science and many others that if you want to have a very comprehensive model you need to expose it to all the knowledge not just on our religion um that being said um you know i think that's also one of the benefits of open source is that you know it's hard to say you know what is the right religion or what is the right way to respond about a certain religion you know do you take a position or do you not um and different nations religions should have the control if they desire you know to develop their own model let's say that answers questions about their own religion and shouldn't be a single owner uh like i say one specific big tech company that decides what everybody's religion or what's the you know the right answer to many of these topics so that's why and arguably it may be even hard for them because there may be even some small religions that are not represented on the web very well that's why for the sake of inclusion the best way is if those technologies are open and you democratize and make it easy for anybody to build the tools for their own language dialects for their own religions for their own uh you know uh a more specific uh context i think jan lecun uh had uh mentioned in the past actually quite a few months maybe maybe a year already that uh uh that why he thought also open source is better than closed source which is like the equivalent of wikipedia that the best way to capture all of humanity's knowledge and inclusively without having single person decide what is the right thing and then leaving out everybody everything else that he doesn't consider important enough is to do it in a wikipedia style to do it in an open way so i think it was very very um you know thoughtful um comment he had made so anyway yeah i see i see the i see it the same way uh and again that's going back to me that's what we hope to democratize should be easy for anybody if they have access to their own data to continue pre-trained post-trained get the model that is desired for their domain because yeah for some these things, it shouldn't be a single tech company deciding for everybody else what the right answer should be.

44:41Yeah. Is it possible to train a foundation model to purely on the linguistic structure of language without any knowledge so that then you can add the knowledge in the fine-tune stage. Yeah. You could. I mean, there's been quite a few discussions and research lately that, you know, what would be best ways to, I don't know, maybe your question's getting that way, to decouple the reasoning of a model from the knowledge because a model that can reason but is liking the knowledge could be smaller, more compressed, and then the knowledge could be external. And then if you're smart enough, you're able to reason about stuff, but then you can access the knowledge when you need it.

45:28It will have smaller models that can tap into specific knowledge when desired. So to some extent, that is possible. And arguably, you mentioned before, RAG. That's almost what people do with RAG. I mean, you train a model that has some basic understanding and knowledge, but when it needs more specific knowledge, it knows how to understand text, it knows how to understand concepts, and can access this external knowledge. bring it in and use it to respond to something almost pretending to be an expert on the specific domain to be honest it's not as effective as fine tuning that's something sometimes people don't understand they think that drugs solve problems and what i tell them you know rag is like having a high school student that has access to google search you know they can quickly go find information process it yeah they have some common knowledge respond to you but it's not the same example let's say if you are a legal company an attorney having somebody who has the experience and they're able to reason for how people should reason the specific discipline while still accessing the knowledge to respond for example if you're an attorney you don't just respond based on some information but i would argue again hopefully no attorneys get offense here that usually they're risk averse that's my understanding again i hope i'm not ignorant that they may respond in a way to help cover their clients across all corners or all cases.

46:56And there's some reasoning, there's some specific rules perhaps that go into their training about how they think about the information to craft a recommendation. And that's something that models can learn with pre-training or post-training. And that's why, for example, just accessing knowledge may not be sufficient. because it's like giving a high school student access to the book. He can read it, but it doesn't mean that he has built the right intuition. In some cases, you need that to be able to give the best possible answers. Yeah. And is anybody – well, this is getting off topic, frankly, but I'll ask this last question, and then we'll get back to Umi.

47:40But are there efforts to train a foundation model more carefully with more highly curated data rather than just, you know, willy-nilly scraping the Internet? Yeah. Yeah. I was going to say that even Meta, even though they only have open-way models, their report for Lama 3 was actually reasonably detailed and helpful. And for Zoom, they make it very clear that they put a lot of effort in cleaning up the data on the web. But actually, one thing that may be not clear to many people is that when we say cleaning up, it doesn't mean that you necessarily need to remove all the bad data, like offensive data, because for somebody to understand what is bad and what is good, you see both the bad and the offensive and the good yeah so it stops them to see all aspects but anyway it's still there's still a lot of careful data curation so that we can give all the most useful tokens the most useful data to those models to be trained so yeah data curation is already happening it's very important um and arguably one of the best ways to continue improving these models is to be more careful and thoughtful both about the type of data you expose them to but also in what sequence you something that's called curriculum training, how do you start training a model on some basic data, and then perhaps you move it to more harder domains, or then you introduce later some capabilities, like being able to understand longer context and things like that.

49:11Yeah. Okay. So UMI, a model builder can go on to UMI and connect to data sets to train a model, either write code or, you know, pull code in from another model. Can you also connect to, you were talking about the cloud, you can connect, is all of that in there that then you can train right in UMI or do you have to export the code? Yeah, so yeah, as you mentioned before, let's say somebody wants to train a new model like the agenda model we released or some other model. Many of the most common data sets that are currently in Hagen face, we provide adapters so they can all be standardized and brought into the same right format and existing recipes that should work very well already out the box so that you can almost right now with again like almost a single command you have already installed umi reproduce training llama on some common benchmarks of your own choosing almost uh you know it's a single command a single configuration file you create um so the the basic you know functionality it's already very very easy to do um now if you have some uh reasonably capable laptop you could do some of that on your own compute and then it just takes some changes to what we call a deployment configuration not writing code not writing these things where you just say here's my credentials let's say for gcp and here's my credentials and here's how many gpus i need that i have already created the vms go run it there so it doesn't take even any code to run it just changed the configuration for how UMI can deploy while running the commands on your laptop.

51:17They can deploy to the VMs you have allocated on GCP. So it's very easy. And again, I tell people, you know, it shouldn't be any harder than that. And again, if you have an HPC, again, it takes again a single command. You can run the UMI code on the HPC itself and you can deploy to hundreds, thousands of nodes GPUs if you have access to. Can you talk about the, you said there was an UMI-powered model that was published on HuggingFace recently and outperformed some benchmarks. Can you talk about that model? Yeah, yeah. So there was, you could say, a gap in the existing models. And Emre, working with the Leica and Geek and Turfam, you see, like, you know what?

52:07It would be good if we have an agentic model that is very good at doing multi-term conversations and at the same time is very good at tool use, which means, with simple terms, you can call external services to access information and then use that to perform a task. And through research that they did in the data curation, they managed to create a recipe that showed that actually you can develop a model that beats GPT-4O in many, actually stacks higher. In some benchmarks, it beats GPT-4O. And in the leaderboard, it currently think it's above Mistral, Gemini and many other models across the board, across all benchmarks.

52:56And in some specific ones, even better than GPT-4O. and the thing that was actually the thing that i'm very excited about this because i tell people you know and i don't want to diminish this excellent research done by the research team in some ways it's a drop in the bucket it's like all these little drops that everybody one everybody of us in the community needs to do and will do because there's so much interest to fill up the bucket of innovation to to make those open models better and the good thing again is because they're already in the same standardized platform, for somebody else to combine M-RES work with somebody else's work, it's very easy.

53:36And then they can again go beyond it and make it better and better, and then the next person can reuse their work and make it one step further. Yeah, that's exciting. And that model's been up there for how long? I think it's now 10 days. Are people downloading it? Are there any metrics around how the community is using it? I think it's been already, at least in the Hacking Phase, over 100 downloads, maybe many more. I haven't looked at it recently, to be honest. But overall, the reception was very good. One way how we know this is, the day the model was published, we got quite a few people reach out to us and say, hey, I want to volunteer and help in UMI's research efforts.

54:18How can I do that? And we see a very big update, which means it resonated with a lot of people. Yeah. And the licenses attached, does UMI offer a selection of licenses? How do you... Yeah. So quite often for a lot of research that happens, because people, they may use LAMA, or they may which has some limited constraints but i think still doesn't block for the most part research or enterprise use unless you're hitting a specific scale and quite often researchers when they produce data or advanced research sometimes they may use data sets that come from sources that restrict them from them releasing their own models or improvements in the fully permissive licenses but by far our preference and what we tell to many people in the research community is that if you need to do that because that's the only way for you to advance the research do it because it helps still educates the community but the best way is to release things with very permissive licenses because then not only fully enables the research community but also the enterprise to use them and then you can have material impact to help enterprise give them a better alternative compared to a black box that you know you get from all the other closed source providers.

55:39So yeah, permissive licenses are definitely preferred wherever possible. I mean, before I ask the next question, is there anything I've missed in talking about, Umi, I mean, that you're waiting for me to ask or that you want listeners to know? I think we covered a lot of ground. The only thing I would say is that I'm very hopeful and optimistic that as a platform, we should help people do this work end to end and also help other people build up on their work, which, again, as I mentioned before, it's critical for open source succeed. But the best way for any of these succeed, for the benefit of all of us, for the benefit of the sciences, enterprises, humanity overall, is if we all work together to make this happen.

56:26if people say you know what I know I was supposed to doing my own small thing I really want to contribute something bigger pursue my own independent research but also in a way that helps everybody else build up on it and the whole community to move forward that's the key thing I mean we all need to work together that's the best way to make this work and I do think that we can make it happen it has happened with many complex technologies like Linux and Unix we can make it happen yeah is this only for, Izumi, only for foundation models or can you build other kinds of software, open source software?

57:00Yeah, great question. So currently it's mostly around foundation models. It's possible to think about Gen AI generative models. We plan to introduce capabilities, which we have internally, but we haven't introduced them in the platform yet, to train classifiers and to perform more classic machine learning tasks because FindOcean models are great tools to even perform these more classic tasks. So we plan to do more and more of that. There's plans that we have to develop tools across the whole ecosystem. How can you better review and improve data or contribute to data improvements? There's all across the ecosystem that we plan to do more work to improve overall the tooling that you need across end-to-end.

57:43And of course, contributions from the community are more than welcome. We want these tools that if it's anything that helps the core community to advance foundation models for the benefit of all of us, we want this to be fully open and uninhibited. Yeah. And I can see that increasingly, that's why I was asking about are people building models from scratch? because there really is a need for a completely open source, open data foundation model that's high performing that people can work with so that there aren't all these questions about data bias or whatever. Where do you see the open source landscape in, say, five years?

58:38and presumably you're you're hoping that umi will will play a big role in shaping that but but do you um do you see that that there will be these fully open source foundation models meaning open data so that uh people can can work with them uh and you i mean as you said and a lot of people have been saying there are going to be these foundation models all over the place, or at least applications built on foundation models. Do you see that there are going to be a handful of foundation models, open source foundation models that people work with, or are there going to be thousands or tens of thousands of differentiated foundation models out there?

59:40Yeah. So I think the, I was going to say the foundation of foundation models, the general foundation model is going to be the starting point for many people to start building more domain-specific or task-specific models. That general foundation, the general foundation models, I think that will increasingly be open and I would for sure hope, but I would also expect, you know, not just wishful thinking, but I would expect that the majority of those huge cases in the future are going to be powered by open models. I have seen that I was expecting that trend over a year ago. I have seen it accelerating the last months, especially weeks.

1:00:17Dipsick had a big impact on this. Where more and more enterprises are reaching out saying, you know what, six months ago when we started, I wasn't sure if we wanted to use open models, but now I see what you were saying. I want to use more open models. The quality is almost there. Many of them are already telling us. The moment we customize them, they're way better than the models that comes out of the box from OpenAI or Google or anybody else or Anthropic. So it makes sense for us to use an open model and customize it and get better quality with the full privacy, security, flexibility, at lower cost.

1:00:45It's just a no-breder. Yeah, it takes a little bit of cost to train our own model, but it's not as hard as it was before. And that's why we built UMI. so given that again i would expect that the the the core the general foundation models are increasingly open if anything uh the existing closed source ai closed source model providers they're going to have as we discussed before increasing pressure and less and less innovation and less and less lead and for them to remain relevant as a viable business they may have to move more and more to add-on applications as opposed to core differentiation on the base foundation model.

1:01:22I think easily that benefit is going to diminish. Actually, it's already happening. And Open is going to be just a better alternative overall for most enterprises. So that's the future I've foreseen. As you mentioned before, I surely hope that UMI plays a key role with this. As I was telling people early on, I was telling them, you know what? even if UMI fails, I really hope somebody else with the same strategy as we have right now succeeds. Because it needs to happen for the benefit of the sciences, enterprise, humanity overall. That being said, I surely hope that UMI is the one that helps achieve this.

1:01:59But again, the biggest thing is we make sure that this happens for the benefit of all of us. Yeah. Okay. Well, that's all fascinating. And if somebody wants to find Umi, give us the URL or where do they go on GitHub? Yeah. So somebody can go to www.umi.ai. And there we have also the links to our GitHub and our Discord where they can join and help contribute. We already have quite a few GitHub issues, as they're called, and efforts and also some projects we have started in our Discord. So, yeah, we would more than welcome for real contributions. As I mentioned before, the best way to make this happen is if we all work together.

1:02:48So I'm looking forward to seeing more and more people coming on board. AI coding assistants like GitHub Copilot, Google Gemini Code Assist, and Amazon Q Developer have quickly become essential tools for developers. They generate code with remarkable efficiency, significantly boosting developer productivity. But the widespread use of AI-generated code brings its own set of challenges. Bugs, vulnerabilities, and suboptimal code can inadvertently enter production, leading to issues with maintainability, stability, even costly outages. How can organizations minimize disruption and risk while maximizing productivity and innovation with these AI tools?

1:03:39With Sonar, the code quality and code security leader, organizations can amplify developer productivity in concert with AI assistants. They can improve developer experience with streamlined workflows and prevent quality and security issues in all code from reaching production. Sonar's AI Code Assurance feature found in their Sonar Cube solution offers a thorough validation process for AI-generated code, including the first in its industry to automatically detect and review AI-generated code from GitHub Copilot. The AI code assurance workflow helps to build trust in AI-generated code, assuring companies that proper due diligence has been performed and code is ready for production.

1:04:35Their AI code fix feature speeds up productivity and issue resolution with instant AI-generated fixes for issues discovered by their code analysis that developers can review and apply directly within their workflow. Sonar is helping redefine the software development lifecycle by use of AI and AI-agentic systems with tools that streamline the verification of AI code and provide guided remediation of issues, supercharging developers to build better applications faster. Join over 7 million developers from organizations like the DoD, Microsoft, NASA, and MasterCard who use Sonar. Visit sonarsource.com slash ionai to request a free demo and see how SonarCube can help your team with machine-generated software development.

1:05:41That's sonarsource, S-O-N-A-R-S-O-U-R-C-E dot com slash ionai. That's E-Y-E-O-N-A-I, all run together. and request a free demo. sonarsource.com slash IonAI

From the publisher

This episode is brought to you by Sonar, the creators of SonarQube Server, Cloud, IDE, and the open source Community Build. 

 

Sonar unlocks actionable code intelligence, helping to redefine the software development lifecycle by use of AI and AI agentic systems, to continuously improve quality and security

while reducing developer toil. By analyzing all code, regardless of who writes it—your internal team or genAI—Sonar enables more secure, reliable, and maintainable software. Join the over 7 million developers from organizations like the DoD, Microsoft, NASA, MasterCard, Siemens, and T-Mobile, who use Sonar. 

 

Visit http://sonarsource.com/eyeonai to try SonarQube for free today.

 

————————————————————————————————————————

The Future of AI is Open-Source | Manos Koukoumidis on UMI & The AI Revolution

Is closed AI holding back innovation? In this episode, Manos Koukoumidis, CEO of Oumi, makes the case for why the future of AI must be open-source. OUMI (Open Universal Machine Intelligence) is redefining how AI is built—offering fully open models, open data, and open collaboration to make AI development more transparent, accessible, and community-driven.

Big Tech has dominated AI, but UMI is challenging the status quo by creating a platform where anyone can train, fine-tune, and deploy AI models with just a few commands. Could this be the Linux moment for AI?

What You’ll Learn in This Episode:

  • Why open-source AI is the only sustainable path forward

  • The difference between “open-source” AI and true open AI

  • How OUMI enables researchers and enterprises to build better AI models

  • Why Big Tech’s closed AI systems are losing their competitive edge

  • The impact of open AI on healthcare, science, and enterprise innovation

  • The future of AI models—will proprietary AI survive?

The AI revolution is happening—and it’s open-source. If you care about the future of AI, innovation, and ethical tech development, this episode is a must-watch.

————————————————————————————————————————

 

This episode is sponsored by Thuma.

 

Thuma is a modern design company that specializes in timeless home essentials that are mindfully made with premium materials and intentional details.

 

To get $100 towards your first bed purchase, go to http://thuma.co/eyeonai

 

————————————————————————————————————————

 

(00:00) The True Meaning of Open-Source AI  

(02:15) The Open vs. Closed AI Debate  

(07:54) Why Open AI Models Are Safer 

(10:34) Defining Open Data

(13:21)Beating GPT-4-O with an Open AI Model  

(16:36) Open AI in Healthcare

(19:31) Why Open Models Will Dominate  

(23:07) How OUMI Makes AI Training Fully Accessible & Reproducible  

(28:44) UMI’s Collaboration with Universities  

(32:29) The Shift Toward Open A

(36:41) Can We Build Truly Open AI Models from Scratch?  

(40:20) The Role of Open AI in Eliminating Bias

(45:02) Will Open AI Replace Proprietary AI Models?  

(50:19) How OUMI Works

(54:44) The Open AI Revolution Has Begun

 

More from Eye On A.I.

All 266 episodes
#240 Manos Koukoumidis: Why The Future of AI is Open-SourceEye On A.I. · 1 h 6 min
Listen in VO