In short
Podcast Notes: Eye On A.I. Episode #304
Episode Overview
- Title: Matt Zeiler: Why Government And Enterprises Choose Clarifai For AI Ops
- Host: Craig S. Smith
- Guest: Matt Zeiler, Founder and CEO of Clarifai
- Sponsor: Oracle Cloud Infrastructure (OCI)
Episode Description In this episode, Matt Zeiler discusses the significance of inference in AI operations, clarifying how it has evolved to be more critical than mere training. The conversation explores how Clarifai optimizes AI deployments across various environments and addresses the challenges faced by organizations, particularly in regulated industries.
Key Topics Discussed
- The Evolution of Inference in AI
- Inference has become a focal point in AI, shifting from training models to deploying them effectively.
- Clarifai began with a strong emphasis on inference, inspired by successful developer products like Stripe and Twilio.
- Clarifai’s Approach
- Clarifai offers a unified compute orchestration layer that enhances AI capabilities for government and enterprise clients.
- The platform includes user-friendly interfaces for AI model training, allowing analysts to train models without heavy involvement from technical teams.
- Key Features of Clarifai's Platform
- Flexible Deployment: Supports cloud, on-premises, and edge environments.
- API and User Interfaces: Provides APIs for developers alongside intuitive UIs for data labeling and model training.
- Inference Optimization: Focus on improving token throughput, reducing GPU costs without compromising model performance.
- Government and Enterprise Use
- Clarifai’s services are employed by various sectors, including law enforcement and military applications, such as Project Maven.
- The shift in governmental AI projects reflects a growing reliance on AI for enhanced decision making and operational efficiency.
- Performance Metrics
- Importance of metrics such as time to first token and token throughput, which are crucial for real-time AI applications.
- Clarifai's benchmarks indicate significant improvements in speed and efficiency, attracting interest from potential clients.
- Future of AI and Inference
- Discussion about the integration of AI into everyday applications, with a focus on enhancing user experience through faster responses.
- Exploration of the edge computing landscape, where low-latency applications are increasingly prevalent.
- Competitive Landscape
- Clarifai aims to distinguish itself with a long-standing reputation and robust performance metrics in an increasingly crowded inference market.
- The company plans to broaden its accessibility through partnerships and integrations with popular developer tools and platforms.
Key Takeaways
- Inference vs. Training: Inference is now seen as the critical phase for AI applications, emphasizing the need for speed and reliability over just training.
- API-first Approach: A strong API framework alongside user-friendly interfaces is pivotal for client engagement and satisfaction.
- Government Collaboration: Partnerships with government entities highlight the importance of AI in national security and law enforcement.
- Edge Computing: The future will see more AI applications running on edge devices, improving response times and application efficiency.
Conclusion Matt Zeiler's insights provide a comprehensive overview of the current landscape of AI operations and the pivotal role inference plays in the future of AI technology. Clarifai's focus on performance, flexibility, and user engagement positions it as a leader in the competitive AI inference market. The episode emphasizes the need for continuous innovation and adaptation as AI technologies evolve.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00It started with just inference back in 2014. The reason we started there and having an API platform was inspirations like Stripe and Twilio, where they built a great developer product and people just built amazing things on top of that. And they became kind of the engine for that. And we saw that there needs to be this engine for AI. And so that's the same product on the commercial side as public sector side. We intentionally built a whole suite of UIs that make the system easy enough for anybody to use. And because of that, we've had even intelligence analysts train models themselves, doing all the labeling, picking a training template, looking at evaluation metrics of how they perform without our teams being involved.
0:43That's key to making the API successful, which is ultimately what you're going to use. Once you have a good model, you're going to use the APIs to fire through lots and lots of data for inference. But to get to that model, the UIs are really important. In business, they say you can have better, cheaper, or faster. But you only get to pick two. What if you could have all three at the same time? That's exactly what Cohare, Thomson Reuters, and Specialized Bikes have, since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds.
1:40How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 % less for compute, 70 % less for storage, and 80 % less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com slash IonAI. Eye on AI, all run together, E-Y-E-O-N-A-I. That's oracle.com slash eye on AI. I'm Matt Zeler and founder and CEO of Clarify. And my history in AI goes way back even into undergrad.
2:44That's kind of where it started. A little bit by luck. I was at University of Toronto and in a program where you have to decide if you want to take, you know, many different options of engineering. and I was deciding between the computer option and nanotechnology, and that's when I happened to run into one of Jeff Hinton's PhD students, a guy named Graham Taylor, and he happened to be my resident advisor on the floor I was living in. So that was the luck of it all, and he showed me some of his research, and back then he was generating videos that looked realistic of a flame flickering, and he said it was all done by AI.
3:20And so fast forward to today, everybody calls that generative AI. This was 2007. So we've been doing this stuff for a long, long time, well before ChatGPT, where everybody kind of realized its value. And so that got me hooked back then because I learned how to program before that. And there was no way I could program in a traditional sense to generate a video like that. So I decided to dive into the computer option, took Jeff Hinton's course, and then I ended up being Jeff Hinton's first undergrad thesis student. that was quite the honor to work with him and I did a thesis on actually generating motion capture data for biological studies of pigeons walking around and how they behave so that was a cool undergrad and by that point I knew that I had to you know focus on AI it was going to change a lot of things in the world and I didn't know enough coming out of undergrad so I decided to do a PhD and I looked at all the different schools and chose NYU to study with people like Rob Fergus who is my PhD advisor and others like Yan LeCun in his lab.
4:31So I got to work with both of them and focused on computer vision for my PhD applying neural networks to understanding images and videos and that was between 2009-2013. During that time I also spent a couple summers at Google Brain and worked under Jeff Dean on the Brain Project. At that time, it was like 20 or 30 people. So it was really fun to work there. And that's where I learned how to scale up AI at the large scale that Google operates at. But I left my second internship a couple weeks early to come back to New York and start working on building a company called Clarify. Because what I realized is the work I was doing in my PhD was actually working better than what Google had internally for understanding images.
5:18And I saw that as an opportunity to start a company around this. And it's been a childhood dream of mine to start a company. It's kind of since undergrad, I knew it would be in the AI space. And so this was kind of my shot at that, having really good research. And I incorporated in November 2013. And about three weeks after incorporating, ended up winning ImageNet with the Clarify name. And that was the biggest computer vision competition at the time and this was the year immediately after AlexNet won ImageNet which kind of put deep learning on the map as as actually being an effective method in the first place so it was great to be the first in the world to reproduce those results and then surpass them and it was a great way to kick off the company because that got inbound interest from day one from investors and and customers so that's kind of the origin story for Clarify and I can go into more about what we do, but I'll pause there.
6:14Yeah. And on that ImageNet competition, was it you and a team that you'd pulled together for the startup or were you still working through academia? So good question. I did actually two entries. One was with my PhD advisor, Rob Fergus. We call that ZFNet or I think other people named us that for ZELR, Fergus. and that got third place in the competition. Somebody else got second and then Clarify, which was only me at the time, got first place. There was a few different techniques that I was doing as part of the company separate from my work at NYU. Yeah, that's amazing. And then from there, you very quickly, Clarify, very quickly grew into one of the first on one of the most influential image recognition companies, and you began working with law enforcement and the Pentagon on Project Maven.
7:22Can you talk about that? Because I remember at that time,
7:29or around 2017, 2018, when I really started focusing on AI, I wanted to talk to you guys, and you were not available to talk. And I don't remember what it was. I think that there had been some sensitivity around the law enforcement applications or something, and you guys just didn't want to be out in the media. Can you talk about that period? Sure. Yeah, so kind of between starting the company at the end of 2013 to that period, we built a lot of products around computer vision, really a horizontal platform. We had our first APIs up in 2014 serving inference traffic. We started by building what people call foundation models today for computer vision, trained to recognize tens of thousands of things in the world so you don't have to train them yourselves.
8:28and we built a few of those for different use cases like wedding photos food our general one is the most popular still to this date general purpose like social media consumer photos all that kind of stuff content moderation filtering drugs nudity weapons another big popular use case and then in 2016 we kind of allowed the platform to let people train their own models label data and train their own models. And so people could fine-tune it for their particular use cases. And shortly after that is when the government actually approached us to be part of Project Maven because they wanted to see, can they take a platform like that and have it fine-tuned on intelligence data?
9:11And they got a demo of our product. I still remember this. It was maybe the only customer meeting I had on a weekend. The leaders of Project Maven came up to New York to our offices on a Sunday, carrying a military bag and walked in and had me demo for an hour. And after that, they're like, yeah, you definitely need to be part of this program. And we were under contract within a month of that. Still our fastest government contract since, actually. So it was a great experience with the original leaders of Maven because they took the approach that an enterprise takes on the commercial world. They evaluate different alternatives.
9:51They do their research. They find the companies that are doing cutting edge stuff. They talk to them, evaluate them, and then pick and get them under contract quickly so that they can see the value from it. And so that was our experience with Maven. And, you know, working with the military, it did raise a few questions internally. And there was other companies, famously Google, that was part of Project Maven, that had similar conversations that happened internally. And I still remember, it's probably like 2018 when I had an all hands that I just told the company that what we're doing in projects with the government like Project Maven is ultimately going to save lives using artificial intelligence to make decisions faster and in a more intelligent way than humans can do it.
10:40And because of that, we're going to be focused on this as a company. And some people didn't like that. and one or two quit after that all hands. But what it did was stop all of the gossip, stop all the debating internally and got us all aligned. And a lot of people after that all hands, many more than the number that quit, said, oh, thanks for doing that. Finally, we're like past this. I'm like really excited to be working on these kind of programs. And I totally agree. It's going to have a big impact. So that was a big, big catalyst to us continuing. and in Project Maven a lot of people think of it just as vision enabled drones but it was much broader than that right can you describe Project Maven and what happened to it I've lost track of it it's uh it's still going I was literally just drove over from a meeting uh with a customer right before this and um I can't go into all the different things that it does but it broadly became the first AI program within the government and definitely the first successful one.
11:48And because of that success, it gathered other lines of effort, all different types of sensor data you can imagine the military having. They want to apply AI to, again, accelerate those decisions and make better ones. And so it started with drones, but it's gone to many different directions from there. Yeah. And it's, it's visual data or is it anything signals intelligence or. Yeah. I mean, it started with visual, but now there's discussions about every, everything you can imagine at this point. And that includes, you know, with, with Gen AI in the last couple of years, becoming very popular and useful large language models.
12:28What can you do with them in the intelligence scenario? That's part of the discussions with, with Maven, but also with many other programs across the, the DOT and intelligence community. Yeah. And, and the output from, from the Maven platforms is for what, what's the primary use? I remember early on there was a lot of talk about targeting and, but, but then there was a lot of sort of, pull back from that by the government saying that it wasn't targeting. It's really for intelligence analysis. Where has that gone? I mean, where does it lie now? Is it still with the Army? Is it or with the joint, what is JAIC?
13:31I forget what the A was, but yeah. Joint AI Center. Yeah, yeah, yeah. Joint AI Center. Is it sort of a joint program across the services? And yeah. No, it's actually transitioned to NGA. So it's part of the intelligence community now. And that maybe hints at kind of the purpose of it. Less so, you know, automating target recognition, but providing just better intelligence. Yeah. And so that's ongoing, Clarify. And then you also had products for law enforcement. Are there discrete products or were these, was it a platform? Did Clarify have a platform or APIs that different agencies would employ and cobble together a system on their side?
14:27Or do you have a product that agencies, whether it's law enforcement or not, can use to scan faces, for example, or whatever the inputs are? Yeah, so everything is one platform. And as I kind of mentioned, it started with just inference back in 2014. The reason we started there and having an API platform was inspirations like Stripe and Twilio, where they built a great developer product. And people just built amazing things on top of that. And they became kind of the engine for that. And we saw that there needs to be this engine for AI. And so that's the same product on the commercial side as public sector side.
15:22We intentionally didn't want to build two different products as a small company. And what that's allowed us to do is, in order to do that, we had to focus on making the overall platform really flexible in its deployment options. So when you deploy with a lot of these government agencies, they typically want to go on-premise, even into air-gapped environments, on classified networks that aren't connected to the internet. So all these different parameters go into our builds now so that they're flexible enough to go in those environments all the way up to AWS, Google, Azure, etc. And that really helped us build a much better product at the end of the day in handling all those things.
16:08But it also focused us on having just one product to take it to solve a given problem. there's usually collaboration with the customer and our professional services team who can customize things in the product, customize the AI models with the customer. And the platform's easy enough for the customer to do this as well if they choose to so that the AI is learning how the customer cares about recognizing or generating what they need. So flexible enough to deploy everywhere and flexible enough to kind of cater to the problem at hand. Yeah, but it's a developer's product, an API into your system that they build into solutions.
16:53Yeah, exactly. And on top of the API, we have really simple to use user interfaces. So, for example, data labeling is a core component to building AI systems. And that requires a lot of user interfaces to draw bounding boxes or draw per pixel masks of objects. And then there's UIs for the task management. Like, what do you want to label? Who is the workforce that you want to label it with? Is it complete? What's the progress? All those kind of experiences, they're not really good as an API only. So we intentionally built a whole suite of UIs that make the system easy enough for anybody to use.
17:31And because of that, we've had even intelligence analysts train models themselves, doing all the labeling, picking a training template, looking at evaluation metrics of how they perform without our teams being involved. So that's key to making the API successful, which is ultimately what you're going to use. Once you have a good model, you're going to use the APIs to fire through lots and lots of data for inference. But to get to that model, the UIs are really important. Yeah. And then you sort of shifted into inference as a service. Is that right? So I would argue we started there and now we're, doubling down on it again.
18:16Yeah, but in... Yeah, can you talk about that? Why is... I mean, I've had Rodrigo Liang and Andrew Feldman, you know, that's Seminova and Cerebris. I haven't had a guy from Grok on. We missed appointments a few times, but that's on the hardware side. speeding inference through hardware and you're doing it through optimization in the code is that's right is that right and so that you can beat these inference chips or these specialized accelerators without any specialized hardware just simply by optimizing the code first of all yeah why is inference such a bottleneck uh because it feels fast to uh to consumers and then uh yeah how how did you go about doing this without and and and does this would would your inference uh strategies just accelerate inference on a platform like cerebris uh that much faster that much more?
19:35Or is a fork in the road either you go down code optimization, algorithm optimization, or you go down the hardware route where you can't have both? Yeah, great question. So inference is a good segue from the last conversation. It's kind of the production use case. Training the model is like a precursor to the inference. Really what you want is an intelligent system that you're going to run over and over, get accurate results, get them quickly, or get them at large scale of batches of data, and do that in a repeatable, reliable way. And that's really where the company started back in 2014. With our first models, we didn't let anybody train in the platform, didn't let them label or anything.
20:22It was just inference. And back then, even going into my PhD, we were one of the first people writing CUDA kernels for AI. I started getting into GPU stuff definitely in 2012, maybe even 2011. But at that time, it really started out at Stanford. There was some work from Andrew Ng's lab. Then Toronto picked it up with Jeff Hinton's lab and people like Alex Krzyzewski writing some really good kernels. And then there was only four labs doing AI back then. There was those two, NYU, where I was at in my PhD, and Montreal with Yoshua Bengio's lab. And so all four of us were collaborating on the use of GPUs.
21:06And once Alex had some really good kernels, we all shared them. And so I remember distinctly in my PhD, I first adopted them for my experiments. And overnight, they were 30 times faster. And that just changed the game. Instead of waiting for results for a day, I could go for lunch and come back and have my experiment done. And so thinking back, I was probably one of the first 20 people in the world actually writing CUDA kernels for AI, which was pretty cool. And that kind of deep expertise was the genesis for Clarify. We had to build toolkits before PyTorch and TensorFlow existed. We had to build GPU operators in Kubernetes before NVIDIA was building operators to support GPUs within Kubernetes.
21:53So kind of all the layers up, we have expertise within Clarify on how to build these systems. And so that realization in the last year and a half was that everybody's still struggling to run these big, large language models that have become very popular since ChatGPT. So why don't we take a lot of this learning we have and a lot of the internal tools that we've built over the years to run models efficiently and at very large scales and open them up as a new product line. And so that's what we launched recently as AI compute orchestration. And we call it compute orchestration, not just inference, because inference is just the first workload.
22:32You'll see more coming for training, for agentic tasks that run in the background for a long period of time, all coming soon. but inference we view as the most important one because of the scale it operates at and because of a lot of the optimizations we can do it benefits inference really really well and your second question was do those optimizations work only on gpus or can they work on other things they can work on other things depending on the optimization so we have a mix of low-level things like CUDA kernels that we optimize. We have things like moving Python code to C++. There's a variety of different tweaks we do, but there's also tweaks that will work across accelerators.
23:16We're predicting usually many tokens in advance using good speculative techniques, and that should work on any kind of accelerator. A Google TPU, a Cerebrus, a Wafer, whatever you may throw at it. So it would be a mix. And we found that it depends a lot on the use case that you're using these models for, how well some of these accelerations can work. So this year, especially, there's been a lot of excitement about reasoning models and agentic tasks. That's really what we built the optimizations within the Clarify reasoning engine for. because if you think of how these models reason, they're actually repeating themselves many times.
24:03They're referencing a lot of stuff in your large context a lot of times. And so there are certain optimizations you can do for those kind of repeatable agentic tasks that perform really, really well. But they won't work as well on every general task. Yeah. And to unpack some of that, I get lost in these conversations and I have so many of them, I'm familiar with the jargon, and I don't want to slow guests down and asking them to explain things that the more advanced listeners already know. But let's explain to people what a CUDA kernel is and why it's important. Yeah, go ahead. Yeah, so CUDA is NVIDIA's programming language for their GPUs.
24:55I don't actually remember what it stands for off the top of my head, but it's been around for like 15, 20 years almost at this point. And it lets you, you know, there used to be languages for gamers to program GPUs to make games, but they're really low level like shader languages. And so CUDA became a general purpose language, just like you can program any software on a CPU. You've been doing that for many decades. This was like the equivalent for a GPU. And GPUs just unlock much more parallelism than a CPU offers because they have thousands of cores instead of tens of cores. And that gives a lot of advantages for neural networks in particular because when you're doing the matrix multiplications that are common in a neural network, it's highly parallel.
25:44So GPUs are a great fit. Yeah. But then kernel is a function. Yeah. Yeah. And so the kernels are an example is a matrix multiplication kernel. So multiplying one matrix by another, that would be executed as a single kernel on a GPU. And so when we're talking about CUDA, we're talking about NVIDIA. AMD has an equivalent now of CUDA called ROCCM. So there's ROCCM kernels, there's CUDA kernels. And that's kind of the lowest level. And it builds up from there. A lot of people today don't have to operate at that level because NVIDIA and AMD and Cerebris, Grok, etc., whoever's making the hardware is usually making really optimized kernels for their hardware so that people can just use the software on top, like PyTorch, for example, which implements the best kernels for NVIDIA GPUs, for AMD GPUs, and other hardware providers.
26:44but when you're trying to push the limits of optimization sometimes you can find techniques at the kernel programming level that are even better than what the hardware providers writing themselves yeah and i mean this is essentially what deep seek did right uh yeah yeah yeah deep seek did a lot of really clever engineering down to the kernels all the way throughout The systems engineering, everything. And how much of what... So the hardware producers like Grok, Simba Nova, Cerebris, who are focusing on inference, their point is that you don't need to go in and tinker with the kernels in whatever the programming language of that particular hardware is.
27:40We've done that, and we provide the super fast inference that you can access through API into our cloud or on promise if you want to do that. In the DeepSeq model, it's not the hardware. You could use any hardware, presumably, but they're optimizing inference on the software side. how influenced were you by DeepSeek? Was DeepSeek just the first to make a splash in, you know, the zeitgeist was already working in that direction? Or did people see what DeepSeek did and say, yeah, that's right, we can do that. Why aren't we doing that? well we were working on computer orchestration about a year before deep seek came out just because the general trend was that open source is uh catching up to the proprietary models and i think deep seek was just a huge catalyst to open everybody's eyes in that same direction and they did a lot of great work on the deep seek team and they published a lot of it too open source GitHub repos, good technical reports.
29:01So kudos to them. And they're not the only ones. There's been a lot of great work out of the Alibaba teams for the Quen models. There's the Kimi K2 models, another good one out there. A lot of great open source models are coming out of China. The one we've focused the most on is OpenAI's GPT OSS. And the reason for that is, one, it's built by an American company and we do a lot of work with public sector so we can deploy that model on premise without any question marks but even more important it's a really efficient model from a compute standpoint you can run it on a single GPU and get really high throughput it and because of that you can offer it for really competitive pricing and it's very intelligent so kind of that combo of cheap fast and intelligent is like the the trivecta i guess uh that everybody looks for and i i don't think anything's even close um from those dimensions uh some may be more intelligent but then they need eight gpus uh and that costs a lot uh so it's uh it's a great model and and we're the fastest in the world at running it because of all the focus we put onto it Yeah.
30:15And so you're offering it as a service or I mean, yeah, it's not open source. The code is not open source. Correct. Yeah. So the way our computer orchestration works is we have some of these popular models like GPT OSS offered as a serverless option. so you can pay per token. But if you, and this is what most customers end up doing, if you like a model or maybe you have your own fine-tuned version of it or your own custom foundation model and you want to run with some of the optimizations we have, you can deploy into dedicated node pools, we call them. And these are pools of the same type of instance that can scale up and down dynamically based on your traffic needs.
31:00And they're dedicated to you. So you don't have other customers making requests in there. You're at full control to deploy whatever you want into them and control how things scale up and down. And so we have both the serverless token-based pricing and the dedicated, which is priced by GPU hour, essentially, options. So lots of flexibility in how we offer up these kind of models. And that dedicated GPU pool or whatever, is that on your cloud? Do you offer a cloud or is that in any cloud and you have ways to ring fence the GPUs? Yeah. So we actually offer both now. So in our VPCs, you can provision, clarify, compute.
31:48and they're in our overall cloud VPCs that sit across multiple clouds. You can also deploy into your own VPC or even your own bare metal Kubernetes clusters. So wherever your compute may live, you can run. And the big advantages of that are you get that kind of flexibility. So maybe you start on premise because it's going to be the cheapest if you have the upfront capital to buy machines. But you're kind of by definition at a fixed size when you're on-premise. And so you need a way to spill over to maybe a NeoCloud as the next option. And then even NeoClouds have a limited capacity. If you're really taking off, you may have to spill over to a hyperscaler or multiple hyperscalers.
32:32And so you can do that all within the Clarify platform without having to retool or re-engineer things for every cloud or on-premise differently. You get the same kind of performance on chips wherever they may be. and that includes not just nvidia chips but amd and uh we're working towards support for google tpus uh kind of as we speak so there'll be other accelerators offered in the near future as well yeah and you were saying that that this can run on any uh hardware so it could run on saminova or cerebris but does it accelerate their inference speed the does it accelerate inference beyond the speed that they offer?
33:16I mean, is there a cumulative effect or is it, you know, it's going to be what you guys can manage or it's going to be what Cerebris can manage on its own? It's not going to be some multiple of that. So those platforms we haven't focused on yet. I'd love to explore that with each of them. but in our experience with new chips there's a long lead time for the software that they provide to catch up and for especially for other third parties to implement optimized software and so i think google tpus are at that point now but they've been in the works for like a decade and and other platforms i think are really early on how well the software stack will allow other third parties to uh to write good stuff with cool has been at it for a long time amd's caught up with rockum um but the others it would be hard i think for for somebody like us to squeeze uh performance out and i think that's actually why you know you see people like cerebris they used to sell the wafers um a lot that was their business now they're focused on like serverless we're just going to host the models and be really fast at it and they are they're great but I think they found it difficult for people to write the software for the hardware that they were selling.
Read the full transcript
34:40Yeah, exactly. Um, and so where does this go? Does, how does this change your business? Does it just extend your business? Um, yeah, I think it, uh, accelerates it because I think the, what we found with computer vision is the, and we're still named leaders in computer vision to this day, like Forrester named us leaders last year in their computer vision platform report we even surpassed everybody like google aws so uh we still kind of wear that crown which is awesome to see but with computer vision uh we find that the use case is always very specific and you need some of those professional services to train for a given customer so a good example is we've just signed a deal last quarter with a nfl football analytics company and they want to use our platform to do all the sports analytics that they're doing, but they need our PhDs to help them train the best possible models to do that analytics very accurately.
35:43Whereas with large language models, it's kind of the opposite end of the spectrum where the one-size-fits-all model is so general-purpose that it can be prompted to do so much that the game there, the focus there is if you can build really performant, really reliable systems that like our clarify reasoning engine is doing now it's how quickly can you and cost effectively can you offer it to customers and so we once we launched our benchmark results on artificial analysis showing us as leaders in the gpt oss model we closed the inference deal in five days i don't think we've ever closed a computer vision deal in five days.
36:25So if that's science to come in a repeatable way, we're very excited about that new offering. Yeah. You mentioned token throughput. Is that the key to inference speed? There's a couple important dimensions. And the throughput is one of them, number of tokens per second the other is usually time to first token and there's even measures different measures of that one which is time to first token period or there's time to first answer token and that means with these reasoning models they do a lot of thinking before they respond with a final answer and so the time to first answer token is kind of combining both time to first token and throughput because if you have to go through all the reasoning quickly you have to have high throughput, but then time-to-first token is kind of when reasoning starts.
37:24And we're doing really good at both. And that ultimately gives you the time-to-first answer, the final token of the answer, which is really fast. And we've had some customers, after they saw these benchmarks, they're like, let me try this ourselves. And they compared us against other popular inference providers out there who really focus on similar things, kernel optimizations. They tout themselves as being fastest at inference. And the customers came back and said, we're 65 % lower time to first token, 11 % faster output token per second. And that resulted in 40 % faster time to answer. And they signed up within five days.
38:05So we're really excited to see the real world performance because it's going to depend again on your prompt, your data, the model you choose to run. There's a lot of variables in overall performance. And so we urge every customer to test it out and work with us to see if we can push it even further. Yeah. And again, for the lay listener, a token is a chunk of data that's roughly part of a word. and the throughput is how quickly the model can process those tokens. And that's a matter of computation speed, right? That's right, yeah. So when you're talking about text, a token is usually like three to four characters of text, sometimes bigger, sometimes smaller, depending on the token.
39:00and some of these models have a vocabulary of like 200 ,000 tokens so that's kind of the options that it gets to choose from every time it generates takes one generative step and in terms of our like optimizations with GPT-OSS we're talking about doing like anywhere from 500 to a thousand of those steps per second and that number of tokens are being generated per second so it gives you the sense of, you know, every little microsecond, actually, when we look at our profiles matters and being able to generate that quickly. Because if you have to do something a thousand times every second, a millisecond is all you have.
39:42Yeah. And does that affect how well, because context windows, you know, the number of tokens that that you can provide a model in as input has been growing dramatically, but, but models get kind of lost, uh, when the context window becomes too large. Um, is that, uh, does the speed of token throughput, does that improve the model's, uh, ability to handle large context windows or is it your view? It's, it's the same problem it's just seeing it faster right it would still have the same problem that's more of an architecture choice training choice on how to get that uh they call it the forgetting kind of problem um uh but with some of the optimizations we've done it actually helped a lot on the long contact speed um like we improve long context processing like 100 000 tokens or more by 5x from where we started um so uh the default implementations you get from some of the open source models and open source toolkits that run them is pretty inefficient in those long context scenarios um but if the model is still going to forget it's uh we we're kind of optimizing performance not that would be an accuracy problem i would say yeah yeah well i've been talking to a lot of people about uh sort of uh post transformer architectures to deal with that problem in particular state space models like mamba i don't know if you know the manifest people in new york they're coming out with some interesting models anyway that those um state space models you work with those uh and and these you know sort of post transformer iterations that attempt to deal with the with the context window problem and and forgetting and all of that yeah i mean gpt oss even interleaves different types of attention mechanisms so attention is usually the operation that has to look it can become really slow with the long context so a lot of the focus is there on things like Mamba and others, but GPT-OSS interleaves different types.
42:12Some of the QN-NEXT models that came out recently are doing some interleaved hybrid attention, they call it. So yeah, we're familiar with a lot of those. They help a lot. I think that's why GPT-OSS in particular is so fast. But I haven't personally dove into these state-space models, but I'm optimistic that Transformers is not the final architecture. And the reason I'm interested in this, if you can handle massive context without the model getting lost or forgetting or missing information, then you don't really need like a vector database or these other external knowledge bases. You can put it all in the context and have the model do inferences on the information in that context.
43:11Do you see that as where this is going? Or do you think there are other architectures that will solve those problems? i do think it's where it's going um but it will probably create need for for different architectures along the way um yeah i think the there's a growing need for like think of um use cases like i want a personal assistant who knows everything about me like every conversation every meeting every tool i've used uh that's a lot of data and and it never stopped growing so So that's kind of the direction we're moving towards. And it's going to necessitate a lot of both engineering and novel AI architectures.
44:01Yeah. What are some of the use cases that you've seen for the super fast inference where inference speed really makes a difference? I mean, of course, everybody wants things to go faster, but what applications do you see this being critical for? so time to first token and really the time to answer come in especially with these chat kind of real-time experiences and we've all experienced it whether you're using chat gpt gemini uh clod whatever your favorite is uh the longer it takes the the less you want to use it and i've actually like moved to to gemini flash because it's very fast and uh and very consistently fast.
44:45Google does a really good job at serving these models. And so the response speed for like every day, I'm in the middle of something, I'm like, ah, stuck in code, I want to ask, is this possible or whatever, jump over to Gemini, I get an immediate answer, and I'm back at my task. If you have to wait even a minute, you're distracted, you're contact switching in your brain, and it's going to slow you down. So the latency is important for those kind of use cases. for a lot of the agentic workloads. So I remember having this conversation with our investors a couple of years ago that about six months after ChatGPT came out that I said all the future stuff is not going to be in a chat window.
45:32It's going to be agents working in the background. And this year, I'm happy to see the shift, especially on coding tools to becoming agents that work in the background. and we're using a lot of GitHub Copilot internally to do that and I think their experience is great. I never took the plunge into things like Cursor where I have to actually change my editor in order to use their product. I just want the AI to do the work and I can stay in the editor that I'm familiar with. And in those scenarios where you're chewing through tokens, they do a lot of thinking, a lot of iteration, but they're asynchronous.
46:12the throughput is the important measure there. I don't care about time to first token because I'm asynchronously queuing that off and it's going to sell me a notification in 20 minutes when it's done. So I think those are good use cases for a lot of the token throughput optimizations we've done. So Clarify has launched this as a service and I presume one of the reasons we're talking is it's growing fast or not. I mean, what's the response been and how crowded is the inference market now since everyone seems to have shifted the inference from draining? Yeah, the response has been awesome and especially that validation that customers as they test it out are giving back to us and we're going to be available soon in many different locations.
47:07So on the coding agent side, we're integrating into tools like OpenHand, which is like an open source agent for coding. We have a pull request out for Vercels AI SDK. We're already in Langchain, LAM Index, popular ones like that that you're used to, Crew AI. And we're going to be on some of these open routers as well, like OpenRouter, where they kind of go across all these inference providers and rank by performance and price. So we're expecting a lot of increase in volume because of all these initiatives to basically be available where the developers are. So we're excited for that. It is a crowded space, I would say.
47:51Totally agree on that. But I think when people look at vendors, they look at, you know, do they have a track record of doing this? And we have the longest track record out of anybody having done this since 2014. we were even out two or three years ahead of any of the hyperscalers with APIs for AI. And so they look at that and they build a lot of trust based on that. And then look at the performance we're getting in these kind of public benchmarks, and that excites them as well because they can do it in a cost-effective way. And then finally, they look at kind of the flexibility and being able to run in our cloud, in your cloud, or your on-premise, and you get to choose, you're in full control.
48:37None of the other inference providers are offering that. So we like our position. We're really excited about how this latest product has come together. And now that it's out there, getting that customer feedback to validate it is great. Increasingly, inference is moving to the edge as a lot of this happens on device. And now we've got, you know, self-driving vehicles hitting the roads where you can't be relying on the cloud for inference. Do you guys work at the edge? Yeah. And even some of these government contracts, we've done work on those. NVIDIA keeps changing the name. It used to be Jetson, Integra, and then Oren.
49:23Now there's something else, I think. But yeah, those low-power NVIDIA GPUs are great. and yeah, we can do all the processing local within them. And edge also means a few different things we found as we talked to different types of customers. So one can classify it as that, those kind of low power devices. Another like telecoms, for example, we're working with a few right now where they have a huge number of data centers to power internet and your cell phones. and realizing that that could be a big advantage in kind of serving up AI traffic if you put good compute in those data centers. And we're talking about like thousands of them versus AWS maybe has like five or so regions in the US.
50:11So it's kind of a couple orders of magnitude difference. And so they can be very local to where you are. So really low latency kind of streaming applications. Like maybe you want video analysis of an operating room in a hospital and getting alerts on that or even driving a robotic arm automatically using AI. Like those types of applications, you really need lots of processing and you need it to be really low latency. So that's another definition we hear of Edge. It's not a low power 10 watt device, but it's really close to you where it's needed. Do you think that right now so little of all the inference in the world is taking place at the edge?
50:55Do you think that that will shift as robotics become more versatile? Absolutely. Yeah. I think the next big wave, a couple of thoughts there. So it would be interesting to know how much inference is being done already on Android and iOS devices. I know iOS does a lot. I have an iPhone and I use, for example, dictation functionality all the time. And it works pretty good. It still doesn't get Kubernetes, which is like the most frustrating thing. But one day AI will catch up. The word Kubernetes? Yeah, it always spells Kubernetes. It's like a person's name or something. But at least it's happening on the device.
51:44So there's lots of inference already happening on the edge on mobile devices, and that'll continue. And every time Apple does an announcement and Google does an announcement, they're pushing the limits of how much capability is in their processors on these devices. So Qualcomm's doing a lot in the AI space now, too. So I think we'll continue to see a lot more in mobile. And then to your point, I think the next big wave in AI is going to be the physical world with robotics. I think it's like five to ten years out before we'll be like amazed by robots. Even the most advanced ones we see out of Tesla are not that amazing right now.
52:26I'm sure they're going to get great, but I don't think it's the AI. Actually, I think it's the physical capabilities, the actuators and all that kind of stuff. I have a long way to go to replicate what a human is capable of. so I'm excited for that but that'll drive a lot of the compute towards the edge as well Is there anything I haven't talked about that you'd like people to know about what Clarify is doing or what you think is going to happen in the future whether we're all going to regret this AI moment Yeah, I mean no, I think we covered a lot we're very excited about the new offering A lot of people remember us as computer vision because we started there and continue to lead there.
53:13But we do so much more in the Gen AI space today. So eager to have more people try it out and give us good feedback. And we'll continue to make things even better. In business, they say you can have better, cheaper, or faster. But you only get to pick two. What if you could have all three at the same time? That's exactly what Cohare, Thomson Reuters, and Specialized Bikes have, since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds.
54:07How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 % less for compute, 70 % less for storage, and 80 % less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com slash IonAI. IonAI, all run together, E-Y-E-O-N-A-I. That's oracle.com slash IonAI.
From the publisher
Try OCI for free at http://oracle.com/eyeonai
This episode is sponsored by Oracle. OCI is the next-generation cloud designed for every workload – where you can run any application, including any AI projects, faster and more securely for less. On average, OCI costs 50% less for compute, 70% less for storage, and 80% less for networking.
Join Modal, Skydance Animation, and today's innovative AI tech companies who upgraded to OCI…and saved.
Why is AI inference becoming the new battleground for speed, cost, and real world scalability, and how are companies like Clarifai reshaping the AI stack by optimizing every token and every deployment?
In this episode of Eye on AI, host Craig Smith sits down with Clarifai founder and CEO Matt Zeiler to explore why inference is now more important than training and how a unified compute orchestration layer is changing the way teams run LLMs and agentic systems.
We look at what makes high performance inference possible across cloud, on prem, and edge environments, how to get faster responses from large language models, and how to cut GPU spend without sacrificing intelligence or accuracy. Learn how organizations operate AI systems in regulated industries, how government teams and enterprises use Clarifai to deploy models securely, and which bottlenecks matter most when running long context, multimodal, or high throughput applications.
You will also hear how to optimize your own AI workloads with better token throughput, how to choose the right hardware strategy for scale, and how inference first architecture can turn models into real products. This conversation breaks down the tools, techniques, and design patterns that can help your AI agents run faster, cheaper, and more reliably in production.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI




