From Prototype to Production: The Best of Inside the Algorithm

30 Sep 2026 · 25 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

A curated “best of” episode on taking AI from prototype to production, what breaks when models face unseen situations, and how to build trust and reliability (including human-in-the-loop, scenario planning, and neurosymbolic approaches).

Guest backgrounds

The transcript references multiple experts across academia and industry; specific names aren’t given. One example is Professor Tom Monk (University of Exeter) researching agents for questioning documentation/models.

Key claims

Prototypes stall because they ignore edge cases, adversarial scenarios, security constraints, and real end-user workflows. LLMs hallucinate due to reliance on text correlations without real-world grounding. Forecasting models fail under structural shocks (e.g., pandemics) because they depend on historical correlations. Trust improves via participative modeling/validation with stakeholders and production-grade controls.

Notable examples

“Car wash 50 meters away” meme where models recommend walking despite needing the car; pandemic, aerospace closures, and a volcano causing sudden demand/operations changes; neurosymbolic systems generating Python/algebra for correct calculations (e.g., arranging oranges/apples).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

From Prototypes to Production: Challenges and Insights

0:45 to 4:33

Discussion on the common challenges faced when transitioning AI projects from prototypes to production.

“But production readiness is actually everything else.”

Understanding Model Limitations and Human Oversight

4:33 to 10:16

Exploration of how AI models perform under unexpected circumstances and the necessity of human oversight.

“Next, let's dig into what happens when a model meets something it has never seen before.”

Building Trust in AI Systems

10:16 to 14:00

Examination of the importance of trust in AI systems and strategies to establish it with stakeholders.

“You'll find the links to those in the show notes below.”

Building Trust in AI Models for Healthcare

14:00 to 15:44

Learn how AI can support regional managers in healthcare to trust and understand models.

“The other regions haven't actually had that experience and it's almost that hurdle that I think is harder to bridge than the technical aspects of rebuilding the model for their specific contexts.”

The Evolution of Large Language Models

15:44 to 16:45

Discover the advancements in language models and their mathematical capabilities.

“The strongest results my guests have described over the season come from pairing language models with solvers, simulations and smaller specialist models, with each part doing the job it's best at.”

Neurosymbolic AI and Human Cognition

16:45 to 19:19

Explore how neurosymbolic AI reflects human cognitive processes in problem-solving.

“They would kind of parse it and they would often produce a Python script or an algebraic equation that captured the mathematics.”

The Importance of Simplicity in AI Models

19:19 to 21:46

Understand why simplicity and production-readiness are crucial for AI models.

“I think very soon it will have tremendous importance in terms of articulating business strategies in this space.”

Challenges for Junior Programmers in an AI-Driven World

21:46 to 23:15

Discuss the risks and challenges junior programmers face with AI tools.

“flagging lines of course across, but yeah, it will, it will help a lot on the kind of like subsequent subsequent phases of like, you know, production, productionization and so on.”

The Future of Jobs in AI

23:15 to 24:38

Contemplate the future of jobs in the age of AI and the role of education.

“around agentic or whatever, it's useful to know a little bit.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:03Hello everyone, this is Jeremy, Chief AI Scientist at Cambridge Spark and host of Inside the Algorithm. Over the last few months, we've had some fantastic conversations with leading academics, researchers and practitioners on how artificial intelligence can be taken from theory to practice, covering topics around trust, turning prototypes to production and what happens when a model meets something it's never seen before. On this special episode of Inside the Algorithm, I've curated some of the best bits from these conversations into a single episode, so you can hear how experts from very different fields approach the same problem.

0:39First, we'll start with the simple question. Why do so many AI projects stall at the prototype stage, and what gets the rest into production? So getting too excited with prototypes and almost believing the prototype is the end production-ready solution, So it's always good to understand and contextualize that a prototype has been built within a certain context to cater for specific use cases. But production readiness is actually everything else. All the edge cases, all the unforeseen scenarios and things that might arise, or even adversarial attacks, etc. So that is what production readiness is all about.

1:23Yes, so this is a question which is interesting not just from a scientific and from an engineering perspective, but also in a way from a business perspective, right? Because we started and we completely internally developed this project. So we have to, first of all, demonstrate that it works, having a working prototype, demonstrating that whatever process we develop, they're executable in real time or in near real time. At that stage of the project, you rarely can say, just buy me a platform to do that. We have to use the tools and the environments and the components of architecture that we already have.

2:11Then probably at the next stage of the project where the prototype proves that it works, The algorithms prove that they are efficient. And actually, there is some level of demonstrated business benefit at that point. One of the difficult points was actually, and it took some time, getting an architecture design approved by security. because obviously there are a lot of firewalls and then security measures put in place on the manufacturing side so that external environment cannot be interfaced easily with process-critical systems. So we had to do a lot of work to kind of, even with our prototype architecture, we had to face some of these serious technical aspects.

3:07I was mentioning before about the so what. So I think it's really important to understand, like, you know, who are the end users? How are they going to use it? How is this going to be surfaced? And so on. It's, you know, if the output is going to be just in, you know, in a folder where no one can access, then what's the point of doing the solution? Because then no one is going to use it. And so it needs to have, like, that mindset of, like, actually thinking, like, you need to start talking to end users, like, early on, even, like, you know, at the first week of like you know when you are trying to to see the the and also to understand like whether do they actually need this solution or actually you know the definition of dawn or like what they were trying to predict even like the definition of like the the outcome of the prediction you know like what um is it like passenger volumes or actually is it aircraft number you know like different like these aspects and then i think as well there is there is an aspect that i see sometimes which is like how much people can absorb change or like want to absorb change as well um i think um in some of the projects that you know we have the most success is also like some of the projects where we had kind of like champions within the business within our clients business and kind of like really believe on like you know these are solutions that actually could make a difference into their business.

4:33Next, let's dig into what happens when a model meets something it has never seen before. Where do these systems break? Where do hallucinations come from? And why does a human expert still need a hand on the controls? Yeah, I mean, most interestingly, you know, humans, when they often make statements, they have a sort of world model in mind, right? So, for instance, if you were to ask me what's the weather going to be like in Edinburgh tomorrow, I'm basing it on some data, of course, you know, the number of years I've lived in Edinburgh. But also, you know, I look out of the window, get a sense of the sunshine and so on and make a judgment.

5:13And with the LLMs, it's often the case that, you know, it's a completely linguistic correlation. So we know, for instance, you know, often when people say it's cloudy, we expect rain. If it's sunny, we expect warm weather. So it looks at these correlations and utters things based on that. So to a large extent, it seems, and this is very surprising. I often tell this to people that the fact that LLMs can do what they can is very surprising, given that they're only looking at correlations in text. So that's shocking that it's so good at all. But I feel like that's where also the main limitation of these things lie, because they don't quite have a way to grasp or ground all of their understanding in some sort of real world structure.

6:06And that's when you see very strange hallucinations. A classic example is there was a viral meme going around where they asked a large language model, you know, I have to get my car washed, which is 50 meters away. What's the better thing to drive there or walk there? And many large language models, even frontier models said, just walk there. But of course, you know, without my car, what's the point of going to a car wash? You know, that kind of thing. So it's very surprising where it feels and also impressive when it works, you know. Okay, so that's really good. So you mentioned the pandemic, I'm really keen to come on to this.

6:43Obviously, there's something which was completely unanticipatable, really, when it did strike in 2019, 2020. What happens to a model like this? And how could it be used or repaired, I suppose, in a scenario where you have a change that's so dramatic, that's so catastrophic, really, for that industry? certainly for a few months what happens to that kind of model what place would it have into trying to aid a recovery in that in that situation i think yeah when when something like this happens like it's almost like a structural change right um and i think it completely kind of like breaks the core assumption of of the mathematical model um and it kind of like exposes in a way exposes kind of like the mathematical kind of like foundations and limits of the yeah of the of the forecasting.

7:36I think in a way it breaks the historical correlation like econometrics and machine learning models they are fundamentally built on the premises of what happened in the past and kind of like finding those correlations and patterns between variables so it is proven that a percentage increase on household income usually translates into a certain raise in passenger demand but then when something happens like this, like a pandemic, then kind of like basically it just doesn't drift the behavioral kind of like relationship. It just like breaks everything and it snaps like, you know, and overnight almost.

8:17And again, like I think in the limitation of like, because we are using historical data and at the end of the day, probably this is a situation where the model has never seen kind of like that situation. No, right. Of yeah, the boundary limits, let's say. so I think that's why going back to you what if it's scenario planning I think that's the way we need to think in practice of how these frameworks should be used so it's kind of like it's almost like as well on the people who are building those models and how quickly we can react to the context changing and use those tools that we have regenerated to actually make better decisions when something like this happens so yeah so like use the framework to generate you know like different outcomes based on on different levels of economic variables or even like you know like kind of like a stress like stress testing like your models of like you know like extreme situations such as like pandemic and aerospace closure and there was the volcano as well for example like right like years ago that close their aerospace as well from one day to another and so on.

9:33And I think the other aspect as well is probably like keeping a kind of like human in the loop. So I kind of like almost like accepting that, you know, like, yeah, we build the solutions to kind of like help us to predict the future in a way. But I think we shouldn't be removing the expertise and domain expertise judgment from those tools and from the math. They complement each other. I think ultimately it's not the limit of any forecasting model or any kind of a model, etc. It's just assuming that we will do everything without the human context. so I think there is always for me there should be always like that that human expertise and you know I hope you're finding this conversation as fascinating as I am, if you're getting value from the research and technical depth we dig into on Inside the Algorithm please take a moment to follow the show and leave us a review on Apple Podcasts Spotify or YouTube every follow helps us reach more of the brilliant minds doing this work and keeping the rigorous substantive conversations going You can find Inside the Algorithm on the Data and AI Mastery podcast feed, or watch it on the Cambridge Spark YouTube channel.

10:51You'll find the links to those in the show notes below. Right, let's get back into it. Next, let's focus on the topic of trust, and how do we find it when working with what sometimes feels like a black box of reasoning? Yeah, I mean, I think things have materially changed with GNI. I think with machine learning, with traditional machine learning, the risks were well understood. That's quite ingrained within many of the largest organizations in the world in themselves predicting risk and that kind of thing. But yeah, it's more from a GNI perspective, the non-deterministic nature of the beast, right?

11:36And that's brought in another level of concerns around risk that they might bring or that the use of LLMs might bring into the mix for all the right reasons. So yeah, there's a whole trend and it's probably starting to fade out a little bit around, you know, this is too risky to implement, etc. But that's a natural fear to sort of not understanding what is possible or how you can actually circumvent the non-deterministic nature of LLMs by adding the right controls and setting up the right evaluations and metrics and measuring what really matters, right? But then, as you said, on the other side of the spectrum, there's a lot of excitement around building prototypes.

12:34And even more so these days, it's really, really easy to build a really good looking prototype. You know, you might almost mistake it for something or some people might mistake it for something that is ready for production. And that's also a risk. So I think there's somewhere in the middle where you have to be cautious around, you know, understanding how LLMs actually work and how you have to work with them the minute you integrate them into a certain application or a pipeline or an agent and also ensure that you build for production grade capability. We started in the Midlands, who were the people that commissioned the model, but we've rolled out to London and to the east of England.

13:19And there are other networks now that are coming on board. So we're hoping that we will reach the whole of England in the next year or so. But the problem is when we're going to other regions, they haven't been involved in the model building process. so where the midlands were there throughout the workshops that we did we did a participative modeling approach where that we had workshops where they were with us when we mapped the whole system and they understood that we understood the system so that there was the the trust between the modeler and the the client was effectively built and then they also came with us through the validation process because we had workshops on that to show them this is why it works this is why you should trust it.

14:02The other regions haven't actually had that experience and it's almost that hurdle that I think is harder to bridge than the technical aspects of rebuilding the model for their specific contexts. Do you see a role for AI machine learning in supporting that bespoken process or indeed that trust process in terms of helping the regional managers believe that this is a really complementary service that you're developing? I think that's where I'm really interested in research at the moment, like solutions that can help people understand models, almost question models. So I've had some conversations recently with academics at Exeter.

14:55So Professor Tom Monk's there. he's been looking at potentially using agents to allow people to question documentation and that type of approach so that they could say or even question a model like what if demand for kidney replacement therapy rose by x amount in the next couple of years and then have the simulation run and produce those results for people I think that would be really powerful because I don't think we underestimate the amount of other things people in the NHS have to do so anything that we can do to make it easier for them to interact with tools that can help them is a win and it's worth doing but it does it will it will take development time and I've not got a clear picture of how that's going to look or how it's going to work best at the moment but that's why yeah actively trying to do research here.

15:49The strongest results my guests have described over the season come from pairing language models with solvers, simulations and smaller specialist models, with each part doing the job it's best at. Let's hear how that works in practice. Yeah, thanks, Jeremy. I'll say a bit about this Sweden talk, because maybe it connects to what people have been thinking about for the last few years, which is large language models. And again, as we both discussed, you know, these are very impressive, do amazing things. But if you look at the GPT-4 or pre-GPT-4 era, what you often see is when you ask these GPTs for calculations, for instance, what does your mortgage look like in five years?

16:32It often would make up numbers. So what changed then is that many of these systems from OpenAI, StratGPD to Anthropics, Claude, they started taking this approach where they would take your natural language, They would kind of parse it and they would often produce a Python script or an algebraic equation that captured the mathematics. And suddenly you started getting answers that made sense, including combinatorial problems like, you know, I have a bag of six oranges. You know, how many different ways can I arrange them such that, you know, when I pick six oranges and four apples, how can I arrange them?

17:10Or what do I get if I pick one at random and so on? So there's an algebraic substrate now for many of these systems, and suddenly they started becoming more reliable. And that, to me, is sort of the key manifestation of this neurosymbolic AI, where we're using large language models purely as a textual understanding and parsing unit. But we understand very well that this sits in a sort of a hierarchy where, as a cognitive function, we have to get agents to plan. So maybe you're interested in taking a holiday to Hawaii. And it doesn't, you know, it's not enough to know what people have done before.

17:50You have bespoke constraints. You say, I want to go in July and come back in August. You know, I want to find a place where my kids can go to. So all of these constraints start fitting in. And that's when you see this neurosymbolic split where there's some bit which is purely neural and it's done very well using deep learning or large language models. But there is a bottom unit where this kind of manipulation takes place. We move things around and get to answers. And this, you know, is sort of interesting going back to human cognition. So in the 1600s, bizarrely, Leibniz, you know, postulated that maybe human thinking is a kind of algebraic system.

18:33And also really understanding and observing the progress of the reasoning within these new agenti capabilities. So what is the point to which reasoning can reach in order to give you benefits? And what is the point beyond which actually the large language model can be an orchestrator of many even more classical tools, such as solvers or forecasting engines or more traditional simulation engines. And then balancing the benefits between quality of the solution and automation of the solution. Because understanding deeply the technical aspects of that, I think very soon it will have tremendous importance in terms of articulating business strategies in this space.

19:28So, yeah, these are two topics that I think are quite interesting and quite relevant in this space. I think there is a shift towards small, highly specialized models, for example. So rather than using large language models... For everything. For everything, yes. Maybe the real value actually lays into those highly efficient language models that you can fine-tune for a specific problem. like clean specific kind of like domain data sets. They are faster, they're more cost effective and you know, like then it's better, right? Win-win. Better for production as well. As AI tools take on more of the work, what does experience still give you?

20:11And how does someone early in their career build it? You'll hear what my guests look for when they hire, what worries them about junior engineers and where they'd point someone who's just starting out. And that 80 % accuracy model that it's already running production then it's better than a model that might take six months to develop and gives you 95 % accuracy, but then it's more complex or really working production and so on. So it's always about thinking agile and think about what delivers quick wins and then you can still develop in parallel, but at least you have something that is giving you value already.

20:52And then simplicity over complexity. A lot of, you know, like all our clients at the end of the day want to understand actually how the predictions work. And, you know, it's much easier to add explainability layer into a linear regression model than, you know, neural networks and so on. So, again, like simplicity. And then I guess like that mindset of like production first, kind of like ready mindset. So like, although, you know, we might look at like developing a POC or MVP, et cetera, and it's a bit like more research look like, research like we, you need to develop the code and, you know, you need to think about like, you know, this is going to be in production.

21:40So like, think about like testing, think about like, you know, clean code, modularized code and so on. So it's not just like a blank canvas and, you know, like, yeah, flagging lines of course across, but yeah, it will, it will help a lot on the kind of like subsequent subsequent phases of like, you know, production, productionization and so on. Yeah. Good question. I honestly don't know. I mean, I think we are running the risk that we are not giving an opportunity to the younger programmers to learn the skills of the trade. So, you know, senior people might really bootstrap their productivity to some extent with these tools, but a junior programmer is going to struggle.

22:24And there's risks also from a psychological point of view. So you might be well aware, you know, if you run a cloud code or one of these tools for a few times and you see everything is making sense, you stop checking, right? You just have this sort of cognitive dissociation because you're exhausted by approving, just clicking approve all the time. So it really becomes an opportunity to rubber stamp things that look kind of right, but you haven't bothered to check. And if this happens a lot, we're going to see huge issues in the quality of the software coming out there. Every undergrad is thinking about these questions.

23:05And what I would say from a practical point of view, you know, you can't quite completely ignore what the kind of hype right now. So if there is hype around agentic or whatever, it's useful to know a little bit. But I would say if you're forward thinking and looking after your passions, it makes sense to then think about, again, if you're approaching it from the point of view of being interested in AI, go back to the discussions we had around cognitive science, cognitive function. What does it mean to have intelligent behavior? How do we characterize it? And if you see the history of AI, there's a complex and rich field going to mathematics, game theory, cryptography, agent-based modeling, multi-agent systems, robotics.

23:52And there's so much variety there that you don't have to pigeonhole yourself to one area. And most likely we'll see a deeper integration of these as we go forward, including in a theory of mind and interoperability across systems. So it's useful to keep that perspective and to think about if we don't buy into the idea that all jobs will be made redundant. And I don't think it's going to be that case. And I certainly hope people are not pushing towards that. Then what's the thing that would solve problems with society? And maybe that's where they need to look.

24:33Thank you for joining us on this special episode of Inside the Algorithm. We'll catch you on the next one.

From the publisher

👉 Discover how Cambridge Spark helps organisations build the data and AI capabilities needed to turn strategy into measurable impact: cambridgespark.com

In this special compilation episode of Inside the Algorithm, host Jeremy, Chief AI Scientist at Cambridge Spark, brings together the best moments from a season of conversations with leading academics, researchers and practitioners. Guests from aviation forecasting, manufacturing, NHS health modelling and neuro-symbolic AI research take on the same core questions. Hearing them side by side shows how experts from very different fields approach the challenge of taking AI from theory to practice.

The episode covers why so many AI projects stall at the prototype stage and where models break when they meet something they've never seen. It also covers how trust is built in systems that can feel like a black box, why the strongest results come from hybrid systems, and what experience still gives you as AI tools take on more of the work.

Key Takeaways

  • Production readiness means handling edge cases, unforeseen scenarios and adversarial attacks. It also means securing architecture sign-off and making sure outputs reach the people who need them.
  • Projects succeed when teams understand how a solution will be used and what "done" looks like, and when they have champions inside the client's business.
  • LLMs learn correlations, not the world. Their reliance on text patterns explains both their surprising capability and their strangest hallucinations.
  • Historical data has limits. Events like the pandemic can break the relationships forecasting models depend on almost overnight. Stress testing, what-if scenario planning and human domain expertise are essential safeguards.
  • Trust is built through involvement. Participative modelling earned deep trust with NHS clients in the Midlands, and transferring that trust to new regions proved harder than rebuilding the model itself.
  • Hybrid systems deliver the strongest results. Pairing language models with solvers, simulations and small, fine-tuned specialist models lets each part do the job it's best at.
  • Simplicity and production-first thinking win. A working model in production beats a more accurate one that takes months to build. Explainable approaches often serve clients better.
  • Juniors need the chance to learn the trade. Guests warn about "rubber-stamping" AI-generated code and encourage newcomers to look beyond the hype to the wider history of AI.

Useful Links & Resources

  • Inside the Algorithm on the Data and AI Mastery podcast feed
  • Inside the Algorithm episodes on the Cambridge Spark YouTube channel
  • Cambridge Spark: cambridgespark.com

Connect With the Show

  • Cambridge Spark LinkedIn: https://www.linkedin.com/school/cambridge-spark/
  • Cambridge Spark Instagram: https://www.instagram.com/cambridgespark/
  • Cambridge Spark X: https://x.com/CambridgeSpark
  • Host Jeremy Bradley on LinkedIn: https://www.linkedin.com/in/jeremy-bradley/

Which of these clips landed hardest for you: the prototype trap, the pandemic story, or the case for pairing LLMs with solvers? Tell us in the comments, and let us know which guest you'd like to hear a full episode with again.

Visit cambridgespark.com to learn how Cambridge Spark can upskill your workforce in data and AI.

#InsideTheAlgorithm #EnterpriseAI #DataEngineering #MachineLearning #AIDeployment

More from Data & AI Mastery

All 33 episodes
From Prototype to Production: The Best of Inside the AlgorithmData & AI Mastery · 25 min
Listen in VO