In short
The episode argues that modern LLMs’ main limitation is correlation-only “linguistic grounding,” which leads to confabulation/hallucinations when questions require real-world structure or extrapolation. It explains why neurosymbolic/hybrid AI (LLMs used for parsing plus logical/mathematical solvers for reasoning and constraints) can improve reliability, especially for planning and tool use. It also discusses “theory of mind” modeling for agents, and how code-focused systems (e.g., Claude Code) succeed because code has verifiable structure and engineered guardrails.
Guest
Dr. Vyshak Bell, Reader at the School of Informatics, University of Edinburgh; Alan Turing Institute Faculty Fellow; Director of Research and Innovation at the Bayes Centre. Background: 16 years at the intersection of logic, probability, and machine learning; foundations of ethics/explainability and generative AI.
Key claims
scaling won’t eliminate hallucinations; neurosymbolic delegation reduces numeric errors; theory-of-mind via algebra/solvers can help agents respect user intent; risk remains when questions fall outside training coverage; junior developers may rubber-stamp outputs.
Notable examples
“car wash 50 meters away” meme; weather-in-Edinburgh correlation vs world-model; pre-GPT-4 math “made up numbers”; Claude Code generating Python for combinatorics/mortgage calculations; security flaws like storing passwords as plaintext; theory-of-mind queries like “what does Vyshak think about Jeremy’s preferences?”
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Surprising Power of LLMs
0:00 to 0:37
Explore the capabilities and limitations of large language models.
“I often tell this to people that the fact that LLMs can do what they can is very surprising given that they're only looking at correlations in text.”
The Evolution of AI and Its Challenges
1:18 to 2:10
Discussion on the historical context and current trends in AI development.
“In today's episode, we're examining something that underpins almost every modern AI system, yet rarely gets discussed on its own terms.”
Humility in AI Development
2:10 to 4:28
Dr. Belle shares insights on the need for humility in AI expectations.
“If you had to describe the single bigger shift you've witnessed in that time, not just technically, but in terms of the field itself, what would that be?”
Understanding Logical and Statistical Models
4:28 to 6:41
A deep dive into the interaction of logical reasoning and statistical models in AI.
“And that can come back to us in a very serious way.”
Neurosymbolic AI: Bridging Two Worlds
6:41 to 7:20
Discussion on neurosymbolic AI and its potential to solve complex problems.
“you know, I have to get my car washed, which is 50 meters away.”
Neurosymbolic AI in Practice
7:20 to 10:17
Exploration of how neurosymbolic AI integrates statistical reasoning with logical models.
“Yeah, that's a very interesting question because to some extent, I feel like the history of the field is reflected in the way these areas have come together.”
The Algebra of Thought
10:17 to 14:00
A look into the philosophical aspects of AI and the concept of algebraic thought.
“That's where the field has come from and only relatively recently as it moved into this more probabilistic area.”
Understanding Symbolic Aspects of Thought
14:00 to 14:51
Explore the symbolic nature of thought and its implications for AI architecture.
“And the question is, can we arrive at that calculus?”
Theory of Mind in AI
14:51 to 17:07
Discuss the theory of mind problem and its relevance to AI and large language models.
“And to some extent, you know, it works to a large extent because it can pick up on the patterns in our real world.”
Integrating Algebraic Approaches in AI
17:07 to 17:52
Learn how algebraic expressions can enhance the understanding of AI context.
“We solve this using a solver and we feed it back to DLLM and it ends up doing the right thing.”
Show all 17 chapters
AI's Problem-Solving Capabilities
17:54 to 20:02
Examine how AI identifies and solves specific problems in various domains.
“I love the idea that your description of ChatGPT, essentially it's spotting that there's a problem to be solved in a particular domain.”
User Intent and AI Interaction
20:02 to 22:26
Discuss the importance of understanding user intent in AI interactions.
“because as you can imagine, in many of these care home robots and other kinds of systems where you have rich social interaction, you have to be very mindful of what the intent of the user is.”
Addressing Hallucination in AI
22:26 to 24:20
Explore the challenges of hallucination and confabulation in AI language models.
“Do you think that this model is going to be something that then helps to get on top of that challenge?”
Reinforcement Learning in AI Development
24:20 to 32:55
Learn about the role of reinforcement learning in enhancing AI decision-making.
“But that doesn't quite, you know, get rid of the problem that it's only relying on the training data.”
The Future of Software Engineering
32:56 to 38:08
Discuss the evolving role of software engineers amidst AI advancements and tools.
“So with Cloud Code, for instance, as you know, the Cloud Code's skills was accidentally released by Anthropic a few weeks ago or a few months ago, I can't remember.”
Advising Future AI Practitioners
38:09 to 42:02
Gain insights on how to navigate a career in AI and the importance of foundational knowledge.
“So, you know, senior people might really bootstrap their productivity to some extent with these tools, but a junior programmer is going to struggle.”
Positive Outlook on AI's Future
42:02 to 42:39
The discussion emphasizes the positive impact of AI on society and job roles.
“So it's useful to keep that perspective and to think about if we don't buy into the idea that all jobs will be made redundant.”
Transcript
Automatic transcript. May contain errors.0:00I often tell this to people that the fact that LLMs can do what they can is very surprising given that they're only looking at correlations in text. So that's shocking that it's so good at all, but I feel like that's where also the main limitation of these things lie, because they don't quite have a way to grasp or ground all of their understanding in some sort of real-world structure. and that's when you see very strange hallucinations.
0:36Welcome to Inside the Algorithm, the show that goes beneath the surface of artificial intelligence into the scientific research, technical breakthroughs and academic thinking shaping where this technology is actually heading. I'm Jeremy Bradley, Chief AI Officer at Cambridge Spark, a leader in transformational data and AI upskilling, career development and progression. Each episode, I sit down with the researchers, data scientists and technical experts working at the true frontier of this field to explore not just what's happening in AI, but also the real science and thinking driving it forward.
1:08Whether you're building sophisticated AI systems, leading teams through AI transformation, or driven to understand what's really happening beneath the headlines, this is the show for you. In today's episode, we're examining something that underpins almost every modern AI system, yet rarely gets discussed on its own terms. The reasoning machinery behind the models and why getting that right may matter more than making them bigger. My guest is Dr. Vyshak Bell, reader at the School of Informatics at the University of Edinburgh, Alan Turing Institute Faculty Fellow and Director of Research and Innovation at the Bayes Centre.
1:42He spent 16 years at the intersection of logic, probability and machine learning and has become one of the clearest voices for a more principled approach to AI, one that goes well beyond scale. In this conversation, we get into why scaling won't necessarily solve hallucination, what the success of tools like Claude Code really tells us about how AI generates working code, and what it means for developers changing roles. I hope you enjoy it. Vyshek, you're really welcome. Thank you so much for joining us. Thanks, Jeremy, for having me here.
2:20fantastic um it's really exciting to to have you on uh inside the algorithm today and not least because you've been working at the you know the interface of of research in uh areas around foundations of ethics and explainability and now generative ai um for 16 years or so so So really nice to have your perspective on the podcast. If you had to describe the single bigger shift you've witnessed in that time, not just technically, but in terms of the field itself, what would that be? Yeah, thanks, Jeremy. I suppose technically maybe the answer is obvious. To some extent, it's the Gen AI growth. But I would say sociologically, a big shift has been sort of the humility.
3:14So, I mean, people say that during the first AI boom in the 80s, there was a lot of enthusiasm and promises and guarantees of success. And I feel like this has come back right now. So to a large extent, you know, before we saw the Gen AI bubble, we knew certain problems were hard. So, you know, for instance, classical search, where you look through a tree and try to find a branch that's useful to you, is often intractable under some conditions. So we knew complication was a big challenge here. What changed, however, with Gen.AI to some extent is that, you know, there's the impression that, you know, a large chunk of human common sense knowledge is accessible, readily available.
4:03you can go in there and manipulate things. And with that, I think, you know, the humility has gone away a little bit. I'm not saying it's not warranted in all cases, but it does feel like, you know, people promise, you know, the world to you. But, you know, there's a risk where we don't really know when and how these systems fail. And that can come back to us in a very serious way. That's a very interesting statement. I think what I see myself is that people see a certain level of capability and then they infer a much grander set of capabilities, whether it's large language models or agents or anything that's derived from that technology.
4:57I mean, it's undoubtedly incredibly impressive technology that will, I'm sure, attract a lot of prizes and applaud it in the future. But it's still in development and there's still things that it needs to work on, I think it's fair to say. Yeah, I mean, most interestingly, you know, humans, when they often make statements, they have a sort of world model in mind, right? So, for instance, if you were to ask me what's the weather going to be like in Edinburgh tomorrow, I'm basing it on some data, of course, you know, the number of years I've lived in Edinburgh. But also, you know, I look out of the window, get a sense of the sunshine and so on and make a judgment.
5:43And with the LLMs, it's often the case that, you know, it's a completely linguistic correlation. So we know, for instance, you know, often when people say it's cloudy, we expect rain. If it's sunny, we expect warm weather. So it looks at these correlations and utters things based on that. So to a large extent, it seems, and this is very surprising. I often tell this to people that the fact that LLMs can do what they can is very surprising, given that they're only looking at correlations in text. So that's shocking that it's so good at all. But I feel like that's where also the main limitation of these things lie, because they don't quite have a way to grasp or ground all of their understanding in some sort of real world structure.
6:36And that's when you see very strange hallucinations. a classic example is there was a viral meme going around where they asked a large language model, you know, I have to get my car washed, which is 50 meters away. What's the better thing to drive there or walk there? And many large language models, even frontier models said, just walk there. But of course, you know, without my car, what's the point of going to a car wash? You know, that kind of thing. So it's very surprising where it feels and also impressive when it works.
7:15just to give yourself a a opportunity to to to talk about your your perspective your work sits at the intersection of logic and probability and machine learning for the audience who maybe assumes that that's those are slightly different uh fields how does that how have you bridged those and how has that been a logical pathway for your work and research? Yeah, that's a very interesting question because to some extent, I feel like the history of the field is reflected in the way these areas have come together. So if you think about what we're really trying to do in AI, we're trying to target the question, how do we get the behavior of a system correct when we don't know much about the world?
8:04So in a way, you can see why it's inspired by cognitive science. It's inspired by psychology, obviously human intellect, which is where it starts from going back to Alan Turing. So in a way, logic and probability are capturing different aspects of the way we look at the world. So, you know, one way to look at the world is to see the objects in it and the relationships they have. So a glass of water sitting on my table, if I push it, it would, you know, spill the water all on the table. That's the relationship aspect. And there's a logical and algebraic flavor to that. But on the other side, going back to my example about rain and sunshine, if you, you know, map the weather in Edinburgh across many years, you get a sense of the distribution.
8:51What can we say about the weather in June? What can we say about the weather in July? And so there is a data aspect to it. And a part of this intersection of looking at logic and probability is to say, what's the best way to look at the world in a joint and coherent fashion that captures the relationships and the statistics? And that's kind of where I got started, because I sort of thought, you know, it certainly makes sense that there are certain aspects of the world that are logical and there are certain aspects that are statistical and we need to look at the ways that these things combine. The latest iteration of this is something called neurosymbolic AI where the statistical aspect is completely built by a deep learning system.
9:38It's, you know, broadly we could call this hybrid AI. It's a very broad term and maybe intentionally so because, you know, what we really think is not the dogmatic definition of keeping, you know, fields separate, but thinking about what problems arise in this context for us to behave and act intelligently and what do we need to make it work, right? So it's not, I'm not often surprised to see philosophers and cognitive scientists also at our conferences because it's a question that interests all of us. I think so. And judging by, so I was really interested to see you presented a talk at, I think recently in Sweden in March in the Wallenberg Advanced Scientific Forum on Neurosymbolic AI.
10:27And that was a title that really sort of pulled me in and made me start thinking about the reality of trying to get these two worlds to interact, the world of the statistical language model and the world of the reasoning, the logical, deductive, planning, rich model, if you like, that ironically, I think, was the foundation of AI for over 30-40 years. That's where the field has come from and only relatively recently as it moved into this more probabilistic area. So I'd love to know then from your perspective sort of what neurosymbolic AI is first of all and what you feel its value is and what problems it can really tackle that maybe the two separate worlds would struggle to.
11:27Yeah, thanks, Jeremy. I'll say a bit about this Sweden talk because maybe it connects to what people have been thinking about for the last few years, which is large language models. And again, as we both discussed, these are very impressive, do amazing things. But if you look at the GPT-4 or pre-GPT-4 era, what you often see is when you ask these GPTs for calculations, for instance, what does your mortgage look like in five years? It often would make up numbers. So what changed then is that many of these systems from OpenAI, StratGPD to Anthropics, Claude, they started taking this approach where they would take your natural language, they would kind of parse it, and they would often produce a Python script or an algebraic equation that captured the mathematics.
12:19and suddenly you started getting answers that made sense, including combinatorial problems like, you know, I have a bag of six oranges, you know, how many different ways can I arrange them such that, you know, when I pick six oranges and four apples, how can I arrange them? Or what do I get if I pick one at random and so on? So there's an algebraic, you know, substrate now for many of these systems and suddenly they started becoming more reliable. And that to me is sort of the key manifestation of this neurosymbolic AI, where we're using large language models purely as a textual understanding and parsing unit.
12:58But we understand very well that this sits in a sort of a hierarchy where as a cognitive function, we have to get agents to plan. So maybe you're interested in taking a holiday to Hawaii. And it doesn't, you know, it's not enough to know what people have done before. You have bespoke constraints. You say, I want to go in July and come back in August. You know, I want to find a place where my kids can go to. So all of these constraints start fitting in. And that's when you see this neurosymbolic split where there's some bit which is purely neural. And it's done very well using deep learning or large language models.
13:35But there is a bottom unit where this kind of manipulation takes place. We move things around and get to answers. And this, you know, is sort of interesting going back to human cognition. So in the 1600s, bizarrely, Leibniz, you know, postulated that maybe human thinking is a kind of algebraic system. Maybe we can come up with a calculus for thought. And the question is, can we arrive at that calculus? Because we know there is likely no place in the human brains where we store symbols. And yet certain things we do have a very symbolic aspect to them. So we could really think of symbols as sort of abstraction for thought.
14:21How do pieces come together? How do we compose? How do we build hierarchies, right? Not surprisingly, a lot of this is happening inside our brains in a very complicated way that we can't quite work on. Evolution is giving us a lot of help here. But if we are to implement a system that kind of captures that complexity and nuance, we need to come up with the best architectural choices that reflect that. So one position, of course, is to say, you know, train a neural network, throw more data at it, see what happens. And to some extent, you know, it works to a large extent because it can pick up on the patterns in our real world.
15:05But there are places where it breaks down, you know, like the car example we talked about. So this is a place where keeping sort of an abstract way to think about our thoughts and composing them makes sense. one thing we got interested in recently was the so-called theory of mind problem so this is a problem that came up that comes up or this area comes up you know in philosophy it also comes up in game theory cryptography and the key idea is imagine now we have a bunch of distinct agents in this case the two of us talking and we both have ideas about the world right so So, you know, you live in a certain part of the country, you have certain ideas and local knowledge, and so do I.
15:48And if you're trying to communicate, so for instance, if you say, would you like to come over to where I live, you might give me some logistical pieces of information. And we share this information, so we're communicating. So theory of mind turns out to be the key component of our social and explanatory dimension of us as agents operating in the world, right? This is how we communicate. We make sure we don't crash into each other when we drive cars. So we try to think about what would it mean for the theory of mind to be somehow present or emergent in large language models. And what we discovered is that, you know, to some extent it can pick up.
16:30You know, if you say things like, oh, Jeremy told Wyshak that, you know, he's going away on a holiday next month. I can work out certain things from that. But as you get deeper and deeper, so if you ask, what does Wyshek think about Jeremy's preferences? Or, you know, how does Wyshek know that Jeremy likes to go here and not there? It starts getting more and more complicated for the large language model. And again, the reason this is happening is because it's picking up on correlations. So the approach we took was to come up with a sort of algebraic approach for capturing the theory of mind. so the language model understands the context, gives us the algebraic expressions.
17:12We solve this using a solver and we feed it back to DLLM and it ends up doing the right thing.
17:36reach more of the brilliant minds doing this work and keeping the rigorous substantive conversations going. You can find Inside the Algorithm on the Data and AI Mastery podcast feed or watch it on the Cambridge Spark YouTube channel. You will find the links to those in the show notes below. Right, let's get back into it. I love the idea that your description of ChatGPT, essentially it's spotting that there's a problem to be solved in a particular domain. it's identifying that domain it's it's then finding the technical solver if you like with technical um tool that that maybe best solves that particular type of problem and then and then and then you know invoking it calling it structuring the problem so it can be solved that way i think i think for a lot of people that's quite going to be quite surprising because i think people imagine that the technical expert in whatever fields that they were coming from would be the person who understands when a problem is of a particular sort and what the best solution approach is for a particular problem.
18:49I'm not necessarily talking about that in academia, but obviously it applies there. But in general, a planner, a pricer in a company, even someone who you know just is a market analyst you know they will have a set of instinctive tools at their fingertips and they will go to those tools in order to to solve a problem and what you're saying is here's a version of um an ai tool with a large language model where it will it will also have an understanding of of the world which allows it to pick the right tool maybe of several but not only that it can it can it can take into account the person who's asking it can take into account maybe the preferences and it can take into account the perspective and some maybe some logistics around around you know around what it would mean to solve that problem of getting the car to the car wash for instance or going to the shop to get groceries and and and what what's what's reasonable and not reasonable in those circumstances And I think that sounds to me like a really extraordinary step forward for the field.
19:58Is that how it's recognized? Well, I mean, to an extent, theory of mind is a question that, you know, from a psychologist to computer scientist or roboticist have been thinking a lot about. because as you can imagine, in many of these care home robots and other kinds of systems where you have rich social interaction, you have to be very mindful of what the intent of the user is. So, you know, if you imagine, you know, speaking to a robot in a mall and it keeps asking your name every couple of seconds, you're going to get annoyed with it very quickly, right? So keeping that user intent in mind is a concern that they've had for a long time.
20:45I think, you know, what my sense is that the way they've tackled this has been sort of ad hoc, where they've tried to keep a very simple model for the user intent. And I suspect as we get into more and more complex applications of LLMs, certainly when you go into agentic AI, we may want to think more carefully about what it is that the user is intending to do. Again, it's not so easy to extract this information, right? So if I just gave you a questionnaire of 100 questions and asked you everything, you're going to have a sort of a sense of frustration by the third or the fourth question. So to elicit this information and this preference can't simply be me asking you 100 questions.
21:32It has to be interactive. So I would say there's a lot of human-computer interaction questions to be worked on and how we arrive at that. But at the bottom, it does feel like keeping a good model for what it is the person that I'm speaking to is interested in is going to make this interaction much more meaningful. And do you think this is going to be a way of sort of getting to what people call hallucination? Because I mean, I think, you know, you and I both know hallucination isn't really hallucination. It's just the model selecting what's most probabilistically likely. It just happens to be either demonstrably wrong or silly or just not appropriate for what we happen to know in our sort of logical world.
22:19But it's still doing what it's, you know, coded to do. so it's not actually making things up. Do you think that this model is going to be something that then helps to get on top of that challenge? Or is that going to always be a challenge for language? Yeah, that's a very hard question to answer because, as you say, it's not a hallucination in the sense of I'm not imagining an alternate world where I live in the Lord of the Rings or in the Marvel Universe, but it's really what people are calling confabulation. And mainly what's happening is that it's looking at, you know, again, correlations, and some correlations are less likely than others.
23:07So an example is this, you know, as I said, pre-GPT4, when you asked it math questions, it would often give you nonsensical numbers, mainly because it saw these groups of numbers appearing with the latter group somewhere else and it's training data and it's just producing that. Now, what's happening with the hallucination problem is that obviously with bigger and bigger data and this kind of neurosymbolic split we are seeing in cloud code where the language model is producing a Python script that does the computation for you, what we're seeing is that that kind of numeric confabulation is becoming less and less likely because now they have a very good delegation module within these systems that make sure that if you ask for mortgage calculations or stock prices, it goes to the appropriate module and gives it back.
24:01So I think the coverage of the trading distribution is getting more and more dense. So it could be that, you know, hallucination or confabulation, as we think of it for many factual points may become less and less likely as you get to larger and larger numbers. But that doesn't quite, you know, get rid of the problem that it's only relying on the training data. So in some cases, you might extrapolate your question in a way that's not there anywhere in the training data. And the only thing you can get are the likely outcomes from the training data, right? So I don't have a very good example for this, but maybe that, you know, the car example is one such thing.
24:46You know, you could genuinely imagine, you know, you could say, I'm expecting an asteroid to fall to Earth, and maybe there's no training data on that. What do you expect people to do? And because it has no training context for this, it could give you something completely nonsensical. And the trouble is, we don't know when, you know, when it does this versus when it can confidently say the next likely outcome makes sense given its training data, right? So that's where the risk point is. That's the pinch point. So in a way, you know, this kind of approach of finding a relationship between the language part and some sort of mathematical component that verifies or understands what is being said could be a way to get out of the training problem, sorry, of the hallucination problem, right?
25:37So in fact, you might be familiar that people have tried out things like these retrieval augmented graphs. And what these things do is that they provide a large context in the prompt. This might be relationship triples. It might say, you know, Obama is the president of the United States from this year to that year. And it uses that information to get to its answers. So that could be an approach. But I would still say because it's working on correlations, there's always a risk that, you know, what you get out of it is plausible, but not what you're looking for, or maybe even wrong. And yes, indeed.
26:18And I think we all may be accused rightly of being wrong at some point. I wonder how much actually we're trying to sort of perfect something which is not perfectible, because human beings will often jump to that kind of correlatory linkage and then try and post-rationalize it almost. I think there's a big risk there. I mean, as you say, it's so easy to make that comparison to the fallibility that humans exhibit, but they are distinct, right? So, for instance, let's say, for instance, you know me as a person who borrows money but never gives it back. And if I ask you, Jeremy, do you have any money in your wallet?
27:03You know, you'd lie to me and say, no, I don't. But again, there's intent there, right? Whereas with large language models, there is no intent. So even if it changes its response, it's completely prompt sensitive. So you can't rely on that. That makes it very dangerous. That's a really good example. In that example, I've invoked something from another bit of machine learning. I've invoked essentially a piece of reinforcement learning, which is that over time I've learned that if I lend you money, to be clear, I haven't, but it hasn't necessarily come back very quickly. And so I've associated, in reinforcement learning terms, I've associated a negative reward with that action.
27:51So I'm taking into account my context that's been learnt over time. So now I've got a temporal relationship here with that action. Do we see one of the developments with large language models and indeed AI sort of agency in general as incorporating that kind of reinforcement approach and that kind of learnt reward in its development to enable it to know essentially which tools should I use? reach to to have the best impact or or you know i know this you have i go down that road it doesn't work out very well so i'm not going to do that again no i agree that that's uh i mean so a reinforcement learning as you know has been a very active area of research in machine learning for a long time and certainly the inspiration for this is exactly what you say the behavioral aspects of of intelligence where you provide rewards or punishments as mechanisms to enforce the right behavior.
28:52And to some extent, many of these, you know, large chat, large language model based chat agents do have some sort of, you know, reinforcement built into the training. So an example is, you know, I don't know if it's still possible now, but a year, you know, one or two years before people could say, you know, Jeremy is the leading expert in quantum AI. and if you keep putting that into chat CPD multiple times, at some point if somebody says, you know, unknown to you, somebody says, who's the expert in quantum AI? And your name would come up, right? So I don't think it's possible anymore to do this, but I think this kind of reinforcement learning is an approach to improve its behavior.
29:36And perhaps one place we have seen, I would say, a magnificent improvement in its performance is in terms of code. And the reason we see that is a bit surprising because if you think about code, it's really structured, right? There are certain things that work and certain things are syntactically impossible. So what's working with code and cloud code and codex and all of these tools is that there is a clear sense of evaluating an output to be correct or wrong, right? So if you ask me to give you a poem, I can write something for you. and you could look at it and say, this is rubbish. But I could come back to you and say, no, actually, it's quite good because it has this wordplay.
30:19So there's a lot of ambiguity in how we judge whether a textual output is good or bad. Likewise, if you ask me an essay about Winston Churchill, I could produce something that looks very dry. And you could say, this is so dry, but I could say, this is exactly what we need. We need a factual output. But with coding, because you have, especially if it's trained on GitHub, you have all the change and the commits inside the file history. You have a clear sense of this function, sometimes a function name, and the signature is very instructive. All of these are signals, right? So if you take those signals and feed it back, some of these pre-trained models produce amazingly good code.
Read the full transcript
31:05I'm not saying it's perfect because often it may not be catching things that an expert programmer would catch. A good example is if there's a security flaw that's easy to pick out, maybe it can't work it out because none of the test functions it builds catches that. But an expert programmer might say this is actually very insecure because it's storing everybody's password as a text file. Simple as that, right? But again, because it's trained on so much of GitHub and other open code repositories, it's possible that it's producing very good code. Yeah, but the coding thing is a really good analogy.
31:43And certainly when you started doing that, it was almost a bit of a mystery because even a, I think you said this the last time we talked, even a semicolon in the wrong place or a comma in the wrong place completely destroys the semantic logic and meaning of the program. Whereas if you slightly position an adjective out of place or move a verb around even within reason. It doesn't at all. You've got the same meaning and people might go, oh, that's quite a nice innovative way of structuring that sentence. It might be more poetic that way, that kind of thing. So I suppose it's in a way producing nonsensical or syntactically incorrect code is easy to catch.
32:29And as you say, with natural language, that's not the case. So the signals, again, reinforce this. point. So, right. So, so do we, do we envisage that then probably these tools are applying some sort of post hoc syntax checker just to make sure that they, that the code suggesting is, is, is going to, you know, compile basically or at least parse. So certainly in terms of the, in the output, is that, is that the, is that what's likely going on in that, in that case? That and many other things. So with Cloud Code, for instance, as you know, the Cloud Code's skills was accidentally released by Anthropic a few weeks ago or a few months ago, I can't remember.
33:15But one of the things is that there are lots of permission checks within the code, right? So it's almost that it's creating a technical solution to human oversight. So if I wanted to execute a piece of code on your computer, you might say, first, review the code by me. Check the folder it's working on. And Cloud Code has a lot of these permission checks within it that's mimicking this kind of oversight. And that's the reason that they turn out to be better and better. And again, let's not forget the fact that, you know, the amount of investment that goes into engineering these systems is massive, right?
33:53I mean, it's about the GDP of a few small countries. So it's not surprising. And they have very clever people working in these companies. So it's not surprising that it's doing so many things well. So I suspect, you know, again, I don't quite buy the hype that it's going to replace software engineers. But I certainly think just like we had, you know, ways to call somebody else's code previously in programming. Now we have ways to call, you know, signatures. Right. So I can just say, write me a function that's going to extract this podcast transcript. Give me the bullet points and send an email to all my colleagues.
34:31So that's quite easy to do. And it's probably going to be able to do that. It's just you talk about Claude and obviously they've had lots of publicity around Mythos and Fable of late. But for the brief period of time that Fable was accessible to outside of the US, I heard in very interesting conversations around from the coding community and from the software community saying, I gave it my hardest problem. The thing that would be really stumping me for weeks, months, and it worked through the code and it sorted it out. And in the two or three cases I heard the details on that, it very much sounded like something was going on, which is analogous to what you're talking about in the neurosymbolic world, which is that it felt to me like it was actually doing some kind of dynamic sort of pathway analysis in the code to find specific values of variables or strings or inputs that would trigger issues, security flaws, bugs emerging, which just by looking at the overall sort of stats, sort of map of the code, it probably wouldn't have spotted, or it would have, and would have been very difficult for a human to spot as well.
35:41So I, that's my hypothesis. I have no evidence that that's what it was doing. But I wondered that, you know, whether we're basically seeing them augmenting these, these amazing tools with a lot of the sort of behind the scenes, you know, smart solver technology that we've been working on also, actually, for a number of years. Yeah, yeah. No, absolutely. And, you know, the other thing is, as you say, the pipeline or even the orchestration of producing code and checking for it has so many things that you, you know, an expert programmer would watch out for. And what they've done is they've engineered those things into it.
36:19So, in fact, just this morning, I got an email to say that Fable is going to be temporarily open to everybody, I think, from now on. Not Mythos, but Fable is. And one of the things it says in that email is that, you know, the things where you say check your work is no longer necessary because there's a couple of while loops sitting there which says, you know, whatever you produce, cross-check the outcome. evaluate it back and then only run it, see what it's changing in your file system and come back to it. So those guardrails have been engineered in. And so if you think about that, that is amazing, right?
36:59It's like, you could imagine somebody working at Microsoft or Google for about 15 years to have all of this insight into how to check code, but any introductory programmer wouldn't have that. There's no textbook that would tell you what to watch out for. And now all of that has gone into these models, which is why I think there is so much excitement. I mean, to some extent, a lot of hype as well. But there's so much excitement about them being these robust, but I don't think robust is quite the right word, but inventive programmers, right? And it makes sense. I mean, and I certainly see that trend possibly growing a little bit.
37:39The question is, at which point are we going to get code that kind of makes sense, but has a huge amount of security flaws and whether reinforcement learning can fix those security flaws too. That I'm not so sure about. And the number of matter for one profession, which is the future, if you like, of the developer, the software engineer, the coder, because if the tools are embedding a lot of that good practice, sort of software patterns and software development process in terms of, you know, constraint checking, output checking, validation, testing, even documentation and all of the good things that we used to really try and drum into our development team, teams what's what's the role of the the the programmer and the software engineer in in in the near future and you know is it is it one of you know orchestrator or manager what what how do you see that playing out yeah a good question i i honestly don't know i mean um i think we are running the risk that we are not giving an opportunity to the younger programmers to learn the skills of the trade.
38:48So, you know, senior people might really bootstrap their productivity to some extent with these tools, but a junior programmer is going to struggle. And there's risks also from a psychological point of view. So you might be well aware, you know, if you run a clock code or one of these tools for a few times and you see everything is making sense, you stop checking, You just have this sort of cognitive dissociation because you're exhausted by just clicking approve all the time. So it really becomes an opportunity to rubber stamp things that look kind of right, but you haven't bothered to check. And if this happens a lot, we're going to see huge issues in the quality of the software coming out there.
39:34And I suspect it's already happening now. I mean, now every time, you know, Windows systems crashes or a browser crashes, I often see tweets on the social media saying, you know, this is an outcome of wipe coding inside Microsoft. I don't know if it's true, but this is what people like to complain about. And we might see more of that. Right. So I think there's a risk there. I 100 % agree with what you say about junior software engineers, junior programmers, and indeed juniors in general in a lot of these technical areas. I think finding the right pathway for that upskilling from new to the field to respected senior and subject matter expert is going to be so important.
40:18It may be different, but I think it's going to be so important.
40:27So I've got one final question for you then, Vaisha, just as a quick fire one, if you like. Just alluding to what you said there, actually, for someone maybe who's a bit earlier in their technical career who wants to think seriously about AI from this foundational perspective, rather than just sort of what's the next model that's coming out, what would you say they should get excited about, get interested in? Where would you point them? Yeah, great question. I mean, obviously, you know, every undergrad is thinking about these questions. And what I would say from a practical point of view, you know, you can't quite completely ignore what the kind of hype right now.
41:05So if there is hype around agentic or whatever, it's useful to know a little bit. But I would say if you're forward thinking and looking after your passions, it makes sense to then think about, again, if you're approaching it from the point of view of being interested in AI, go back to the discussions we had around, you know, cognitive science, cognitive function. What does it mean to have intelligent behavior? How do we characterize it? And if you see the history of AI, there's a complex and rich field going to mathematics, game theory, cryptography, agent-based modeling, multi-agent systems, robotics.
41:43And, you know, Now, there's so much variety there that, you know, you don't have to, you know, pigeonhole yourself to one area. And most likely we'll see a deeper integration of these as we go forward, including, you know, theory of mind and interoperability across systems. So it's useful to keep that perspective and to think about if we don't buy into the idea that all jobs will be made redundant. and I don't think it's going to be that case and I certainly hope people are not pushing towards that, then what's the thing that would solve problems with society and maybe that's where they need to look.
42:21Nice. Well, that's a lovely positive note to end on, Vaishak, and I think one that will hopefully inspire and encourage a lot of people to really get into the field and see it as a positive thing to execute in. Thank you so much for joining us today. I really enjoyed the conversation. Thanks so much, Jeremy.
43:09perspective, check out our flagship show, Data and AI Mastery, with Dr. Raoul Gabriel-Urmer. Until next time, stay curious, stay rigorous, and stay ahead of the algorithm.
From the publisher
👉 Discover how Cambridge Spark helps organisations build the data and AI capabilities needed to turn strategy into measurable impact: cambridgespark.com
Large language models are surprisingly good at producing fluent, plausible text. So why do they still confidently get simple things wrong?
In this episode, Dr Jeremy Bradley is joined by Dr Vaishak Belle, Reader at the University of Edinburgh's School of Informatics, Alan Turing Institute Faculty Fellow and Director of Research and Innovation at the Bayes Centre. Vaishak has spent 16 years working at the intersection of logic, probability and machine learning and brings that lens to one of AI's most persistent problems: hallucination.
The conversation traces why scaling alone will not solve reliability, what neurosymbolic AI actually is and why tools like Claude Code quietly depend on it, how theory of mind is being engineered into language models, and where reinforcement learning fits into the future of AI reasoning.
If you work at the frontier of AI research or engineering, this is a grounded, technically rich conversation worth your time.
Follow Data & AI Mastery so you never miss an episode.
If you enjoyed this conversation, you might also like this episode featuring Dr Petar Veličković. Petar joined us on Data and AI Mastery to explore how graph neural networks bring structured reasoning into systems like Google Maps and how AI is being used as a genuine discovery partner in mathematics.
Apple: https://podcasts.apple.com/gb/podcast/bridging-ai-research-and-real-world-impact-dr-petar/id1779783413?i=1000734000576
Spotify: https://open.spotify.com/episode/7qA0AY9MlLS2L9PANqlXNi?si=37d8a8fb43c14cc9
YouTube: https://www.youtube.com/watch?v=GwMUSNidnvE
Glossary Terms
Neurosymbolic AI: an emerging field that merges the intuitive pattern recognition of neural networks with the logical, rule-based reasoning of symbolic AI.
Theory of Mind: refers to an AI’s capacity to attribute mental states to humans or other agents and understand that these states may differ from its own.
Confabulation: In AI, it is the generation of factually incorrect, distorted, or entirely fabricated information presented as absolute truth.
Delegation Module: a software component that allows users or systems to assign tasks, roles, or access rights to others.
Retrieval Augmented Graphs: an advanced AI framework that enhances large language models by grounding their responses in interconnected data networks, such as knowledge graphs.
Reinforcement Learning: a machine learning method where an AI agent learns to make decisions through trial and error.
Dynamic Pathway Analysis: a computational method used in systems biology and bioinformatics to model and simulate how biological processes change over time.
Chapter Markers
(00:00) - Why LLM hallucinations happen
(02:12) - The biggest shift in AI over the last 16 years
(07:42) - How logic and probability shaped Vaishak's path into AI
(10:11) - Introducing neurosymbolic AI
(11:27) - Claude Code, algebraic delegation and the theory of mind problem
(20:01) - Theory of mind in robotics and human computer interaction
(21:54) - Confabulation versus hallucination
(26:42) - Why AI errors are not the same as human dishonesty
(27:24) - Reinforcement learning, reward signals and learned behaviour
(32:36) - Syntax checks and the engineering behind reliable code
(37:52) - What happens to the software engineer's role
(40:49) - Advice for early career AI thinkers
Useful Links
Connect with Dr Vaishak Belle on LinkedIn: https://uk.linkedin.com/in/vaishakbelle
Learn more about Vaishak’s work here: https://www.vaishakbelle.org/
Read Vaishak's report on The Future of Neuro-Symbolic AI: https://ojs.aaai.org/index.php/AAAI/article/view/42130
For more AI insights follow Jeremy on LinkedIn: https://uk.linkedin.com/in/jeremy-bradley
Explore Cambridge Spark’s AI upskilling programmes at https://www.cambridgespark.com




