In short
Eye On A.I. Episode #130: Mathew Lodge - The Future of Large Language Models in AI
Episode Overview In this episode, Craig S. Smith interviews Mathew Lodge, CEO of Diffblue, who discusses the role of reinforcement learning in AI, particularly in code generation. Lodge highlights the potential and challenges of large language models (LLMs) like GPT-4, contrasting them with reinforcement learning approaches.
---
Key Topics Discussed
- Reinforcement Learning and Code Generation
- Reinforcement learning (RL) focuses on achieving specific goals, such as maximizing accuracy in code testing.
- Diffblue utilizes RL to automate the writing of unit tests for Java applications.
- RL can optimize algorithms for efficiency, as demonstrated by Google’s AlphaDev optimizing sorting algorithms.
- Challenges of Large Language Models
- LLMs, while powerful, often prioritize generality over accuracy, which can lead to semantically incorrect code outputs.
- They excel in creating syntactically correct code but struggle with understanding the semantics behind the code.
- Lodge emphasizes that LLMs may not be the ultimate solution for code generation due to these limitations.
- The Future of Language Models and Intelligence
- Discussion on the potential merging of no-code and low-code solutions with AI technologies.
- Addressing skepticism from software developers regarding AI's capabilities in code generation.
- The evolution of programming languages and how higher-level languages impact machine learning development.
- Comparison of AI Techniques
- RL is precise and goal-oriented, while LLMs are more probabilistic and general.
- Examples include AlphaGo and AlphaDev, where RL is effectively used for problem-solving in areas with vast solution spaces.
- Programming Language Evolution
- Insights into how programming languages have evolved and how abstraction affects the way developers write code.
- The impact of tools like GitHub Copilot, which use LLMs but must be critically evaluated by developers for accuracy.
---
Key Takeaways
- Reinforcement Learning’s Importance: Despite the hype around LLMs, RL remains a crucial technique in AI, particularly for tasks requiring high precision.
- Limitations of Large Language Models: While LLMs can produce human-like text, they lack a built-in understanding of programming semantics, making them less reliable for code generation than RL-based approaches.
- Skepticism in the Developer Community: Developers are often skeptical of AI-assisted tools, which can complicate adoption despite their potential to save time and improve productivity.
- Future of AI and Coding: There is potential for a shift towards tools that integrate natural language processing and RL to create more efficient coding solutions that are accessible to non-programmers.
- The Concept of Prompt Engineering: The discussion reflects on the unpredictability of LLM outputs based on prompts and how this challenges the notion of “engineering” prompts for consistency.
---
Conclusion The episode wraps up with reflections on the ongoing developments in AI technology and its implications for the future of programming. Lodge asserts that while current technologies like LLMs are influential, they are not without limitations, and the journey toward more accurate AI-assisted coding tools is ongoing.
For more information and to access the transcript, visit [eyeonai.com](https://www.eyeonai.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00In reinforcement learning, you're going for accuracy. You're conducting this search, you're trying to find the best possible answer you can, the most accurate answer you can. In the case of unit testing, it's how close to 100 % coverage you can get. And in the case of AlphaDev, it's how fast can we make this algorithm? Is it faster than the previous algorithm that I've come up with? And so you're very focused on a goal. In contrast with large language models, which don't have a goal at all, their goal is to be very general. So you've gone from a very specific task in reinforcement learning to a more general task with large language models, where you're going to predict what comes next, and it's okay if it's not exactly correct.
0:37You're trading accuracy for generality. Large language models, the code is usually syntactically correct. We've never seen it generate code that doesn't have the right syntax. What it gets wrong is semantics, because it doesn't have a model for semantics. It doesn't have a model of the language. It just knows what it's seen before and patterns that it's seen before, and that's what it tends to get wrong. I'm Craig Smith, and this is Eye on AI. This week, I talked to Matthew Lodge, CEO of DiffBlue, a company that uses reinforcement learning to automate the writing of unit tests for Java codes. Matthew talks about how reinforcement learning has been lost in the shuffle around generative AI, but remains one of the most powerful kinds of artificial intelligence in any machine learning developer's toolbox, and in fact, how reinforcement learning is the power behind much of generative AI.
1:42I found the conversation with Matthew fascinating, and I hope you do too. Now, here's Matthew. Matthew, it's great to have you. Why don't you start by introducing yourself and then we'll get to some questions. Great, great. And thanks for having me on your podcast. So my name is Matthew Lodge. I am CEO of DiffBlue. We're a generative AI for code company. We're a spin-out from Oxford University in the UK. In terms of my background, I started as a software developer in real-time systems and safety critical. So I worked on code that flew on the space station, the Boeing 777. And later on, moved to Silicon Valley and been in B2B product management for a long time.
2:27Worked at companies like Cisco and Symantec and VMware, as well as startups. You worked on a startup that did software-defined networking 10 years too early, which is the same as being wrong. I was also SVP product to Anaconda, so the machine learning, data science, the Python R distribution, very popular in the open source world, before being CEO at DifBlue, where we are using reinforcement learning to do generative AI for code. So we have a product, our first product, that automatically and autonomously writes unit tests for Java applications. Yeah, and we were talking yesterday about how reinforcement learning or code generation through reinforcement learning has been kind of lost in the buzz around transformer-based large language models, which, in fact, have a lot of problems when it comes to things like code generation.
3:29And I was describing to you how I've worked sort of many hours with AutoGPT and with going back and forth between AutoGPT and GPT-4. Every time I'd hit an error, then I'd plug it in GPT-4. It would tell me what the error is and suggest some changes and go back and forth and back and forth and just never get anywhere. So can you talk about that difference? uh i you were giving a talk yesterday i think about uh about that difference and why people don't uh talk about reinforcement learning as a solution for code generation yeah well i mean large language models have this huge general appeal i mean so gpt has really made this kind of technology accessible to a general audience and that's really the genius of large language models is that it's very easy for everybody to understand what they can do because they can just see it.
4:39Reinforcement learning has traditionally been used in AI for game playing. That's really where it showed its forte. So in the early days of OpenAI, they were building game playing AI, very successful Dota engine that's adversarially trained. And that technology, essentially, you're going on a search. you're trying to find the best move in game playing. So AlphaGo is a really great example of reinforcement learning led approach to game playing. And so the place where reinforcement learning is really useful is where you have a space of possible answers that's too big to check all of the answers.
5:19So people think about Deep Blue IBM's product from what, 30 years ago now, as being this sort of breakthrough chess playing product. But essentially, all they did was brute force search the move space and it just looked at every possible move and they picked the best one and and that's how modern chess software works there's no neural networks involved it's they just look at every move when you get to go you can't do that because number of moves is greater than the number of atoms in the universe right so you have to take a different strategy and that's essentially what reinforcement learning is about it's about uh identifying the areas of the solution space where the best solutions are likely to be found and spending more time searching there and basically ignoring the rest.
6:00So you ignore the areas of the space where you're not likely to find the answer. So you're not guaranteed to find the best answer, but you can find a very good one. And essentially that's what AlphaGo does. It has a prediction of who's going to win the game. It uses that to, it tries different moves. It uses Monte Carlo methods, randomness to you and me, to suggest additional moves. And it tries them and it makes a prediction about who's going to win the game if it makes that move. And it essentially conducts the search that way. So it's always asking who's going to win, who's going to win, who's going to win.
6:31And from that, it can predict, try and move, and then it can essentially follow a sequence of moves and build a playbook for playing the game. And you can do the same with code. So Google has recently done this with AlphaDev, so the DeepMind team at Google. a couple of weeks ago came out with AlphaDev, which is an approach. They took reinforcement learning and they applied it to very common computer science algorithms. So they picked sorting and hashing algorithms, which are every software program today uses those things multiple times a day. So if you can squeeze and optimize, squeeze some time out of those algorithms, the effect is enormous because they are so frequently used.
7:15And so essentially what they did is use reinforcement learning to try different moves. They tried different implementations of the Quicksort algorithm. And this is down at the assembly language level. So they were just trying different instructions like what happens if we delete this instruction? What happens if we simplify or change, modify the code, essentially mutate the code? And in that search, they were able to find more efficient implementations. So for Quicksort, the implementation is about 1.7 % faster for a large array of numbers, which doesn't sound like much until you remember that this is a function that gets calls millions of times.
7:55And so the cumulative saving is enormous. And for small data sets, the improvement is 70%, so much bigger for small data sets, which is probably a more common use of the function. So essentially, they were searching for better code. And Difblue's product does the same thing. We search for unit tests. We make a guess at what a good unit test would look like for a particular piece of code. So a unit test, the idea is you're isolating a unit of software and what you want to do is find regressions using these tests. So these tests, you put in the input to the function, you check the output to make sure it's doing the right thing, you put in enough input so you cover all the different branches inside of the code, you check the answers for all of those to make sure that they're correct.
8:40And what that means then is that when you run the unit test later, after you've modified the code, you can find regressions. So changes in behavior, because a regression could be a bug, or it could be a deliberate change to the software. You've changed the way the software works. And so you would expect the results to be different. But in the case of unit tests, what the search we're doing is we write the test, we look at the signature of the method, we have some analysis of the program as well, so we can make a good guess. and we try it, we compile that code, we run it against the method under test, and we see how it did.
9:13What kind of coverage did it get? Did it exercise the branches? All of those things. Based on that, we can predict what a better test might look like. So we compile that, we run it, we try against the code, see how it does, and we iterate through that approximately 10 ,000 times. And at the end of that process, we've got a set of tests. So it's a very different method to large language models. Large language models are based around the notion of text patterns. When you're training a large language model, it is learning, it's building a statistical model of text. And it's using the context of that.
9:49In the case of natural language, it knows the role that a token plays. A token is essentially a word in large language model. So it knows that this is a noun or it's a verb, what part of speech it forms. So they're using that information in order to construct the statistical model. And so what you see with large language models is that from the prompt, it can do a very good job of figuring out what comes next. And so that kind of technology is kind of like a auto complete, turbocharged auto complete. You know, auto complete has been an IDE feature for years since the first IDEs came out. But it's able to go a lot further because it has a much, it's much better at guessing what comes next.
10:34Essentially what is happening in the two different techniques is that in reinforcement learning, you're going for accuracy. You're conducting this search, you're trying to find the best possible answer you can, the most accurate answer you can. In the case of unit testing, it's how close to 100 % coverage can we get. And in the case of alpha dev, it's how fast can we make this algorithm? Is it faster than the previous algorithm that I've come up with? And so you're very focused on a goal in contrast with large language models, which don't have a goal at all. Their goal is to be very general. So you've gone from a very specific task in reinforcement learning to a more general task with large language models where you're going to predict what comes next, and it's okay if it's not exactly correct.
11:18So you're trading accuracy for generality. And in tools like ChatGPT, if you ask them to write code, or GitHub Copilot, or Tab9, Tab9's product, essentially it's making its best guess at what the next piece of code is. It knows the code that came before, it knows the code after, and that's all it knows. And it guesses, and that's okay because you have a developer sitting there looking at the completion and deciding what to do. Does this make sense? Maybe I can, it's not quite correct, but I can fix it. So, they're very different techniques and it's large language models because of ChatGPT have been taking all the limelight because that's the thing that people can see and understand.
12:00them yeah uh and before alpha dev uh deep mind had alpha go is alpha dev is is not a product it's a research project is that right that's correct yeah so they released a paper about alpha dev but i don't think it it's been released outside of google yeah and uh how does uh that relief to no i'm sorry not i said alpha go i meant alpha code um oh yes alpha code yeah yeah alpha code also a research paper and alpha code was designed to win programming competitions and so alpha code is an example of where you're taking a large language model and asking it to generate lots of different alternative programs to match a specification.
12:53So the thing that you have that is just completely unlike the real world in the programming competition is you have a very good, unambiguous, clear description of exactly what the code should do and what the output should be. And so they take that and they use it to synthesize many different versions of the program. because the other thing with large language models is the stochastic component, the randomization that you can dial that up and down in most of most LLMs. And so it will generate different alternatives essentially. And so the idea is that you have a better chance of coming up with the best alternative that way.
13:35And so in AlphaCode, they take that description, they generate a whole bunch of candidate programs, and then they run cluster analysis on those programs because they They want to find out a lot of those programs effectively do the same thing. The code might not be exactly the same, but they do the same thing. They're semantically equivalent, maybe not syntactically equivalent, so the code doesn't match word for word, but it does the same thing. So they cluster all of those to find out, to eliminate duplicates, and then they essentially cull it down to a couple of candidates and they try those.
14:11And that's how it produces solutions. So it's very good at winning programming competitions, but it's not really a real world thing. Yeah, but it sounds as though for code generation, you could write a system or build a system that does generate accurate code using reinforcement learning. And why hasn't that been done? Well, I think it has been done. It's like the old phrase, the future is unevenly distributed. So we've been doing this since 2016. There's an open source project called EvoSuite. They started in 2012. They built a reinforcement learning research project originally for a number of academics at Sheffield University and elsewhere, they built EvoSuite and they use reinforcement learning to find unit tests.
15:14And they've been doing that for a very long time. So you're starting to see more products based on reinforcement learning while LLMs have grabbed all the headlines because of their broad appeal. Yeah. We talked about the threat discussion with that's, again, grabbing all the headlines in the last month or so. And you were telling me that this is actually a problem that's existed for a long time and has been dealt with for a long time. Can you go back over that? Yeah. So it reminds me of the discussion that was had very early on around software controlled, what are called safety critical systems.
16:05So you can think of a safety critical system as where if it gets it wrong, people die. That's the easiest way to think about it. So we're talking about the software that flies aircraft and other drones, things like that, signaling control for things like train networks, railroad networks. That software has to be absolutely correct. And Boeing, some of the very early stuff that I worked on, so both Airbus and Boeing at the time working on fly-by-wire systems, Airbus did it first and Boeing followed along later. And essentially those systems, the effect of getting it wrong is that people die. And so those systems are very highly verified.
16:49And so the behavior of those systems is very well understood. And they have this notion of the envelope, sort of, if you think of the flight envelope of an aircraft, if you exceed the envelope in some way, then bad things happen. So, you know, maybe you stress the aircraft so much that things, you know, the wings break off or, you know, stabilizer or something like comes off the aircraft. You can't exceed the parameters of the flight for a particular airframe. And so a lot of work goes into verifying those systems. essentially what they do is they prove that they can't exceed that flight envelope.
17:23So there's a lot of work around making those systems deterministic in operation so that you can apply these proofs to it. And you can sort of constrain what the algorithm will do. And it strikes me that's a good analogy for what some of the requests offer around large language models. if you start hooking these things up to real world systems, then you need to understand the envelope of solutions. And that is one of the big challenges with anything that's based on a neural network is that inherent unpredictability. You have a giant statistical model that humans can't understand. It's too complex, too complicated.
18:03There are starting to be companies out there that are, you know, sort of segmenting the input space and looking at how you can test these algorithms and understand, you know, this envelope more. but essentially i think that is the challenge that's the next challenge for the ai industry is you know think about it the way safety critic people do or have done for for decades but but why then uh why then this this near hysteria about uh you know unleashing uh large language models or large transformer-based models and their potential for catastrophic outcomes. Is it because they cannot be controlled with these different constraints?
19:03I'm really puzzled by this debate. Yeah. I don't pretend to understand everybody's point of view in this debate because, as you say, there are wide disagreements between the various participants. I can tell you how I think about it. I think there are really a couple of different distinct things that people are worried about. One is the misinformation challenge. Yeah, it's very easy to, you know, you can ask GPT to write something in the style of somebody else. Right. You know, you can ask it to write in the style of Gordon Ramsay and it'll sound like he does on Hell's Kitchen or something like that.
19:46So you can see the potential for misuse in those areas. I think the more general challenge is the idea that if you start hooking this up to things that impact the real world, I think that's the bigger concern because we just don't really understand them very well. We don't understand how they work. We can't predict them. I think that's where a lot of this comes from. I don't buy into the world is going to end hysteria. And this is a new form of intelligence. It's not a new form of intelligence. It's a statistical model of text. and the work that we do at DifBlue around with large language models in using them to generate code.
20:28What we find is that large language models, the code is usually syntactically correct. We've never seen it generate code that doesn't have the right syntax. What it gets wrong is semantics because it doesn't have a model for semantics. It doesn't have a model of the language. It just knows what it's seen before and patterns that it's seen before, and that's what it tends to get wrong. And in your example with AutoGPT style, it produces some code that maybe doesn't compile or doesn't run. And so you ask it to go fix the code that it generated. It has no idea what the code does. It is just looking at patterns in that code and trying to problem solve by suggesting a different pattern.
21:12And that's the inherent challenge with large language models. So there's limited understanding as large language models. And if you don't have a way to validate the result of what it's producing, and you just take that and you use it verbatim and just assume that it's correct, then you can get into this very quickly spiral completely off the rails because the model is not trying to be correct. It's trying to do the best it can with a single shot. these algorithms are deliberately trading accuracy for generality. Yeah. And going back to code generation in particular, since code generation is so useful for everybody in that there's so much software that needs to be written as the economy continues to digitize.
22:17Why haven't, for example, why hasn't, from your point of view, hasn't Google put more focus on alpha code that writes complete, albeit simple programs uh so that so that we we can use natural language to instruct a reinforcement learning system to write code for us that that compiles and and executes yeah i i well i don't know the answer to that but I do know that they've followed suit around a co-pilot style product. So there are now three of those you've got a co-pilot, you've got Code Whisperer from Amazon AWS and now there's Bard, there's a version of Bard that does coding. And the advantage of those approaches is that they're highly accessible for developers because you can just put something in and it will generate some code inside of your IDE.
23:22And that's great. It's very simple from an interaction model. And I think it's what we've seen in the case of OpenAI, I don't know about how Google thinks about this, but in case of OpenAI, they clearly believe that large language models are on the road to general artificial intelligence, artificial AGI, artificial general intelligence. and so they're trying to push that technology as far as it can go and they're not interested in things that are less general as a result so i don't know if that's true at google or not but but certainly it would be useful if there was uh if there were were systems that could generate code uh that were not general that were more accurate yeah yeah absolutely and i think there's plenty of room for that.
24:14I mean, you can see this very quickly being useful in today's low code products, right? You've got, you know, the problem out there is there's more software to bring and there are people to do it. And so you've got several different approaches to solving that problem. You've got the no code things where, you know, stuff that really doesn't have to have code and could have like a flow chart, you can just build those things. And that's a very well established markets been around well over 10 years. You've got the low code things where you're trying to just simplify it. And people like Out Systems have been doing this for over a decade, but I think there's sort of more modern versions of those products that try and make it even simpler.
24:54And it seems like there's a lot of opportunity to you to marry the two together and get something that can produce useful results with even lower code. I don't know what you call that uh the uh have you used copilot or do you guys use copilot in your work we don't use copilot in our work well it's it's very controlled so part of the issue is we have a lot of code that our customers give us and we agree uh to keep that confidential not share that with anyone so we're incredibly careful about how we use any kind of tool where we would be sending code outside of our organization. So we don't tend to use it for that reason.
25:36And essentially, it's on an exceptional basis. Can we talk about intelligence? You sort of dismiss this idea of intelligence as opposed to statistical models. And it's something that I've talked to a lot of people about and of course it gets into a semantic debate about what intelligence is. Right. But can you sort of give your view on how much large language models really exhibit intelligence? And then this goes on to the whole discussion of sentience and consciousness and that sort of thing. Yeah, I was telling you I had a call with Noam Chomsky and he's very dismissive of what large language models do and does not believe that there is any real intelligence being exhibited there.
26:53Yes. Yes. So, I mean, nobody knows the answer. It's a fun thing to debate. It's like, what is consciousness? It's some kind of electrochemical process that's going on in our brains. But we don't really know how that works. And although neural networks are inspired by the way the brain works, but they're really not a model of how the brain works at all. They're not not close. It's like a Hollywood movie that's based on a true story. So the thing that what you tend to see is blog posts where they say, well, I asked GPT this thing and it seemed to have a theory of mind, right? It seemed to do these things.
27:38And then you can, somebody else comes up with a counter example. So they get one example and they stop. And it's like, well you need more than one example and and you need to be able to relate that back to why why is it doing this and what is the mechanism that's at work here that why do you think it has a theory of mind you have to be able to explain that you can't just like assert it because you got a good result in one particular situation that's not science um it's cherry picking essentially and so um there have been number of claims that we made this this claim of emergent behavior so emergent behavior is things like flocking.
28:16So when birds fly in flocks, essentially, that's some very simple rules for flocking that birds seem to be following, which is, you know, if like two neighboring birds are moving this direction, I should move in that direction. And that's how, and you can build, you know, sailor automata, this is from, you know, 30 years ago, you know, explored this whole idea where you could build these very simple automata, very simple rule sets and they show that they would exhibit these emergent behaviors like flocking. The only problem with this is that nobody's actually been able to demonstrate emergent behavior.
28:48It looks like it might be emergent behavior, but there's actually a really good Stanford paper like how it's not. And so it's this thing where we as humans, we want to see these patterns. It reminds me a lot of Taleb's book, Fooled by Randomness, where he sort of talks about trading and the idea that you could be a successful trader for 10 years on pure luck. And you know, if you're in that situation, you think you're a genius, right? I'm a genius trader, I've made all these fantastic returns for the last 10 years, and it could all just be randomness. And it's the same, we see the same sort of thing going on.
29:24So it's like, GPT is very easy to understand when you see this example, it's like, oh, this must be, it's very easy to get a right, a leap to the conclusion that something must be happening in there that is intelligent. But again, it's like, great, you need to have a theory of where is that intelligence in this process? Because we do know how it works at the macro level. We understand how it's trained. We understand the perceptrons, the neurons that are in there, how they work, the layers, how they're connected, all of those things. Where is the intelligence in that? And you've got plenty of counter examples.
30:02And this is made worse by the fact that the more modern large models like GPT-4 are also trained by humans. It learns human answers. So humans will come up with some of the answers that you're getting here. You're like, wow, that's a brilliant answer. And things that GPT-3.5 maybe didn't get quite right, it has been trained on the right answer for those. So a human being has gone and taken that example, written the correct answer, and it has been trained up, fine-tuned on that as part of the process. That's what they call reinforcement learning with human feedback. By the way, it's not reinforcement learning it's an additional training step where they're fine-tuning the model so it's i that to me is the best explanation of why people think it's intelligent i don't think it's intelligent because i know how it works yeah yeah and uh uh where do you think that research is going to go do you have any any uh opinion i mean uh you know, you, you, you, either you keep, uh, expanding the size of these models and, and looking for, uh, for increased capability or, uh, or you, maybe you, you, you combine them with reinforcement learning.
Read the full transcript
31:20I mean, there already is some reinforcement learning in there, uh, to, uh, to make it more accurate or you know there's a lot of a lot of people are just uh using the large language model as a generator of of syntactically correct language uh but uh calling uh uh vector database with you know ground truth knowledge that that then the the language model formulates into into coherent responses and then there's the browsing models that everyone's come out with so that you're going out to the internet and and citing your sources i mean i'm curious what your view is on the future of that. Yeah, so I do think it's going to be, it's not just large language models.
32:19I think it is going to be some kind of ensemble approach. Jeff Hinton has this really great quote from, actually from the interview that you did in IEEE Spectrum, Craig, where he talks about, he's like, what about all the non-language tasks that the brain does? So he gives basketball, says you learn to play basketball by throwing the ball so it goes through the hoop. You don't learn to play basketball by reading a book. It's got nothing to do with language. So that whole, that part of learning is completely missing from large language models. It's just, and that I think is where there is a difference of opinion.
32:53You see the folks at OpenAI saying, you know, these are not the problems you're looking for. You know, they sort of, oh, yes. You know, language is the root of everything. Well, Chomsky argued that language was innate, right? In the brain, that was one of his, you know, famous assertions that he's made. Lots of people disagree with that. And so you've got that kind of disagreement going on. In my case, I don't think reinforcement learning is the example of throwing the ball so it goes through the hoop. That's exactly what that is. And so I think the answer is not large language models on their own.
33:29The answer is not reinforcement learning. I think it's something else. We don't know what that is. And that's what makes this space fun and exciting. Yeah. And in terms of applications, do you think that reinforcement learning, I mean, presumably there are a lot of people working on reinforcement learning applications that are as useful as Diff Blue's application. Do you think that, I mean, right now we're in a period where there's this tsunami of products coming to market or that will be coming to market in the next year based on pre-trained transformer models. is there a parallel space where products are being developed based on reinforcement learning, or do you think that there will be if, in fact, those produce more accurate results?
34:34Yeah, I think so. I mean, there already are. O 'Reilly published a book on reinforcement learning for developers a couple of years ago now, And that gives you an idea of that. So the author of that has a website where he collects examples of reinforcement learning in the real world. It's, you know, the LLM hype machine is in full gear right now. I mean, if you saw, you know, Mistral AI, the French AI startup, the people quit Meta and Google four weeks ago and they've raised 100 million euros seed funding. um that's the kind of craziness we're in right now and so there's a lot of noise yeah uh but but you you you think that there is there are other people working continuing to work on reinforcement learning for for applications uh yeah absolutely yeah yeah so phil winder wrote the o'reilly book on reinforcement learning a couple of years ago and and he maintains a website with lots of examples of reinforcement learning and you follow his Twitter feed and all kinds of different applications for reinforcement learning.
35:50I think the fact that they work so well for a particular problem is what makes them less popular because they're not as general. And so it's a question of, you know, what are the things that come to our attention and what are the things that we can relate to? And anything that is easy to relate to is going to be much easier, is going to have a much better job of going viral and attracting attention versus more specific solutions. I think that's part of the challenge you have with anything that is a more specific solution. I mean, most general audiences haven't don't understand what AI for code products do.
36:25They don't really don't really understand that. And that's fine. It's not their domain. And the great thing about large language models is that they can write you a haiku or a poem or a funny story. And and that's incredibly relatable. yeah uh i mean you were saying one of the challenges for diff blue is just raising awareness because i was asking if it if if it if it works so well yeah and people spend so much of their time developers spend so much of their time writing unit tests uh and and you're focused on java right you you guys only do uh java but uh but there's certainly a that's a massive space i would think it would spread like prairie fire through the java community everyone would say oh my god thank god look there's this tool that we don't have to write unit tests you you know yeah well there's a lot of skepticism so your program is in general skeptical lot and they they don't like hype and so they are skeptical.
37:33The biggest issue we had prior to GPT coming along is that people didn't believe that our product did what we said it did. They're like, that's an incredible claim. How can you possibly claim that? Where's your evidence? You know, very skeptical and that's how they are. That kind of goes with the territory really as a software developer. And so we've seen that in terms of how our process plays out, how we gain customers is that first of all, they don't believe we can do it. Then we show them that we can do it. And they're like, well, that's great, but you need to show that it works on my code.
38:09And so it's a brand new area. Most developers have never seen a tool like this before. And so they really want to understand what it does and they want to convince themselves. And so we give them, obviously give them every opportunity to do that we have a community edition you can just install into IntelliJ and just try it. Back in 2020, we were doing, we did a Wall Street Journal op-ed. We basically said, you know, it was 10 years since Marc Andreessen's, you know, software is eating the world op-ed. And we said software has ate the world, but now, and software is now going to eat itself, right?
38:45Software will write software. And the Wall Street Journal editors were also very skeptical, and they gave us a really hard time. We had to work incredibly hard to prove to them that this one was real and two was going to be an important trend in the software world and therefore the business world because business runs on software. And to their credit, they did run it in the end. But it's like that. And people really need to believe that this is something that is worth their time and would make a significant difference. Also, I think, frankly, because it's been tried before. before. So certainly over 20 years ago, there was what we call computer-aided software engineering tools, case tools that tried to do this kind of stuff and automatically generate code.
39:31They did a really bad job. Code was awful. Developers hated it. So it kind of got a bad reputation. Yeah. Again, you know, I'm not a coder, but I'm fascinated by software. And one of the reasons I got excited about these large language models is the prospect that I could, you know, write my prompt and the model would write my code and, you know, I could build my own software. And, you know, people have done that at a certain level with GPT-4. But again, why isn't there a product where I can write my prompt for a reinforcement learning program that then would go out and search the space and come up with the most viable code?
40:35code and and uh and then you know I I could be on my way with uh with with whatever app I want to build yeah yeah is there a reason yes it's uh really hard for anyone to write down exactly what it is they want their program to do that's that's what the specification problem so this is a very old problem in computer science this was the whole water the idea of waterfall development in the early days of software was that you'd write this full and complete specification for exactly what the software was going to do at the beginning, and then developers could then take that fantastic specification and we had everything in it that they needed to design and build the software according to the spec.
41:20And the problem is nobody can write the spec. And part of the reason for this is that you're essentially what you're doing is you're asking the people who want the software to do something for them. you're essentially asking them to break it down into how it should be done. And they don't know, right? And that's why you need developers in the first place. Because, you know, the users of the product, the people who are describing what it should do, they're experts on their problem. They know what their problem is. They are not experts on how to solve it. They just know that, like, here's the problem, and this is what I want to solve.
41:53They don't know how it should be solved. that's a very different intellectual exercise to figure out, okay, so how would you do that? So that's the joke about AI generated code from English text. It's like all it requires is a complete and exact specification for that to happen, which means every developer is safe because that almost never happens. There are exceptions. Back to the safety critical world, If you're going to do verification, you do need to have a very, you need a formal specification of what the software should do. And that's why it tends to be used for very small areas of code that are safety critical, right?
42:35The real core stuff. What's interesting is we do a lot of work with Amazon Web Services. They do a lot of formal verification of their code across their code base for really mission critical parts of AWS service, like the code that runs on the hypervisor, secure boot, All of that code is verified for when you power on the server and all of the microcode that runs before you even start the BIOS or start the operating system. All of that stuff is verified at AWS. So it's interesting. AWS is somewhat of an outlier in that they're doing that because they're not in the safety critical world, but they see the benefit of formally verifying all that software.
43:17And so, yeah, I think we're in for an interesting time. I think you can probably get close. I think people are, I am sure there are companies working on this problem right now and that you'll be able to use that. But as you say, even today, you know, I'm not a programmer anymore. Nobody would pay me to write code. But you know, just for things like running my own website or projects where I just need a little chunk of code to do a specific thing, I can go to ChatGPT and it'll give me something that's pretty close. And so for sort of like dabblers like you and me, because I would say that, you know, I'm at the dabbling stage now, I'm not a professional programmer.
43:56It's an incredible benefit to have that available. Yeah. For somebody like you who can look at the code and see whether it's correct, semantically correct, and then correct it. for somebody uh like me who who can't do that yeah i just end up in endless loops what you were describing of the uh you know writing of the specifications there's this whole world of prompt engineering that's uh that's developing for working with large language models. I mean, that's really what you're talking about, right? Is prompt engineering, being able to write the instructions for an AI model, whether it's an LLM or a reinforcement learning model to follow.
44:53Yeah. Yeah. I really don't like the term prompt engineering, Craig, because engineering implies there's some kind of rigor there and there isn't. The main reason the prompt engineering doesn't exist is not a thing, is that the fact that small variations in the prompt can produce massive variations in the output. Because it's an unpredictable model. So in what way are you engineering the prompt? You have no predictive ability. You make a change of the prompt, you have no idea what that's going to do to the output. So it's not really engineering. There are good rules of thumb for a prompt. And in one-shot systems like large language models, putting examples into the prompt really helps the model give you a better answer.
45:37And so the more context you can give the large language model, the better it will do. And that's really, that's the rule of thumb. That's prompt engineering in a nutshell. Just give it as much as you possibly can. And it will be a much better job. Then you're constrained by the number of tokens you can input. Yes. But on a reinforcement learning model, if you can specify your intent with some precision, it could write correct code. So in that sense, and in order to do that, I mean, we were talking yesterday. I mean, at a certain point, if you can specify a natural language that precisely, you might as well be writing code.
46:32Yes. But I can imagine that programmers may be trained in the future to write in that kind of, in natural language, that kind of spesys. I know what you mean. I'm going to have to edit this, yeah. It's too early in the morning. But who are going to be able to write that specifically,
47:08then that natural language, there could be an orchestration layer that then decides which programming language is most appropriate. Yeah, yes. But the skill would be in writing specifically a natural language, your intent. Yes, that's right. And there was a really great Twitter post the other day. It says essentially that's what programming is. It's a way of thinking and a systematic way of thinking, breaking down problems and figuring out how to solve them. That's really what's going on in a program's mind. The typing of the code is sort of incidental to that. And he was responding to some of the Microsoft crazy claims that 61 % of Java is now written by Copilot.
47:55And it clearly isn't because the market for Java programmers would have collapsed if that was the case. And Microsoft's own statistics, there's another paper from Microsoft Research came out last week and shows the acceptance rate varies quite considerably from like 0.33 to 0.55 at the maximum level. So I don't know where the 61 % figure came from. And his point was, it's like, look, saving you time typing is not really what increases programmer productivity. It's that breaking down of the problem and thinking analytically and understanding how to solve it. And so maybe you're right. Maybe that's what programming will look like in the future.
48:32It's like I'm old enough to remember this transition from assembling language to high-level languages. So, you know, Algol 68 and Pascal and Modular 2 and C and C++ and eventually Java, all of those languages. And, you know, assembly language programmers, you know, in the early days were like, well, I can write better assembly code. It's tighter. It uses less instructions. It uses less memory, all those great things. But the productivity benefit of writing a high-level language is just so overwhelming that people, very few people have to write assembly these days. Everyone writes in a higher-level language.
49:04And so maybe that's it. It's the next level of evolution. It's the next step up in abstraction. That's it for this week's episode. I want to thank Matthew for his time. If you want to read a transcript of this conversation, you can find one, as always, on our website, eyeonai, that's E-Y-E hyphen O-N dot A-I. We love to hear from listeners, so drop us a line. And remember, the singularity may not be near, But AI is about to change your world, so pay attention.
From the publisher
Welcome to episode #130 of Eye on AI with Mathew Lodge. In this episode, we explore the world of reinforcement learning and code generation. Mathew Lodge, the CEO of Diffblue, shares insights into how reinforcement learning fuels generative AI.
As we explore the intricacies of reinforcement learning, we uncover its potential in game playing and guiding us towards solutions. We shed light on the products that it powers, such as AlphaGo and AlphaDev. However, we also address the challenges of large language models and explain why they may not be the ultimate solution for code generation.
In the last part of our conversation, we delve into the future of language models and intelligence. Mathew shares valuable insights on merging no-code and low-code solutions. We confront the skepticism of software developers towards AI for code products and the task of articulating program outcomes. Wrapping up, we reflect on the evolution of programming languages and the impact of abstraction on machine learning.
(00:00) Preview & sponsorship
(01:51) Reinforcement Learning and Code Generation
(04:39) Reinforcement Learning and Improving Algorithms
(15:32) The Challenges of Large Language Models
(23:58) Future of Language Models and Intelligence
(35:50) Challenges and Potential of AI-generated Code
(48:32) Programming Language Evolution and Higher-Level Languages
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI




