XBOW CEO and GitHub Copilot Creator Oege de Moor: Cracking the Code on Offensive Security With AI

10 Dec 2024 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Training Data - Episode Featuring Oege de Moor

Episode Overview

  • Title: XBOW CEO and GitHub Copilot Creator Oege de Moor: Cracking the Code on Offensive Security With AI
  • Hosts: Konstantine Buhler and Sonya Huang, Sequoia Capital
  • Guest: Oege de Moor, CEO of XBOW and creator of GitHub Copilot
  • Description: The episode discusses XBOW’s AI-powered offensive security system, which outperforms human penetration testers and rapidly assesses security vulnerabilities.

Key Topics Discussed

Introduction to XBOW

  • Company Overview:
  • XBOW specializes in automating offensive security.
  • Its AI system conducts security assessments in minutes instead of days.
  • Aims to protect software systems amid increasing AI-generated code and cyber threats.

The Importance of AI in Cybersecurity

  • Growing Need:
  • The rise of AI in code generation has made software development accessible to many, increasing the likelihood of security vulnerabilities.
  • Cyber attackers are also leveraging AI, leading to more sophisticated threats.
  • Automation in Penetration Testing:
  • XBOW automates the penetration testing process, which is traditionally skilled and time-consuming.

Performance Results

  • Benchmarking:
  • XBOW’s AI scored 75% on established industry benchmarks, later improving to 85% on original benchmarks created by the team.
  • In a comparison, a professional pentester took 40 hours to solve challenges that XBOW's AI completed in just 28 minutes.

Market Implications

  • Disruption of Traditional Security Models:
  • The episode discusses how AI will fundamentally change the offensive security market, which is currently small due to reliance on skilled human experts.
  • AI allows for continuous security assessments rather than infrequent testing.

Technical Insights

  • How XBOW Works:
  • Uses advanced AI models and tools to simulate attacks and find vulnerabilities.
  • Incorporates guardrails to prevent harmful actions during testing.
  • Comparison to Human Pen Testers:
  • Initially, the AI's approach resembled human methods but is expected to evolve, enabling it to discover vulnerabilities that humans may overlook.

Challenges and Considerations

  • Managing AI Risks:
  • The importance of ensuring that AI tools are used ethically and not for malicious purposes.
  • The need for safeguards and validation processes to ensure the efficacy of findings.

Future Directions

  • Expectations for XBOW:
  • Plans to enhance the product’s capabilities further and make it available for broader use.
  • The potential for significant transformation in web security over the coming year.

Notable Quotes

  • "This is a very large financial institution that everybody watching this podcast would have heard of, high confidence."
  • "We believe that the market will grow enormously."

Mentioned References

  • Semmle: Oege’s prior startup focusing on code analysis, acquired by GitHub in 2019.
  • HackerOne: Company known for its bug bounty programs.
  • The Bitter Lesson: Influential post by Richard Sutton discussing the value of general methods in AI.

Closing Thoughts

  • The episode emphasizes that AI is not just a technological advancement but a necessity in the fight against sophisticated cyber threats. Oege de Moor's journey from academia to commercial innovation underscores the urgency and potential impact of AI in cybersecurity.

Additional Resources

  • For more insights, listeners are encouraged to explore works by Richard Sutton and Dario Amodei, and to stay informed about the evolving landscape of AI in technology and security.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Because we now have AI code generation, everybody can create code. But not everybody knows about security. The models that generate the code have been trained on all public source code. There's a lot of vulnerabilities in all the public source code. And so we generate much more code with more security problems. On the other hand, attackers are already using AI to make their own work more effective. So we also have a greater threat. So more code, more attack, more attacks. That kind of makes the automation what Expo is doing absolutely essential.

1:00Today we are excited to welcome Uhe De Moore, founder and CEO of Expo. As the creator of GitHub Co -Pilot, Uhe has helped push the boundaries of modern AI. Before GitHub acquired his last startup, Summall, Uhe was a computer science professor at Oxford. His new company, Expo, is one of the most exciting AI native companies to launch this year. They're able to automate offensive security with an AI penetration tester. It's one of the best examples of AI services of software that we've seen. We're excited to talk to Uhe about the breakthrough results of Expo and what's next in AI.

1:45Uhe, Expo now matches the capabilities of the world's best hackers. Is this one of the first industries that's going to be completely disrupted by AI? Absolutely. It's going to completely change the way application security is implemented in the enterprise. Really, it's an example of service as a software. people will be able to replace a lot of routine human work with complete automation. And that will free up the humans to do the truly creative work themselves. So we could tell us about some of the results that you announced recently because I think they're really quite striking. So when we first built the first version of our product, we decided to try it out on renowned industry benchmarks.

2:46These are challenges that human hackers use to hone their skills. And we got these from a bunch of commercial providers, including Port Swigger and Pentastel app. On these benchmarks, our products scored 75%, which was amazing. In fact, it was so good that my first reaction was that surely there's something wrong here, but actually happening is the probably the benchmarks are so well -known that they occur somewhere in the training data and the mobile is simply regenerating the artist. us. Sir, we created a new set of benchmarks, completely original, guaranteed to be not in any training set. And those that scored even better, 85%.

3:34Wow. So then, the question is, so how good is that really? To answer that, we got in five professional pentastasers from reputed firms. And we asked them to solve exactly the same set of 104 challenges. One of these people is really at the top of the game. The very best type of pentaster, the kind of person that you'd asked to secure a multi -billion dollar hatch fund. And he scored the same. He scored the same as the AI, however, the human took 40 hours and the system took just 28 minutes. Yeah, that is striking. And when we first partnered with Uhe, it was science. I got to say we did not know if the AI would even be able to perform remotely as well as humans.

4:28And then when UHIT called and said, hey, Constantine, we've got some results to share that will blow you away. It certainly did. It certainly did. What do you think was your over under back in January, February, when it was still science as to whether an AI could perform at the level of these 20 -year season penetration tester experts? So at that time I didn't think that it would be achieved so quickly. I thought it would take at least a year to reach a reasonable level of proficiency. And even then, I would expect that it would work at the level of a mediocre human pentathlon, not at the level of the absolute top.

5:13In fact, since we announced these results, we've been working quite closely with a bunch of early design partners, and at one of them this morning, we found an incredible critical vulnerability of very surprising and the way it worked, if you look at what the AI is doing, It was called the web app and then it found some source code written in PHP and this source code was intended to access another another host But it used an insecure Signing algorithm in order to make that connection And so X -Bubble is able to get to the other house generate to generate links and access that. Nothing interesting found there.

6:07So then it continued crawling around and found another roundpoint and decided to try and use the same trick as I previously discovered. Didn't quite work. Needs another parameter. No problem. Browses around, find some more source code, I'm in JavaScript. Seize the number of candidate parameters, tries them all out, finds one that works, and now it has access to an endpoint and when it explores that it turns out it's intended to download PDF files. But not only could you could you could you have download PDF files, you could actually download a password file. So this is quite quite serious. And what I find fascinating about this type of example is that the AI is exploring like human pentastory.

7:01It's taking quite interesting creative terms that would be hard for most human experts. So just to summarize what you just said, this is a very confidentially confidentiality, obviously. This is a very large financial institution that everybody watching this podcast would have heard of, high confidence. And the AI was able to find a very advanced vulnerability. This is the type of institution that has human penetration testers constantly targeting it and trying to find vulnerabilities, a massive budget on security. It was able to find a whole file full of passwords. That's what just this morning.

7:48That's right. We have something like that every day, every old day. Wow. Congratulations on the results. Maybe can we take a step back and for those who aren't that familiar with this specific market, I've heard you and Constantine talking about pen tasting and I think Constantine called them hackers. I don't know if that's the same thing. Like what is the offensive security market? And you know, I guess how do you define the market that you're going after and what is Expo? Thank you for taking a step back. So offensive security is currently the best way to secure a software system. You invite external experts to call and simulate attacks against your systems and the report of whatever they find so that it can be fixed before the bad guys get happy.

8:41Now, this is a highly skilled activity. People, people meet years of training to do it. And it's expensive and slow. Typical cost of a so -called penetration test is something in the order of $18 ,000. Because it is expensive and slow, people only do it once or twice a year. But that doesn't make sense because there are systems evolved much faster than that. And so there will always be periods of time that insecure systems are out there. And what X -Bode does is automate this process, this highly skilled activity of launching simulated attacks and trying to find vulnerabilities. and because it automated, you can now run it continuously instead of just a month or twice a year.

9:44What drew you towards this market? I think Constan mentioned your background in founding SEML and having seen GitHub co -pilots. What drew you towards this specific market? Because it feels like there's a dozen teams going after AI coding. You're the only team I've met that is taking this specific approach to offensive security. So it was kind of the natural thing to do. So my previous company called SEML also was in security, but finding flaws in source code. And at SEML we had an offensive security team which reduced our product in order to find potential vulnerabilities and then our security researchers would find exploits and we would tell the world about what we found.

10:33Even at that time, it was kind of embarrassing to me that that last step of finding the exploits was done manually. Then when I was at GitHub, GitHub acquired our company at GitHub, I had the opportunity to find the co -pilot project. And so it was natural to now take my new found interest in AI and apply it to the challenge of automating offensive security. It was very lucky. One of the star researchers at Summer was Nico Weisman. He joined me in creating and one thing I'd love to ask you about, I think Expo is such an interesting case study for this brother thesis we have that, you know, AI is actually changing markets of yesterday that weren't as interesting, AI is actually really expanding and dramatically changing the nature of those markets.

11:34And I think this is a really interesting case study, so I'd love to dig into it a little bit more. The pen testing market is, you know, relative to, say, endpoint security or network security, It's a relatively small services heavy market today. And so to your point, offensive security is so important, and it's the gold standard, but it's a relatively small market. How do you think AI is going to change the nature of that? So first of all, it's small because it is powered by a small group of a highly skilled human experts. I think AI is going to change the market fundamentally in a couple of ways.

12:18So first of all, because we now have AI code generation, everybody can create code. But not everybody knows about security. The models that generate the code have been trained on all public source code. There's a lot of vulnerabilities in older public source code and so we generate much more code with more security problems. On the other hand, attackers are already using AI to make their own work more effective. And so we also have a greater threat. So more code, more attack, more attacks, that kind of makes the automation that XBOY is doing absolutely essential. So we believe that the market will grow enormously.

13:10We hear one of the, and Sonya, one of the analogies that I think about with this market is, is frankly, the adversarial nature of conflict, a human conflict. Cybersecurity is an adversarial game. You basically have two sides that get better and better equipment and they fight each other and it's a little bit of a game of cat and mouse, not completely unlike war. And physical conflicts and human history and one of the reasons why we think that this market is particularly interesting is think about how frequent war games are played in the military. In the US military or in any military broad war games, rent teaming, in fact rent teaming has been an initiative in most militaries for decades and centuries where you actually simulate a war game simulation.

14:04So this is a level of national importance. And really what you have built is in my eyes, the first ever AI cyber warrior. I mean, this is, I described it as a hacker because this is an AI cyber warrior that can do things that no software has been able to do before ever. And when you'd launched these results, I know with confidence because we talked about it, a bunch of people from DC called us up and said, whoa, wait a second, this is very consequential from DC and all over the West Coast. This is highly consequential. And I'm sure it didn't go unnoticed by adversaries to the West as well. and that they have probably been working on issues like this.

14:48So my question is, how do we stay ahead of the competition, true competition as a nation state competition, not business competition? How do we stay ahead of it and how do we make sure that Expo is a force for good in this massive adversarial cybersecurity game? So first of all, we stay ahead by moving very fast. and at Expo we were very lucky to work closely with several of the creators of big foundation models which are ahead of the rest of the world. We are also extremely cognizant of the potential of dangerous users of our technology. Therefore we've decided to make it available only in the cloud by making it available only mean a cloud and not in some downloadable form of software.

15:46We can actually control what scope is being used, it is being used on to launch attacks. And so we can recline our from our customers that they prove to us that the scope they wish to have tested is actually legitimately there and it's not being used to attack someone else. AI security warrior. Constantine, I think you're going to, you're new expose CMO. The head is incredible. Oh, hey, I'd love to learn about, you know, the, you know, how, how the product actually works and how the models work. How much of the magic of what you've built is you mentioned you work with from the major foundation model companies.

16:29How much of the magic of what you've built kind of exists in the foundation models versus things that you are building on top? So most of the magic is in fact in on top. We work with several of the foundation model providers and we are we're very happy that they are in stiff competition and they're playing hopscotch, they run pulls ahead, the other one pulls ahead and every time the foundation model gets better it benefits us but the true magic comes from the security team at Expo. We built some of the very best hackers in the world working for us and that domain knowledge is what informs how our product works.

17:16Can you double click a little bit into how it works? Is it prompt engineering? Are you fine tuning the models? I know that you probably want to keep your cards close to your chest as well in terms of how it works, but I'd love to hear the high level how you've built it. Sure. So I've already talked about these benchmarks that we use to evaluate our our product at the beginning. And that is absolutely key. Benchmarks, benchmarks, benchmarks, it's a live blood of a company of a product like this. And so we've organized these into kind of curriculum to teach the model how to solve cybersecurity problems better.

18:10And the benchmarks are critical to evaluate all the other changes that we make. And the other big components of our proprietary technology are the tools that we give to the LLM in order to forge easy tax. Human pantheroscopic has a toolkit of a bunch of things that they use in order to do attacks. But here it's a bit special because we want these tools to work well with our lands. For example, since we were focused on web security initially, we need a web browser that is driven by the alarm. you need to click around, you need to fill out forms, answer them, answer for it. And so we created a special browser to do that sort of thing.

19:06Thirdly, and this is pretty important, we need guard rails. Maybe first try to try to our product on some of these benchmarks. It struck me like an over -eater, super brilliant teenager. I would do lots of attacks and find something and then it got very excited and it goes, I did a sequel injection. Let me show you what I can do. Drop table. This is for the Strophic. If you use that at a customer, this is a big thing about our pen testing services. you have to make sure that you do not actually do the harm that a real hacker, an adversarial hacker would do. So we've been building guardrails to carefully watch over the shoulder of this brilliant teenager and stop it when it is not doing things that might be unsafe.

20:12Then there's an initial phase of attack surface discovery. So what we have is a fantastic exploit finder, but you have to point it at the right endpoint to begin forging an attack. And so this is running a bunch of tools and prioritizing where to go first. And then finally, as you already mentioned, those of course, prompt engineering, three of sorts, prompting to keep it on track and make sure it finishes one goal and when that finishes a goal, it goes on to the next, as of course. You described the technology as a brilliant teenager who sometimes overeager and maybe finds an exploit and actually drops that table.

21:00Some places in the world there are actors that don't have the same discretion to add those guard rails. What do we do to stay ahead of those actors and make sure that Expo can protect those that are doing good against them? So first of all we need all the obvious safeguards in in place. We need firewalls and also in that type of technology AI will also play a role. But first and foremost, we have to make sure that we find the vulnerabilities and the exploits before the bad guys do. But what expo is all about? How do you deal with hallucinations? I hear about people saying if my LM does 50 % or 60 % gets it right 60 % of the time, I'm good.

21:55I imagine security is one of those fields where that is insufficient. How do you deal with managing around the stochastic nature of the unpredictable nature of these LMs? So fortunately because it's automated, you just have to run it many times. Going back to your earlier question about the foundation models, what we do see is the better the foundation models get. the less attempts we need to make in order to find exploits. So it's kind of interesting how the influence how you are deployed and package in price or a product like this. Very much like humans if you if you get a human to perform the service for you, you actually pay for the time.

22:55How long they tried? How many things they tried? So we are thinking about doing the same kind of thing, charging our customers on the one hand subscription license. But on top of that, you can pay for attack hours. If you want to do a really thorough test and make sure that you absolutely find everything you can pay more and obviously that would then pay for the inference time on our side. I went into the pricing and packaging a little bit later because I'm very curious about that and I think you are one of the first kind of examples of services of software and so you are really paving the way in terms of how these things are priced and packaged.

23:42Before we get there you mentioned inference time compute and you know I think we're broadly very excited about what's happening as more and more of the compute is shifting from pre -training to inference time. What do you think the impact is going to be in your market? For us it can only be good when the value that we deliver remains constant. The price for delivering it becomes goes down. We see this even over this very short time that that X -BOW has media in existence, so we only expect that to continue. Uh, uh, on the probabilistic nature of these LMS, just revisiting that concept for a second, my mental model of what's going on is you have a state space with billions of possible states, the actions that the hacker can take, the actions that this penetration tester, this AI, penetration tester can take billions of possible states.

24:41And you've introduced this really intelligent, surestic as to the directions to go. You in theory could execute all possible states in perpetuity if you had infinite compute and infinite time. But reality, you have its constraints. And so I'm wondering, is that a reason why this might be the first or one of the first markets to enable full AI automation as in the stochastic nature of it? And the fact that even if you find one exploit, it's extremely valuable and you don't have the expectation of a complete exhaustive search. You do want to be it to be sure, but you find everything that a very skilled human being would find.

25:22And so this is why we, we've kind of exhausted our first set of benchmarks. People do these offensive security exercises not only to find availability, but also to have the peace of mind. That it's not easy to find stuff they didn't know about. And so we do have to make that. We have to present the evidence to our customers. that we do find everything that's built human beings would find and people will insist on having that reassurance. And when it is found, is it verified by a human or by the machine? We have a validator that automatically validates that the report is correct and reproducible before it goes to human.

26:28But of course, in the end, a human will have to take a look at it and fix the problem. Makes sense. O 'Hare, I'd love to dive a little bit deeper into the results that you've attained so far. So you mentioned, you're at 85 % on your current benchmarks at the level of the best human pentestres in the world. What have been the most surprising things that you found as you dig into the nature of those results? But the thing that I found most surprising was that we originally we only had benchmarks with particular instructions. So it would say something like you're going to test a web app for managing medical descriptions, try to log in and access the the prescriptions of another user.

27:23And it would do that successfully. But then we ran another task where we took the instructions away completely and just said, here's a web app, go explore. And the AI was able to find exactly the same vulnerability because it was able to read to us on the web pages and say, oh, this is biomedical of prescriptions, but we it's not a good idea that one user can access the prescriptions of another and so it would go and find a true mobility completely autonomously. I think that that's part of the reason that this technology is so exciting compared to all these security tools that came before because these LAMs have an understanding of the of the real world, it actually can assess what is important to go and test.

28:16It doesn't have to do this complete exhaustive search of all the possibilities. It can interpret what is important for this particular application. That's really cool. That's really cool. And then does the way that the AI system kind of reach its results. How does that compare to the way that a human pen tester would go about approaching the problem? I'm kind of thinking of, you know, AlphaGo and move 37, just, you know, very different from how we as humans would think about it. What is the model doing? So it's early days. Today it's very similar to what a human being would do. I completely agree though, but we have to be We already hear of rich shut -ons, bitter lessons in the end, because it learns continuous firm data on benchmarks, on more and more examples, it will start finding attacks that were unimaginable from a U .M.

29:19perspective. Which is a good thing. You say cautiously, I'm curious as to why, isn't that a great outcome? Yes, it's a great outcome. So I'm merely saying that today, when you read the traces, absolutely. This is what you would expect in good human to do. I fully expect that we'll go beyond that in a couple of months, certainly within years. Oeh, where do you think the biggest remaining room for improvement lies? And I'm curious, you mentioned looking at the traces of these models. Would you say that they are reasoning already today and is the furthest further improvement remaining in the reasoning area?

30:08Or how do you think about that? I think about that. But that's clearly the case. But most of the improvements will come from more data, more reinforcement learning on on particular examples and as we do a set, we all lead to a similar improvement to games like Loggo. How do you get more data that's just running more simulations? I imagine you've used a lot of the data there is. A couple of different ways. We have quite a few contractors, security experts, who create more benchmarks for us. There's also the opportunity of mining open source. So we've only recently started doing this. Just letting it lose on a large number of images on Docker Hub.

31:10And finding, just let it go. Every time you find something, that becomes a new thing that is can learn from. And so it might find it by doing a hundred attempts. And so in practice, if you had to do a hundred attempts, that probably wouldn't work at a customer because you would already get shut down because there's too many things, too many attacks happening clearly, clearly that shouldn't happen. But because it's open source, we can run a little ourselves. We can do 100 attempts. But now we have the data to try and make them ultra -batter to find it more quickly. You mentioned open source and Docker Hub.

32:00And so that obviously gets me thinking about GitHub. And you hear for those who don't know was the creative brain and creator behind GitHub Copilot, one of the most widely adopted AI applications in the world. Was there a moment when you were developing co -pilot or productizing it that you realized this AI is going to get so good that it's going to automate entire processes, what people now call agents. I'd actually take these actions on entire processes. And was there a moment where you said, hey, security is actually a very relevant area for this to happen? So, I wrote a memo in the summer of 2020 where I sketched what would later become a copilot.

32:51But also, we were already speculating that perhaps it will autonomously be able to fix bugs. just look at the issue, take a turn and now we see this about functionality emerging. So yes, I think that was pretty clear from the very beginning. I think the moment where I realized that would happen was I took a set of exercises, interview questions that I normally ask, of people at Hoxford and asked the world to solve them. And if you just give it one attempt, didn't do it. But if you give it a hundred attempts or even a thousand attempts, it would do most of them. And at that moment it was pretty clear that as the models get better and they need less attempts, they will be able to do these types of things.

33:50And one of the things that we also hoped it would be would be doing security analysis. So, admittedly, I didn't have offensive security on my bests in the summer of 2020 just yet. Any other lessons from productizing GitHub co -pilot that you think are relevant to sharing here? I actually think the most interesting thing about GitHub co -pilot was that it was done by such a small team. When we launched, we were only 10 people, something like that. And it's just a testament to how fast you can move. We're, say, a dedicated team of people that believes. How big was Expo when you launched the results?

34:34We were from 13 people. So actually quite big. Well, 13 really brilliant people. Since you were part of the co -pilot journey from the very beginning, I'm curious what you think of the current market for co -generation. A .I. Startups, it seems like it's one of the most crowded categories competitively right now. Do you think there's a path to building a company there and can one of these start -ups beat the incumbent GitHub that already has so much distribution? I agree. So I like a lot of what's going on, particularly at my other work, cursor, factory, but it's really difficult to compete with the distribution of a journal like a guitar.

Read the full transcript

35:21I do think that there may be an opportunity to go after a different market. So Gidepp is a braining supreme among professional developers. If you go after people who do not code for a living, there's an opportunity and replicate this quite well, for example. How do you think coding will transform in the future? Do you think the market that Replic serves doing village is be a dramatically larger and more important market as AI kind of continues to take over the world or how do you think coding changes? Yes, so I think the biggest change is going to be that many, many more people are are unable to create their own software.

36:08So that's a big transformation. But even for professionals, it will be much more about the conceptual ideas about sketching an architecture and then having the AI fill in the details. Longer term, I believe that we may be moving away from code as we know it today. The artifact that you make, I'll say developer, is the conversation with the model and so that is what you should store because that records what the code is supposed to do, rather than the details in a particular coding language. English is the coding language. So it's right. There's transformations that. Yeah, so the English is decoding language, perhaps with some diagrams, you explain it better, but it's just the next step in moving up in abstraction.

37:12Originally, it was all in machine language, and we had higher level programming languages, and now we're going to national language and images. So you talked a little bit about education and coding. I'm going to go down a little the version for a second. Because one of the amazing things about your life is a URA professor for much of it. And a very, very good one at that. So for context, Uhe was a computer science professor at Modeling College in Oxford. And Modeling is one of the most prestigious colleges at Oxford. He was one of the most amazing computer science professors. I got to study abroad at It's one of those incredibly serene places where they've got the deer park in the thousand year old buildings and the British man who tells me that the door at his entrance is older than my country.

38:03And all of the things that you would expect from one of the most prestigious academic institutions in the world, including Uge who was a professor and could walk across the grass whereas I, a mere student, would only be able to if I was holding his cape with his permission. And you left all of that to come into the commercial world with Semmel 15, 20 years ago. Can you tell us a little bit about your personal journey from leading academic at highly prestigious institution to commercial CEO redefining the cybersecurity I actually built into computer science because I loved coding. My very first program was a word processor to show my dad who was a professor of Semitic languages could type his manuscript on his computer.

39:03So when I started sitting computer science, I got totally taken by mathematics and the foundational theories. And so that's what I pursued as an academic initially. Then when I became a professor, I wanted to go back to my love of coding. So I started a new research group in programming tools, which eventually led to the spin -out that was that was semil. Well, I love the serenity and the peace and quiet of a place like Modern College. In our field, speed is incredibly important and speed can only be achieved with small teams that have a profit motive. It's just different from trying to event something because you have a paper deadline for an important conference.

40:10Or you've got to invent it because otherwise this important customer will not sign up. And I actually loved that additional excitement and pressure. And that's what led to me leaving Modeling behind and going full in on Samo. Love it, a great advertisement for capitalism.

40:43So, you know, I mean, if you do a really foundational work, a university, a university of Oxford, that's probably the best way to do it. But as soon as you start doing applied stuff, there's no place like a startup. I love it. I guess on that capitalistic note, I'd love to understand how you think about generating profits at Expo. And, you know, since you are one of the first agents, first, services as a software company, I think you're really going to set the precedent for how these types of agent applications are priced and packaged. So maybe can you just expand a little bit about how you're thinking about how to do that with your offering?

41:22Sure. So we would like to, we would like our product to run continuously as part of engineering processes. I mean that is the main value proposition that instead of doing append test ones or twice a year, you run it continuously after every change and immediately fix problems before they even reach production. So if you think about it like that, The most obvious pricing model would be based on the size of the engineering team, very much similar to products like GitHub Advanced Security. However, there is a different dimension here. Are we touched on a little bit earlier in the conversation? Some customers will want to do a super thorough test, really making sure that they exhaustively eliminated every possible of exploitable vulnerability.

42:32And in order to serve such customers we should have a service component to our pricing where you pay for if you pay more you get a more thorough test and the way we talk about this is in terms of attack hours. How many hours of attack do you get? And so if you buy a normal license, it's based on the number of engineers in your organization. And that comes with a fixed number of attack hours suitable for your environment. But then if you want to go more thorough, you pay that extra service fee in order to go deeper. That's super interesting. So you are really tapping into services like pricing models and budgets, but on the back end, you have the gross margin profile of software.

43:32I, right. And I, but I, I think that enterprise software is moving more and more towards a consumption based model as well. And here there is a, a very clear correlation between attack hours and the benefits to the customer. And I think that that's correlation between resources you consume and the benefits against as a customer that has to be very clear for a pricing model like that. My other takeaway from your modeling story was, I mean, you've always said this impact, interest and impact. You got into development tools because they touched people. You got into this because you know that this is going to change the world.

44:16I mean, you're highly confident that this type of technology is going to change the world with cybersecurity, whether it's us or someone else. I would put it a little differently. We absolutely must create a X -Bell because if we don't do it, the bad guys will get there first and so for sure. I mean, we do it because it's interesting and we think that it's a great commercial opportunity, but it's also an imperative. It's an imperative for the free world that we actually create this thing to protect all the soft earned free world. That has been so clear from minute one of meeting that that is the driver behind you and this brilliant team that you've assembled of academics and builders and technologists that are incredible.

45:08The other thing you mentioned was in the model that started with speed, the ability to move fast. Let me say you have moved really quickly, you and your team. What should we come to expect from Exbo in a year? What do you think will be the product that's focused on the product and technology impact? what will be happening from a product and technology and capabilities perspective a year out. You're going to replay this to me in the next board meeting aren't you? Four board meetings, Bellary, four. So I believe thought in a couple of months, so we're currently in a phase where we very carefully try out the product.

45:55with a select few early design partners. And the reason that we do that is because it needs this human supervision in order to control the brilliant teenager that we discussed, that we discussed before. Once we are over that, but face and we are confident as we can let it lose without any supervision, I think everything is going to move very fast. Our part of the reason is that this type of product is very easy to deploy. You can just point it at an existing service and immediately find results. So I would expect that by next summer we have significantly transformed the web security the state of web security, hopefully by demonstrating our work on our episodes, but also on platforms like HECK1.

47:01Okay, this has been one of my favorite episodes so far. Thank you again. Shall we wrap up with a quick lightning round? Go for it! Okay, awesome. Number one, favorite startups, other than Expo. So now, I love the way that you can just type in a few words. You get a completely original song. It's spine -chilling to me. The other start of my lack of latte is harmonic, applying AI to mathematical reasoning. Are you making soon those songs about coding in security? I know. There's a great deal. I said my wife a new song about sitting on the balcony at home in my house. There's a hard metal one about the AI cyber warrior.

47:55Well, toilet. I think we're actually going to need that. Okay. Perfect. At our annual AI event that we throw, we had Mikey from Suno there. soon over there and we crowd created an AI hot girl summer song. It was actually very catchy. That was great. That was great. Uh, who was there? Um, what other markets do you think AI is going to disrupt with this service as a software model in the short medium and long term? So in the short comment, but this is already, uh, this is already happening. Everything related to customer support is clearly going to be impacted by this type of technology. I think this is not exactly service as a software but I think that much of the problems we currently see with social media could be mitigated using this type of technology.

48:55I mean, you read all these reports about how social media is affecting the mental health of children all over the world. AI has the power to help with this type of problem. And long time, I think health and biology or the areas where this will make the biggest impact. What advice do you have for other startup founders? Focus on only one thing. Move as fast as you can. If you do those two things, then it will all come all right. Love it. One last question. And an optimistic note. What do you think is the best possible thing that can happen with AI over the next decade? I already touched on it. The opportunities in health and biology and to significantly expand health outcomes everywhere in the world is amazing.

49:59Dario Moda wrote this essay, Machines of Loving Grace, and I think he laid out very beautifully what the potential benefits of generating AI for all of us. Okay, thank you so much for joining us. This has been absolutely fantastic and we're so grateful to get to work with you and for the fact that you're building this on behalf of the right players, the people that are trying to do good in the world. Thank you very much. We'll be in a pleasure to be here.

51:35you

From the publisher

Oege de Moor, the creator of GitHub Copilot, discusses how XBOW’s AI offensive security system matches and even outperforms top human penetration testers, completing security assessments in minutes instead of days. The team’s speed and focus is transforming the niche market of pen testing with an always-on service-as-a-software platform. Oege describes how he is building a large and sustainable business while also creating a product that will “protect all the software in the free world.” XBOW shows how AI is essential for protecting software systems as the amount of AI-generated code increases along with the scale and sophistication of cyber threats.

Hosted by: Konstantine Buhler and Sonya Huang, Sequoia Capital

Mentioned in this episode: 

Semmle: Oege’s previous startup, a code analysis tool to secure software, acquired in 2019 by GitHub

Nico Waisman: Head of security at XBOW, previously a researcher at Semmle

The Bitter Lesson: Highly influential post by Richard Sutton

HackerOne: Cybersecurity company that runs one of the largest bug bounty programs

Suno: AI songwriting app that Oege loves

Machines of Loving Grace: Essay by Anthropic founder, Dario Amodei

More from Training Data

All 110 episodes
XBOW CEO and GitHub Copilot Creator Oege de Moor: Cracking the Code on Offensive Security With AITraining Data · 52 min
Listen in VO