In short
Eye On A.I. Podcast Episode #204 Summary
Episode Title
Peter Guagenti: How AI Will Change the Way Developers Work (Tabnine’s Vision Explained)
Host
Craig S. Smith
Guest
Peter Guagenti, President and Chief Marketing Officer at Tabnine
---
Overview In this episode, Craig Smith converses with Peter Guagenti about the impact of AI on software development, detailing Tabnine's innovative contributions to AI-assisted coding and exploring the future of AI in programming.
Key Themes
- The evolution of AI in code assistance
- Privacy and security in AI applications
- The collaborative role of AI in software development
- Future trends and challenges in the industry
---
Detailed Insights
- Peter Guagenti's Background
- Transitioned from web developer and engineering lead to leading Tabnine.
- Experience in digital transformation and enterprise software.
- Tabnine's Origins and Vision
- Founded: By Iran Yahav and Jorv Weiss, who aimed to leverage AI for simplifying software development.
- Significance: Tabnine created the first LLM-based code assistant in 2018.
- Current Focus: Supporting large engineering teams with private deployments and a “protected model” ensuring user data privacy.
- AI Code Assistance Innovations
- Enhancing code generation and automating repetitive tasks.
- Protected Model: A unique approach ensuring compliance with licensing for enterprises.
- Personalization: Adapting AI tools to better fit the context and workflows of individual engineering teams.
- Autonomous Code Generation
- Spectrum of Capabilities: Ranges from simple assistance to fully autonomous code generation.
- Acknowledges the necessity of human oversight to ensure quality and relevance.
- Discussion on current tools achieving varying levels of autonomy, with emphasis on the need for iterative processes and real-world feedback.
- Misconceptions About AI Replacing Developers
- AI is more likely to empower developers rather than replace them.
- Focus on creativity and complex problem-solving as human strengths that AI can complement.
- Future of Software Development with AI
- AI can streamline maintenance tasks, allowing developers to focus on innovative design and advanced coding.
- Emphasizes the ongoing complexity in software development that requires human insight.
- Challenges in AI Deployment
- Need for skilled coders to manage and curate quality training data for AI models.
- Addressing the trust crisis in AI, where companies need to be transparent about their data usage and training practices.
- Performance and Productivity Insights
- Tabnine's tools reportedly enhance productivity by 20-25%.
- Example of a hedge fund using Tabnine for accelerating tech debt reduction and feature delivery.
- Code Quality and Review
- AI can assist in ensuring code quality through automated code reviews.
- Discussions on the potential pollution of training data due to low-quality generated content.
---
Conclusion Peter Guagenti emphasizes that the future of AI in software development is not solely about automation but fostering a collaborative environment where developers can leverage AI to enhance their productivity and creativity. The focus should also be on building trust and maintaining quality in AI systems to ensure their successful integration into the software development lifecycle.
---
Call to Action Listeners are encouraged to reflect on the implications of AI technology in their work and remain engaged in its evolving landscape. For more in-depth conversations and insights, follow the podcast on social media and subscribe for updates.
---
Sponsor Message
This episode was sponsored by BetterHelp, offering online therapy designed to fit individual schedules and needs.
Stay Connected
- Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00There's a spectrum of capabilities inside of all these tools. And it's not just code generation, it's break fix, it's other generation, like things like documentation, it's fully autonomous testing, it's all of these things. And we think about the spectrum as running from an AI assistant at one end to a fully autonomous AI software engineer on the other. I think the most important thing to note, if you want to understand how this stuff works, is that the model is only the foundation of the house. The actual interaction with those models is where all of the actual success is. Hi, I'm Peter Guagenti.
0:29I'm president and Chief Marketing Officer at Tab9. I've got a long history in enterprise software, including having built and scaled companies like NGINX and Cockroach Labs. But actually, I spent most of my career as a web developer and engineering lead myself, actually, in management consulting and in digital agencies. And I worked on over 100 different company projects over 15 years doing what we now call digital transformation. And tell us about Tab9. You know, I've heard a lot of people refer to Tab9 before this call. I did a little bit of research, but I'm not entirely clear about the company's history and its direction.
1:12So can you tell us about Tab9's origins and what it does and what it's looking to do in the future? Absolutely. So Tab9 is actually the originator of the AI code assistant category. So the company was founded by two very good friends, Iran Yahav and Jorv Weiss. Iran was actually a professor at the Technion, which is a computer science program in Israel. And they've actually been focused on how to leverage AI and machine learning to simplify software development. And this was actually, they started this work almost a decade ago and had been making very good advancements. And then LLMs emerged. And back in 2018, the team actually released the very first LLM-based code assistant for Java, and that was Tab9.
2:02And so fast forward to today, the company spent in that time actually building out the capabilities, ironically, before CIOs or engineering leads really appreciated it. So it was mostly just selling directly to developers. There's now over a million users of Tab9 on a frequent basis. Millions have tried it and used it over the years. And about a year or so ago, the company really pivoted and started supporting large engineering teams. As a significant player in the category, we had engineering managers coming to the company and saying, okay, we want this capability. And they were drawn, unlike some of the cloud provider tools, they were drawn to a few things for us.
2:44One is we've always done fully private deployments. So there's no code sharing. There's no user data sharing. it all either runs in your private self-hosted environment, either VPC or on-prem, or in a fully private SaaS environment, which we operate, where you have completely private individual environments. And they also were drawn to our model flexibility. So we created from the very beginning what we call the Tab9 protected model. So we have our own LLM that sits under both our code completions and our chat agents that's actually trained only on permissively licensed code. We do support other LLMs.
3:22We actually have a great model that we developed in partnership with Mistral. We work with all the large model providers, including Anthropic and OpenAI. But the protected model gives highly regulated or highly sensitive organizations an option to only use a model that they know they have rights to use the code that are underneath it, which is pretty critical for many of these folks. And then you asked the question around sort of where are we today and where we're going. We've really been focusing on innovation around AI code assistants and other software development AI agents that are highly personalized individual engineering teams.
4:01We have a thesis, right, which is that generative AI is great. But, you know, if you look at, say, the code assistant category and use generative AI, it's like hiring a seasoned software engineer off the street, right? You're going to end up with someone who has very strong knowledge and really understands the principles. But if you ask them how to do something, it'll do it in a way that's academically sound and correct, but not necessarily a fit for you. And we think that this is one of the things that's holding back wide scale adoption of generative AI is fit to purpose. And so we've really focused in the last couple of years and continue to double down on leveraging context, RAG, even customization to models and other connections to other systems to really inform the AI platform.
4:45So instead of behaving like a software engineer off the street, it behaves like somebody who's operated in the company for the last five, 10 years, right? A fully onboarded, fully experienced software engineer. And we think that's the future, right? We think the more context and personalization we can provide around these AI agents where they really do fit you and how you work and your team's behaviors and standards and guidelines, then the more value that it's going to deliver. Yeah, you know, everybody has been following, and I've done a number of episodes on automated code generation. A lot of the things that have been coming out from AlphaCode to, you know, certainly Copilot and CodeWhisperer.
5:27And most recently, I talked to the founder of Julius AI, and, you know, everyone's waiting for Devin to be for their general release. The question that's being asked about all of these is, at what point are they fully autonomous? At what point can they generate code that you can trust without having them operating as a pair programmer and without having them make suggestions? uh where it is i know you guys came out with this atlasian jira i don't know if it's a plugin but integration uh that i think is is fairly autonomous can you talk about that and then what what's the process to get to that full autonomy uh julius of ai who i had on fairly recently is focused on tabular data, on creating an autonomous agent that you can work with to manipulate tabular data.
6:41Is that where this is going, that there will be these narrow use cases that are refined enough that they're autonomous? or at what point does it break out into more general coding? Yeah. Well, look, the way we think about this is there's a spectrum of capabilities inside of all these tools, right? And it's not just code generation. It's break-fix. It's other generation, like things like documentation. It's fully autonomous testing. It's all of these things. And we think about the spectrum as running from an AI assistant at one end to a fully autonomous AI software engineer on the other end. And I think I would be careful about the hype that you see around some of the tools that are out there that are showing really fancy demos and things like that.
7:33And I can make a really fancy demo as well, right? I can make something that looks like it's, you know, going to completely replace a software engineer, but it's not realistic, right? It's not realistic. And I think to answer your question, this is going to sound slightly conflicted, but it's important to say it is, yes, there are things we can do fully autonomously today. So we released around the Atlassian Team 24 event. We're an Atlassian-backed company through Atlassian Ventures. We're also a huge Atlassian partner. And we want to get people excited in showing them a straight Jira ticket to code tool, which actually does straight app dev.
8:09It's not a, you know, here's some code and maybe you can go fine-tune it. It was taking a ticket and actually creating a full node application. And that is possible today, right? It's just the question is, where are the constraints, right? And where we're very focused is we're trying to take as many of these discrete tasks as possible that a software engineer is working through today and take it to as full completion as possible. So an example of this, obviously, the first one was documentation. We do complete documentation generation across multiple touch points. That's basically fully autonomous now.
8:48You don't have to worry about that. Yes, we believe that we should always have a human in the loop on these things. We think that there is significant value in having a human in the loop. So even with that full autonomy, we want human review and validation, just like we have code review and things like that from architects and managers today. That's what a software engineer should be doing as AI works. and a lot of that capability exists pretty deep in the product. It's for break fix. We're seeing it. We have almost fully autonomous testing now available. That's very strong. We've got a lot of this capability fully developed.
9:23I think to the point of where we get to where there's effectively autonomous app dev, I think we are not far away. And I think the way we're going to get there, however, is probably through an iterative process, right? And a slightly different user experience than we have today with these AI platforms, where we actually have AI working in conjunction with a software engineer to ask further questions and ask for clarification and to propose options and get feedback, which is actually how we work today, right? It's how we work as humans today. And so, you know, we're probably not very far off. We actually have in the lab, you know, pretty advanced versions of a lot of this stuff already.
10:05But I think what's going to happen is we're going to start building up that layer cake a piece at a time by creating fully automated versions and autonomous capabilities for pieces of it. Now, I want to say one really critical thing in this, because I think there's this, the people who are here getting most excited about this are people who are not software engineers. And frankly, the people who get most excited are people who don't usually really understand how software engineering works. So the idea that these AI bots are going to come and we're not going to have software engineers anymore is foolishness.
10:36Anybody who's ever worked in a microservices, data-intensive, modern application will tell you that the level of complexity in a modern application is such that it is highly unlikely AI is going to completely replace humans anytime in the foreseeable future. And what I suspect will happen is as we automate more and more of these tasks and more and more of these processes, the apps themselves are going to get more complex. It's going to unlock the ability for us to do even more interesting things. And you're going to see an up leveling of the software engineer themselves to get out of the doldrums, to get out of the sort of rote, repeated tasks that tend to eat up a lot of our time and more into design and evolution and really advanced thinking.
11:21and I look forward to that day, right? I mean, if you go back 20, 30 years ago, we used to do things like memory management and you just have to figure out how to write to disk and all these other crazy things that we did when we did app dev. As we automated more and more of those things, think about how advanced the applications got. I can't even imagine what the applications are going to look like 20 years from today. Yeah, although if there is this iterative process where you're talking to an LLM or to a generative AI system and refining your intent to the point that it can code something that's executable or compilable, that does open up development to non-coders like myself, who may not be creating extremely complicated apps, but can create apps that the code assistants aren't able to now.
12:23And do you think that's important or do you think that's kind of ancillary to the main direction of where all of this is heading, which is to enable engineers to write increasingly complex apps? I think it's an interesting question because I think there is a question here on what is a software engineer and who makes applications. That's really what you're saying, right? Because the only thing that the AI assistant really replaces is the need for us to handwrite code, right? Because at the end of the day, you think about a modern microservices-based application with lots of nodes and inflows and outflows of data and specific experiences that you're creating.
13:13You do have to know how to architect and engineer that in order to do that well, even if someone else is writing the code. I started my career as a developer and designer. Started out doing both user experience, designing the system, and then actually hand coding it. Very quickly got out of the code. and then eventually even got out of having to do all of the flows, right? And however, we're still architecting the actual product and the actual application. And, you know, just we kept moving up the stack, right? And all those things. And I think to imagine a world where someone who's a complete layman and doesn't understand how the system works to actually go and quote-unquote write an application is probably unrealistic, right?
13:52Because you actually still have to understand the interplay and interconnectedness of the data and the systems and the components around you. I'd imagine though that this starts to look 10, 20 years from now more like the way an architect builds a skyscraper, right? Where they don't have to go and plan for every piece of material and they don't have to plan for every component that goes in, but they have to have a vision. They have to understand how these systems work and interact with each other, at least if you want them to perform well and to be successful, right? I think you're seeing it already, probably sneak previews of it today, even without AI.
14:26functions as a service, drag and drop componentry. I spent a significant portion of my career working in open source frameworks like Drupal and WordPress and Magento. And at the time you could sort of Lego snap together applications, but you still really did have to understand how these things worked. And I'd imagine that, and this is pure speculation, obviously, but I'd imagine if you fast forward 10, 20 years from today, the people who are still going to be most successful building applications are the ones who understand the total system and how these systems work and pull together the right componentry, I just imagine a lot of it is going to be generated automatically by AI for us.
15:01And Tab9, you started as a Java tool. Are you in other languages now? Yeah, no. The very first LLM we built was basically just trained on Java exclusively, right? And that was, you think about, this is back in 2018. So this is now six years ago, seven years ago when we first started doing this. The LLMs were very limited. The context windows were very small. We had to be very focused. You fast forward to today, however, and we support over 80 languages and frameworks. We support anything that really any software engineer is typically going to work with. And then on top of that, Tab9, unlike most of our competitors, we actually do fine-tuned models.
15:46So we work with mostly very large engineering teams who are our enterprise customers. We have sort of individual developers and small teams who use our SaaS product, but our enterprise customers tend to be big banks. They tend to be defense organizations. They tend to be pharmaceutical companies or software companies or even hardware manufacturers. And so with them, a lot of times we're actually going and doing fine-tuned versions of our models that are trained on their code. And that starts to get really interesting when you do that, because not only does it make the behavior of the model stronger for your specific use cases, but actually it allows us to close the gap on obscure languages and frameworks that tend to be not a lot of great public training data for.
16:28You think about machine code for very specific hardware companies or obscure software languages that are used in financial services. You see a lot of that. And so that 80 doesn't even cover it. I feel like most of the time these days, when someone asks us for language coverage, we just immediately go in and start testing against against their use case and start immediately seeing that it's actually super performant. The powerful thing about these LLMs is these LLMs that basically know how to write software. So as long as they have the visibility to the structure of a programming language, they're incredibly successful with that language.
17:04Yeah. But then exactly that point that they know how to write software, That would argue that eventually they would be able to architect a piece of software, or at least guide a user in how to architect a piece of software beyond the coding. Can you talk about developing the fully autonomous JIRA tool? How did you train it and what its capabilities are? And then talk about how do you then take that process and broaden it to something else? So I think the most important thing to note, if you want to understand how this stuff works, is that the model is only the foundation of the house. The actual interaction with those models is where all of the actual success is.
17:57And a perfect example of this, go and use, you know, something like ChatGPT directly and ask it software development questions. And then go use tab nine connected to GPT 4.0. and ask the same exact question. And you're going to get a radically more evolved answer. And that's actually because the LLMs themselves are really just a font of knowledge. That's what they really are. And there's this whole other layer of work we do around context and what information actually gets sent into the LLM when you ask it a question in order to give it appropriate context and understanding, plus the prompt engineering itself, plus even just the UI UX.
18:37like, you know, things we do around helping to structure the question or ask follow on questions and things like that. And the way I think about this is and this is important context to explain the autonomous generation. If, you know, you and I met on the street and you asked me, you know, I'm looking to buy a car. What car would I get? Right. And I knew all about cars. I would give you 100 answers that are all potentially relevant to you. But if you're my best friend and I know you, you know, you've always wanted a convertible and you live in warm weather and you are, you know, you just had a windfall of cash and you really want something fancy and you want to be known.
19:15Like, I'm going to start directing you in a very specific case versus, you know, somebody who's new family. They've got three children. They, you know, they do the school run and they need something that's going to be safe and practical for that. You know, it's a totally different context. And that understanding is really key to be able to operate autonomously, because it's never enough to be able just to answer the question. You have to answer the question appropriately for the use case, for the context of the organization it's serving, for the users it's serving, for, you know, 100 parameters.
19:50And so what we do today natively in tab nine is we use local code-based awareness, global code-based awareness, and then integrations to non-code sources of information on top of all of our prompt engineering, all the work we do in order to make sure that the recommendations come back really strong. So local code-based awareness is literally looking at every open file. It's looking at all this, the things that are accessible from there, you know, things like the errors and other things visible from within the IDE. global code-based awareness is just that. It's a Git connection into whatever your repos are.
20:23And then non-code is things like Jira and Confluent or whatever. So why is that relevant to the autonomous? Because when you start giving it design parameters, it's not enough to have all the design parameters. It has to understand what else it's interacting with and how else it works, right? And so the Jira autonomous agent that we built is actually very simplistic, right? It's really just looking and saying, okay, this is the design specification. Now let's actually go and generate the application. It makes a bunch of assumptions around language. It makes assumptions around data storage. It makes a bunch of assumptions that are built into that.
21:02That if you have full context, then those aren't assumptions anymore. Now they're actually parameters that are being fed in. And so if you think about where this has to go from here, and we had this even in, if you watch the video of the Jira agent we built, it actually does ask clarifying questions, right? And we think that that's the future of this is taking that spec and then identifying what it doesn't know. And that's a really, really challenging problem, right? Because AI is not self-aware, right? So how do we figure out what it doesn't know and and what it needs to know in order to solve that problem.
21:42And that's really where this next wave of innovation is going to come from, is being able to resolve that, usually through a combination of what additional context we feed it automatically, what additional context we ask for, and then what clarifying questions we need to resolve. And I think that's going to get us to a place of where these agents are behaving more like Jarvis in Iron Man and less like your chatbot at the call center where you're just trying to get tracking information for your package, right? And that context, the global context you were talking about, how is that fed in to the LLM?
22:21I mean, do you have, do you load up a vector database with all of the, with the code base around the problem at hand, or have you just fine-tuned the model on a specific? In our case, it's actually a vector database, right? We're using retrieval augmented generation. What we have done, however, is we have a number of proprietary things that we've done around what specifically we're using as context and in what way. Because that's really the art, right? The art of getting these things to respond the way you want them to is all of the effort you do around that, exactly. And so what we have is basically a vector database that gets built.
23:03It gets preloaded with the parameters that we believe are appropriate. it. And that's both based on our local code base awareness. So if like you're one of our credit card based customers, it's whatever we can access with your permission from in the IDE. If you're an enterprise customer, then it's what an administrator gives us access to in their Git repo. That vector database for those customers actually does live in their environment, not in ours. And so does the model, actually, if they use one of our private models, it also lives in their environment. And so that's all self-contained. And we've basically have been testing and tuning this for years now around what is relevant and what is not relevant when we think about the prompt engineering around this and then what context goes in.
23:43LLMs themselves are not smart things, right? I think this is where we find them remarkable because of this idea that we can ask these sort of freeform questions. But if you actually start using them for very specific generative use cases, you think about there's a great article recently in Fortune about the rise of generative AI in law, right? What you're discovering is you go and ask chat GPT legal questions and it hallucinates like you wouldn't imagine, right? And it'll even tell you things that it'll tell you citations that aren't real, right? And it's already caught a bunch of lawyers out in cases live, right?
24:21And, you know, but then you go and look at what LexisNexis is doing and these folks who really understand it, they're doing some really advanced things once again to use their knowledge to check the knowledge. and eventually, and we're already starting to do this with some of our code review agents, which we'll roll out later this summer, we have models checking the models. So we have models sort of holding each other accountable, right, in order to eliminate any sort of hallucination or identify any issues. And I think that's the kind of effort that's going to be required for generative AI to actually be useful.
24:53Yeah, and on the vector database, presumably you have that workflow automated. that you plug into? Yeah. Recently, I was talking to people at Boston Consulting Group who are way ahead of the other groups on AI and generative AI. And they have been experimenting with loading, instead of using a vector database, just loading, you know, as these context windows get larger and larger loading uh you know a couple hundred pages of uh content of text or code or whatever it is into the prompt or into the i'm not sure if it's it's not really a system prompt well maybe it is a system prompt but the and then when you ask a question the inference comes out of that i mean it's like rag but instead of having uh to first create a vector database you just put all the stuff front loaded in the prompt.
25:59And understandably, that doesn't really work for a lot of use cases because it creates a lot of latency issues. But would that be an easier way to do it? I mean, this show is sponsored by BetterHelp. When your schedule is packed with kids' activities, big work projects, and more, it's easy to let your priorities slip. Even when you know what makes us happy, it's hard to make time for it. But when you feel like you have no time for yourself, non-negotiables like therapy are more important than ever. Therapy is no longer something to be hidden. I've benefited from therapy. Many of my friends have.
26:41If you're thinking of starting therapy, give BetterHelp a try. It's entirely online, designed to be convenient, flexible, and suited to your schedule. Just fill out a brief questionnaire to get matched with a licensed therapist and switch therapists anytime you like at no extra charge. Never skip therapy day with BetterHelp. Visit betterhelp.com slash IonAI today to get 10 % off your first month of therapy. That's betterhelp, B-E-T-T-E-R-H-E-L-P.com slash IonAI. IonAI all run together, E-Y-E-O-N-A-I. Go to.com slash IonAI for 10 % off your first month of therapy. You'll enjoy it. No, I mean, it's easier from an engineering perspective, but it's actually like the brute force method.
27:39I mean, I would argue that whoever is doing that doesn't actually understand how context works, right? Because if you're just loading that in, And you could do, an individual can do that with ChatGPT today and just copy and paste an entire catalog of information and ask it questions about that information. I think the thing that they're missing in that is some of that information is relevant and some is not. And it's not just about latency. Yes, it'll dramatically increase your latency if you're feeding in that much information for it to process in every single query that you're putting out. But also it might give you irrelevant information.
Read the full transcript
28:12You know, maybe that's good if you're trying to say, ask it to synthesize or summarize information, right? Or asking it to make conclusions about a specific body of knowledge. That's a good use of feeding that in. But when we're talking about what we're doing with software development, you know, there are things that are relevant and there are things that are irrelevant. There are good examples and there are bad examples, right? And you can feed it too much information and then get back mediocrity, right? And so I think when we talk about this, we're talking about relevance and quality. That's the things that we're really focused on.
28:46And what are the biggest weaknesses with generative AI today? Relevance and quality. And so when we're talking about context in this place, we're not just using the technical concept of context, but what is conceptually and intellectually appropriate context to answer the question well. that's really where all of the fine tuning of us as AI code assistants, that's where we're spending our time right now. And I think we're the pointy end of the spear for generative AI applications in general, right? I think this approach that we're taking is going to be leveraged by everyone. We're seeing it already with image generation, for example, where they're being very, very specific about what parameters end up getting sent in in order to make sure that the right imagery comes back.
29:31And that also includes constraints and controls, right? And that's a whole other area that we haven't even, we've only seen the tip of the iceberg yet, right? Which is generation is one thing, but what about constraining that generation in order to specifically fit certain use cases, right? And certain parameters. That's another layer of context that you're not going to achieve by just dumping in a bunch of material. You have to be really selective about what gets sent over. And this RAG approach, that works for specific, very narrow use cases. How then do you broaden that so that your Gen AI system can do more than that?
30:19I'm going to get out of my technical depth very fast on this. I am not a modeler. I am not a math PhD like our engineers are. Let me wrap it up in this way. It's not just RAG, right? I think what we're talking about here is the complexity of how we manage the whole experience, right? So it's RAG, it's semantic memory, it's prompt engineering. It's even just how the queries are posed and what information is sent over with the query in that prompt, right? So there's a lot of components that go in and what we've discovered is it's not an A or B or C, it's really an A and B and C and really tuning each of those for each of the use cases.
30:58So one of the things that we've we innovated and we're and we're proud others are picking up now is these sort of AI agents right in the chat right so things like you know explain this code to me, onboard me to this project, fix this code, find security vulnerabilities, you know those sorts of things. And each of those, the reason why they're structured as simple UX buttons or prompts, and they're structured the way they are, is because they have preloaded capability all behind them, right? And once we know your intent, right, this is what you're trying to do, then we can actually go and give the best possible answer based on an understanding of that intention, right?
31:35And so there's a lot of complexity in getting this right. There really is. And I think it's very easy to to not like to sort of understate the importance of these things and then at the same time be frustrated with the output when you're just asking unformed questions and not getting a good answer yeah but but to go from fully autonomous uh uh you know handling of uh jira tickets to i don't know what another uh obvious use case i mean where are you guys going with full autonomy? What's the next? We're working through the entire SDLC. We actually think that code generation is exciting and it sounds interesting, but where do we actually spend all of our time?
32:19We spend all of our time in maintenance. We spend all of our time in break-fix. We spend all of our time in things like refactoring, performance tuning, security issues. And actually most of the issues in software development, the stuff that eats up a lot of the release time is actually code review. So where we're going is working on building in as much autonomous resolution of each of those tasks as possible. And frankly, we're moving out of the IDE, right? We're moving into code review. We're moving into security review. We're moving into all these other areas. The other thing that we're really focused on is making sure then that the things that are being generated actually suit the company's standards.
33:01So, you know, later this summer, we're going to be releasing a new set of capabilities that allow you to constrain the AI. So, you know, been very focused, obviously, all of us have been very focused on making them able to do more. You know, now we're at a point, though, of where we're introducing, with all this AI-generated code, a potential risk with, you know, poorly trained software engineers or unmanaged or lightly managed software engineers now introducing code that just doesn't get accepted. Right. And doesn't make it past the pull request. So things like simple things like we follow the Google Java coding standards.
33:35Okay, then let's make sure that all of the code that is being generated or being refactored fits those standards. So they make it past the pull request. You know, and that's a very simple example. But the bigger examples are, does it use our functions? Does it use our APIs? Is it structured the way we write those things? Does it follow our security parameters? All those things. And that's a layer of AI capability that is generally untapped today by most of the tools. But us as the first to market, we're leaning into those now, right? And really saying, okay, how do we make this so it's even easier?
34:13And frankly, so we serve more people, right, in the software development team. Yeah, I mean, that's interesting on the code review because, you know, I'm a journalist, but not a practitioner. but I've been talking to people about, specifically on code, but on all generated content, the more generated content there is on the internet, the more that generated content is going to slip into training data sets. And if that content, if what's being generated is not the best that can be written, either in text or in code, you kind of pollute the training data. and with all of these people using code assistance, it's speeding up the writing of code and it's increasing the volume of code being written, but it's not necessarily increasing the quality of code being written.
35:25So does the code review take care of that problem to go through and say, wow, this is spaghetti. It needs to be refactored. Yeah, absolutely. It absolutely does. I mean, the things that we can do right now are things like making sure that it fits the standards and the structure, making sure that it actually runs well, right? Runs correctly. Making sure that it follows performance and security best practices. And getting a human in the loop also allows us to make that code review agent even smarter, Because if they push back or ask for clarification or reject something that is a recommendation of the code review agent, then we're going to make that quality even stronger.
36:10And I think you raise a really important point, which is where's the training data? What is the training data? Was it good? And I think a lot of the general purpose LLMs trained on anything they could find. And it actually made them weaker. It didn't make them stronger on these things. And so, you know, we've really heavily focused on really curating what training data went into our models. And then even when we fine-tuned, you know, we use, if you don't care about the copyright and license compliance, which about two-thirds of our customers don't care about license and copyright compliance, a third really do.
36:46The two-thirds who don't, we took the off-the-shelf Mr. All Open Source and we fine-tuned it because we discovered that it needed a little bit more guidance around how to perform certain tasks and what that looked like. So we used our curated materials to do that. I think as more and more content, whatever it is, be it code or imagery or text gets generated by AI, we're going to have to get more and more strict about what it is we're actually training on. right and you know i think there's a there's a debate around what volume of trading data is required for high performance you know sam malton very famously went in front of british parliament said oh we couldn't build these things without you know accessing all of the wealth of human knowledge in the world and then mistral comes out and releases a model that's that's a direct competitor in high performance that didn't right so so i think we're still only even learning now what what is required to make each of these each of these models perform the way we want yeah and And how do you curate the training data?
37:43I mean, that's because of the volume of data required. That's kind of a monumental task. It is a monumental task. We've been at it for years. So, you know, and frankly, a lot of these languages don't change. We're doing things like Java and Python and others that have been around for a very long time. So, you know, we've put in the work already on the foundation of these things. and we continue to add and curate and pour through a lot of that stuff as well. And it's not, there's no like, you know, silver bullet to that. You actually have to go and put in the work to be able to make sure that you're using good, high-quality sources.
38:23And does that mean having experienced coders read code and, you know, block? No one or trustworthy sources versus not. Yeah, I mean, Jensen Huang famously sort of poo-pooed the need to learn to code for future generations. But you're going to need people that can do that curation, if nothing else, so that you have clean code going in to these models. Look, I'd be careful about responding to things that sound good in marketing that aren't actually based on reality. Right. I mean, I think I and I say this as somebody who cut, you know, cut my teeth building brands. Right. And building building companies around around a reputation.
39:15Right. And the reality is, is that, you know, you are absolutely going to have to know how to write code. you are right you're gonna you're gonna have to at least know how to read it and know that it's performing the way you want it to and test it right all of those things and you know i think the reality is that curating the training data alone means not just that you have to understand the code you have to be able to actually assess its quality and capability right and i'll give you a real real example of where we do this today you know you look at open source projects that's a great example of this.
39:48Open source communities are filled with hundreds or thousands or millions of people issuing pull requests. Somebody still has to go and make sure that that code reaches a certain standard. And depending upon the criticality and the expectations of each of those projects, that bar may be pretty high. I operated in two very different open source communities for a lot my career so the drupal community you know lamp stack you know web content management and then nginx right and in the drupal community it's not mission critical for the most part it's just web experiences you know there is a standard for software development but they also really encourage more contribution from people who are not computer scientists and so the bar to get into that code base you know it is i wouldn't say low but it's but it's definitely approachable and accessible Nginx, I can count on literally on my fingers and toes the number of humans who've ever been allowed to contribute code to that project.
40:48Because of how high performance it is, how mission critical it is, how small it needs to be. It runs 80 % of the internet. So you actually have to really be strict about what gets allowed in and what doesn't. And so I think this training data and the models themselves and what they're based on, it's going to have a lot of those same super high standards. right it's going to have to in order for us to be be secure that and and and trust that it's actually going to be able to work the way we want it to yeah what kind of volume is required um of of to train a an autonomous uh code agent i mean is it is it uh something that a team of curators can come up with in a year or is it?
41:37I couldn't even answer that question. I think the reality is that we have some autonomy today. We're still figuring out where the edges are and still figuring out what the gaps in knowledge that are within that. The thing I would remember is LLMs, why were they successful with software early on? Why was that the first use case that actually worked? right? Because what LLMs are best at are things that are incredibly well-structured and are finite and well-understood. And what is more finite, structured, and well-understood in language than programming languages, right? So when we see the performance of these models in software, I think we're already seeing, okay, we've got more than enough training data, right?
42:26The LLMs are not going to be where the advantage is anymore. And we're seeing it. I think we assumed that there would be some level of commoditization on the llms we are already seeing it right we we have switchable models within tab 9 so you can run open ai you can run anthropic you can run you know the hour models you can run mistraw the different mistraw models we're going to roll out some others and we're already seeing that the relative performance between them when given the appropriate prompts is very very minimal difference very minimal difference and so So give it a reasonable period of time.
42:59The LLM is no longer going to be the thing that differentiates the application. It's going to be, what are you doing about it? What are you doing to actually shape it? If the LLM is knowledge and insight, how are you actually interacting with that knowledge and insight to actually generate something that is useful and appropriate for your use case? And this has always been the case, right? You think about all of software before this doesn't look that different from what we're dealing with with AI. It's a three-legged stool. You know, what's the model, what's the data, and then what's the application and use case in order to make it perform, right?
43:31The models are becoming commodity. The data you're using now to shape that experience becomes even more critical. And that's where we lean into things like context and institutional knowledge inside of an enterprise and an engineering team. And then there's all of the work, I think, when we talk about full autonomy around what then is that experience that somebody who's trying to build the application, what is it that they're doing? What are they having to ask? What additional information is required? What is the back and forth with the AI agent in order to be able to generate that application?
44:01Yeah. Although if you, when you say the models are becoming commoditized, those are the pre-trained models. But if you trained a model from scratch, I mean, the architectures are commoditized. But if you took the algorithms and trade them on highly curated code, you presumably would not need to then add fine tuning or connect it to a reg. It would have that. No, that's incorrect. That's incorrect. I mean, the thing that I think we as users really seem to struggle with is not understanding the value and importance of the prompt. What questions do you ask and how you ask those questions? And what your expectations are and what incremental insight you provide as you do that, that is what makes a generative application successful or not.
45:10It is that simple. And I challenge you once again, if you're unsure about that, go and use one of any, not just us, any of the generative AI applications to accomplish a task and then do the same thing through the chat UI straight into the LLM and go look at the difference. And you can even do this, for example, in Tab9, you can do it today. So we have the ability to turn on and off the workspace awareness. So you can just go free trial, tab nine, you know, 98 free trial, go and use it, hook it into your IDE. And then we have a toggle, workspace on or off. Open up a project, ask it a question with workspace off.
45:49You'll get an academically accurate answer that may or may not be hyper relevant to you. Usually not, right? It's going to be a generalized answer. You ask to write a function, it'll give you a good function. Then turn on workspace awareness. And then you'll all of a sudden see, oh, I get it. you are writing a function in this way, you do it this same way 10 times already in this application against this API. So that's what I'm going to now recommend, right? Take that to the next logical step for writing an application, right? Unless you are writing the code itself, you are, it's going to have to make a number of decisions and have to follow certain assumptions around you and what you're trying to achieve and who your user is and what it's for and how it's structured and what data sets to use and what data sources to rely on and where to read and write things.
46:37And there's so much that goes into that from an autonomy perspective. Where do you think that's going to come from? Right? It's not going to be native in the LLM. Right? Personalization and having these agents be aware of you, that is the future of generating. Yeah. And again, the way you guys are doing it is through RAG. RAG plus semantic memory plus prompt engineering plus, I mean, there's a bunch of components I'm going to technically, yeah. And I will butcher it if you ask me to explain it. So great conversation, follow up conversation with our CTO or our model builder. Are there any case studies that you can talk about of where tab nine has improved productivity by some particular method?
47:20Yeah, we've seen it. I mean, even just looking at how long we've been in use and how many developers we've had and how much code automation we've had, we've probably written 1 to 2 % of all of the world's code at this point. So it's a staggering number when you actually look at how many users we've had over the last six years and how much has been done. I'll give you a few statistics that are interesting and notably. I'll give you some that are tab nine and some that are just category, because it's actually very interesting when you look at the category. We generate somewhere between 30 to 50 % of all of the code on the projects that we touch for our customers today.
47:58Once you get to a certain level of maturity and using it, that's the rough range that we're at. The productivity impacts on that are very individual to individual teams. And so rather than explain sort of what I see at tab nine, let me give you Carnegie Mellon's view of this. Let me give you McKinsey's view of this, or IBM Research. There's a bunch of really phenomenal third-party research that's been done around this, where they've actually done true laboratory testing around how much automation are we seeing for what tasks in software development. And then where does a developer usually spend their time?
48:35Because it's not like we think about, all the vendors talk about 50, 60 % autonomous code generation at our peak, but how much time are we really spending writing code versus reviewing code versus maintaining code? So I love these third-party studies. According to the studies, and they all sort of land at roughly the same number, regardless of which tool they're using, which is actually really interesting because all of us are very similar in performance as tools. 20 to 25 % real world productivity savings for software engineering teams, 20 to 25%. I mean, you think about how much effort we would put in to get 5%, 10%, to be able to save one fifth or a quarter of someone's time and be able to redirect that towards your tech debt, redirect that towards improving the application as opposed to maintaining the application.
49:26I think this is incredible. When I first joined Tab9, I've only been in the company now seven, eight months. Very first thing my sales team asked me for was basically an ROI calculator. They want to understand what it was. That's how I ended up going deep on all these third-party research studies. And it was staggering when I did the ROI calculator because Tab9 is$39 per user per month for our enterprise product,$12 per user per month for our pro product, right? The productivity savings we're talking about are$50 ,000,$60 ,000,$70 ,000 per engineer for a product that costs you hundreds. So it becomes kind of a joke when you look at it where you're like, this is like a fundamentally better way of working.
50:09The ROI is sort of clear. The question is not ROI, but where and how do you deploy it? And then what do you do with that time savings, right? Where do you reinvest that energy? Yeah, yeah. Because it could be, as you said, in code review or something like that, but it could also be in writing just more code, right? That's exactly what we're seeing. So we work with a hedge fund. As you imagine, hedge funds and these folks are incredibly competitive, like incredibly, incredibly competitive. Anytime I think we're competitive in software, I can then go look at financial services and I realize we're nothing.
50:46The reality is they've been using us for quite some time. They've got it in use across their entire engineering team. That's what they've been doing with it, right? So they've actually, what they've reported back to us is they are increasing the burndown of their tech debt, increasing their time to market for new features and for new capabilities. And that's what we're all looking for. We're all looking for that little bit more competitive edge, that little bit more capability and convenience in what we do every day. Or, frankly, with the layoffs we saw last year and tech debt increasing and all the pain that's come, if you're an engineering manager, you're looking at your teams like, my team's already working 10 hours a day.
51:26They're getting burned out. Quality's dropping because of that. I'm losing my best people. You know, I look at these agents as coming at just the right time to be able to actually take some of the weight off of some of these people. So it's not even necessarily that you're going to take that 20 % savings and reinvest it. You're going to have your team, at least in startup land, you're not going to have your team working 10, 12 hours a day. Okay, we're up to about an hour. Is there anything I didn't cover that you would like to have listeners hear? You know, I think there's a, you've done a great job, Craig, with covering the innovation in the space.
52:03I think your interrogation of the data side is very interesting because I think there's a lack of understanding around how this stuff really works, which I think is leaving us in a place of where there's not, there's almost unrealistic expectations with what the LLMs can do out of the box. and that risks of putting us into a trough of disillusionment with these things, right? Because we have inflated expectations out of context, right? And so I think that's really good. I would throw one other thing out for you, which is actually I think the biggest existential crisis in generative AI today is actually we have a trust crisis.
52:40And I think this is a real issue. And we hear it in a very different way, right? I think most of the time when people look at our category, they think of Copilot, right? They think of GitHub. GitHub came in after us, but with a giant megaphone and Microsoft's deep pockets, they've been able to go and insert themselves into engineering teams all over the world. But at the same time, what we hear when we talk to Fortune 500 companies is they don't trust what those companies are doing. They don't trust their behaviors towards training data, and they don't trust what's happening with how the models were built.
53:15They don't trust that the CTO of OpenAI can't explain in front of a journalist what it was trained on. I can tell you, I can give you a list today of every single piece of code that Tab9 was trained on. So why are they not answering that question, right? And, you know, then you look at some of the behaviors. There's a great New York Times piece where they actually exposed the big tech brands changing their terms of service so they can go and exploit the data. They showed them actually violating each other's terms of service. There are active lawsuits against OpenAI and GitHub and others for what they've been doing with regards to everything from journalist content to imagery to code.
53:56You know, there's a ton of this stuff out there. And I think there's a problem with all of this, right? There's a problem with all of this, which is we are risking stunting our own future as technology brands by not doing more to build trust, by not doing more to respect our users, by not doing more to respect our creators, right? And so, you know, I think that's something as an individual, like forget even my role at Tab9 for a minute. I've been in technology now for 30 years and I've been in multiple waves of innovation, the birth of the web in 94, you know, the birth of an explosion of open source in the 2000s and all of that.
54:34And I've always personally felt like I had two responsibilities. I had a responsibility as an entrepreneur to build a great business, but I had a responsibility as a citizen and technical thought leader to see that we're doing things the right way, right? And that we're doing things in a way that actually betters us all, right? And I think we are really at risk right now with generative AI, really benefiting a very small few and hurting the others. And frankly, that I feel like is short-sighted thinking that ultimately hurts all of us. It's a tragedy of the commons. And so I'm hopeful that my peers and others wake up and realize that building trust, not breaking trust, is a critical part of what we're doing.
55:15Building relationships, not exploiting relationships is a critical part of what we're doing. and that the real success of AI will come when we're building opportunities for everyone to benefit from it. And I know that's certainly the attitude that myself and TapNine take.
From the publisher
In this episode of the Eye on AI podcast, we sit down with Peter Guagenti, President and Chief Marketing Officer at Tabnine, to explore the role of AI in software development.
Peter takes us through his journey from web developer and engineering lead to leading Tabnine, a pioneer of AI code assistance.
We delve into the innovative ways Tabnine is pushing the boundaries of AI, from enhancing code generation to ensuring privacy with its Protected Model—offering enterprises fully private AI solutions tailored to their specific needs. Peter discusses how Tabnine is addressing the challenges of fit-to-purpose AI, making AI tools more context-aware and personalized to the workflows of individual engineering teams.
Peter also sheds light on the future of AI in software development, addressing the pressing question: Can AI truly replace developers, or is it destined to be a powerful collaborator?
Learn how AI can elevate software engineering teams, helping them overcome the repetitive tasks that slow down progress and focus on the creative aspects that push the industry forward.
Don’t forget to like, subscribe, and hit the notification bell for more in-depth conversations on the latest AI advancements.
This episode of Eye on AI is sponsored by BetterHelp.
If you’re thinking of starting therapy, give BetterHelp a try. It’s entirely online. Designed to be convenient, flexible, and suited to your schedule. Just fill out a brief questionnaire to get matched with a licensed therapist, and switch therapists any time for no additional charge.
Visit https://www.betterhelp.com/eyeonai today to get 10% off your first month.
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(00:38) Peter Guagenti's Background
(01:20) Tabnine’s Origins
(03:49) Innovating in AI Code Assistance
(05:10) The Path to Autonomous Code Generation
(07:49) Human Oversight in Autonomous AI
(10:17) Misconceptions About AI Replacing Engineers
(14:15) Future of Software Development with AI
(17:04) Autonomous JIRA Tool and Broader Applications
(22:36) Leveraging Vector Databases for Context
(27:34) Balancing Contextual Data for AI
(29:54) Expanding Generative AI Use Cases
(34:16) Ensuring Code Quality with AI
(37:17) Curating Quality Data for AI Models
(41:17) The Need for Skilled Coders in AI
(42:59) Future of Generative AI Beyond LLMs
(47:00) Case Studies: Tabnine’s Impact on Productivity
(51:49) Conclusion: Building Trust in AI Technology




