In short
Eye On A.I. Podcast Episode #160 Summary
Episode Details
- Podcast Title: Eye On A.I.
- Episode Title: #160 Atul Deo: The Future of Generative AI at Amazon
- Host: Craig S. Smith
- Guest: Atul Deo, General Manager for Amazon Bedrock
- Air Date: [Insert Date]
- Sponsor: Shopify
---
Episode Overview In this episode, Craig S. Smith interviews Atul Deo, who discusses his professional journey and shares insights into Amazon Bedrock, a platform aimed at democratizing access to advanced AI models. The conversation covers the evolution of machine learning, the role of generative AI in software development, and the challenges faced in creating AI applications.
---
Key Topics Discussed
- Atul Deo's Background
- Transition from engineering to business roles.
- Experience with Yahoo and Amazon, focusing on M&A and product development.
- Amazon Bedrock
- Introduction to Bedrock as a service for building generative AI applications.
- No UI support; primarily an API for developers to interact with AI models.
- Access to a variety of high-performing foundation models.
- Generative AI and Code Whisperer
- The development of Amazon Code Whisperer, an AI tool for code generation.
- Current limitations of AI in coding, emphasizing the need for human oversight.
- Discussion on the evolution from simple autocomplete features to more advanced code generation capabilities.
- Retrieval Augmented Generation (RAG)
- Explanation of RAG as a technique combining generative AI with retrieval of specific information.
- Importance of keeping information up-to-date, particularly in customer service applications.
- Challenges in AI Development
- Discussion on the GPU shortage and its impact on AI scalability and application deployment.
- Differentiation between using GPUs and AWS’s custom chips (Tranium and Inferentia) for AI tasks.
- Foundation Models and AI's Future
- Overview of foundation models and their varying capabilities.
- The potential of multi-modal models that learn from diverse data sources.
- Predictions for the future of AI and its integration into personal and enterprise applications.
- Building Generative AI Applications
- Tools and resources provided by Amazon Bedrock for developers, including a sandbox environment and training courses.
- Customization options for models via fine-tuning and continued pre-training.
- Ethical Considerations and Guardrails
- Introduction of features for managing output, including guardrails to filter inappropriate content and PII redaction.
---
Key Takeaways
- Democratization of AI: Amazon Bedrock aims to make AI accessible for developers without deep technical expertise in machine learning.
- Generative AI's Evolution: Significant advancements have been made, but human involvement remains vital in overseeing and refining AI-generated outputs.
- Challenges Ahead: Issues like GPU shortages and the need for better AI models are current hurdles. However, AWS’s investment in custom chips is providing an advantage in this landscape.
- Future Vision: Atul Deo envisions a future where generative AI becomes integral to daily operations, enhancing productivity and allowing individuals to focus on more complex tasks.
---
Conclusion The discussion highlights the transformative potential of generative AI technologies within Amazon's ecosystem and the broader implications for developers and businesses. Listeners are encouraged to consider the advancements in AI and their impact on various industries.
---
Additional Information
- Follow Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
- Eye On A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
- Listen to the Podcast: Available on various platforms, including Spotify and Apple Podcasts.
---
Note: The singularity may not be near, but A.I. is fast changing your world, so pay attention.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Bedrock is kind of the service that people come to to build their application. So, but it doesn't offer any UI support or whatnot. It's for educational purposes. But the core problem with large language models being used and certainly in any safety critical application, but in any application for which ground truth is important is hallucinations. Now we make our own chips here. Chips which gives us a tremendous advantage in these times where, you know, there is shortage of GPUs. Hi. Good tech solves problems you know about. Great tech solves problems you haven't even thought about. What can the commerce platform trusted by millions of merchants do for you?
0:44It's time for Shopify, the commerce platform revolutionizing millions of businesses worldwide. Whether you're a garage entrepreneur or IPO ready, Shopify is the only tool you need to start, run, and grow your business without the struggle. Shopify puts you in control of every sales channel, so whether you're selling satin sheets from Shopify's in-person point-of-sales system, or offering organic olive oil on Shopify's all-in-one e-commerce platform, you're covered. Shopify powers 10 % of all e-commerce in the United States, and Shopify's truly a global force, powering Allbirds, Rothy's, and Brooklyn, and millions of other entrepreneurs of every size across over 170 countries.
1:37Plus, Shopify's award-winning help is there to support your success every step of the way. Sign up for a$1 per month trial period at shopify.com slash IonAI. That's Shopify, S-H-O-P-I-F-Y dot com slash IonAI. That's E-Y-E-O-N-A-I, all run together. Go to Shopify.com slash IonAI to take your business to the next level today. Excuse me, sir. I couldn't help but overhear. Did you say Shopify? Oh, Shopify.com slash I on AI. Oh, carry on. Hi, my name's Craig Smith, and this is I on AI. I just got back from Las Vegas, where I was attending Amazon Web Services' re-invent conference, a cloud conference for developers.
2:43And I had the opportunity to speak with Atul Deo, who worked on developing Amazon Code Whisperer, a generative AI application to help write code, and now heads Bedrock, a service that gives developers a choice of high-performing foundation models to create generative AI applications. He talked about RAG, retrieval augmented generation, and building agents that execute tasks. Atul gives a clear and clear-eyed overview of today's tech landscape within the context of Amazon's ecosystem and beyond. I hope you find the conversation as interesting as I did. Hey, this is Atul Dave here. I'm the general manager for Amazon Bedrock.
3:34I manage, as part of my role, I manage product and engineering for Amazon Bedrock. Prior to AWS, my background was in computer science. I'm an engineer. I wrote code for a few years after an undergraduate degree. And then I basically realized I was just in a cubicle writing code without understanding the implications of what I was actually writing code for. And I realized I should probably get educated on some of the business aspects. Went to business school, went to USC for that. And then after business school, I joined Yahoo. This was back in 08, 09, when Yahoo was going through some interesting transition with the Microsoft acquisition offer and whatnot.
4:14And then I worked there for the next few years, and I transitioned from doing finance, business operations, and then moving into corporate development, where I focused on mergers and acquisitions for Yahoo towards the end. And then I moved to Amazon in 2014 and joined Amazon's corporate development team where I focused on mergers and acquisitions again for AWS. And as part of that role, I got to work deeply with some of our product leaders, product and engineering leaders in AWS. And I really, instead of just being on one side of the table where I am evaluating companies to acquire them, I wanted to kind of actually be a founder myself.
4:54And I wanted to do that in a very low-risk way. And I realized that AWS is probably the perfect kind of place to do that. And this was around 2016, 2017 time frame when machine learning, the team at AWS was being formed. It was kind of thriving under Swami. And I decided to join the team. And since then, I've been involved in a bunch of different projects. I helped launch Amazon Transcribe. Then I launched a number of AI-powered capabilities for Amazon Connect, which is our contact center software services. Specifically, they're infusing machine learning into customer care. So things like I launched a service called Contact Lens for Amazon Connect, which essentially helps transcribe customer conversations.
5:37I'm sure you've called a customer care scenario, like you called your airline company or your credit card company with some complaint. And these companies want to analyze kind of what conversations happen between their customers and their customer care professionals, right? And essentially, contact lens allows you to extract a bunch of different insights from those customer conversations. What are the sentiment? What are the key topics that customers reached out? What are the trends? Did the agent, the customer care agent, did they actually behave properly? Did they follow all the rules that a company had laid out for?
6:08And essentially, here we use the power of AI to do this analysis at scale and help the companies understand this better. I think from that point onwards, I was like, I think I had invested enough in infusing machine learning into other products within existing products within AWS. And then generative AI, the term was not really coined back then in 2020. but we had started to see rise of large language models and these models were materially different than what we had seen in the past. So I embarked upon development of Amazon Code Whisperer, which is our ML-powered code generation service. And I helped launch it in Preview.
6:49And then since then, I've been involved in development of Amazon Bedrock. So Code Whisperer is highly specific to code. But at that point, I basically realized that this thing is way beyond code. The generalizable nature of this technology could essentially allow people to build a host of different applications just outside of even code and helping developers kind of get productive. So that's basically the short deal. Yeah. So Bedrock is a service to help customers build generative AI applications. You were saying earlier that you guys don't really use the term platform. It's not there. There isn't necessarily an interface that you're working in.
7:33Before we get to Bedrock and Gen AI, I want to well, Code Whisper is generative. So I wanted to ask about where you are with the companies with Code Whisper. I'm not a coder. And two or three years ago, I was talking to people who were imagining the day when you could use natural language to talk to a computer and it would write code for you. So when Copilot and then Code Whisperer and the other code generation tools came out, I got very excited. when AutoGPT came out on GitHub, I got very excited. I thought, great, now I can talk to a machine and it can code something for me. And there was a plethora of articles in the tech press about, you know, I coded this app with CodeWhisperer or something.
8:37But when I tried to do it, it really is an auto-completion tool. It's not a code generation tool. How far are we from having end-to-end code generation tools that then you could pass the code to a review AI that would spot problems and fix them? So I would just kind of expand on that, right? So code autocomplete has been around for a long, long time. Like you had things like IntelliSense. I mean, I used to write code, as I said, in my early part of my career. I wrote a lot of.NET and C Sharp and Visual Basic code. And obviously, you do the famous kind of IntelliSense recommendation would give you kind of the next one word.
9:31Or kind of probably gave you, after a few years, it started giving you a little bit more. But that is materially different than what is happening now. So with these large language models, it's not just a tiny bit of autocomplete. At the very least, you're getting a quick one, completion of the entire line of code, or you're getting next several lines of code. And you can just even get, just say in plain English, that, hey, I want to write a function to do so and so. And based on the context of your prior code or rest of your code in your kind of project, you'd get next several lines of code. In fact, it'd complete the entire function for you.
10:05That is materially different than what we had in the past, which is just simple autocomplete of one or two words. So the key difference, though, in kind of where you want it to be versus where things are today, you still have to have the developer's mind where you need to decompose your broader problem into a set of smaller problems and then basically tell kind of the whatever tool you're using, let's say Code Whisperer, I want to write a function to do and so and so. But usually, as you know, real apps don't only have one function. It's usually a combination of many things. So the developer today still has to put together some sort of a plan as, hey, here's kind of my logical flow, and here are the various things that I need to do.
10:50So that's basically what is happening. That's the current state. Is it possible to use an LLM like GPT-4 to give you the steps to break down a problem into these sub-problems? So that's basically the right way to do it today, where you can use some of the reasoning capabilities of the model to basically say, I have a big problem, kind of how do I actually come up with the right architecture? And because the model has seen a lot of code in the past, it will generally tell you kind of, hey, what are the best practices? And if you are building a specific thing in mind, here are the kind of key kind of classes that you may have to create.
11:32Here are the key functions, key methods, and kind of here's the general structure. And it might even tell you here are the next set of functions, and it might start giving you code. But I think the key point, though, is ultimately the human in the loop still needs to roughly understand and verify. It's not that it's going to give you an app that's going to be perfect. It might do that for simpler apps, but inevitably, if you are a builder, nothing plain vanilla is ever acceptable to you. You want to put your own twist and turns on the thing that comes out of the model. I mean, that's always been the case.
12:05Builders want to be creative. Yeah. And the code review tools, if you generate an app that combines multiple functions that have been written by CodeWhisper based on the roadmap given to you by GPT-4 and it doesn't execute, are the code review tools strong enough to go through and say, well, you have an error here? So two parts. One is I'll tell you the inherent feature of Code Whisperer is it identifies security vulnerabilities in the generated code, and it is not just in the code that it generated because usually it's a combination of what a developer wrote and what was generated by the model, and it will give you kind of a list of security issues in the overall code base as in the current window.
13:07But separately, even code reviews, we already launched a service called CodeGuru, I think in 2019 already, which is even before the advent of generative AI, where it could analyze lines of code and just do code reviews using power of machine learning. Wow. Okay. And then, so I'll play around with that. But Bedrock, describe then what Bedrock is. Sure. So Bedrock is a simple API that allows any developer in a company to access some of the best large language models or foundation models without the developer having to understand nuances of machine learning or having to manage any underlying infrastructure.
13:50They don't need to know kind of what GPUs, what instance, none of that. All they need to know is what is the application that they are building. And they just need to be able to interact with the model in plain English, give them prompts, and the model gives you a response. Yeah, I was saying to you earlier at the expo part of reInvent, which is where we are in Las Vegas, there was a booth for something called Party Rock, which is built on Bedrock. and it's a simple user interface to build Gen.AI tools or applications. The guy was showing me that you can pick your LLM. So to me it looked a little bit like an orchestration layer where you're picking your LLM and then you're talking to that LLM.
14:53Is that right? That's exactly what happens in Bedrock. So as I said, the developer has a choice of different LLMs in Bedrock. We support a number of them from some of the best startups in the company in the world, and including Amazon ourselves. So we have developed a set of foundation models, which we call as Titan. It's a family of models. And developers can pick any of these models because, you know, different developers have different needs for different models based on their specific use cases. They may have certain requirements on accuracy, latency, cost, and whatnot for their use case. And depending on that, they will basically choose the right model.
15:28Right. And then in PartyRock's case, it has a UI that, you know, you put your prompt in, and then it builds a very basic app based on that. is one thing that I was asking and maybe you could answer. The problem with, and Party Rock is at this point kind of a sandbox for people to play in. It's not producing production. It's for educational purposes. Right. But the core problem with large language models being used in any, certainly in any safety critical application, but in any application for which ground truth is important is hallucinations. Is there a companion service within AWS? Yeah. So it's part actually we launched a feature within Bedrock.
16:40We call it Knowledge Basis for Amazon Connect. So essentially, as you may know, there is a very common technique in the generative AI world that has evolved over the last few months. It is called as Retrieval Augmented Generation or RAG. Right. So essentially, the large language models are pre-trained on a broad set of public data. And they don't really understand two things. They don't understand latest information because they were trained up to a certain point. So the second thing they don't understand is they don't understand the context of your company and your particular documents. So even within the company, the documents can sometimes be changing because maybe the CEO of a company decided to change the return policy.
17:23The return policy was certain on November 30th, and it is a new return policy on December 1st. And obviously, if a customer comes and asks the question about return policy on December 1st, you don't want to be giving it the wrong answer, right? So sometimes you want the latest and the greatest answer. But you want also at the same time the capabilities of the generative model to give you the right answer. So essentially, with retrieval augmented generation, you extract the right snippet of information from your company's data. And you include that as part of the context for the model when you pass your request.
17:58And the model takes into account two things. It takes your actual specific request. For example, I may say, summarize this doc. You give a doc to it. But here you may say, hey, here's a bunch of information. Here's the question. And here's the pertinent information. The customer is asking about return policy. And here's the exact updated snippet from our internal documentation, which talks about this return policy, which was changed yesterday. And here's the thing. Now the LLM is able to take that into consideration while it crafts its response for the end user. Right. And maybe you can talk a little bit about RAG because I haven't delved too deeply in it.
18:38Yeah, just describe the architecture, what kind of database you're using. Sure. So let me just describe RAG first in more kind of technical detail, and then I'll explain what our specific solution for that is. So essentially think about a company. So the company has a bunch of documents. Now, the first step they have to do is they have to break those documents into smaller chunks and then pass those individual chunks into what we call as an embeddings model. Now, this embeddings model is not really a generative model. It takes text as input and it generates mathematical representations or vectors for those chunks of text.
19:19Now, these mathematical representations capture semantic meaning of the text, input text in a way that it can understand, like the model can understand better. And then, essentially, the company has to, let's say it has a million documents. It takes this million documents and, say, creates 10 million chunks. And then takes those 10 million chunks, passes them through the embeddings model, and now it has corresponding embeddings. Now, these embeddings need to go in a specialized store, which is called as a vector database. Now, this is like a new class of related products that has emerged in the last few months.
19:58And this vector database is optimized for storing embeddings, but also retrieving embeddings. Now, let's say the query comes in from our example before. I'm trying to ask about a return policy. Now, I can't compare text with embeddings. I have to compare embeddings with embeddings. So when a query comes in asking about return policy, I again pass that same query to the same embeddings model, and it generates the corresponding embeddings for the input query. And now the vector database does matching between what it has in terms of its corpus, the embeddings corresponding to the 10 million chunks, and the embeddings corresponding to the input query.
20:44and then it's able to quickly find out, hey, based on this end user's query, I was able to find that right chunk which corresponds to the return policy. And it then returns me back the specific text. It says, you know, the new return policy as of December 1st is so and so. So now I was able to extract that. Now the next step is I've extracted this information. Now somebody has to do the work, particularly a developer has to do the work to then take this in real time, run time, and then pass, append it to the prompt. So the original question to the model or to the application was, what is the return policy?
21:23Now, the developer has to take this information from the vector database, append that to the, before it passed, makes that request to the large language model, and then the large language model generates the response based on that final compile. It's called, basically, that's why it's called as retrieval augmented, because you're augmenting that output from the vector database into the prompt, and that's why you're getting the generated response after augmenting it. Right. I mean, I've had on the podcast, Adonuf, who's from Datastacks, which is a vector database company, and I've had Aido Liberty from Pinecone.
22:05Yes. But they didn't use the term rag. is that because they're only building the database part? So now the interesting thing is vector databases have been around for a while, right? But they did not have this popularity with the generative aspect. RAG has this retrieval augmented generation. So that generation part only became popular in the last, I would say, 12 months since the advent of the popularity of large language models. So now let me just switch gears. So we've discussed about what RAG is. Now let me just talk about what Bedrock does for RAG. I think the answer is pretty obvious. There's a lot of undifferentiated heavy lifting here.
22:48The developer, as I said, has to develop, let's say the developer has a bunch of documents in S3. S3 is our place where people store their data. And now from that point onwards, they have to take these documents, chunk them, pass those chunks to an embeddings model, then store them in a vector database which means they have to go sign up for some vector database, create an account manage permissions and then they have to manage this runtime workflow of hey when a query comes in I have to then create the embeddings for that input query get the retrieved answer, append it to the prompt and get a final response there's too much going on so with knowledge bases a lot of these aspects are automatically managed but at the same time while they're managed, we give control to the developer.
23:36So the things like they get control on what kind of chunking strategy they have to use. There are different ways of chunking. I use a very simplistic example of saying 1 million documents, 10 million chunks. But there are different ways in which people can go about it. They can also pick a different embeddings model. They may want to pick a different vector database. So we are natively integrated with Pinecone, with Redis, in addition to our own vector database, which is the OpenSearch serverless engine, vector engine. So in fact, with OpenSearch, customers don't even have to go to the OpenSearch console.
24:15If they say they want to use OpenSearch, we actually create the resources on their behalf. And they get automatically, the embeddings get created for them automatically, and they can go and access it. And the vectorization of the documents that you want to put into the language model, does Amazon also have that end of the workflow where you just upload a million documents? Exactly, exactly. So the only thing a customer needs to do here is upload a bunch of documents in S3. After that, Knowledge Basis for Amazon Bedrock does all of this heavy lifting for you. And it also has a query API. So you can simply pass your end user query, which is, hey, what's the latest return policy?
25:08And all the sausage making in the back, all they get is the answer, based on the latest return policy documentation. And they get an answer. And then they can do many things with it. There's also a generate API, which they can just pass this to the specific model and then get the final response. So all that management is automatically taken care of for the developer. Yeah. Okay, I have another question that's kind of related. I've been talking to people a lot and hearing people complain a lot about rate limits, about the quota allowed an individual customer in terms of queries per minute or tokens per minute is not big enough to build heavy enterprise level applications.
26:03And I've been having kind of a debate with a CEO of a company who I know pretty well who says, oh, that problem is being solved by being able to compute inference on CPUs instead of GPUs by distillation, by using sparse models and that sort of thing. But then when I talk to people, other people, one guy in particular at a company called Nomad Data says, no, this problem is fundamental. It comes down to GPU availability, which comes down to silicon starts, which comes down to$20 billion fabs. You just need more fabs. So where's the truth? Yeah, so I would say a lot of what you see in terms of limits, this is a true problem.
27:05If you look at almost every service out there, people are bottlenecked on GPU availability. Like suddenly we had this massive spike. I mean, a few things happened at the same time. We had COVID, supply chain crisis, rise of large language models, massive demand for GPUs. And now not this time for crypto mining, but for LLMs, particularly in inference where people are seeing the value in what generative AI brings to the table. and everybody wants to infuse generative AI. But at the end of the day, there is a compute shortage. So that is basically what is happening. But one of the things that differentiates AWS here is years ago, in fact, this was exactly the time, almost 10 years ago when I joined Amazon, we had started investing in our custom silicon chip development efforts.
27:53And we acquired a company called Annapurna Labs. It's an Israeli-based... I'm sorry, what's the company? Annapurna Labs. So it's an Israeli startup. And while the initial focus was mostly around networking, because AWS, that's the core. We are good at infrastructure. So that was the initial focus. But quickly, the team started investing in specialized chips for machine learning, and specifically machine learning training and machine learning inference. And that gave birth to what we call today as Tranium, which is chip optimized for training, and Inferentia, which is chip optimized for inference.
28:32Now we make our own chips here. So of course we are still a great partner with NVIDIA, as you saw in Adam's keynote today. But we also have our own chips, which gives us a tremendous advantage in these times where there is a shortage of GPUs. So what I've heard from the people who are saying rate limits are a bottleneck that's here for a few years. They say that big companies are experimenting, building pilot products, but no one is putting major generative AI products into production on an enterprise scale because they just can't get the token throughput. Is that true? I don't think that is the case.
29:33I mean, we have plenty of customers who are starting to go production and do very interesting things. As I said, I think if the company, if they are using certain cloud providers where they are only solely dependent on GPUs, then that is probably going to be the case. But as I said, in our case, we are diversified here. where in addition to GPUs, we are also using a lot of tranium and inferentia. Yeah. What about GPT-4, which seems to be the preferred model for the time being? Does the constraint exist there? So, GPT-4 today, GPT-4 is not available on Bedrock. It is only available on Azure OpenAI.
30:16So, I'm assuming that Azure OpenAI is probably running on GPUs. Right, right. But so maybe this is a problem that exists primarily with GPT-4, but other foundation models, depending on how their architect or which ships are using. I think in your framing of question, you touched upon a lot of interesting topics. So things like quantization, things like distillation. And so I think instead of going into that specific nuances of that, I do want to say that, hey, you know, people have different problems and you don't need the most expensive, your most powerful tool, your most compute hungry tool to solve every problem.
Read the full transcript
31:02And that is the approach which we have taken with Bedrock, where we have a range of different models of different levels of capabilities. So, for example, we have Titan Lite, which is a very small, nifty, fast model, which gives customers tremendous cost advantage, and it can be fine-tuned for their specific use cases. On the higher end, we do have Cloud, Anthropics Cloud, which is a very powerful and capable model, which customers find it very comfortable or better in many cases than GPT-4. Right. If you're using RAG in a vector database and you're only using the model to compile the language, the natural language, or to recognize the natural language on the input side, do you not need as powerful a model?
31:57I think RAG, it seems very simple on paper, but there are different flavors of RAG. If you have a very simple use case where all you're doing is just retrieving the context, appending it to the prompt and generating a response, then the answer is yes. But if it involves multi-step kind of reasoning and kind of thought processes, hey, where do I get this information? Do I need to execute some API calls? Maybe I need to understand the customer's history. Oh, the customer's history is stored in the DynamoDB table. Let me go get sad history. Oh, and then based on that history, I need to do so and so.
32:34So if there is a lot of complex reasoning involved in generating that final answer, because all of this, you know, what happens in this multi-step process is as I'm retrieving kind of new information, I just keep appending all of it and including it as part of the original request. and what started out as a very small, simple request, you know, obviously some simple questions can be pretty, simple sounding questions can be pretty complex in reality. So, and you may have to go kind of do multiple things to get to the final answer. So that is, if the answer, if it's a simple, single step kind of retrieval, augmenting, generating, then you should, in fact, use a simpler, smaller model because that is going to give you the best bang for your buck.
33:17But if you are going to require some complex reasoning orchestration, then you should use a more capable body. Yeah. So Bedrock, how long has Bedrock been around? So Bedrock, the service became generally available in September, end of September. Yeah. And as you said, it's a service. It's not a platform. So how does someone use Bedrock? So, again, I think the nuances of what is a platform we can go into, we don't usually call anything in AWS as a platform. But Bedrock offers a range of different capabilities, right? So, as I said, the most basic capability is giving access to the models via an API.
34:05But in addition to the plain models as an API, we also allow companies to customize the model. So, there are two ways of customizing. predominantly one is called as fine-tuning, which is the most common way of customizing and essentially involves use of label data. So essentially me saying, hey, this is a resume that I like, this is a resume I don't like. And if I'm a recruiting firm, I give 100 examples of that and I fine-tune this model. Next time I give a new resume, it will exactly know what's my preference and how to classify it based on my like or dislike. So that's an example of fine-tuning.
34:43There's another example of a way of customizing the model, which we've launched this morning, is called as continued pre-training. Essentially, it is domain adaptation. So it doesn't require any label data. But all you're saying is, hey, I'm a company in a particular industry, and we have a lot of jargon, company-specific. Here are a million documents. Just understand more about me and my company and what we do. and that way you can better answer my question instead of giving me some generic canned responses. And the advantage of this is, you know, RAG is good when you have a question and you need to inject context which is specific for that question.
35:24But sometimes you just want to give it kind of a general kind of education, the model general education on you. So this is the domain adaptation capability that we've added. So that's just recapping, right, in terms of what Bedrock does. One is plain access to the models via the API. Second is customization via fine-tuning and continued retraining. We also allow the managed rag experience. And then we've added a capability called agents. As I was saying, oftentimes when you ask a question, like, for example, I'm just going to give a simple, did you buy the groceries? I don't know. Let's take an example.
36:03The question can be simple, but you may actually have to go and buy the groceries or do some action as part of that question. So what happens with agents is the companies can provide the context of their data stores. They can give context of their APIs and basically say, hey, based on the end user request, go take some action, execute some things and make use of the data sources to answer the questions or execute their commands. And agents does all that without kind of, without developer having to manage the context of the session back and forth, without having to write complex prompts. So that seamless experience is provided by agents for Amazon Bedrock.
36:46And the last thing I would add is we've also launched something called guardrails. So today, the models themselves have some inherent guardrails, which they don't allow certain topics to be discussed, or they don't give responses to certain things. But companies also want to create certain rules or guardrails for their specific applications. And they may vary based on individual application within their own company. And they may want to apply multiple guardrails for their application. So we introduced or launched this new feature for Bedrock, which allows companies to basically say, hey, here are some denied topics.
37:24If a question from the end user comes on this, give the scan response. Or basically say, hey, there's certain topics like violence, hate, sexual in nature, and I want certain filtering. So there is a default filtering. You can set it to low, medium, or high. You can change the filters for inputs and outputs. We also have an upcoming capability here called as redaction, PII redaction. So you can redact the inputs going to the model, PII inputs, and you can redact the PII outputs coming out of the model. And also, last but not least, which is like a conventional block list. I don't want anytime a certain word is said, basically kind of either mask it or kind of basically say I'm not going to answer questions if a certain word comes up.
38:10So those are the various controls. In addition to that, we are continuously, this landscape is evolving so rapidly. So our goal is obviously giving customers access to the best large language models, foundation models is the starting point. But what we want to do is give developers a set of tools so that they can build their generative AI apps easily. And to access this, all you need is your AWS login? Exactly. And there's a tab or something for? Bedrock, yeah. Bedrock. And it has all the various functionality. And you can, I mean, you can even, there's a playground within Amazon Bedrock. You can just select a model and start interacting with the model by asking it some questions.
38:57And you can copy-paste some context and you can see what the model responds. In the playground itself, there are interesting configurations. You can set a temperature to, say, 1. Temperature 1 means the model is going to give more kind of wilder response. Temperature 0 means you want the model to be more conservative or kind of have less variations in its response. kind of like sometimes even humans, we tend to have our own implicit temperature on a given day. Right, that's right. And how do you charge for this? Is it by throughput or usage? Yeah, it's usually per thousand tokens. There's different pricing for the input tokens and different pricing for the output tokens.
39:42So typically what we see in production is customers tend to have a lot of input and outputs on the model tends to be more smaller. Right. Is it for a typical application, and again, I'm thinking personally here, is it prohibitively expensive for someone who just wants to build an app for themselves? It's not prohibitive at all. I mean, you can do it in tens of dollars. You can experiment and kind of quickly get a lot of value from it. Yeah. And how much support is there? Because AWS is famously self-serve. You can configure machines and everything on the website. But is it simple enough that someone can navigate on their own?
40:32It is simple enough. I think there are a number of things that we have done here. The first thing is we have a lab. So you can basically do zero-cost lab. You can just kind of go and do it in a sandbox environment and you can play. And you can actually make use of Bedrock to implement a few things based on the particular project. We also have a number of training kind of courses that we have launched on generative AI and specifically on Bedrock. And we're going to launch some more advanced courses over the next few weeks. And then for a lot of enterprise customers, we also have what we call Gen.AI Innovation Center, which essentially for companies who don't have a lot of experience in building Gen AI apps, we have a set of experts on our end who work closely with the right POCs in the company to build an app and help them get started.
41:23Yeah. And we started out talking about Code Whisperer. Is it integrated with Bedrock? Code Whisperer, so the way I would think about it is we have a three-layer stack, right? So the bottom layer, you have the infrastructure layer where you have the chips that I talked about. Then we have the middle layer, which is bedrock, which is kind of the, everything is built on top of bedrock in the Gen.AI world. And then at the top layer, you have the application layer. So Code Whisperer is an application. I see. So developers don't have to know about what model, what prompt. They are just doing their job.
41:54And what we also launched the top layer this morning was Amazon Q, which is, you know, it's an application for both business users and for developers. Right. Where do you think – we have a few more minutes. Where do you think this is going? You know, I talk periodically to Yan Le Koon, and he says the day will come soon where everybody has generative AI, personal generative AI assistant that they're interacting with throughout the day. and certainly companies, there's tremendous promise for companies and operating their businesses, optimizing their businesses. Do you have a vision of where this is going within, say, five years?
42:48So I think a few things are happening. I think, again, for me, usually I see history repeating itself. It's just different kind of intensity of kind of what the repetition is happening with. So in the past, we've had the industrial revolution. We've had, obviously, arrival of calculators, computers. And this is, for me, this is just the next generation of that disruption where today we can't imagine a world without our phones, without our computers. And with generative AI, with large language models, essentially, we will be able to communicate, get far richer responses on a personal productivity level.
43:25We can also offload a lot of repetitive tasks, multi-step tasks with things like agents. I think this is all happening very fast and furious. And I think the main thing is, though, that these things are often 98 % good today. And the last 2 % is where we'll really have to kind of, the space is moving rapidly. So my biggest kind of thing in 2024 is, how do we close the gap on the last 2 %? And the closing of the gap on the last 2 % is going to happen with better models and also with better tooling around the model. So things like the concept of RAG emerged very rapidly over the last 18 months. So there are going to be the next set of RAGs and other techniques that help us get closer to that 100 % mark.
44:18And that basically will be, you know, the dawn of a new era where we will be all more productive because we'll not be doing boring stuff. We'll be doing more interesting things. Yeah. I'm fascinated by the work on world models. I've had Yanlo Kunan talking about it, a guy named Alex Kendall with a company called Wave AI. They have a model called Gaia One. Are you guys working on world models as opposed to language models, models that learn directly from the environment or from video or other forms of unstructured data? So again, as I said, we are working, we have launched a set of 1P models, which we call Titan.
45:04And we are continuing to develop Titan in a host of different ways. Multimodal inputs is going to be critical for us because, you know, there are only so many things that can be captured just purely by text or by images or by video. You do need real world understanding. Because ultimately, if you look at our, like my, I have a five and seven year old, and I see kind of how they learn. They learn basically based on the real world. That is, they don't have to be fed the entire internet to get smarter. So they are learning rapidly from social cues from the real world. So I do expect models to learn a lot from the real world and physical world.
45:39Yeah, it's an exciting time. I mean, it must be crazy working at one of the big tech companies like Amazon. You know, even RAG. Nobody was talking about RAG that I was aware of six months ago. How do you keep up? For us, one of the best things we have going is an amazing set of customers. Since yesterday, already I've been in so many customer meetings where customers are relentlessly hungry. They keep coming to us with asks. and the dissatisfied customer is the best customer because that helps us keep us on our toes and helps us innovate. So as long as we continue to listen and capture that feedback, I think that's basically, I'm anticipating how we will stay up to date and kind of keep innovating.
46:31Hi, good tech solves problems you know about. Great tech solves problems you haven't even thought about. What can the commerce platform trusted by millions of merchants do for you? it's time for Shopify, the commerce platform revolutionizing millions of businesses worldwide. Whether you're a garage entrepreneur or IPO ready, Shopify is the only tool you need to start, run, and grow your business without the struggle. Shopify puts you in control of every sales channel. So whether you're selling satin sheets from Shopify's in-person point of sale system, or offering organic olive oil on Shopify's all-in-one e-commerce platform, you're covered.
47:19Shopify powers 10 % of all e-commerce in the United States, and Shopify's truly a global force, powering Allbirds, Rothy's, and Brooklinen, and millions of other entrepreneurs of every size across over 170 countries. Plus, Shopify's award-winning help is there to support your success every step of the way. Sign up for a$1 per month trial period at shopify.com slash IonAI. That's shopify, S-H-O-P-I-F-Y dot com slash IonAI. That's E-Y-E-O-N-A-I, all run together. Go to shopify.com slash IonAI to take your business to the next level today. Excuse me, sir. I couldn't help but overhear. Did you say Shopify?
48:18Oh, shopify.com slash I on AI. Carry on. That's it for this episode. I want to thank O'Toole for his time. If you want to read a transcript of today's conversation, you can find one on our website, IonAI. That's E-Y-E hyphen O-N dot A-I. And remember, the singularity may not be near, but A-I is fast changing your world. So pay attention. you
From the publisher
This episode is sponsored by Shopify. Shopify is a commerce platform that allows anyone to set up an online store and sell their products. Whether you're selling online, on social media, or in person, Shopify has you covered on every base. With Shopify you can sell physical and digital products. You can sell services, memberships, ticketed events, rentals and even classes and lessons.
Sign up for a $1 per month trial period at http://shopify.com/eyeonai
On episode #160 of Eye on AI, join Craig Smith as he sits down with Atul Deo, General Manager for Amazon Bedrock.
This episode delves into Atul's journey in the tech world, from his beginnings in computer science and business to spearheading major AI-powered projects at Amazon.
Atul shares his insights on the development and impact of Amazon Bedrock, a platform designed to democratize access to advanced AI models. He discusses the evolution of machine learning, the integration of AI in customer care through Amazon Connect, and the emergence of generative AI technologies like Amazon Code Whisperer.
Learn about the challenges and breakthroughs in creating AI applications, the significance of large language models, and Atul's vision for the future of AI at Amazon.
Don't forget to leave a 5-star rating on Spotify and a review on Apple Podcasts!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(03:30) Exploring Bedrock and Atul's Professional Journey
(07:14) Diving Deep into Code Whisperer and AI-Powered Coding
(10:54) The Future of AI in Code Generation and Review
(13:23) Unveiling Bedrock's API and Its Ecosystem
(16:36) Retrieval Augmented Generation (RAG) in AI
(21:51) The Evolution and Integration of Vector Databases in AI
(25:29) Confronting the GPU Shortage and AI Scalability
(29:32) The Landscape of Foundation Models and AI's Future
(33:27) Broadening the Horizon: From Language to World Models
(38:33) Closing Thoughts and Future of AI




