In short
AWS Podcast Episode #754: Accelerating Healthcare Decisions with AI Agents
Episode Overview
- Hosts: Jillian Ford with guests Gigi Yuen (Chief Data and AI Officer) and Kenji Fujita (Staff AI Platform Engineer) from Cohere Health.
- Release Date: March 20, 2026
- Main Topic: The role of AI agents in enhancing healthcare decision-making, particularly through the use of Cohere Health's AI-powered solutions.
Key Themes Cohere Health's Mission
- Objective: Improve collaboration between payers (insurance companies) and providers (healthcare entities) to streamline healthcare processes.
- Challenges: A significant portion (20-30%) of healthcare spending is wasted on administrative tasks, hindering timely patient care.
AI Implementation and Trust
- Trust and Transparency: Essential for AI solutions, especially in healthcare where decisions affect patient care.
- Collaboration with Experts: Importance of involving clinical experts from the beginning to ensure accuracy and reliability in AI outputs.
Cohere Review Resolve™
- Functionality: An AI copilot that assists in making medical necessity reviews more efficient and accurate by analyzing structured and unstructured data.
- Expected Outcomes:
- Reduction in review times by 30-40%.
- Faster patient care access and improved patient outcomes.
- Increased accuracy in clinical determinations.
Detailed Discussions Challenges in Healthcare Automation
- Complex Decision-Making: Balancing automation with human oversight in high-risk scenarios. Key decisions must involve clinical expertise.
- Avoiding Hallucinations: Establishing strict guidelines to reduce errors in AI outputs.
Design Principles for AI Systems
- Evaluation-Driven Development: Defining metrics for success and continuously assessing AI performance.
- Human Oversight: Implementing systems that ensure human review for high-risk decisions while allowing automation for lower-risk tasks.
Technical Implementation
- Amazon Bedrock AgentCore: Chosen for its ability to meet stringent security requirements and facilitate quick development of AI agents.
- Ease of Development: Mentioned benefits of using AWS tools, including rapid iteration and reduced operational concerns for developers.
Metrics and Data Utilization
- Importance of Ground Truth Data: Collecting and utilizing accurate data to inform AI decisions and model evaluations.
- Model Evaluation Framework: Establishing a leaderboard to assess the performance of various AI models regularly.
Key Takeaways
- Integration of Domain-Specific Data: Understanding the "why" behind incorporating specific data into AI systems helps organizations tailor solutions effectively.
- Human-Centric Approach: Engaging clinical experts and understanding user needs is crucial for successful AI deployment.
- Continuous Learning and Adaptation: The AI landscape is constantly evolving; companies must remain flexible and open to testing new models and approaches.
Advice for Companies Entering AI Development
- Evaluate Your Metrics: Define what success looks like for your AI applications.
- Invest in Team Diversity: Create cross-functional teams that include domain experts, developers, and data scientists.
- Start Small with Testing: Use AWS resources to prototype and iterate on AI solutions before full-scale implementation.
Conclusion The conversation underscores the significant potential of AI in healthcare, emphasizing the importance of trust, expert collaboration, and a clear evaluation framework. Cohere Health's approach serves as a model for other organizations aiming to harness AI effectively within their operations.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Cohere Health's Mission
0:45 to 3:42
Gigi Yuen discusses Cohere Health's role in streamlining healthcare operations.
“Gigi Yuen and Kenji Fujita from Cohere Health.”
AI Solutions in Healthcare
3:42 to 5:16
Gigi explains the importance of trust and transparency in implementing AI in healthcare.
“There's just a lot here, I think, especially from healthcare, but I think other folks who are in different industries can even see some parallels to some of the challenges that they're thinking about.”
Incorporating Domain Expertise in AI Development
5:16 to 8:00
Discussion on engaging domain experts during the AI development process.
“Also, there are a lot of important experts' opinion we have to take into account.”
Monitoring AI Performance and User Personas
8:00 to 9:23
Exploring the importance of monitoring AI performance and understanding user personas.
“You know, I grew up in an era where we talked about test-driven development.”
Incorporating Domain-Specific Data
9:23 to 12:43
How companies can effectively incorporate domain-specific data into AI systems.
“So maybe you can tell us how companies can really think about incorporating their own domain-specific data in terms of AI.”
Human Oversight vs. Automation
12:43 to 14:00
The critical balance between automation and human oversight in healthcare AI.
“That is such a good call out because I know I've seen companies already start going down the route of maybe using a specific data set for an example.”
Human Oversight in AI Healthcare Decisions
14:00 to 14:47
Learn how human oversight is critical in AI healthcare applications, especially in high-risk scenarios.
“But we've made a clear decision, regardless of regulatory and which geography we operate, AI will never use to deny a patient's care or to even deny a provider's request.”
Prior Authorization Automation Examples
14:47 to 17:20
Discover practical examples of how AI automation is applied in prior authorizations for different medical procedures.
“But I think it's in a high risk scenario where human oversight is, you know, in terms of every single case and every single nuance detailed.”
Building AI Platforms with AWS
17:20 to 19:55
Understand the considerations for building AI systems using Amazon Bedrock and AgentCore.
“Let's get into how this has actually been built.”
Transitioning to AgentCore
19:55 to 22:38
Learn about the challenges and advantages of transitioning to AgentCore for AI implementation.
“how aws was really able to to help you with everything you were saying earlier like the technical challenges, the business challenges that you had, and to be able to actually build it in production today?”
Show all 20 chapters
Evaluating AI Models in Healthcare
22:38 to 25:47
Explore the importance of evaluation driven development when choosing AI models for healthcare solutions.
“I want to go back to something you were saying earlier, because I think this will resonate with a lot of folks who are listening.”
The Importance of Data in AI Development
25:47 to 28:00
Understand how data quality and expert involvement are crucial for successful AI deployments in healthcare.
“So I'd love to hear how you think about looking at evaluating different models and the term I love that you chose earlier, evaluation driven development.”
Importance of Metrics and Data Labeling in AI
28:00 to 30:00
Learn how metrics and data labeling influence AI decision-making in healthcare.
“So that, I think, is keeping an open mind.”
Evaluating AI Models for Specific Use Cases
30:00 to 32:10
Discover strategies for assessing and selecting the right AI models.
“There is a lot to really unpack there that I think every single listener can get a takeaway from.”
The Impact of AI on Operational Efficiency
32:10 to 34:20
Understand how AI systems can improve productivity and job satisfaction.
“I think having the platform, investing in the platform is key to this style of development, both the evaluations framework and the gateway.”
Key Considerations for AI Implementation
34:20 to 36:55
Explore essential steps for companies to successfully implement AI agents.
“So 85 % of decisions are made within minutes.”
Navigating Challenges in AI Development
36:55 to 39:50
Learn about the challenges and best practices in AI agent development.
Advice for Starting AI Projects
39:50 to 42:01
Get practical advice for organizations looking to begin AI projects.
“I think these are all prerequisites that people really need to think about.”
Advice on Building AI Agents
42:01 to 43:56
Listeners receive valuable advice on initiating AI agent projects.
“I've got one last question for each of you.”
Masterclass on Development Strategies
43:56 to 45:11
A deep dive into evaluation-driven and AI-driven development practices.
“wanting to dig into agent and so forth is think hard about the team you assemble to to make this real think um So agentic work truly does require a new profile of developers and a development team.”
Transcript
Automatic transcript. May contain errors.0:00Jillian:This is episode 754 of the AWS podcast, released on March 20th, 2026. Welcome everyone to the AWS podcast. I am your host, Jillian Ford. And this episode today, I am super excited about. I think there's going to be something for everyone here. I know agents is really top of mind for, I mean, let's face it. It's like every single person on the planet is probably thinking about this right now. And you get to learn from two people who have been in the trenches at a company that has not only just been thinking about this, but actually has business critical applications that are using agents today.
0:44So I'm really excited to talk to Gigi Yuen and Kenji Fujita from Cohere Health. So Gigi Yuen is the Chief Data and AI Officer, and Kenji Fujita is the staff AI platform engineer at Cohere Health. So there's something here for everyone, whether it is you're someone who is thinking about agents and how do you apply it, maybe you're in a highly regulated industry, and maybe you want to understand how to use it in AWS, some advice from these two. We're going to cover all of that. All right, let's get started. So So Gigi, I'd love to understand first if you can tell our listeners about what is Cohere Health and the specific business problems that you were thinking about within the healthcare industry.
1:38Well, first of all, thank you for having us on your podcast, Julian. What a privilege. So Cohere Health, we are a clinical intelligence company. And our mission is to streamline the payer and provider connectivity and the collaboration. so just take it a second when i say payer and provider what do i mean by that payers are insurance companies or non-profits or government entities that finance health care services whereas providers are entities like a doctor's office hospital systems provider groups who provide the care so i guess providers provide care and payers pay for the care so We have millions of providers in the United States and hundreds of payers.
2:23So you can imagine with these two entities, in order to support the whole healthcare ecosystem, there are a lot of transactions, a lot of administrative tasks, right? Pile off, claims processing, payment, quality, care coordination, and unfortunately, the fraud ways and abuse that comes along the way when you have all these back and forth. And if you look at different studies, most recently health affairs published an article where about 20 to 30 percent of health care spent in the state are spent on administrative tasks, 20 to 30 percent. And some of them are necessary. Some of them are avoidable or I would say low value.
3:07And depends on which studies you read, it's about half of this 20 to 30 percent are low value administrative tasks. So that's about half a trillion dollar. So that's the problem space Coher Health is set up to solve. We want to eliminate the waste. When you do that, what does that mean, right? Patients can get the right care faster. Providers can actually focus on what they do best. And the payer can really, really do a good job financing their care. So that's the nutshell. Wow. There's just a lot here, I think, especially from healthcare, but I think other folks who are in different industries can even see some parallels to some of the challenges that they're thinking about.
3:53So when you were addressed with these challenges, how did you think about it in terms of implementing AI solutions? yeah um it is a big problem but it's also a very personal problem a healthcare is very personal even when you're simply talking about getting a bill that you don't understand why or having to wait a few weeks to get an answer for to get imaging it's very very personal so as we think about ai solutions it really has to do with trust and transparency it is the utmost important thing is the trust. And remember, I was just reading to my kids Bernstein's books, Bernstein's famous books.
4:39Love those books. Right? And there's this book that talks about once trust is broken, you can't get it back. So I think as we as technologists and think about building and rolling out these AI solutions, we have to get it right the first time, which is very unforgiving in this notion of stochastic and deterministic world. So that's something that is top of mind for us. So what does that mean? We need to make sure what we do is reliable and consistent. There's no tolerance for hallucination, none. Also, there are a lot of important experts' opinion we have to take into account. It has to be clinically sound.
5:22A lot of literature that we lean in, a lot of guidelines that we can count on. But the funny thing is, when you put two doctors in this room for the same case chances are they don't agree on everything right that's why we love to get a second opinion right so as we think about building ai we have to take all that into consideration like what do we mean when a system is performant based on whose opinion is that and what are the nuances that we have to account for where we really need to say human needs to be in the loop yeah um i think i think last but not least is um the security and privacy aspect.
5:59And we can get into some of that as we think through when we kind of get our hands dirty in building a system. That's a lot of technology safeguards, but also process safeguards that we have to consider.
6:13These are some themes that I think a lot of businesses are thinking about, regardless of what industry they're in, especially with AI. I think the bar just keeps getting set higher and higher, which is great because now what customers want to be able to implement to be able to serve their end customers is going to be an even better, more accurate, more performant application for them. So I'd love to understand, how are you thinking about no hallucinations, ensuring that it really is a safe and reliable application? I think let's put it in terms of the software development lifecycle. Yeah. So when we are doing design and kind of the reference architecture, it's easy to have the technologists in the room.
7:08I think, especially in healthcare, and I imagine in many, many nuanced domains, we must have the experts in the room on the get-go. It is common for many healthcare startups to talk about performances, talk about clinicians in the loop, but I think there is a difference when you engage your domain expert in the beginning versus at the end when we simply ask them to do validation. In career health, every single development project to have clinicians on the team and in the loop as opposed to waiting until we've already built the prototype or already about the launch of solution asking them to validate.
7:52And I think that domain expertise is key to ensure that we're measuring the right things. I think that leads to my second point. You know, I grew up in an era where we talked about test-driven development. And I think now, especially with a lot of this agentic solution, we really have to move to the mindset of an evaluation driven development where eval comes first what are the metrics that are important how are we going to track we talked about no no tolerance for hallucination it took us a month to iterate on exactly how we quantify hallucination in particular clinical settings right so all those upfront work is more important than ever I think that's key.
8:36Once the solution is launched, now we are at the monitoring and tracking phase. Obviously, having 24-7 monitoring is key, potentially using the Gen AI itself to help as a judge. But I do believe in the importance of human audits. Having that regular, intelligently sampled human audit is really, really critical. And I'm going to say one last thing, and Kenji, you may have something to add too, because we've been working together on this. It's we're learning that there's different personas that will use the AI solutions, especially since healthcare is such a personal space. And even learning about how do you roll out to different persona groups over time help us build a more trustworthy and useful applications.
9:32Yeah, the one thing that I would add there is the key focus for us at a lower level is having strict guidelines and standards so that all patients get unified care while allowing for some level of user preference when it comes to how that care is received.
9:54there's so much to impact so i'm glad that we've got more time i'm going to ask so many different questions but let's dive into that what kenji was just talking about that uh user you'd use the words i think user standard care but really like focusing on the end user and that sounds like you also need to be able to have that domain specific data in order to be able to bring it back to providing the best care or the best experience for that end user. So maybe you can tell us how companies can really think about incorporating their own domain-specific data in terms of AI. It's a great question. I think technically there are many ways, like case-specific context or pre-training of models, the spectrum is wide.
10:45But I think instead of talking about the how i want to talk about the why and the what just for a moment right like why like ask ourselves why do we want to incorporate domain specific data is it because i want the solution to run faster run cheaper be more accurate be more contextual um is it a matter that you want more code control over your system's output but especially given the industry operating and to Kenji's point, where there's a standard of care that we want to maintain. Like, I think understanding the motivation behind using these domain-specific data will help us pick the right data and at what point do we incorporate them, right?
11:30So there are instances that would make sense to incorporate more of a knowledge system, right? Like in our world, it's more around medical society guidelines, standard of care, ontology framework, incorporating them will allow us to have a more consistent framework in how the AI operates. But then when the goal is to provide more contextual, accurate response in every single interaction, then we need to have our AI be able to access very specific individual case notes. And they are relevant and they are important. It goes back to to why yeah and um last but not least put a plug into the notion of data use rights is really critical right when we entrust it with patient data and when we entrust it with business sensitive data what rights does cohere health and our partners have to use the data for what purpose is really really critical and it's quite nuanced right are you using it for operations versus are are you using it for learning or are you using it for educational training?
12:39All that nuances have to be accounted for. That is such a good call out because I know I've seen companies already start going down the route of maybe using a specific data set for an example. And then it's the engineers who get really excited about the problem. I want to find out later on that like, oh, sorry, we can't use it because of like the rights to the data for maybe it's like compliance, licensing, whatever kinds of reasons. So that's a such a good call it that I think will definitely help a lot of the listeners. Same with what you were saying earlier about having those experts really from the early stages of the process instead of the human in the loop just being the end part.
13:24So maybe you can help us really share your thought process of how do you decide when to actually automate versus actually having that human oversight that's part of the process? It's the million dollar question, right? It's risk and reward. I think human in the loop makes a lot of sense when we're working with high risk, high reward cases. And like, for instance, in our case, we have a good number of solutions that target prior authorization automation, trying to get the patient to the right care faster and with less paperwork. But we've made a clear decision, regardless of regulatory and which geography we operate, AI will never use to deny a patient's care or to even deny a provider's request.
14:20Anytime there's any doubt that we cannot say yes, we will always have an expert of the same specialty to review the case and even sometimes have a conversation with the requester and bring in that really important human conversation in the loop so so to us like we draw a clear line of when we will not we'll always have um let me take that back and to cohere health we always draw a clear line on when we will not automate is when we have to say no to a patient's care but kind of going back to your question earlier julian about human oversight even in the cases where we choose to automate i think oversight is still essential i think it's just a matter of at what point does it come into play yeah um oversight should always happen with the system design reference architecture and how we design the eval and oversight should always happen without it.
15:19But I think it's in a high risk scenario where human oversight is, you know, in terms of every single case and every single nuance detailed. And figuring out the middle ground, I think is the key to figuring out how to scale. Yeah. Maybe you can give us like an example that I think can help the listeners maybe visualize what that could kind of look like in their own business? Yeah, sure thing. So for instance, going back to the prior authorization automation example, we could look at a variety of clinical areas, right? Like we've, a lot of us have experienced getting a prior off for imaging, trying to figure out what's going on, right?
16:04A diagnostic reason, or you go get a prior off because you had to get a knee surgery, right? depending on the clinical use case, the risk tolerance is quite different, right? It's something when you will need to bring me to your operating room and kept me open versus, you know, getting an MRI, which is, you know, taking half an hour of your day and with minimum radiation exposure. So really considering that I would, you know, the way Cohe Health approach it is for a knee surgery decision, we better be absolutely confident before we automatically, you know, say yes to the case without a human review.
16:38whereas for a diagnostic imaging we will likely say um as long as there's no contraindications as long as patient risk is considered as long as it's covered by your policy so there's no financial risk we'll go ahead and say yes right without a human intervention so that that nuance it's um it's that's why the human the human experts are domain experts in the loop as we design the system is so critical.
17:08I love that. Really having a framework for, with the experts assessing the actual risk and then using that to be able to design how the human in the loop, human oversight is part of the entire process. Let's get into how this has actually been built. So Kenji, I'd love to understand really, what are some of the factors that your team was thinking about that ultimately led you to choose Amazon Bedrock and Amazon Bedrock Agent Core? Sure. Yeah. So I think timing was a huge factor for Cohere. We, this past year, have invested a lot of time and resources into building out a platform around our AI.
17:59And a lot of that incorporates how do you scale with the new agentic services that are out there in the market. And so we were attending workshops with AWS for AgentCore. We were evaluating the different components. I think the key thing that stood out right away was the speed of innovation here, the ability to build these agents quickly, effectively, and safely. A couple key components for us that I think most developers can get held up on are memory and MCP. And so it was clear to me that these were paramount for the service teams at AWS when they were building out AgentCore because the tenancy concerns that we have in a highly regulated space like healthcare are covered with some of the components of memory client out of the box.
19:01And implementing them only takes a couple lines of code, which to me is a huge benefit. The gateway is another thing. So we had started building out our own MCP servers, but with the identity built on top of gateway and the different targets that the gateway provides, we've been able to at least iterate on our research and scale out the potential use cases for our agents. Because, again, it only takes a couple lines of code to implement an entire MCP server target.
19:42and that speak it says a lot that you're saying you're able to build it quickly effectively and safely because bedrock agent core is relatively new so it sounds like there must have been some folks at aws that really helped you so maybe you can share some of the some uh how aws was really able to to help you with everything you were saying earlier like the technical challenges, the business challenges that you had, and to be able to actually build it in production today? Yeah, I think what helped us was a little bit of handholding around understanding our use case, right? Our top concern is always the security of the data when it comes to implementing a solution like this.
20:29And so the first thing we brought to them was how do we transition our short-term memory to agent core so that we have tendency separation?
20:42The workshops that we went through with some of the Jupyter notebooks that they had available had everything that we needed right out of the box. And so going through some hands-on experience in a test environment was super helpful. And then understanding the documentation was also key. And I think the namespaces that the AgentCore memory client provides are very intuitive to set up.
21:13I'd love to know, based on your experience between the workshops, the documentation, is there anything else that kind of stood out to you that can help folks who are on that journey of implementing agents? Yeah, so I touched on it a little bit, but I think the amount of experimental research that we've been able to do has far exceeded what we thought we would be able to do by this point. And a lot of that is due to the ease of development here. So I'm trying to think of a good example to share. There are two different use cases that we have. One was transitioning an existing service over to agent core, right?
22:03So we had spent all of this time setting up and like evaluating and setting up memory for a chat agent and transitioning over to the memory client was a relatively trivial task for our team to implement. And a lot of that was due to the documentation and the workshops that they were able to attend. But then there's also net new development. And what's clear to me is that AgentCore really takes away a lot of the ops concerns from the MLE. So the developer on the ML side, who typically wouldn't be a DevOps expert, doesn't have to evaluate those concerns as heavily because a lot of it is handled by the AgentCore service.
22:55I want to go back to something you were saying earlier, because I think this will resonate with a lot of folks who are listening. MCPs are really a huge hot topic right now. And a lot of people are thinking about building it themselves. So I'm curious from your experience when you were at that stage of you had started building it yourself and then you would start then use Bedrock Agent Core. If there was any other learnings that you had from that experience that you can help someone else who's really thinking about building it myself or should I use Bedrock Agent Core to make that easier? Sure.
23:33Yeah. So Cohere has been around since before a lot of this technology existed. And so some of our applications have tendency built into them that the agents need to follow, right? So we want to make sure that the agents are following the same patterns that we already had in place pre-agentic implementation. And so we were building out our MCP servers and ensuring that we had auth proxies to send through the same sort of approach that we would follow on the core application side. But the agent core gateway handles this implicitly with the identity provider. So it's one of those things that with a couple lines of code, like I was saying, you can just pass through the same authentication method that we would typically set up on our own with the configuration enabled by AgentCore Gateway.
24:33way. So I'm curious now that you went from before you started to really like go down a path, building yourself, then started with Asian core. Did that change at all? Maybe like your timeframe of when it was that you were able to put into production? Your any of any other areas of like your velocity? Oh, 100%. Yeah, we were talking about this all week. I think going into this quarter, we had maybe an agent or two that we had planned to develop. And looking ahead at Q1 and the rest of 2026, it's full of agents. It's full of agentic systems, multi-agents. And I think Gigi touched on this a lot, but the evaluations come first.
25:17So a lot of our focus has been on building out the evaluation framework this quarter so that we can scale and continue to build new agents that we know are going to be successful on the first pass using AgentCore.
25:33I've got a few questions that I want to get both of your opinions on. So Gigi, I'll start this one with you. So model choice, this is a super hot topic that I know a lot of businesses are really thinking about. So I'd love to hear how you think about looking at evaluating different models and the term I love that you chose earlier, evaluation driven development. And based on your experience, maybe some insights that you can share on the business value of that approach. Yeah. Actually, it could be helpful for me to go back a few years of history. Kenji talked about how Coher Health has started a few years ago before this agentic revolution.
26:20And in fact, we were using our own transformer models to do NLP months, if not quarters, before ChatGPT came out. So the company was on already an accelerated trajectory on adopting this cutting edge tech. So I think we have a unique perspective because it's always the question bill versus buy um or do we tune like so i think the decisions between do we continue our own journey in building our own model from scratch versus adopting one of the frontier model with prompt tuning or maybe go down the path of fine tuning slash maybe pre-training so so we've been having these um internal healthy debates for for quite a few months i think it's a unique experience that i'd love to share more widely and um i think one one thing is change is constant and the only the best thing i could do as a leader for these amazing technologists is to set very clear metric right accuracy cost latency reliability and and honestly how much eval data is needed for each approach um we cannot not push any AI out without publishing eval data, especially in the industry we're in.
27:42So really thinking through all those metrics, and I can be honest with you, Julian, earlier this year, it's still leaning very heavily to early self-training. And more recently, it's leaning more and more towards fine-tuning or perhaps using one of the frontier model, at least at the get-go, so that we can get really good coverage. So that, I think, is keeping an open mind. what has really helped us is once we agreed on these are the business and operational metrics that are important to us and setting up a leader leaderboard so that we have the ability behind the scene not as part of the constant spring planning that we have you know a way to keep monitoring and tracking which models are winning and what makes sense for us to make the search.
28:29And I think what makes mine and Gigi's lives a lot easier is the investment that the company has made in experts that help us label our data too. So Gigi touched on it earlier, we do have a ton of licensed clinicians that not only help us when it comes to evaluating our decisioning modality, authoring policies and those components, they also label data for us so that we can use that in our evaluation data sets. So on top of an automated LLM as a judge approach, which we have in place, we also have clinicians labeling the data for us behind the scenes. 200 % Kenji. And change is a constant. Having a leaderboard to keep watching against metrics is so helpful.
29:19And what we've learned too is as this ground truth data keeps coming in every day, for us it's 50 ,000 to 60 ,000 labels a day the answer will change because when you have more and more ground truth data your ability to fine tune something or self tune something changes well so I think being clear on let me take that back so for those who's editing probably get rid of what I just said the last five seconds I think going forward is being really cognizant on how we collect grand truth data so that we can have that use case specific insights to make these decisions.
Read the full transcript
30:05There is a lot to really unpack there that I think every single listener can get a takeaway from. Metrics, having a leaderboard, I know a lot of businesses out there that I speak to don't have a model valuation framework. they're usually just sticking with one large language model and they stick with it until maybe there's a reason not to but I love that you're really assessing all these different options that are out there I think and obviously it's very clear that you're looking because you have all these different metrics that you've defined ahead of time you're able to then and you've got this process, you're able to have the best cost possible at the lowest latency, at the best performance, which at the end of the day, that's what companies are all looking at.
30:57They are just don't have a system that can be able to help them get all of the benefits that they're really looking for. So I would love to hear your advice for a company that right now they're using a single large language model and they're curious of, there's probably other models that there are definitely other models that are out there, but how do they go from one to being able to assess others so they can pick one or more that are going to be best for their use cases? I think you summarized it well, right? Let's make sure you know what's important to you. What are your metrics? And then investing in that automated eval framework both human in the loop and using large language models so that these decisions can be made with data-driven decisions i think that's a big part and when it does show that there may be value to switch my personal experience says that it's not always one size fits all it's not that you switch from one frontier model to another for all your use cases or even within the use case with every single piece of your pipeline so it goes back to architectural conversation like working with the architect and thinking through how you design an architecture that allows you to have that composability the ability to for some in our case certain use cases rely more heavily on smaller models that we host internally and certain use cases we rely more heavily on frontier models but if your architecture doesn't support it and every times and you build, then you mix the adoption costs very unbearable.
32:48Yeah, Gigi touched on it. I think having the platform, investing in the platform is key to this style of development, both the evaluations framework and the gateway. I think agent core gateway is meant to be a gateway for the agents. I think having an LLM gateway or an AI gateway is also valuable to put on top of this framework so that you can iterate quickly, change targets, and evaluate these at a much higher pace.
33:22So I'm very curious about really the business impact that you've been able to see in AI. I know there's still companies that even struggle to be able to measure the ROI of AI. And so hearing, I think from your experience will certainly be able to inspire them. I know in addition to everything else you said earlier, that definitely has, if not already. Yeah. Oh, well, let me go back to the example I started earlier in this conversation. In the prior of space, we've worked with clients who have to deploy dozens and dozens of nurses and MDs and doctors to review these cases in order to just meet the volume.
34:07and meet the turnaround time requirement. It's a highly regulated industry. We have 14 days to decide, but starting in the new year, you only have seven days to decide. So it's easy to just try to fill bodies of the problem. But with the AI system that we've built out with our clinician in the loop, we're able to achieve 85 % automation. So 85 % of decisions are made within minutes. and for those 15 % of cases that require high touch human review with our agentic system we've seen about 30 to 40 % improvement in productivity so and um and something that's tough to measure but we're hearing feedback from the users is it makes that it helps with the job satisfaction because they spend their time making clinical decisions as opposed to trying to figure out where the information is or trying to dig through all the requirements, right?
35:04Everything is surfaced. All the deep research is done on their behalf and they can really just focus on the area expertise. So yes, we save time, but also I think we have better retention.
35:19that is definitely a testament to the operations that your team has done i mean to be able to get to 85 automation your end customers being even happier with the solution um that really speaks to i think for all the listeners who are thinking about how to be able to get there and and they're especially those I know who are really in a monolithic type of architecture right now where they have one LLM and they're not able to maybe experiment with a number of different models that are out there to be able to maybe have a certain use case that's for the, they can get away with like a lower cost, one that'll give them better latency, all those different factors that you were talking about.
36:07So I think I just love those metrics because I think it just shows others who are in the early stages of their journey what's possible. Okay. Some other areas that I think I'd love for your opinions on. All right. So Gigi, based on what your experience, what are three things that all companies should really be doing when they're in the early stages of being able to build agents before they actually push into production? I think I'm going to know your answer but maybe i'll be in a little context right i've been doing this line of work for 20 years so agents are not right the the movement from a proof of concept or a pilot to a large scale solutions i think mckinsey says that 95 percent of ai prototypes and poc don't even go into production and only half of the ones that go in production actually stay in production so this has been an age-old problem regardless agent or not i think agent does create actually a pressure because on one hand you can innovate and experiment faster but on the other hand there's more unknown whether you have to manage so you actually exemplify the challenge and you're right we're going to start with evaluation driven development if there's one thing you're going to take home from this podcast that's the one line that i will really encourage us and as we think about these metrics to kenji's point having the right experts to label provide a data provide a grand truth even if it's weak grand truth right they're all valuable make sure that they tie back to your business and operational metrics so they're not pure functional non-functional by eval but tying back to the overall company or your client strategy the other piece i've personally seen a lot of struggles going from poc to scale is not having that not investing that time let me put it the other way actually put it in a positive way let me try again another part i've seen successes in taking from poc to large scale deployment is taking the time to define and letting everyone know what must be true for the agent to be successful at scale because the nature of POC is to simplify right is to not consider certain edge cases and it's to assume certain level of integration and operational efficiencies so in order to kind of flip that switch we must be very very clear on what are the important criteria for it to be successful and they're not usually ai related they're usually about the people who are going to be using the tool are you going to give them the right training are you going to give them the right transition plan are you going to bring in the right advocates because as i mentioned earlier everyone look at technology differently they're different personas right who do you bring on board to help you evangelize and oftentimes things fail because of processes, right?
39:13Because if you don't change the processes, but you get the new tech, you're not going to see the benefits and you can quickly fold, right? The solution can quickly, the impact can quickly be minimized. And then I think the third thing is, it's like classic, right, system build, right? We always assume data integration is easy and it's never easy. So I think that's the piece that we always have to kind of take a step back and say amazing AI system. Let's make sure the people, the process, and the data already and having that clarity so that everyone's on the same page and marching to a single. It's key.
39:53I think you just saved people a lot of time on that because people get so excited about AI and already start thinking about like, oh, what LM, for example, are we going to start using, but getting the domain experts really at part of the process, making sure you really understand, like that you were talking about earlier, like that what you're building is going to change the entire process for your end customers and really having clarity on what that means for them. Are they going to even use it? Even the data ingestion part as well. I think these are all prerequisites that people really need to think about.
40:29I've seen that as well as they often become overlooked and then you have to take two steps back before you can go ahead. anything you would do differently if you were starting over today?
40:52You stumped us with that one. Let's see.
41:00I think that Jerry is still out. We don't know yet. We don't know yet. One thing that I'm still thinking through is, is it more effective to have a small team that focus on the agentic development in a larger org? Or is it better to plug the agent development into every single development team? And let me try to say it the other way. Is it better to have a centralized agent development team, a gender development team, and let them drive the innovation and then dissimilate what they learn, is that more effective? Or is it more effective to have each development team to start adopting and doing more of a federated model?
41:46Jerry is still out on that one.
41:50That sounds like that'll be a part two episode. We will let you know, Julian. I don't know, Kenji, anything you may want to say? No, that's a good point. I don't have anything there. It's a tricky problem.
42:09All right. I've got one last question for each of you. Kenji, I'll start with you. So we've got listeners here that are in all different industries and all different stages of their journey within AWS. So for those who are considering building AI agents, what's one piece of advice that you have for them to get started? I think AWS has so many resources out there right now to get, to get, um, to make it self-serviceable. So I would, my, my best advice is to just start testing out and trying the tools. And at least in my experience, asking the questions early and often to the service teams, to, to, our account managers.
42:56It's helped, you know, uncover the solutions that we would have spent more time trying to figure out ourselves.
43:07Kenji is right. Being able to have hands-on experience is the best thing to make good decisions. And I guess one more thing I'll add is don't let FOMO get in the way. like just because everyone seems to be deploying and benefiting from agent systems really always start with the why you'll be surprised there's still a class of problems that may not be agentic right and there's a class of solution that can be very well solved with a single lom and just really kind of going back to the business and success metrics to make these decisions is key and i'll add one last one probably not in the right order so if the person who's editing can help us out here um i think my my advice to the folks who are wanting to dig into agent and so forth is think hard about the team you assemble to to make this real think um So agentic work truly does require a new profile of developers and a development team.
44:24We mentioned earlier about the importance of having domain experts in the loop on day one. Those who are technology savvy domain experts, they are gold in the team. And we've had amazing success seeing the hardcore platform developer working side by side with a data scientist and working side by side with a clinical MD and a nurse, that combo really allows us to iterate super quickly. And that's a new framework. I don't think we've had that kind of dependency in terms of diversity and skill sets, being in the same room at the same time before the agent or order. Wow. This was seriously a masterclass on, I mean, evaluation-driven development, AI-driven development.
45:20There clearly was something here for everyone. Thank you so much, Kenji, Gigi. This was phenomenal. Really appreciate both of you spending time here with me on the AWS podcast. Thank you for having us. Thanks, Jillian. This was awesome.
45:42Thank you.
From the publisher
In this episode, Jillian speaks with Cohere Health®, a clinical intelligence company focused on strengthening payer-provider collaboration as well as improving the speed and accuracy of clinical decision-making, both pre- and post-care. The company’s clinically trained AI helps accelerate access to patient care, improve patient outcomes, reduce administrative burden for providers, and improve healthcare economics across the care continuum.
Using Amazon Bedrock AgentCore, Cohere Health built Cohere Review Resolve™, an AI-powered copilot that optimizes the accuracy and efficiency of health plan medical necessity reviews. Cohere Review Resolve™ analyzes both structured and unstructured data–such as clinical records, patient notes, and faxes, quickly identifying and surfacing evidence to validate the medical necessity of a requested treatment.
This conversation reveals why Cohere Health chose Amazon Bedrock AgentCore for their first production deployment of agentic AI in a highly regulated healthcare environment.
Cohere Health expects that Review Resolve™ will reduce review times by 30-40%, helping them to meet critical, mandated turnaround times. For patients, quicker decision-making will accelerate care access, increase adherence to therapy, improve outcomes and reduce costs. With Amazon Bedrock AgentCore, Review Resolve™ also improves upon their existing 90% automation rate while helping health plans by further increasing the accuracy of clinical determinations, thereby improving medical expense savings and patient outcomes.
