Agentic AI: Redefining How We Interact with Technology

6 Aug 2024 · 47 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Talking AI Podcast - Episode Summary

Episode Title

Agentic AI: Redefining How We Interact with Technology

Host

Matt Paige

Guest

Amir Behbehani, Founder and Chief AI Engineer at Memra

---

Episode Overview In this episode, Matt Paige engages in a detailed conversation with Amir Behbehani about the future of artificial intelligence, specifically focusing on the concept of agentic AI and the obstacles it must overcome to be effectively integrated into enterprise environments.

Key Themes and Discussions

  1. Understanding Agentic AI
  2. Definition and Purpose: Agentic AI refers to AI systems capable of performing tasks autonomously, rather than merely assisting humans. This raises questions about the future of jobs and workforce dynamics.
  3. Memory Management: Emphasizes the need for both short-term and long-term memory in AI to improve the context of interactions, enhancing the effectiveness of AI applications.
  1. Challenges in Current AI Implementation
  2. Data Access: Many AI systems struggle with accessing proprietary data, limiting their effectiveness in enterprise settings.
  3. Hallucinations in LLMs: Current language models (LLMs) can produce incorrect or irrelevant outputs, known as "hallucinations," which needs addressing for reliable enterprise usage.
  1. Emergent AI Stack
  2. Five-Layer Framework: Amir introduces a five-layer AI stack:
  3. Layer 1: LLMs and foundational models (CPU equivalent)
  4. Layer 2: Data storage, including Vector and Graph Databases (like a hard disk)
  5. Layer 3: Context management (short-term memory)
  6. Layer 4: Agentic Layer, which orchestrates tasks
  7. Layer 5: Application Layer, where user interaction occurs
  1. Agentic Capabilities and Workflows
  2. Master and Subordinate Agents: Discusses the hierarchy of agents where a master agent delegates tasks to subordinate agents, maintaining workflow integrity through memory management.
  3. Feedback Mechanisms: Importance of feedback loops to mitigate errors in agent behavior, preventing "rogue agents."
  1. Future of Work with AI
  2. Impact on the Labor Market: Amir speculates about a shift in job roles as agents take over tasks traditionally done by humans, potentially leading to increased entrepreneurship and alternative employment models.
  3. Use Cases for Agentic AI: Highlights potential applications in matching algorithms for job placements or dating apps, enhancing the efficiency of interactions.

Key Takeaways

  • The Importance of Memory: Memory and context preservation are crucial in improving AI reliability and effectiveness.
  • Opportunities for Automation: Organizations are encouraged to explore automation within existing processes before pursuing new customer-facing solutions.
  • Transformation of Workforce Dynamics: AI could redefine job structures, transitioning individuals from traditional roles to more entrepreneurial endeavors.

Conclusion Amir Behbehani shares a visionary outlook on how agentic AI can redefine technology interactions and work landscapes, emphasizing the significance of enhanced memory management and contextual understanding in AI systems.

---

Additional Resources

  • Memra Website: [Memra](https://www.memra.co/)
  • Connect with Amir Behbehani on LinkedIn: [LinkedIn](https://www.linkedin.com/in/aimlengineer/)
  • AI Opportunity Finder: [HatchWorks AI Opportunity Finder](https://hatchworks.com/ai-opportunity-finder/)

---

Feedback Listeners are encouraged to subscribe for more discussions and leave feedback or topic suggestions for future episodes.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If we extrapolate this out, and this actually does start to manifest into a reality where these agentic agents are able to execute tasks. You logically need less humans in the loop. And if it gets to the point where they're not just executing tasks and specific workflows, but entire job functions, what happens then? Welcome to the Talking AI Podcast, where we talk AI with both experts in the field and early adopters. I'm your host, Matt Page, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI. Generative AI and the power of LLMs has amazed us all over the past few years.

0:38past couple of years. But when you get down to brass tacks, there's some major gaps and issues that need to really be resolved for this technology to fully integrate into the enterprise, into our works, into our lives. LLMs still hallucinate. They don't have access to all of our proprietary data. And at the core, they're lacking memory in a lot of senses. But ultimately, there's a gap in the operating system of today and what will require for future LLMs to take full advantage, a new emerging operating system, if you will. And that's why today's guest and topic of the Emergent AI Stack is going to be super relevant and timely.

1:17And today we're joined by Amir Babani, founder and chief AI engineer at Memra. And Amir is an AI research scientist with previous exits to both Google and Meta. And Memra is making agentic AI real with their platform that helps build, deploy, and automate your workflows through AI. But welcome to the podcast, Amir. Pleasure. Thank you. And did I hit Memra correct? How would you define Memra and what you all do? Well, I think the root of the word, there's a cogby that has to do with memory. And it also has to do with the word word, actually. So we're kind of playing, we're riffing a little bit, a little bit of a pun there.

2:05And the idea is that we're working on both a memory management layer for both long-term and short-term memory to help contextualize interactions with LLMs. Memory is such a big component that is lacking in some ways with how we're working with LLMs today, a huge opportunity area. But, you know, what other gaps or problems do you see with how we interact with LLMs today versus like where you see it could go in the future? What are the main things kind of holding this technology back from really being mainstream in the enterprise and in our everyday lives? The main thing that I think about is instead of having human beings interacting with GPTs to surface information, maybe making the human being marginally more effective at whatever it is they do from writing code to writing a document, etc.

3:09The way I think about these things is if the automated systems are leveraging the GPTs to then engage in performative actions. So, for example, reading from a data store and then writing via an API to then perform a set of actions that result in a set of tasks being performed on behalf of the human. R &B. So I sometimes think about this when you're engaging a GPT, is the conversation being by and large in the interrogative. But when you're engaging with, for example, agentic AI, the conversation is mostly in the imperative. So if you think about there's the imperative, the declarative, the exclamatory, the interrogative, the interrogative, you're asking questions.

4:01with the imperative you're giving a command now when we talk about memory and we talk about reasoning and we talk about agentic ai there is sort of an interrogative dialogue that's happening behind the scenes you can think of it as a socratism however that's happening by the agents not by the human being engaging in that socratic dialogue to sort of fine-tune the thought process The fine tuning of the thought process is happening automatically. The human is just saying, for example, write my documents to a database. Of course, there's a lot of questions that that invokes. What database? What's the user's password?

4:40What's the path to that database? What data do you want me to write to the database? What document? But of course, if you have agents behind the scenes that can answer, first of all, that can ask those questions, then can answer those questions and maintain in memory as context the correct answers to those questions, then you have something, in my opinion, pretty powerful. And that I sort of refer to as agentic reasoning. Yeah. And that's an interesting topic there. You're talking about this concept of agentic and the agentic nature of AI and where it could go. And today, it very much is kind of, you almost have to have a human in the loop.

5:22I almost equate it to today, we're very manual in the way we interact with AI. It's kind of a one-to-one relationship. I kind of see this future-looking view where it is very much autonomous in nature, one-to-many, many-to-many. It could be many agents interacting with many agents. Maybe there are humans in different areas of the orchestration. But the big gap you mentioned, say you're interacting with in AI today, whether it's Chappie, GPT or whatever, right? It doesn't have access to your data. And if it does, it's on a very manual type of scale. I'm like connecting to some document within the frame of the context, right?

6:04So quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact, tailored AI use cases based on your business, your goals, your pain points, and your industry. No fluff, no generic use cases, just real ideas that fit your business and the ranked by ROI potential. It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder.

6:44You mentioned orchestration and you know, I almost like to shift over to the emergent AI stack, right? So I'm actually going to pull this up for, for folks that are, uh, watching live. Uh, but yeah, the emergent AI stack, it's got five different layers as you define it. Uh, but let's, let's start with the bottom layer, LLMs and other models. That's the foundation in essence, start, Start by taking us through. And if you want to set some context for the emergent AI stack, feel free as well. I wrote this paper about a year and a half ago, and I think the thoughts were still sort of culminating in my mind.

7:29I was trying to reason through, you know, is there going to be some sort of AI stack that is sort of differentiated from existing stacks. And this was an attempt kind of piecing together sort of disparate concepts and disparate items. But I think a lot of it is sort of relevant a year later. I certainly think that maybe there's a metaphor here as well. So in this sort of emergent AI stack, as we call it, there is sort of layer one, which we liken to, again, metaphorically to the CPU of a computer. Layer two, we liken it to the hard disk or the store of long-term information, long-term memory. Layer three, we think that is sort of more of the short-term memory or the context window and things that accompany the context window.

8:25And then layers four, or layer four, I should say, That is the agentic layer, which I should also say that each of the layers, as you move up the stack, there's sort of accumulation where layer four benefits from layer three, layer three benefits from layer two, et cetera. And then ultimately at layer five, there's going to be an application layer. I've seen some hints at some of the applications. There aren't that many of them. I think the majority of companies with whom I've come into contact so far with regard to this AI stack, they're all engaging at the infrastructure layer. But that's not the same with the application layers, not around the corner.

9:19It probably is. And it'll benefit from the other layers about which we've spoken. So that was sort of the beginnings. And maybe take us through this data layer, which is an interesting component. Because, you know, LLMs are trained on a ton of data, you know, huge corpuses of data, the entire internet, but they don't necessarily have access to your proprietary data, data within your enterprise and whatnot. But as you can see there, it's connecting to, you know, your own relational databases, you have CRM historical quotes. This was kind of hitting off an insurance related use case, but it's almost translating it into these vector databases and knowledge graphs.

10:03Talk us through that component of it and potentially even differences between the vector database knowledge graph. Is that an either or, or do they benefit having both? Yeah. So I think layer two, you may think of that as sort of a rag stack, right? So the vector databases are involved there. Graph databases are involved there as well. And what we have found sort of in our work is that you store documents or other forms of information in these vector databases. So you can turn documents into embeddings. You store the embeddings in the vector database. And you can make a call to the vector database and retrieve some sort of set of similarities that you then augment your LLM with.

11:00So now you're taking effectively the private data on which the LLMs were not trained. So they have now sort of through this RAG stat, they have access to the private data. But we found that if you can categorize the density of information, then there are different stores of memory that you can then use to augment the LLM. So, for example, what we found is with vector databases, they're very good at providing answers to general inquiry or inquiries, general inquiries. Whereas graph databases, they're much better when the density of information is higher. So, for example, if I'm looking at a document, I may say, what is this document?

11:54It may be a contract. It may be an invoice, et cetera. And a vector database is very good at providing context to answer that question. But if I want to know, for example, who are the parties to the agreement? What's the value of the contract, for example? Is the document signed, et cetera? Generally, the density of that information is much higher. And if that information is stored in a graph database and then retrieved and used to effectively augment the queries, we get much better results. Now, the problem with the graph databases is that after a while, they get so large that effectively they'll overflow the context window.

12:44So in our particular case, we vectorize the knowledge graphs as well. so it's sort of like we get the best of both worlds we get the the the benefits of that sort of vector similarity search but we also kind of benefit from the retrieval of very dense information and we get very accurate responses very precise answers i should say so that's been our experience at that layer two, layer three and four we can talk about. But I will say with regard to layer four, given that's the agentic layer, it's benefiting from whatever's taking place at the RAG layer. So the agentic layer is not necessarily interacting with the documents directly.

13:33They're interacting with a representation of the documents, the RAG layer. So you want to make sure that it's not garbage in, garbage out, so to speak. You want to make sure the information with which you have to function at that rag layer, in that embedding layer, is the best it can be, number one. And then number two, you want to make sure that you can retrieve from that layer as precisely as possible. That's why categorizing the information is important and then retrieving information based on the category of information. And that category is really about density. So how dense is the information at the rag layer is going to define how effectively you're retrieving information at the rag layer to then allow the agents to then do their thing, do whatever they have to do functionally.

14:30Why is the density important? Is it just from a computational standpoint, dealing with more dense versus less, or how does that factor in? I think the identity has to do, I mean, again, probably it has to do with the limitations of similarity search. So cosine similarity is, you know, sort of how these vector search engines work. They may use cosine similarity or dot product, Manhattan distance, Euclidean distance. these sort of these measures of distance, I think they're all very approximate. Yeah. So they're going to get some approximate region of that vector space and then use that to augment your query.

15:18Whereas in the graph database, there's still, if you're vectorizing the graph database, those sub graphs, you also have that issue about the similarity metric being kind of potentially a limiting factor. But I think on the edges of the graph, of the knowledge graph, there's just context that seems to benefit the retrieval. So that might be a reason. So as we move up the stack, so we get into layer three, the context layer, and really that layer four, this kind of orchestration, which you have as the OS layer, this is kind of the new novel piece of this. and I mean there's several components that are new and novel to this but this concept of agents and having agentic capabilities is a whole new way of us thinking about it like today it's really humans kind of executing these different tasks playing these roles tasks building up to workflows building up to you know processes building up to entire work functions effectively Right.

16:29But talk us through that component and what's critical here, because I got to imagine the agents have to know what type of data they can access, when they should access it. You know, all these different components that play into their role that they're designed to do. But when the human is interacting with the GPT, they're maintaining context in their own mind. And they're winnowing down what information they want to elicit from the GPT as a function of the context that they're sort of maintaining. We have our own short and long-term memory in our mind that we're leveraging, right, as part of this.

17:12Precisely. Precisely. Now, when you have these agents engaging and there's no human in the loop, if memory is not maintained and sort of context is not preserved as precisely and as accurately as possible, you have parasitic effects. So a parasitic effect is when these error terms are being fed back into the reasoning process. And then you're effectively, you're exponentiating, or at least at a minimum, you're sort of multiplying these error terms. And you kind of get these agents that go awry very quickly. Rogue agents. Rogue agents, that's right. And so their ability to perform complex tasks where complexity is defined by either a serial process that has a lot of intermediate steps or each sort of task that comprises the overall workflow has a high degree of interdependence.

18:24And in those particular cases, the agents sort of just go off the rails. So preserving memory and preserving the ability to know where to retrieve as needed, that seems to keep these things, these agents on the rails, so to speak, for as long as needed to complete the task. Yeah. So are there feedback loops you have in place that are kind of hitting these negative scenarios to keep the system in check in essence, or is it not working in that kind of way? Or I know you've also talked about kind of master versus subordinate agents. Is there an interplay there as well between the different roles of the agents and what they're effectively doing?

19:25With a master agent and a subordinate agent, memory comes into play again. So what does a master agent do? You give a particular function to the master agent, write my documents to a database, for example. That master agent has to effectively take that function and break it down into constituent tasks. And then the subordinate agents go and do each of those tasks. So it's almost like a map reduce process where the master agent is mapping to the subordinate agents the relevant task. And then maybe it's summing the sum, so to speak, at the very end. and how do you do that? Well, of course, memory, you have to store all the, you have to store the state of the tasks of the workflow, of the tasks being done pursuant to the overall workflow.

20:24You have to store that and manage that. And that becomes sort of a memory management problem. So again, memory management comes into play there as well. Is the memory management, is it taking place across all of these different layers and components in essence, or is there one layer where memory is coming into play? Like, is there memory effectively associated with each individual agent where it has its own, you know, shorter long-term memory? Or how does that conceptualize, I guess, across this stack? I was, I don't know. I've been thinking of it, everything in between the application layer and the foundational layer is like a bus.

21:06That's how I've been thinking of it. Yeah. Yeah. Okay. Interesting. And so if we go to the very tip of it, you know, the app layer, that's going to change as well. Right. So effectively today, how we integrate with or work with applications, to your point earlier, it's very much kind of declarative. we're having to hunt and click and pick and basically execute the whole workflow ourselves via this GUI. Where do you see that going in the future with these agents? Is it a completely different user experience, user interface that we're interacting with? Much smarter applications. So the simple use case, for example, is let's say you and I want to schedule a meeting.

21:52There's a lot of hunting and pecking just to do that. It's kind of onerous it's not particularly enjoyable um and maybe an assistant can sit in the middle of that interaction and um and facilitate that transaction but uh and that's sort of how it's done you know with humans these days but using an application that benefits from agentic reasoning now you don't have to necessarily send a calendly link and i have to go kind of click it and find a time and it just it just eliminates that necessarily that all the reasoning is happening behind the scenes and there are already companies um that are doing that i think there was a company uh block it i think it was called they're backed by sequoia that's an example that's a company at the application layer that's benefiting from magentic reasoning.

22:54And it's a very seamless interaction. You send an email to somebody and the agents behind the scenes are sort of resolving that double coincidence. And this is effectively focused on that use case of scheduling meetings and that not having to have this kind of step in the middle of either us talking. Calendly took it to the next step where at least one person can kind of view a calendar and select it. Now, this is basically pushing that responsibility to the agent and that solution. Correct. That's right. On the topic of use cases, what other use cases are you seeing, whether they're ones you're doing at Memra or ones you think are interesting and kind of ripe to be picked off first?

23:46Because I have a feeling there's going to be this group of use cases that are like easy, low-hanging fruit, and it's going to kind of progress from there. It's a harder, more difficult thing to get the human out of the loop effectively. I mean, the use cases that come to mind, and maybe even in the general case with AI, a lot of the use cases are where you have two-sided marketplaces. So let me speak to the general case, not focusing on agentic AI, but AI more broadly. So if you think about Uber or this sort of the shared economy applications, one of which being rideshare, you have two demand curves, right?

24:36One of those demand curves is the demand for the riders. And then the other demand curve is the demand for the drivers. And there's a recursive cross-product price elasticity of demand function there, right? Adobe Acrobat Reader, Adobe Acrobat Writer, that's an example of sort of like platform economics. And in the case of Rideshare, there's a human driver. And that creates the bipartite demand curve. Now, if you replace the human driver with a self-driving bar, suddenly that bipartite market just becomes a single demand curve all over again. So now taking – of course, you couldn't do that 10 years ago when Uber was brand new.

25:26So, of course, they had to foster a marketplace. Now, if you now – so that's more of the general case with AI. if you now kind of focus with regard to agentic AI, one of the use cases that comes to mind or one of the set of use cases that comes to mind, matching algorithms. So maybe on a dating app, as an example, you use machine learning to foster a match or with job apps, job sort of marketplace applications. Maybe there's an algorithm, them keyword matching as an example that matches the employer to the prospective employee.

26:09And machine learning is okay. It's certainly a worthwhile approximation. But how about if agents are doing it? Yeah. It's vastly more dynamic. And maybe even the agents can leverage machine learning. I'm not saying they can't. but certainly those marketplace interactions come to mind as a set of use cases. That's interesting. I think hitting on the double-sided marketplace, I think I saw Tesla teasing the other day with their new fleet of autonomous cars coming soon. So we'll see what happens with that. But to your point, that kind of takes at a whole different scale the human out of the loop completely on that whole side of the supply and demand equation there.

26:58But if we extrapolate this out and this actually does start to manifest into a reality where these agentic agents are able to execute tasks, you logically need less humans in the loop. And if it gets to the point where they're not just executing tasks and specific workflows, but entire job functions. What happens then? Where do you see kind of the future of work going if this does start to manifest over time and extrapolate out to where you have whole job functions effectively being done by AI agents? What's the role of humans at that point? That's a harder question. A labor market transaction has two sides to it.

27:49There's someone who's selling their skills and there's someone who's buying the skills. To the extent that you're selling your skills and you're on the sales side of that labor market transaction, Yeah, potentially you're competing with AI. And so you may view AI as maybe a threat, maybe a dampener to sort of the wage. Or maybe you might even view it as to the extent that you're leveraging it, although there's probably a governing factor to what extent are you leveraging it. It may be improving your quality of life as you perform your particular job. But from the buy side of that transaction, AI is really powerful.

28:39Yeah. Yeah. Right? So now the question then becomes is if AI becomes a catalyst to sort of change the extent to which one group of people are on the sell side of the labor market transaction, maybe they start moving to the buy side. Maybe they become more entrepreneurial. Maybe they decide I'm not going to go to work at some company, but I'm going to become a vendor. That which I was doing at company XYZ, I'm going to strike out on my own and do that job as a vendor. Of course, it was like Ronald Coase that talked about why teams of persons come together within a firm. This whole theory of a firm.

29:36And sort of the idea was what he termed at the time when the marketing costs were prohibitively high. Of course, we call those marketing costs these days transaction costs, right? But if AI is coming and sort of reducing those costs, then I don't know, maybe potentially it creates an environment where the person who would otherwise need to be part of sort of a cohesive entity to reduce some of those prohibitive transaction costs, they may not face that barrier and will perhaps do something on their own. I mean, that's a, it's a hard question because it requires a lot of variables that go into it.

30:18There's political variables and there's a secular nature to this. So it's hard to prognosticate. Yeah. So on the shift in topic, that's become a common term now in the new age we're kind of living in. at what point do you feel like that that you know probability of either something being wrong something uh you know pretending it's right and just making stuff up along the way become small enough to where we can effectively trust the output because i find myself at times where if it's you know if i'm having a discussion around you know brainstorming a strategic conversation or something where I don't need actual, you know, it needs to be correct, then it's perfect, right?

31:09But if there's an element where I need the right answer, I find myself going back and fact checking. Do you think there's going to be a point in time to where that probability of error effectively so small to where we can effectively trust the output? And is that kind of going back to this emergent AI stack and the operating system to where that becomes reduced so much to where you just trust the output, especially when you're working in some of these enterprise use cases where messing up can mean very bad effects on the back end? Well, I mean, in our particular case, we're taking data. And this is maybe answer the question to the best of my abilities from my experience with working with these elements.

31:56In our particular case, we're taking data that exist in documents, which are relatively unstructured, and we're writing them to a structured relational database. And it's one thing that we write the answers totally wrong. So that's one element of it. But there's another element where we actually get it right. But every time we run the process to write another record to the database, even though it's the same process, therefore it should be the same record, that record is written differently. So if you do this over time, you multiply it across the entire matrix, you have a bunch of dates as an example, and they're totally formatted differently.

32:52Even though it's the correct date, It's just one is mm forward slash dd forward slash yyy. But the second one has the correct date, but it's just formatted. It's different. So what we found is that if we have some sort of manifest or instruction file, so this is in addition to categorizing the information density. if we have some sort of manifest then we're able to

Read the full transcript

33:28again it's not just reducing the hallucinations because this is not hallucinatory it's getting it right it just doesn't know exactly how you need it to be formatted so we've been able to reduce the heterogeneity of the data that are being written to the databases down to zero. Really? Yeah, there are techniques that you can use. And again, ultimately, these techniques boil down to preserving context and leveraging memory. And in so doing, you are, at least in our particular case, you're eliminating the hallucinations and eliminating the sort of the variability of the results as well. And that's very, it's really important, actually, in our work.

34:25I found it interesting, too. I was looking at one of the demonstrations, I think it was on something posted on LinkedIn from you using Membro. What was interesting is at the AI agent level, it had specific workflows that it was designed to execute on. And if it came across one that it was not tailored to use, it kind of had that response back like, you know, I'm not I'm not suited for this or whatever it was, which I think is also impactful, too, because there's this element of today. If you're just working with, you know, chat, GBT, whatever, you know, LLM, it's going to try and do it. But it had instructions that this is what I'm supposed to do.

35:06If I can't do it and I'm not going to do it. maybe there's another agent that can, which is another interesting element, I think, that you all kind of started to build in within the system as well. Yeah, that's at the master agent level, right? So can you imagine the master agent hallucinating a workflow? I mean, that's off doing its own thing. Yeah. That might be quite bad. So we're able to sort of control the master agent quite easily. But some of the things that I'm speaking about in terms of the manifest, in terms of the the embedding space, et cetera, that all has to do with the subordinate agents.

35:42And the subordinate agents, they're sort of, I like to think of them as, as sort of runners in a relay race. Yeah. And they're kind of handing a baton to the next runner. And if they're not preserving context of memory, they're going to forget who's the runner from whom I have to grab the baton. And, oh, did I have to grab the baton? And it's just, that's a mess. What is this thing I'm holding? Yeah, what is this thing I'm doing? So those are some of the things that, yeah. So in today's world, I'm assuming you are effectively building, architecting, engineering, all of these master agents, subordinate agents.

36:21Is there a point in time in the future where, you know, there's a master agent actually creating subordinate agents on itself in kind of an agentic autonomous nature or solving workflows in just a better way than a human could define it? Is that something that could be potentially there in the future? I think it's already being worked on. Really? I think you could, you know, the base case is a master agent with planning capabilities. So sometimes you can think of, it was like Frederick Taylor and his scientific management, principles of scientific management. He was talking about taking a particular function and breaking it down into constituent tasks.

37:11Of course, I guess that's the basics of division of labor. these agents effectively, the master agent is doing something very similarly. It's breaking tasks down into bite-sized pieces that effectively the LLMs that are powering the subordinate agents have a high probability of getting right. So the governing dynamic here, as opposed to sort of like the division of labor or human beings instead of it being comparative advantage. The governing dynamic here is maybe information entropy or some sort of measure by which the LLMs are just getting it wrong with some probability. So if you can control that, then I think that is sort of a master agent that can take on large, complex tasks, break them down into their constituent, and then allocate subordinate agents to take care of those particular tasks.

38:29Yeah, I feel like over time, there's going to be this whole new slew of companies that are coming out that are effectively doing with just a few people plus agents, what massive companies are doing today. And those massive companies have a whole different, you know, P &L and just, you know, all the operating expenses and human labor and everything that goes into it. So I think it's going to be very difficult for those to adapt because it's not just like the innovators dilemma of, you know, past generations. I feel like it's a whole new level of the innovators dilemma, in essence, that may be upon us with a lot of the, you know, kind of existing companies that are out there.

39:22Yeah, yeah, it's a good point. Also, it's not just the paper pushers, so to speak, within an organization that have some added competition vis-a-vis agents. It's really the explicit knowledge worker. Yeah. And sort of the explicit knowledge workers, you know, these are software engineers. These are mathematicians, statisticians, data scientists, et cetera.

39:53Yeah. So on the topic of software engineering, just because that's in the area that we play in large part, where do you see that going? Do you see engineers completely being replaced? Do you see it's kind of the more experienced engineers being able to leverage AI in a more exponential nature versus somebody that may be more junior in nature? Where do you see that discipline going? in the future as this starts to progress? I think what I see right now is that the efficacy of a software engineer, vis-a-vis the GPT models, that's increasing. And then to the extent that you're leveraging agents, yeah maybe you're leveraging agents to help you with your software engineering tasks although ultimately i think that agents are being developed to help the applications reason not necessarily to help the engineers reason um personally i think from you know many years i was working in the machine learning world and there you have sort of deterministic output that you're trying to use to train a stochastic process, right?

41:25Because there are events, they've actually taken place and you're using that as the training set. And then your model, your predictive model is outputting results that are stochastic. With the AI stuff, it's sort of the opposite. You have stochastic output and you're trying to make it as deterministic as possible. So I sort of see the role of the engineer as doing that now, building the infrastructure to do the latter, just as sort of the ML engineers in years prior have built the infrastructure, data pipelines, et cetera, to do the former. At least I find myself doing that. Yeah. So less in the weeds of actually writing the code and you're kind of setting up the systems architecting, that aspect of it becomes more valuable, more important, unless, you know, even intensive to do it, I suppose, as well.

42:22So the future, last thing I'd be curious of, as much as you can share, where do you see the future of Memra going? Are you all building this future operating system? Do you see it in line with that? Is it a platform that you're looking to evolve and build into that people are leveraging? Where's the future of Memora going? Future of Memora has a lot to do with the customers that we have benefiting from the agentic solutions. I mean, we have probably the same trajectory as any other startup wherein you You have to identify product market fit and lay the scaffolding to realize that feedback loop, right?

43:09Or the set of feedback loops wherein your customers are benefiting from the product and the services. But in the way that I see it, I think I would love to help companies that otherwise have manual processes. and sort of see a lot of those processes as automated as much as possible. And in a way that the human being sort of commands these agents to do that. So when I think of sort of like the decade prior, again, the machine learning, where machine learning was sort of coming into its own, it was sort of a decade about leveraging predictive models and outputting predictions and then consuming those predictions to do things.

44:00Now I see it as one where it's more about automation and then probably leveraging optimization after. So probably when Memor will evolve into sort of servicing companies along those trajectories. That's really interesting. I think that's a foundational area that a lot of companies need to start. It's like start with your existing processes. Where can you begin to automate those? And it may shift your entire workflow altogether. You know, I think sometimes people see the shiny object and they want to build some new, you know, customer facing solution right out of the gate. You know, I'd say start internal, see what you can do there, start to build your muscle up over time and then see where you can evolve it as it may relate to completely new revenue streams, products, solutions, all that, all those kind of things.

44:59I think so. And I think with regard to this conversation, as we've kind of spent a good amount of time talking about memory and infrastructure to preserve context and leverage it, that's all bottoms up. Yeah. So when you're talking about a particular use case, a lot of that is kind of top down. But my thought process is that unless you've given a lot of thought to sort of this bottoms up reasoning, you're going to get a lot of errors, you know, just sort of kind of doing top down reasoning only. So maybe you kind of want to do the top down reasoning from the perspective of the use case and GTM, et cetera, but bottoms up reasoning to make sure that you're supplying demand as accurately and sort of correctly as possible.

45:42Yep. Well said. Well said. Well, thanks, Amir, for being on the podcast. Where can people find you and find Memra as well? What's the best way to get in touch with y 'all? So I'm Amir at Memra.co, and that's M-E-M-R-A dot C-O. And the website is Memra.co, M-E-M-R-I-C-O. M-E-M-R-A dot C-O. Yeah, and check out the Emergent AI Stack. We'll have that in the show notes. Really good piece that kind of talks through this new Emergent AI Stack. Well worth the read. Thanks for being on, Amir. Very much. Thank you, buddy. See ya. Thanks for listening to the Talking AI Podcast. If you enjoyed the show, give us a follow or subscribe on your favorite podcast platform.

46:26And don't forget to leave us a review. We love those. For more info on Talking AI, visit TalkingAIPodcast.com.

46:36The single biggest mistake we see companies make with AI is they don't properly train their teams. We see it all the time. Companies roll out AI tools and expect people to just figure it out, but using AI effectively requires a totally different mindset and skillset. And that's exactly why we built training for every level of your org, from AI training for teams and executives to training engineering teams on our generative-driven development methodology. Or if you've already identified your AI use cases and want to just prioritize where to start, we offer an AI roadmap and ROI workshop to help you build a quick plan.

47:09It's all about going from we should use AI to actually driving real value with it. Head over to hatchworks.com to learn more.

From the publisher

Welcome back to the first episode of season 3 of the Talking AI podcast, formerly known as the Built Right podcast!

We’re kicking off with Amir Behbehani, Founder and Chief AI Engineer at Memra, a platform that helps build, deploy, and automate workflow through AI. He and host Matt Paige look at some of the hurdles AI needs to overcome to fully integrate into enterprise, such as lack of proprietary data access and memory. Amir delves into the science behind these challenges and potential solutions. He introduces the concept of ‘agentic AI’ and the emergent AI stack, made up of five layers, aimed at enhancing AI functionality through memory management, data orchestration, and agentic reasoning.

They also discuss the role of vector and graph databases, the importance of memory and context, and exploring future use cases in a marketplace with agentic AI. Amir also speculates on the transformative potential of AI in the future labor market - hinting at more entrepreneurial endeavors as individuals leverage AI for improved productivity and automation.

Enjoyed this episode? Don't forget to subscribe to our podcast for more insightful discussions on the latest in AI technology! Follow us on social media for updates and exclusive content. Your feedback matters—leave us a review and let us know what topics you'd like to hear about next!

Key moments:

  • Introduction to Amir and Memra
  • Challenges with current LLMs
  • Emergent AI stack overview
  • Layer-by-layer breakdown of the AI stack
  • Agentic AI and memory management
  • Future of work with agentic AI
  • The future of Memra

Key links: 


Mentioned in this episode:

AI Opportunity Finder

Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/

More from Talking AI

All 84 episodes
Agentic AI: Redefining How We Interact with TechnologyTalking AI · 47 min
Listen in VO