In short
Podcast Summary: Talking AI - Generative AI Is Not an Island: ML’s Core Principles
Episode Overview In this episode of *Talking AI*, host Matt Paige engages with guests Simba Khadder, Co-Founder & CEO of Featureform, and Omar Shanti, CTO of Hatchworks AI, to explore the relationship between Generative AI (GenAI) and traditional machine learning (ML). The discussion highlights the misconceptions that GenAI operates independently from established ML principles and explains how both fields share foundational concepts and techniques.
---
Key Themes and Discussions
- Generative AI vs. Traditional Machine Learning
- Misunderstandings: Many believe GenAI is entirely new, however, it builds on established ML principles.
- Continuity: There is a clear continuity between classic ML and newer AI technologies, emphasizing that GenAI is not an isolated innovation.
- Unique Aspects of Generative AI
- Prompting: The introduction of prompts distinguishes large language models (LLMs) from traditional transformer models.
- Use Cases: Emerging applications like copilots and agents are changing organizational structures, often requiring new teams unfamiliar with traditional ML.
- Challenges in Implementation
- Pilot vs. Production: Transitioning from pilot projects to full-scale production remains a significant hurdle for both GenAI and traditional ML.
- Data Dependency: Both GenAI and traditional ML are heavily reliant on data quality, making data wrangling a critical task.
- Evolution of Retrieval Augmented Generation (RAG)
- RAG Defined: A method of combining retrieval mechanisms with generative models to enhance performance, requiring careful attention to how data is presented.
- Future Directions: RAG is evolving beyond simple implementations, emphasizing the need for contextual data and metadata filtering.
- ML Lifecycle vs. GenAI Lifecycle
- ML Lifecycle Steps:
- Data: Involves touching all aspects of data management.
- Training: Developing and training models using the prepared data.
- Serving: Deploying the models for real-time use.
- Evaluation: Continuous monitoring and assessment of model performance.
- GenAI Lifecycle: Similar yet distinct, focusing more on fine-tuning pre-trained models and managing the complexity of deployment.
- Team Dynamics and Organizational Structure
- Cross-Functional Teams: Successful implementations require collaboration among data scientists, data engineers, and ML engineers.
- Democratizing AI: Tools like Featureform aim to empower data scientists to independently build production-grade pipelines without extensive engineering support.
---
Key Takeaways
- Continuity of Principles: GenAI is deeply rooted in traditional ML practices, leveraging lessons learned from earlier models.
- Data is Fundamental: Both GenAI and traditional ML share a heavy reliance on data quality and management processes.
- Implementation Hurdles: Moving from pilot projects to production is a common challenge, necessitating robust MLOps strategies.
- Feature Engineering: Understanding the relationship between traditional feature engineering and embedding techniques in LLMs is crucial for effective model training and performance.
- Effective Team Structures: Cross-functional teams that integrate data science, engineering, and business perspectives are vital for maximizing the value derived from AI initiatives.
---
Closing Thoughts Throughout the podcast, Simba and Omar emphasize the importance of learning from the history of machine learning to effectively navigate the evolving landscape of generative AI. They underscore the need for organizations to adapt, understand their data, and structure teams that can build and maintain AI solutions efficiently.
Additional Resources
- [Featureform Website](https://www.featureform.com/)
- [Connect with Simba on LinkedIn](https://www.linkedin.com/in/simba-k/)
Final Notes The episode concludes with a reflective note on the necessity of practical AI training and the importance of developing a structured approach to deploying AI solutions within organizations.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00When I talk to MLOPS teams, one of my first questions is, is your ML team successful? Are they driving real business ROI? If not, you're just making zero value faster with MLOps. Welcome to the Talking AI Podcast, where we talk AI with both experts in the field and early adopters. I'm your host, Matt Page, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI. There's this impression that generative AI is completely different from everything else, and that it lives on its own island, and everything about it has to be new. that's actually a bit of a fallacy.
0:36And lucky for you, we're joined today by two experts in this space to help break it all down for us. Simba Cotter, CEO of Feature Form and returning guest Omar Shanti, CTO of Hatchworks AI. And for those that are not seeing us on YouTube, we all got the hoodie vibe. We did not plan this. And Simba, you win the award for best hair of any guests we've had so far. So if you're on like Apple or Spotify, go to YouTube. It's worth the look. Your hair game is strong today. uh if for context simba is a former google engineer head of machine learning at triton prior to founding feature form but welcome to the show guys thank you so much for having to be here awesome well to hit on that topic i'm gonna riff on this a bit because you know gen ai it hit the the market everybody was freaking out and you know there's this idea that it's this new novel thing which it is to a certain extent but it doesn't mean everything about it has to be new it doesn't have to live on this new island with a completely new approach from Curious Shell's perspective on that, having been working in both the kind of traditional machine learning space and now with large language models as that's begin to proliferate.
1:48Sounds great. Yeah, I mean, I think I'd like to start by maybe talk about what is different. Because I think before we get into what's the same, I think it's interesting to really point out what actually is different. So one thing that's different is prompts. So Transformers is a concept that my architecture is not new, or it depends on how you define it. It's definitely not. It's definitely way pretty. It's like 2016, 2017 era. Yeah, it was Google. When were you at Google? Were you there when that actually came out? Just curious, do you actually remember that paper coming out about Transformers?
2:24I do remember a lot of papers coming out about Transformers. One of the more pivotal ones for me was a paper called Almo, which I think really got into doing transformers for NLP. And yeah, I do remember it coming out. I wasn't at Google working in that specific domain. But I do very much remember reading all those papers. I think another pivotal paper for me was a recommender system paper that YouTube put out, which is using embeddings for recommendations. But anyway, yeah, so I do remember it coming out. But yeah, I mean, one thing that's different from those old transformers and stuff like GPT, LLMs, is the idea of a prompt.
3:05But that's kind of a new concept of like prompting a model. We didn't really have that before. We would really apply models to text, but the chatbot style prompt thing, that workflow is different. And part of why it's different is the LLMs have way more ingrained knowledge. Like we did this, do a better job. The line between believable and not believable is pretty stark. If you don't hit that line, like if you were to go to an old version or kind of hack with some of the old transformers even from three or four years ago and make them do what GPT does, you'd kind of get pretty garbage output. So one thing that's different is prompts.
3:46The other thing that's different is the use cases. A lot of use cases that we're seeing are things like copilots, even agents. and they, because they're so different, they're kind of coming under different teams, different organizations, especially large enterprises are taking the approach of, wow, this is huge impact. We don't really have an LLM expert in-house. I think a lot of reasons why people are doing this differently is because they're kind of spinning up new teams, sometimes teams that have never even done traditional ML before and are very much just built in this AI era. and those people are going ahead and kind of building from scratch.
4:24And I think there's some reason for that. But I think what we're learning now is there's a lot of lessons, so many lessons that we learned from traditional ML, so much workflows, concepts, procedures that would so much benefit LMs. And I'm seeing the LM people relearn a lot of lessons that recommender system people learned the hard way many, many years ago. Yeah, it's like the grizzled vet, which is the ML and then the new scrappy person on the scene, intern maybe, that's just like ready to tackle the world and they can learn a few things. But Omar, curious from your perspective, it was kind of interesting thinking through organizations that are adapting it now versus those that may already have this ML practice and thinking about, is it a separate thing or is it kind of integrated together?
5:16Yeah, it really is a great question. And I think locking in on this answer sort of helps demystify this sort of novel field that a lot of people invest a lot of time in kind of constructing a complexity around. I think ultimately, a lot of the principles are fairly straightforward. They're fairly simple. And that's why it's fairly easy to build a pilot. It gets difficult when you try to go to production with it, but in that way, it's similar with ML. But I think ultimately, Gen AI comes in this history of AI broadly, specifically in this history of machine learning. We now kind of characterize ML into sort of two camps.
5:53You've got the generative AI and then sort of the more classical ML, the more discriminative AI. And we emphasize the difference. But sort of as Symbol was saying, the two are fairly similar. They're built on the same techniques. And it really is sort of a porous border understanding where embeddings used for product recommendation ends. and where a generative AI, more conversational interface built on that same embeddings technology begins. Quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder.
6:25It's a free tool that helps you uncover high impact tailored AI use cases based on your business, your goals, your pain points, and your industry. No fluff, no generic use cases, just real ideas that fit your business and they're ranked by ROI potential. It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder. So rather than sort of complicate the picture, if we sort of remove one or two levels back, then I think we can see a more or less similar architecture, Not all the way, but I think at least on the conceptual level, there's a lot of similarities, which we'll get into later on.
7:11But then beyond the architectural side, I think there's some attributes that they both share. I've hinted on one, which is both are fairly easy to do in pilots, but harder to do in production. In practice, both are mostly built on and bottlenecked by data. So a lot of the job of building a production-grade application, both of them is wrangling with data. and then to Simba's comment around sort of the grizzled vet I mean both are increasingly abstracted away wrapped behind services so you don't really need to see the underlying engine there's a question of maturity which I'm really excited to get to together but that's the trend in both as well as being deployed on edge so ingrained into your day-to-day life whether you've got Gen.AI in your IDE or whether you've got sort of a classical ML model that is you know baked into your word processor or something like that.
8:05And I think one place where you can really see the continuity between the two is conversational experiences. So chatbots, drive-through agents, digital assistants, et cetera. The exact same infrastructure and workflows that we're serving what we're now calling older ML, which is really like deep learning LSTM models. So it's not really that old or classical, is now serving generative models. So it's sort of a slight paradigm shift, at least in the way we think about it, but it's the same infrastructure serving both of them. Yeah, there's like five different rabbit holes I want to go down now. But I think one nuance, and I think one of you mentioned it, but do you almost view like LLMs as a subset of ML versus two completely distinct things?
8:49Or are they distinct in nature? Is one a part of the other or is it like that whole thing where this type of dog's a dog, but I can't think of the analogy there? My view is not that there... I think of data as this core thing. You have data. You can just drop everything. And you have different kind of attachments you can add to your data to make it useful. One attachment could be just as simple as basic analytics pipelines, a BI tool like Looker. is this thing you snap on and boom, now you have dashboards and that solves a certain problem. Chirisal ML, predictive ML solves a never set problem.
9:28So you think fraud, recommendations, a lot of other use cases are well solved by that kind of plugin. And then there's this new plugin, which is LLMs, which again, you apply to your data and it has a different interface, different experience. It's very oriented towards outputting text, obviously. It tends to work really well when interacting with users, like co-pilots. And that's how I view it. So I think they're separate tools, but they're both tools that are essentially appendages to your data. That's perfect. Data being the core kind of foundational thing. And I think one other just understanding point for the audience that would be good, when you think of ML, you hear of the topic feature engineering.
10:08I remember back in the day when I worked on this advanced analytics team, by these super smart data scientists and they talked about feature engineering. And in my head, used to think of features like product features, right? But it's different. Feature engineering is kind of, you know, manually creating these different features based on your data. You're kind of changing them in some way, whether it's categorical, numeric, attribute-based, whatever it may be. But you're doing it intentionally versus the LLN, that's where you get into embeddings and kind of vectorizing the data to create these embeddings, which in a sense could be viewed as synonymous to features.
10:49But I think this is like an interesting topic as well when you think of ML versus LLMs. Curious, you know, how you would equate those two and how you would speak to those two. Yeah, my view. So what is a feature first? So for me, when I think of what a feature is, it's just an attribute. It's a description. Let's tell the user, right? Like my age is a feature of me. In traditional ML, we have these models. Let's say you have like a, you know, random forest. You have these inputs. They have to be numerical and you feed in all the numbers and it will have an output. There's this view, and I actually think that it's very quickly going to get outdated, where RAG and vector databases are synonymous.
11:35There's this kind of naive workflow that we have where we take textual data, we chunk it up, we embed it, throw in a VectorDB. And then on query time, when the user asks a query, we just make that into an embedding somehow. There's ways to do that. And then you grab relevant snippets, throw it into your prompt. Boom, that's RAG. And every single place you learn RAG, it kind of gives this exact same workflow. So there's this idea of like, that is RAG. I think it's wrong. I actually think it works pretty well, but it's very naive. And I think very soon we'll look back and like, it's funny how we thought that this was, that was all that RAG was.
12:16It was just like a nearest neighbor lookup. What RAG really is, is you have a context window, certain size. Your goal is to maximize the amount of useful information in that context window to provide for that alum. And what we're finding already is that people are actually finding that traditional search, not vector search, but like just like literally like what Elasticsearch does, can sometimes actually be, you know, like a vector DB. We're also finding kind of knowledge graphs and these other concepts also working really well. And let me give you like a very simple example of this, of a time where you could use the exact same features for a traditional model and NLM, and they completely make sense in both contexts.
13:00Let's say for me, you're a bank. For me, you have features like my age, my net worth, my income per year. And you could imagine a few other features about my financial data. If you're doing fraud detection, you could totally imagine, hey, look at the transaction. Here's somebody, here's how much he makes, here's how much he spends a month, et cetera, as it's fraud. That totally makes sense in that traditional model. In a chatbot, let's say I ask the question, hey, how should I invest my money? You would also want to know my age, my net worth, how much money I make. It's the exact same data you want.
13:37So really, if you start to change thinking of RAG from this super tie to this very specific implementation of RAG to conceptually, how do I fit the right information into the prompt? I think companies that have done this will find better results, will be able to iterate faster. and be able to unify ML and LLMs for the benefit of both. Yeah, and for those that aren't familiar, RAG is retrieval augmented generation. We got a lot of good content on the Hatchworks AI side about that. I'm almost feeling like, Omar, we need to do a joint piece here with Simba at some point in the near future about exactly what you're talking about.
14:14That's a really interesting kind of insight from that perspective. Absolutely, Maden. If I could sort of dovetail there, I share Simba's way of thinking about this as RAG is a pattern as opposed to a specific implementation. And right now, there seems to be this sort of equation of the two. I think we can even go back one step further. So as Simba's saying RAG is about maximizing signal, we can go back and say prompt engineering in and of itself is an optimization problem where you're trying to maximize your signal, however that's defined. This is like the output that you get from the model based on a finite context window, which the trend is to continue increasing.
14:55And then RAG is figuring out sort of how to balance external data. So external data in the sense of not instructions, but actual context to feed into the model. And then as Simba mentioned, like really succinctly, that doesn't necessarily have to come from an embedding store. where it can be text to sequel, can be... We were earlier talking about papers. One of my favorite papers that has sort of come out in the field is sort of spelling out reactive agents, where the notion of a tool telling us what information... The notion of an agent telling us what information it needs from the tools that it has is almost like, rather than giving the agent a book, giving the agent the Dewey Decimal System and the library catalog for it to go tell us, hey, go get this book and come back to me.
15:45so between things like text to sql between running other apis other models reactive agents really break out of the limited implementation that you know some people focus on and think bigger simba i'd love to hear from you as well like one of the big things one of the big use cases that i've seen where rag has needed to evolve a little bit beyond that static implementation is when you You need to factor in things like metadata filtering and keyword matching. So Proximate nearest neighbor is not only enough. Now you need a series of hybrid approach. Could you maybe speak to that a little bit? Yeah.
16:23I mean, a simple example might be like e-commerce. Like if I'm asking about if I have all these embeddings, maybe like descriptions. Let's say I have descriptions of items in my vector database, but I know I'm only looking for shoes. then you're going to want to make sure that nearest neighbor lookup is actually limited to shoes. And again, this is something we face in recommender systems for a long time when we use similar techniques. Another example that doesn't even use text that a lot of people are I'm sure familiar with is GitHub Copilot. If I have GitHub Copilot and I'm looking at a function and I'm deciding what to generate.
17:04You probably would get way more signal by following all the functions, you know, like almost building a graph of like, okay, I'm on function A and then, you know, it's calling B, C, and D and it has these parameters and here's what it's been called. Pulling as much code as you can to build almost like a mini module and pushing that to GPT will probably get you way further than just like embedding that function and finding similar functions like that doesn't even really make sense when you think about it um there are some places where it would obviously where like you're writing the same flux it's been written before but if you're really trying to understand the code base you kind of want to give metadata or an even context that actually doesn't benefit from embeddings so yeah um there's many many examples hybrid search um most of them have just there's a very basic like only filter to this type of data and then do nearest neighbor search on that.
18:01The problem is that most approximate nearest neighbor indices don't actually support metadata filtering. Every example of it is kind of hacked on top with some examples of people actually trying to like natively menta didn't. But originally a lot of, it was a really hard problem to solve. Yeah, it's interesting. So I'm more of like a product guy. I'm not a developer, but I've been going through this interesting experiment, playing with agents to actually build kind of simple web apps. And my wife had this idea and I'm like, this is a perfect idea to kind of go test and do that. And I've gotten into this flow.
18:37I used Replit's new AI agent to kind of get everything set up, which is super helpful for me. That's been my biggest blocker. But once I got to that and I hit my limit, I kind of shifted over to ChatGPT and I was doing kind of what you just said. I would say, okay, this is the feature or the bug we're trying to fix. and I would give it the context of the, whatever it may have been, the JavaScript or HTML or Python. And it had that context and it was like very minimal issues or bugs that I'm running into. And it would say, hey, this is what we're going to do. And then I'm lazy. So I'd say, great, here's my current JavaScript.
19:15Now give me the complete code back with your recommended update. And I'm just like in this process of like just hitting these small little, incremental features and bug fixes, it's working insanely well. So that's been this like kind of aha moment, but what you were just talking about there from like the context in that piece, that's been super good versus just like some kind of blanket statement out and just expecting it to have that context, right? So one thing Simba, I'd love to talk to you about is the ML life cycle and the Gen AI lifecycle. And I think one of the sort of the big challenges here is that people don't necessarily build the same things when we talk about ML lifecycle versus Gen AI lifecycle.
20:01When we talk about ML lifecycle, we typically associate it with building, deploying, maintaining a model versus Gen AI lifecycle. Very rarely is actually, I mean, there's like a few kind of buckets. Most rare is you're going to build and deploy your own sort of small language model or a large language model, though that happens. Prerequisite$5 billion, right? Exactly, yeah. Some other cases, it'll be like maybe fine-tuning with some chance for learning or other cases that might be sort of building like a series of orchestrated prompts or like an agenda of architecture. So sort of building different things, but maybe one way forward would be let's contrast the process of building a traditional ML model versus the process of deploying a generative AI pilot.
20:50Because I think businesses are approaching both of these in a similar vein of saying, we want to take action on our data to unlock memorable, differentiated experiences, and so on and so forth. So I think probably that's the most apples to apples comparison. And this could sort of easily be like a hundred level course on like what an ML lifecycle is. So Simba, feel free to like sort of gloss over the things that aren't interesting and focus in on the things that you're most passionate about, perhaps such as like feature engineering, model serving, detecting drift, you know, things of that sort.
21:26Yeah, I would classify or yeah, I would break the ML life cycle down to four parts. So I'll start there and then I'll contrast LMs. The four parts are data, which is broad, but it's just anything that's touching data and has nothing to do with models. So that includes feature engineering, analysis of data, et cetera. So there's data. Then there's training. So training is, okay, I took data, I built training sets that I think have valuable features and my label on it. I'm now going to train the model. Cool. Once the model is trained, you're going to have to deploy it. and in some cases if it's like a batch scoring model that might be very simple in a lot of situations it's like touching a user if it's fraud it's recommend their system etc it's more real time then it's you know it's it's much harder so then they're serving finally there is evaluation and monitoring which is okay so i built my data sets i trained my model i served my model.
22:28Now my model is in production. I need to make sure it's behaving well and that it continues to behave well because data changes over time. Things change. You want to make sure that the model is continuing to behave well. So those are the four parts. Now, and just to be clear, it's not linear. You can obviously go linearly, but it's very common to jump back and you're going to be training lots of different models. You might not deploy a model that you trained. It's very common. You might deploy multiple models in the AV test and you have kind of like special evaluation tools for that. So it's definitely more complicated, but the four steps tend to be the same.
23:06Let's look at LLMs. So data. If you're doing RAG, which means you're taking external data, it's similar. Like you have these data sets. It could be like a dump of chat conversations you've had. It could be a dump of PDFs. Who knows? You're processing them. You're doing feature engineering on them. You see, obviously, more embedding-based techniques. Even though those are used in traditional ML, you see them more in LMs. And you also see traditional analytical features, as we talked about, from age to average to graphical-based features, et cetera. Then there's training. Well, LLMs are pre-trained.
23:52So you can skip training, but the equivalence of training would be fine-tuning. And when you fine-tune, and in my view, it just makes sense to fine-tune. Like, it's not really a... If you can fine-tune, there's almost never a negative of fine-tuning. So if you're doing RAG, you can rag and fine tune. A lot of people view them as one or the other, but you can very easily do both. To have external data, you do that. You fine tune. Now, there's similar problems here about experiment tracking. I'll have all these different fine-tuned models or fine-tuned in different ways. How do I track those? But I need to deploy those.
24:28If you're using ChatGP or OpenAI, it's behind a model. But I'm seeing a lot of people deploying like Lama and other things, especially some of the bigger companies so they have full control over it. And that's hard. It's actually harder than deploying traditional ML because much bigger model. You need to be much more deeply in tuned in how to make your GPUs work well, how to scale those out. GPU sharing isn't really a thing, so you have to figure out how to do that, et cetera. Finally, there's evaluation. Same thing. You have this model in production that's serving. You want to be sure you're collecting user responses.
25:08even some of it's like not explicit, like implicit, like, oh, they read this or they left immediately or they asked another question, which kind of implied that I didn't answer the question. So then there's evaluation. Now, those are four things are the same, right? It's data, training, surveying, eval. In both traditional ML and in LLMs, the things that people spend most of their time talking about are data and eval. So even the hard parts are pretty much the same. Obviously, training and fine tuning are hard. Serving can be really hard. But 90 % of papers and things I read are not papers, but rather blog posts and things I see people talking about in industry are almost always about how do I get my data to do what I want it to do?
25:55How do I handle my data? And also, how do I make sure my model is behaving correctly in production? so many good insights there i feel like that's a part that i'm going to like re-listen to several times when i go back to this um but omar curious if you have any additional thoughts on that firstly i'm loving this i'm having fun at this conversation i wish it was like hours longer there's so many potential rabbit holes to go down um but i think what makes serving so i think what makes productionalizing models hard is probably the most um compelling content for our audience. So I want to echo the sentiment that of the four kind of nodes that Simba provided, data and eval are hard across the board.
26:40Now, when it comes to generative experiences, so far we've talked about sort of just pre-populating a vector store. There's quite a few things. So companies who are using RAG, typically using RAG in the sort of architectural pattern that we mentioned is not equal to RAG, but who are using that sort of vector store pattern, often encounter a number of hurdles around things like chunking, indexing, potentially involving in metadata filtering as well, as we sort of talked about keyword search and so on. And building production-grade pipelines that bring the data where you need to get it is really crucial.
27:19And Simba will talk about how Feature Form stands as a major sort of problem solver in that space. and how sort of feature forms architecture can alleviate some of that issue around data, specifically around building production-grade pipelines to empower your applications, as opposed to sort of bottleneck them. And then regarding evaluation, we're producing some fairly exciting content. So stay tuned, audience. Coming out of our labs group here, we should have some information or like some sort of guidelines on how to test generative AI, ideally using generative AI. We're starting off a little bit narrow in scope, focusing primarily on chatbot-based interfaces, so conversational chat-based as opposed to single-turn things.
28:07But one thing that we've been seeing great sort of progress with is using language-based models to assist us with contrasting syntax from semantics. So when we test chat-based interfaces, a lot of what we want to do is figure out, is it semantically conveying the correct information? From there, we can test some of the syntax for things such as tone, whether it's friendly, professional, or so on. But this has historically been a challenge because when you're dealing with stochastic operations, such as a generative bot where you're getting different sort of responses, you can't set up unit tests in just the same way.
Read the full transcript
28:45So building a testing suite that you can run at scale is based on similar practices as the past because always there's continuity. but there is an element that differs as Simba sort of led us with at the start there. I think we've got so much to talk about, so I'll probably draw that there. And going back to sort of what makes ML models hard, you know, Simba, we always hear the usual figures floating around 80 % of models don't make it to production. 15 % of ML projects don't really succeed. They don't sort of pay themselves off. And half of the ones that make it a successful prototype don't make it into production as well.
29:27So varying signals all sort of pointing to similar drop-off in productionalization. And I'd love to just talk about why that is. I wish, well, Matt, maybe we can do some magic. Let's plop in that image over here around the MLOps sort of toolkit where you've got ML code as just this very small black box in the middle and then configuration, verification, monitoring, all of these other nodes. This is like a very famous graphic in the space, comes from another very seminal Google paper years ago. Simba, I sort of think that the key to why a lot of ML projects don't make it into production is because as the name suggests, the focus is on ML as opposed to the focuses on the whole ecosystem around it.
30:13and MLOps is sort of a corrective step there. Could you riff on that idea, please? Yeah, I would love to. Yeah, I mean, I think one of the first bits is, I guess, even for MLOps, when I talk to MLOps teams, like, hey, how can we be successful? One of my first questions is, is your ML team successful? Are they driving real business ROI? If not, you're just making zero value faster with MLOps. So the first step is really making sure, And this is going to take time and experiments. Can you find the places and the levers in your business that actually drive value of ML? If yes, it makes everything else easier, right?
30:52Because you get more buy-in, you can make more things happen because the potential of the ROI is worth the cost, especially if you're dealing with building things from scratch where you haven't really done much before. So I think step one is make sure you hit those high ROI use cases. Once you do, where I've seen pain points, is people start hiring a team of data scientists scientists and no one around them like okay cool like go do ml well how do i get access to my data okay like i see it in this data lake i have to wait a month to like get all the right approvals to be able to touch it and you i have to deploy my own spark cluster now and tune it and and and and and it's like there's all this stuff like you said that comes around it um from access control to scale.
31:39Even if you can get the kind of basics of that working and you have a notebook working, let's say, which is usually how a data scientist is going to start. I'll grab a sample of data. I'll work in pandas. I'll get it working in a notebook. Then there's getting in production. That jump, that chasm between my working notebook to a production grade model that has SLAs and uptime and latency constraints and it has to be monitored in a certain way. And if it's an LLM, especially, it's like it can't do really dumb things. How do we make sure it doesn't do really dumb things? And how can we feel confident about that?
32:16And a legal might have to check off because we're a bank. And because we're a bank, we need to make sure that everything about this model is well-documented. You start getting through all that and you realize that this data scientist who's a PhD in stats and just like this super smart person is actually spending 95 % of their time not doing stats and actually dealing with infra, organizational management, being a PM, like all this other stuff around it. And I think realizing that at the beginning, having that right, the person that you hire, once you have models in production, and you're trying to make them better and tune them.
32:54And the person who like gets them from scratch is very different. One is more like a startup hacker. Like I'm going to just like break through walls to get my mom production. It might not be the best mom in the world, but it's going to work and it's going to be a great baseline versus the hey i'm the professional who um an analogy would be if you're trying to see get from one point to another on a flat track go get a sprinter if you're trying to go from one point to another through a messy forest of vines go get like a wilderness man and not like you know a sprinter because all their benefits of being a sprinter are completely taken away by how much they have to fight through so i think there's like this organizational complexity that really is the killer.
33:36And it's always going to be there, not having a good way to describe ROI and to have that kind of carrot that makes it all worth it. It just makes it impossible to get through. Yeah. It's like, we got to make the unsexy parts of everything sexy again, because that's what actually makes it work at the end of the day. We're curious, Omar, like, what do you think is that ideal team? I know it's nuanced, obviously, but I think to your point, Simba, a lot of folks just think, oh, data scientists, They need a whole team of those, but they need support so they can actually do the thing they're best positioned to do.
34:10Yeah, spot on. If I could maybe summarize in one line what Wisdom Symba just dropped on us, it's build business-centric models iteratively with a cross-functional team. And what constitutes that cross-functional team? we've got the data scientist but as we mentioned the minority of the actual work is producing data science code most of it is going to be sort of data collection verification quality imputation feature engineering like that kind of data manipulation stuff so the data scientists can do that brilliant there's the data engineering task which ideally the ml team is empowered to self-serve to so building either production grade pipelines or building queries into a virtualization layer that allows them to sort of self-serve to data as close to the source as possible, as true as they can.
35:04And then once the data scientist has got that data, cleansed that data, everything, and built the model, then I would say sort of the role of the ML engineer comes to the fore where this engineering role is responsible for taking this artifact and building a production grade, scalable, secure system around it. I think I left the data engineer off of the hook too early there. There's sort of the pre-computed features as well as the real-time features. And the data engineer has got to make sure that at whatever cadence the model runs, the data is ready to support it. So that might involve real-time sort of pipelining, which is very, very difficult.
35:40And then, you know, once you've got this model in this production-grade application surrounding it and so on, you need to monitor it, as Simbo was saying, and then continuously improve and communicate business metrics to sort of make sure that there's further investment for future uses. And I really think that that's a wonderful segue. I wish that someone had built a tool that just helps you sort of from more or less A to Z over here. I wish that person was on the call with us. Oh, wait, we've got Simba. So Simba, I've been sort of very candid with my appreciation of Feature Form as sort of an entire holistic product suite.
36:24For folks who don't know, actually, I'll give, I won't butcher my sort of description in front of you, Simba. So I'll pass it over to you to sort of introduce Feature Form. But I think between the breadth and the depth of features, as well as some key decisions around what I loosely call virtualization, but I'm sure you have a better word for it. It sort of feels like it's lowering a lot of the barriers to making this ML Ops team of the future, something a lot more realizable and within reach for a lot of like data scientists and AI and data orgs. So I'd love maybe if you could talk about how the team of the future gets, well, firstly, what is FutureForm and how much more simple is the team of the future with a tool like FeatureForm in your stack?
37:13Yeah, I would love to speak on this. So for those who don't know, FeatureForm, the category that we're in is a feature store. So FeatureForm is a type of feature store. The name's kind of a misnomer. It implies a place you store features. We really view ourselves as a place where features are built. So it's a workflow to API where features are built. The name FeatureForm comes from Terraform because we kind of wanted a declarative way to define and deploy your features. We spend a lot of the time, even in this conversation, talking about how data is critical. What we see in a lot of orgs is one of two things.
37:49Either data scientists try to productionize their own data pipelines. You'll often end up with situations where it never makes production. It's not the right level of SLA that you would require from prod. It's using stale data because setting up streaming, managing streaming pipelines at high velocity becomes impossible. So it kind of becomes a non-starter for a lot of data scientists. On the other hand, what we see, and this is actually more common, is data scientists will pass their notebooks to data engineers, whose job it is now to productionize these data pipelines. Now, data science is inherently a science.
38:28You throw away a lot of things. There's a lot of R &D. So for a data engineer, if I have a stack of tasks and there's never enough data engineers at the company. And I have three tasks, which are build dashboards for these three execs who really need them. And 75 tasks, which is from those darn data scientists who always want more features in production because they're always iterating. And most of those features aren't even used in a few months. Which ones are you going to do? Probably the exec-based dashboard ones. So data scientists... Yes. Yeah. It makes sense. When you think of ROI, if I'm a data engineer and I think of ROI, those dashboards will probably get more ROI than a specific feature pipeline.
39:12So what we need is a way to democratize the way so that data scientists themselves can go and build production-grade pipelines without having to know what Iceberg is, without having to know the intricacies of tuning Spark, without having to do all that stuff. Feature Form kind of provides a harness where if you give us some SQL or data frame code, we will run it on your infrastructure. It will all kind of look very native to a data engineer. It will look very similar to the pipeline that they would build. But you as a data scientist don't have to worry about all that. You just give us SQL. It's ready to go.
39:46You can train your model. You can use it in production. So my view of MLOps and an ML platform that does well, to answer your question about the team, is democratizing it so that data scientists can do most of the work themselves and these specialists come in to tune things where they need or maybe set some configurations that are specific to their business. But the generic bit, building, streaming, feature pipelines, it kind of looks the same across most companies. And so we've kind of taken that pattern and we've turned it into a harness. And that's a huge part of what Feature Form is, democratizing feature engineering so data scientists can take it to production themselves.
40:22I appreciate that overview. I think for folks too, if you're like, A, this all sounds really cool. I need help. Or B, you just uncovered so many unknowns I didn't even know about. Definitely can find Omar or Simba. Let's chat through it. I'm curious, any last rabbit holes we want to dive down before we wrap things up? Anything top of mind that we would have gotten into earlier before we wrap up this episode? I think we've covered all the main concepts. I think if I had to give a big takeaway, it is understand and learn from what we've done historically and leverage those things to do what is new better.
41:09Bring in the right experts or ask the right experts to make sure that you are really doing this well because otherwise you're going to run into the same traps that all those experts have run into in the past of just one now. It's much more similar when people tend to make it. I echo that as well. I could spend a lot of time talking about feature form, but I think ultimately one of the themes here is that complex systems are made up of simple things done well. And we know from solid principles, from Unix philosophy, from all of these sort of foundational software engineering and even AI and data philosophies, how to do simple things and how to structure them into more complicated sort of systems.
42:02We can look at the debates that we've had around things like ORMs, around abstractions in machine learning, successful cases such as scikit-learn. And we can sort of build on the back of the great work that has come in the past to move forward. There's all kinds of, lately I've been talking a lot about abstractions and sort of how in software engineering terms, you want the abstraction to go with the direction of stability. So something that is more stable should not depend on something that is less stable. Otherwise that thing becomes less stable. And right now in this area where the use cases and usage patterns around generative AI are still being sort of drawn out.
42:48I think it's very important to keep that wisdom in mind, and it's something that is all too often forgotten. If I could just go back then to Featureform and just share some closing comments. So it does feel that with Featureform, the promise is to unlock data where it is and to let your infrastructure work for you. Simba in your comments it sort of makes me think through the development to getting to the promise of the data mesh. 60 % or so of data lake implementations fail. The siloing of knowledge leads to so many handoffs and loss of context going through there. So the more end user self service we can enable up the stack hopefully the better outcomes that we achieve.
43:36And I think whether we're building ML models and we want to empower folks in the MLOps lifecycle, or whether we want to enable self-serve analytics, whether it's through generative AI or through a dashboarding technology, that is sort of the drum that we are all marching to. Would you agree there, Simban, Matt? Very much so. I think the cool thing here, especially if you're in this space, is there's this huge network of people like Omar and myself who are pretty much committed most of their time entirely on moving the space forward and making everyone's life easier and getting the stuff in production and providing an end more value to end users.
44:15Yeah, it's such a great spot. I think to wrap it up and like I got in my head, I think I just stopped listening. The quote you had, Omar, complex systems are made up of simple things done well. That's just so powerful. Andy on the post-production side, there's your quote for the episode. But let's wrap it up there, guys. So folks know where to find Omar by now. I think LinkedIn's the best spot. Reach out to him. If any questions triggered from this episode, hit him up. You can hit me up as well. It probably just won't be as interesting or insightful. But Simba, where can folks find you and find Feature Form?
44:51Yeah. Best place is also probably LinkedIn. I'm very active on LinkedIn, so you can find me there. If you want to send me an email, you can just find me at Simba at Feature Form dot com. Pretty easy, pretty easy to guess. And you can always send me an email and I check my own email. So, but it's another place to reach me. Awesome. Thanks guys. Thanks for talking some AI with us today. Thank you so much for having me. Thanks for listening to the Talking AI Podcast. If you enjoyed the show, give us a follow or subscribe on your favorite podcast platform. And don't forget to leave us a review. We love those.
45:21For more info on Talking AI, visit TalkingAIPodcast.com.
45:28The single biggest mistake we see companies make with AI is they don't properly train their teams. We see it all the time. Companies roll out AI tools and expect people to just figure it out. But using AI effectively requires a totally different mindset and skillset. And that's exactly why we built training for every level of your org, from AI training for teams and executives to training engineering teams on our generative-driven development methodology. Or if you've already identified your AI use cases and want to just prioritize where to start, we offer an AI roadmap and ROI workshop to help you build a quick plan.
46:00It's all about going from we should use AI to actually driving real value with it. Head over to hatchworks.com to learn more.
From the publisher
Generative AI may feel like the new kid on the block, but it's not as different from traditional machine learning as you might think.
In this episode of Talking AI, we’re joined by Simba Khadder, Co-Founder & CEO of Featureform, and Omar Shanti, CTO of Hatchworks AI, to discuss how generative AI builds upon established ML principles and infrastructure. We explore the misconception that it exists as its own little island, when in reality, we see a clear continuity between classic ML and newer AI technologies.
Simba and Omar break down the unique aspects of GenAI, such as prompting and large language models (LLMs), while also highlighting the shared foundations with traditional ML. They also touch on the evolution of retrieval augmented generation (RAG), the ML vs. GenAI lifecycle, and what MLOps teams should consider when it comes to driving value in a business.
Want to stay ahead in the fast-moving world of AI? Tune in to Talking AI where we break down the latest trends and tools in Generative AI, MLOps, and more. Subscribe now, and don't miss out on practical insights from industry experts. Let’s transform your data into real business impact!
Key moments:
- The similarities and differences between generative AI and traditional machine learning
- How prompts differentiate LLMs from earlier transformer models
- The challenges of implementing AI in production versus creating pilot projects
- How RAG has evolved and what it could mean for the future
- The importance of data as the core foundation for both traditional ML and generative AI
- How feature engineering in ML relates to embeddings in LLMs
- How an ML lifecycle differs from a GenAI one
Key links:
Mentioned in this episode:
AI Opportunity Finder
Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/
