In short
Practical AI Podcast Episode Notes
Episode Title
Broccoli AI at its best 🥦
Episode Description
In this episode, the hosts discuss "Broccoli AI," a concept representing practical AI applications that yield tangible benefits for businesses. Bengsoon Chuah, a data scientist in the energy sector, shares insights on developing and deploying NLP pipelines, specifically in environments with limited resources and high risks.
---
Key Takeaways
Introduction to Broccoli AI
- Definition: Refers to AI that is beneficial and practical for businesses, focusing on real-world applications rather than hype.
- Context: Discussion centered around the energy sector, where traditional infrastructures often lack cloud services and data science resources.
Guest Introduction
Bengsoon Chuah
- Data scientist in the energy sector.
- Experience with machine learning and NLP in traditional industrial settings.
- Interest in how to implement AI within legacy systems while overcoming limitations.
---
Discussion Highlights
Unique Challenges in the Energy Sector
- Limited access to cloud services due to security, legacy, and connectivity issues.
- Importance of using on-premises solutions to deploy AI.
Data Availability and Structure
- Data Collection: Traditional industries like energy gather vast amounts of data from sensors and operational reports.
- Unstructured Data: Significant amounts of unstructured data (e.g., comments in reports) hold valuable insights but are often overlooked.
- Analysis Gap: Historically, there has been no effective way to analyze this unstructured data at scale.
Implementation of NLP and Active Learning
- Initial Steps: Bengsoon had to justify the data science role in his organization by demonstrating the value of data analysis.
- Proof of Concept: Conducted a pilot project focusing on unstructured safety data.
- Labeling Process: Developed a system where users would label data, which involved significant collaboration and iteration with domain experts.
Bootstrapping Labeling Process
- Users acted as labelers for the model, which was crucial given the lack of prior labeled datasets.
- A voting system was implemented to resolve labeling disagreements among users.
- Used Arjila (an annotation tool) to facilitate the labeling process and gather feedback.
Model Development and Deployment
- Initial goal was to achieve about 60-70% accuracy to provide visibility into operational safety data.
- The model was developed using Sentence Transformers, appropriate for NLP tasks.
- Active learning was used to improve the model continuously, incorporating user feedback on predicted labels.
---
Technical Tools and Frameworks
- MLflow: For model registry and versioning, essential for maintaining model history and tracking performance.
- Prefect: Used for orchestration of workflows, allowing for scheduled tasks and easier management of data pipelines.
- DuckDB: Employed as a lightweight SQL database to manage and clean data pulled from SharePoint, streamlining the data preparation process.
---
Insights on Future Directions
- ML Ops Evolution: Emphasis on adapting ML Ops practices to specific contexts without overcomplicating processes.
- Embedded Databases: Anticipation of broader adoption in the industry, particularly with advancements in LLM (Large Language Models) and Gen AI.
- Focus on Accessibility: Excitement around making AI technologies more accessible on varied devices, shifting from large-scale computational requirements to more practical implementations.
---
Closing Remarks
- The episode highlights the importance of practical AI applications that bring real value to businesses, particularly in traditional sectors like energy.
- Bengsoon's experiences serve as a case study in navigating the complexities of deploying AI in environments with resource limitations and a focus on user engagement and feedback for continuous improvement.
---
Links and Resources
- Show Notes: [Link to show notes](https://github.com/thechangelog/show-notes/blob/master/practicalai/practical-ai-280.md)
- Guests:
- [Bengsoon Chuah on GitHub](https://github.com/bengsoon)
- [Daniel Whitenack on GitHub](https://github.com/dwhitena)
- Tools Mentioned:
- [MLFlow](https://mlflow.org/)
- [Prefect](https://www.prefect.io)
- [DuckDB](https://duckdb.org/)
- [Agrilla](https://argilla.io/)
---
Conclusion This episode of Practical AI underscores the significance of implementing AI in a manner that is not only innovative but also practical and beneficial to business operations, particularly in sectors facing traditional constraints.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.
0:44What's up, friends? Intel Innovation 2024 is right around the corner. Accelerate the future. Registration is now open, and it takes place September 24th and 25th in San Jose, California. This event is all about you, the developer, the community, and the critical role you play in tackling the toughest challenges across the industry. Ignite your passion for AI and beyond. grow your skills to maximize your impact, and network with your peers as they unleash the next wave of advancements in technology. Here's what you can expect. Understand the emerging innovation and trends in dev tools, languages, frameworks, and technologies in AI and beyond to empower you and the solutions you're building.
1:28Get in-depth technical experience. Join hands-on workshops, labs, meetups, and hackathons to collaborate and solve problems in real time. You can explore featured partner and Intel solutions. They have partners there, startups there, customers there. And Intel is showcasing the latest in products, services, and solutions across keynotes, tech sessions, and the show floor to help you meet your development needs. Collaborate with experts, learn and have fun, engage in interactive sessions to connect, get certified, gain unique ideas and perspectives, build long-lasting networks, and of course, have fun.
2:05And get inspired, hear from leading industry experts, technologists, startup entrepreneurs, and fellow developers, along with Intel leadership, CEO Pat Gelsinger, and CTO Greg Lavender, as they take you through the latest advancements in technology. Don't miss this chance to be at the forefront of innovation. Take advantage of early bird pricing right now until August 2nd. Register using the link in our show notes. Or to learn more, go to intel.com slash innovation. Once more, that's intel.com slash innovation or go to the show notes and click that link.
2:46Welcome to another episode of the Practical AI Podcast. This is Daniel Whitenack. I am founder and CEO at Prediction Guard. And this is a pretty special and fun episode for me because I get to kick back with an old friend of mine. We went to the same university, for those that haven't heard of it, Colorado School of Mines in Golden, Colorado. Of course, a shout out to all the ore diggers out there that are listening. But yeah, we have with us today Bingsan Chua, who's a data scientist now and working in the energy sector. I was really fascinated to talk over the years with Bingsan about all the things he's doing, and in particular, his kind of approach and learnings around active learning and NLP models.
3:42And yeah, I wanted to invite him on the show to talk through some of that and learn a little bit from him. So welcome to the show. How are you doing? Hi, Daniel. Thanks for having me. Yeah, it's been a while since the days in Colorado School of Mines in Golden. Good old days. And now you're working as a data scientist in the energy sector and also working in Asia, which is super cool. I'm wondering if you could give us a little bit of a sense of some of the unique things about doing data science and machine learning type of things in the context of the energy sector, in the context of an actual enterprise, real world kind of situation.
4:24because we talk a lot about, recently we've been talking a lot about all of these Gen AI models and APIs and such, and that is super cool. But also there's a lot of on-the-ground work going on in data science that maybe looks quite a bit different than that. Yeah, thanks. So, I mean, I work in the energy sector, and it's pretty much a traditional type of sector. A lot of the companies, as you go around, at least in Asia, or at least over here where I'm at, we do not actually even have things like cloud services or subscription and stuff like that due to different reasons and stuff. But at the same time, there's an appetite for machine learning, AI, data, and all of those things.
5:11You see people talk about Gen AI as well. But I think the ones that I've noticed, at least for me personally, that has really brought a lot of values are kind of what you guys were talking about in the previous episode of Broccoli AI. Broccoli AI. I love it. Yeah, not so sexy, but still really important, really brings value. Particularly, I guess what we're going to talk about is active learning in the context of NLP, natural language processing. And so I think that's a pretty exciting place to be in, to do. I mean, it kind of translates into Jenny and to a certain degree. But at the same time, I think to at least the context that I'm in as well, in the sector that I'm in, we do not have things like health services ready for us.
6:01And so you have to figure out ways to kind of bring that about in an on-prem server VM. So how do you work around that? And how do you actually bring in cloud native modern technologies within a traditional kind of structure? So, yeah. Yeah. And from at least my impression, even though there's sort of not this, whether it be for security reasons or legacy reasons or just connectivity, there's not the type of connection to cloud services like you're talking about that others might be working with. But at the same time, at least my impression is that this sort of sector and maybe others, there's other related verticals where they have been sort of data driven to some degree for some time.
6:55And I don't know if you could speak to that, like the types of data that people, you know, have been processing or storing or are available in those contexts. Yeah, I mean, that's a good point because a lot of these traditional industries, at least in the energy sector that we see, we have sensors that are constantly flowing in data all the time. The data is there. I mean, we are collecting data, right? And then going a little bit deeper, then you find that a lot of us have been collecting a lot of unstructured data too. And so at least within, you know, where in my experience, at least what was happening was when I came in, I was pretty much the only data scientist and so I had to make my existence justified in a sense.
7:46And so I knew that I had to do something about like bring value within this organization and like I'll be able to have proof that like, hey, look, data science actually does work and does bring value. But the quickest, I guess, low-hanging fruit that I've found within, at least in the context of where I'm at, is the whole unstructured data. So we have been collecting thousands and thousands, hundreds of thousands of unstructured data. But there has never been a way to really analyze that at scale. So people have been analyzing it. They've been able to do some sort of a human analysis on it. but there's never been someone who's able to like say, hey, what's been happening in the past 10 years that we've been collecting all this data?
8:32What's it telling you? Nobody has really been able to do that. So I thought, okay, you know, maybe that could be one thing that we could actually bring in is that like you have all this data that's ready for you, that has all the insight, that has all the information that is locked up, right? Waiting to be unearthed pretty much, waiting there for us to just extract it or mine it. So I did a quick POC for one of the department's company and said, hey, guys, like you've been doing a Power BI, Tableau Power BI kind of thing. You've been able to do stuff with your structured data, right? And you've been able to plot it on beautiful graphs and stuff.
9:15I could be able to analyze in that sense. What about all the unstructured data that you've been collecting over the years? And they said, there's no way we could do it, right? I mean, like, you've got thousands, I was like, oh, you can't, there's no possible way for you to put it in RBI. And so I said, like, well, maybe we could explore ways that we could actually get into machine learning to actually help you to scale that analysis from an open standpoint. Thankfully, they bought it. And so we took off from there. And it was pretty cool. I mean, it was a journey for sure. A journey to learn. Yeah.
9:48So like when you say unstructured data, give people a sense of like the kinds of files or not maybe the specifics, but, you know, I'm imagining a file store with some type of files in it that contain something. Give people a sense of that. Yeah, for sure. I mean, I guess unstructured data typically would think about it as like text, right? It could be text or it could be something that's just out of structure. And in like Microsoft docs or... So we have been storing all of those data within SharePoint, Microsoft SharePoint, right? And so what I've seen is these unstructured data, it usually comes in a tablet form, right?
10:31You have a table that is collecting all of the structured data, but alongside with it, there's always a comment or that additional things that you have to actually tell the story of what you're actually collecting. And those are the ones that I think it's always locked up. And you would see it's very typical in any kind of industry where you have a tablet data that collects sensor data and everything or reports of what's been happening. And those are structured. But then there's also a column that would bring in some sort of remarks, observation or actions taken and whatever it is. Comments. Comments, yeah.
11:15But more often than not, we just kind of like play through it and kind of like not really put too much attention to it. At least within the data set that we were looking at, we had found that there were a lot more insight in those data than the structured data that they've been collecting. I'm talking about like safety data. So we've been collecting like safety reports every single day, a couple of hundred sometimes. or attends them every day over the years. And so these data has some sort of insight, right? And it brings in an insight of how is the safety condition operations. And so with that, people usually see the categories or what the reports are being reported for.
12:05But then when you look at the categories, you realize that sometimes it doesn't really jive with the stories that they're actually trying to bring in in the unstructured data and so that's where the part where we felt like it's gonna bring in additional insight to the unstructured data that we see yeah and so like there's this uh we've talked a little bit about like the data that you was there the potential insights in that around safety and maybe other insights as you kind of came into this industry and kind of were getting an understanding of like let's say that you built the coolest data science app that was out there and and had a cool model in it to do some analysis like what is the reality of how that would have to run like what is in production mean uh or in in your context yeah so many things i mean just start from the data we don't even have labeled data at the start but say you want a classification app or model.
13:07You need to have some sort of a label, right? And we don't have that. And so you have to figure out ways to bootstrap your labeling process to start off and stuff. And then all the way down to like, of course, training the model is pretty easy these days, right? But then you have to think about things like, hey, what does that look like to put it on the server? Most of our servers are running Windows server. And so I've had experiences putting production some apps on Windows server and that was painful. And so like we have to figure out ways that like work with the ITs and stuff and said, hey, can you deploy a Linux server for us instead and just work it up from there and set it up from there.
13:51Like being said, like that's just overall picture of it. And then you get into details of like, how do you actually store your model. We've got to have some sort of infrastructure to kind of hold that which in our case felt like MLflow is pretty good model registry experimentation tracking and stuff to keep track of what type of models that I'm using and stuff like that and then so many things honestly and then like how do you actually put it on an orchestration kind of service you could use cron jobs but then you know it may not be so flexible then you kind of need to work something out and so you have to get some sort of orchestrator to spin it up and and kind of like make that a service for your infrastructure as well i love this discussion because it uh i think it fits the theme of the show so well around being practical and the fact that yeah i'm sure that there's actually a good number of listeners out there who are really wanting to do machine learning, AI, data science type of things.
15:02And they are sort of in a similar situation in their company because actually I think probably more of the majority of companies are in this sort of situation than sort of infrastructure wise, extremely modern and just cranking everything out on Kubernetes in the cloud and that sort of thing. So yeah, I love this. So you were talking a little bit about the sort of problem of, so you have all of this tabular data with extra comments and unstructured data and certain things you want to do like extract insights or classify maybe some of the unstructured data. But then also nothing has been labeled over time.
15:50It's just unstructured data. So talk a little bit about that bootstrapping problem and how you've thought about that in terms of, I've got all this stuff, I want to create a model, but I have no starting point. When we were going through that whole labeling process or data preparation process, it was pretty interesting because we really didn't have anything, no labelers or anything. And I didn't have a budget to get an external labeler and a data for me. Just hire thousands of people online? Right. I know. I could just do that, maybe. But then again, even that, I've thought of it, and our data is sensitive to start with.
16:28But at the same time, it's so nuanced. And it's so nuanced to the context of our company. And so a lot of data that every company has is just so nuanced to their own context. you know and um at the same time too one of the new ones is this is the way that like um these texts are being written and so a lot of code switching happening you know um which means code switching which means like in asia a lot of times we speak in english but we will kind of put in some you know native languages that we know or that we grew up with and so it's just kind of like you go back and forth back and forth and it's just kind of kind of a common thing especially especially in Southeast Asia.
17:09And so you can't just hire somebody online and just kind of label it for you because you just can, I don't even know how to bring in the context of these guys, right? And, but thankfully, I mean, I would say that the part that really helped was, I happen to have a really good sponsor for the project. And these sponsors, they were super on board. They were not technical. They were SMEs in their own departments. And they knew this stuff, but they know enough of machine learning and data and AI that, hey, it's kind of a model that predicts not necessarily making 100 % accuracy get-go. And so they understand those nuances in a sense.
17:56And so they kind of supported that and understood the kind of things that we have to go through, you know, as a practitioner that we have to go through. They kind of understand that part of it. So that was really helpful for my part because having a good sponsor means you get really good support for the project. But that also means that there's a product that I'm working on, the app that I'm working on is for their people, their subordinates and their people, and people are reporting on them. And what happens is they said to them and say, hey, guys, this is your app. I want you to help Bing Soon out with building this app.
18:38And so having the users themselves on board from the foundational bootstrap level really helped us. So because we didn't have any labels, they had the guys to actually be the ones to label for us. They were the labelers. And so the users themselves were the labelers. So honestly, I was really blessed to even just have that kind of worked out together. I think that kind of worked so much in my favor. Yeah. And in that labeling process, how did you develop your sort of set of instructions for like how you like explaining the problem to them or helping them define the problem and the categories, for example, in a classification model?
19:25How was it for you? Because I've also had the experience personally, probably, and been burned a couple times where I'm like, oh, this problem makes sense to me. I set up the labeling thing. I release a bunch of labelers in there and either the instructions don't make sense or I've biased this in some way or, you know, likely because I, like you had mentioned, I wasn't super close maybe to the users in that situation. But yeah, any learnings from that experience? I think it's a lot of iteration with them. I had so many times like I would travel to see them. They work in our operations. I literally traveled there to see them in person.
20:07And I said, we would just go through, hey, these are the labels that we want to label. We kind of get a general idea of what they are. but when you get into the weeds of it like you when you get into the details of it you're like you would think in this situation that label should be here should be number one but then now somebody else said no it's two so um the way i worked it out was there will always be contention i noticed no matter how tightly knit your labelers or your team are there will always be contention and you just got to work around with it at least to in my experience um i just had to work around with it and the way I worked around with it was I just kind of had a voting system you know and um I I set up an account so the technical side of this is I I could have just given them excel sheets and they could just label them on an excel sheet but I find that um you know they are doing for me a favor and a thing but I want them to have a really good user experience instead of just going through Excel sheets and stuff.
21:13And so I used Arjila back in the days and when they started. Pre-hugging face days. Yeah, yeah, exactly. And I noticed that Arjila is amazing in the sense that it allows you to set up different users and you could, I mean, even other kind of interfaces I've used labels to as well. But Arjila was able to, I could use the API and just kind of like set up each user, right? And then for each user, I would just kind of sample the same data for them. So I had to actually go through the first round, the same number of set of data for them to label, say 500 of them, I think I remember. And they would all label within two weeks.
21:57And at the end of that two weeks, I'll collect them and I'll find which are the most contentious ones. And so the ones that are the most contentious, the ones that have the least percentage of the majority, I would pull it up and I said, hey, guys, what do you think about this? This is contentious for you guys. Why is it contentious? And you work it up from there, right? Because chances are you're going to see the same kind of a label again or the same kind of a data again. and if you talk it out, hopefully when you see a similar thing and you said, hey, we already talked about this, we all agreed that we're going to go with this.
22:34And so that's the first round. We had the same labels, the same data set for everyone. It's a bit inefficient to a certain degree, but I think it's important to actually get into that space so you understand the contentions of each person. And then the second round, I have everyone just kind of like label that scale pretty much. and yeah we collect that from there pretty much interesting yeah so you did in this process you did a initial sort of offline bootstrapping of of labels right and did that so like scale wise like what sort of scale when you're solving so here we're talking about an nlp problem creating a classification model on some labels of this unstructured data.
23:22Of course, this would vary by domain, but sort of what scale of labels did you shoot for when you were doing that initial, trying to get to that place where you could start up and train your first model? We did just about 1 ,800 to 2 ,000 labels of rows of data, basically. And then we start training our first model. That probably means you're not training like a 400 billion parameter model with 1800 samples. You don't even have an infrastructure to be able to train that. Yeah. No, we don't. So what does a broccoli AI model look like? I mean, this is all text, right? And so we were going with like the simplest you can find on Hugging Face at that time.
24:09And I think at the time, sentence transformers were really making it big, you know, for different reasons, like whether it's topic modeling or classification. And at the same time, too, I remember they came up with the set fit model, which is fine tuning the sentence transformer, which was honestly revolutionary for me. Amazing. And I thought it was amazing that like you're able to do something that was meant for similarity. but then you could actually fine-tune it for classification and with pretty good performance. And it's supposed to be something that is a few-shot classification model, a few-shot kind of fine-tuning.
24:49And so I thought 2000 should be enough for me to start somewhere. And in fact, when I trained that, I tried some other models, but I think Sentence Transformers were the ones that actually gave the best performance out of all. It still wasn't that good. you know talking about like 60 something 70 percent kind of thing in terms of f1 score but when i talked to my sponsors about this i said hey guys like you're okay with like me deploying this at like 60 70 percent and i said no actually that's fine right because um the objective for this was number one to bring visibility of these reports to the users is one of the pain points that they said was for us to be able to know what people have been reporting at least in the past 24 hours, they had to get on SharePoint and just different hoops and loops to try to find out filtering and stuff.
25:43But to be able to get that sent out in the email with the classification was already a win. And so I thought, okay, let's do that, but let's not stop that, right? I mean, we should actually create a pipeline and that's where the active learning comes in. And it really helped because I'm glad that I actually used Argyla to start with the bootstrapping of our data set. And having that Argyla, which means our users are already used to the interface and they already have an account. And so I was able to kind of hack around with Argyla as a Python API. And basically, I was able to create a loop where pretty much what this model does every day, it will bring in the new data that people have been reporting for the last 24 hours and make some prediction on it at about 60, 70 percent F1 score, accuracy, whatever it is, and then send it out to the users.
26:42And these users will see it. And at the end of that email, they will say, hey, I don't think this is that signal. It should be this signal. I want to give a feedback. And at the end of it, they're able to click on a link that brings them to their profile in Arjila that will allow them to give a feedback for the particular day's data set. And so over time, now it's in production every day. I would get, from time to time, I would get people giving their feedback. And we've gotten like close to 4 ,000 data sets now labeled from this active learning. And so we would train a model periodically, not on what.
27:21I could have done it automated, but I didn't really want to like just, I didn't feel the need for it yet to put it on automation. but then at the same time like you know you're just collecting an event and we're just training it from time to time basically
27:45hey friends outshift cisco's incubation engine merges innovation with the art of possible a launch pad for transformative emerging tech outshift blends startup agility with corporate strength to develop next-gen technologies from the ground up in AI, quantum technologies, cloud-native, and more. Their newest AI innovation, Motific, addresses a critical challenge in the rapidly advancing world of Gen AI. Bridging the gap between concept and deployment, this model and vendor-agnostic solution supports the entire Gen AI journey. From assessment and experimentation, Motific accelerates deployment from months to days while safeguarding against gen ai security trust compliance and cost risks all while empowering business function and it teams to rapidly configure end user assistance powered by organizational data motific provides advanced customizable policy controls to prevent unauthorized access to sensitive data and helps ensure compliance throughout the entire process With deep visibility into operational and business metrics, Motific enables you to track ROI, optimize costs, and make informed decisions.
29:02By offering a centralized view, Motific deters shadow AI usage and empowers teams to innovate responsibly. So move beyond the traditional constraints of AI implementation, utilizing AI deployment that is both responsible and is revolutionary. ensuring your projects are not just quickly launched, but built on a foundation of trust and efficiency. Visit motific.ai. That is M-O-T-I-F-I-C dot A-I.
29:52so bingsun uh it's super interesting to hear kind of how the you were able to engage the the users of the application through this like reporting process essentially that they were you know had some of the right incentives in place to to respond and to give you updated labels And you mentioned also the model repository, saving models, getting them out with MLflow in the context of you deploying your model on-prem, you updating the model. You just mentioned kind of retraining the model. What does that look like for you right now in terms of that cycle of when you would want to push out a new model after gathering this data?
30:37How you would judge that to be worthwhile or useful in any sort of testing that is relevant to that cycle of getting in new labels, retraining, evaluating, that sort of thing? What does that look like for you and how do you kind of put in the right or how have you thought about the right metrics to understand when to update the model? At this point, honestly, we keep it simple. We just kind of like periodically do it at a cadence, you know, a couple of months or two. But I did think about like, what does it look like to actually measure the drift of the data and stuff like that of the model predictions?
Read the full transcript
31:16That could be one of the ways that we could do it too. But what I'm seeing is actually the model is doing its job fairly well, well enough to actually solve the business problem. And so we don't see a need to actually implement more sophisticated monitoring unless we need to. You know, that's where we're at with it. Yeah. And when you push your model, sort of like you update it, you had mentioned the model repository. How are you shipping your model out to the application? Because I think, like you had mentioned, you only have so many resources. I think there's also a lot of people out there in your situation where I think it was Kristen Lum on a previous episode, she had talked about kind of that data scientist out there that is maybe one of very few or the only data scientists in a potentially a large organization and having to like do all of these things they're not like an ml ops person they're not a model trainer they're not a observability person they're doing all of that right so there are limitations to you know how much sophistication you can put in place.
32:32And I think that some people go way too far and they're like, oh, I'm going to implement all of this stuff. And it actually makes their life as a practitioner less happy than otherwise. So yeah, how have you found that balance? And what does it look like for you to do these cycles in terms of tooling and the things maybe that you, like you say, you mentioned you thought at some point, maybe it's relevant to implement some of this of observability stuff, but maybe not yet, or there's other priorities. So what does that look like for you in terms of how you decide what level of sophistication is right and how you push things out?
33:12I'm 100 % with that, honestly, because it is a matter of priority. My customers are happy, I'm happy, and I'm not going to change what's good. I don't want to break what's been working, right so to speak but that being said like i you know when it comes to like all of that i think i have a general idea of what would be the minimum thing so now i'm working on some other things as well like anomaly detection and stuff like that which needs to be deployed so having gone through that that kind of like set up a like a pseudo infrastructure for me to know what kind of infrastructure that i'm going to be looking for for whatever else that i'm working on And at the bare minimum, I think model registry is super important.
33:54And being able to call the different versions that you've been training and being able to track that and being able to call it through an API, through a function, you know, MLflow has this great Python connection with it. And so being able to do that is just amazing. I mean, it keeps my life sane, right? I don't have to like figure out where I store my model pretty much. So I would be doing exactly the same things with whatever I'm working on next, which I've since moved on from that project and I'm just kind of maintaining it. That project's now in maintenance and now I've moved on to a different project to solve a different part of the business in that sense.
34:34But that project kind of set, like I said, set the foundation and knowing what kind of things that needs to be done. So, sorry, I'm kind of going ahead of myself. So one is MLflow being the most important thing for me in terms of this sort of scenario. The other one is orchestrator is also really important. Having a really robust orchestrator. So for me, I think Prefect was perfect for me. And I was able to do different things and stuff. The types of things that you're orchestrating are what types of things? You could do it real time with Prefect at the same time. You could also be listening and stuff.
35:11But at the same time, you could also just running on schedule, calling different functions, sub functions and things like that. So that was really cool to be able to have that. That's pretty much what we do right now. We're not really going into real time monitoring yet. Until we do that, then we'll have to figure out something else more sophisticated. Yeah. And are you just shipping your models sort of as part of a Docker container or something like that? Pretty much, yeah. We do use Docker containers just so that we can keep it contained in that sense. Yeah, that's awesome. I think you had mentioned in one of our conversations something about DuckDB.
35:53Where does that fit into some of this? So the raw data that we get from is from SharePoint. But if you have anyone who has any experience with SharePoint in terms of wrangling and data stuff, it's so painful. So I thought that would be good to actually have some sort of a middle layer, mini lake house of data lake kind of thing. And I didn't want to bother my IT guys too much. So I thought DougDB is a great thing for it. I don't need a VM for it. And you can have an embedded SQL service that you can use. So that's being pulled every day, pulling the data into DuckDB. And DuckDB will be the one that actually cleans up the data, preparing the data to send it to the model.
36:40And that becomes like a pipeline for me to be able to work around the whole complexity of SharePoint, really. yeah I personally found a lot of use for for DuckDB even in the past uh yeah even in the past year on the even on the more gen AI stuff where you're doing sort of like uh text to SQL or like queries and that sort of thing and every company we're working with has different crazy sets of data or different configurations of this or that and that layer of having a kind of unified analytics layer, but also not that sort of, you know, easy to pull in to Python, easy to spin up, easy to test with locally and then deploy with.
37:28Yeah, that's been really useful. I remember you talked about LensDB for Rack and things like that. And it's the same thing. I love embedded database. I think it just works well, you know, and it's kind of scalable eventually, you know, and I think I really like that. I think there was one blog post I've always referred back to because I also went through the, you know, you and I were at Mines at the same time. And then like there was like data science hype. And then there was like the big data period where everybody was in Hadoop and Spark and all this stuff, which I know a good number of people still use Spark.
38:07But there's a blog post by the mother ducker company. Yeah. But I think the title is big data is dead or something that basically goes through some of the discussion around like, Hey, we all thought we had big data, but like the actual query problem, like the types of queries that we need to run, these aren't like big data problems. What's needed is different. So yeah, for those shout out to whoever wrote that blog post, cause it was really, really good. If you ever want to come on the show and talk about it, that would be awesome. Yeah, well, as you kind of look back on this process and some of the things that you've learned, like, what are you looking forward to in terms of like the future of the process of your own work or of the things you're learning?
38:53Or maybe like, as you go into this next phase, it sounds like you're working on some new things. You'll want to reuse some of the tooling and kind of process that you have used. But, you know, What's different or what are you excited about for this next phase in light of what you've learned over the past years? Generally speaking, I think MLObs is just so nuanced in different contexts. Everyone has a say of what should be done. And I think if I learned something from this was nobody really knows everything. So you kind of have to figure out from there and you kind of take a risk on certain things that you decide in terms of your system design and stuff.
39:37What I'm excited for is actually to be able to take this and see what it looks like for other things, right? And in other applications, like whether it's anomaly detection or whatever it is. In a broader sense, I think I'm excited to see things like embedded database, you know, getting more and more mainstream, especially in the context of LLM and Gen AI and stuff. I love to see that getting more and more mainstream as well. One of the things I'm always thinking about is scale is one thing because a lot of the applications that we talk about today, especially in the context of Gen AI, we always talk about the bigger compute and bigger scale.
40:24I would love to see that getting smaller, which it is happening now, getting more accessible on different devices and stuff, being able to do more cool stuff on device and advanced stuff. I'm excited for that, too. Yeah, I think there's a lot of people excited for that and sort of this new phase of AI where people talk about AI everywhere or this sort of thing, which in reality, you know, there's been machine learning and data science sort of everywhere for some time. But that sort of wave of these newer generation of models kind of being runnable in more practical scenarios is exciting. But yeah, thanks for joining Beansoon to talk about a little bit of your broccoli AI.
41:11It's been fun. Love it. Thanks for indulging me. Yeah, yeah. You and I can hype the broccoli AI and I'm sure we can get Dimitrios to help us hype it too. I don't know if he trademarked that term. He's got it in his hype cycle now. I love it. Yeah. Thanks so much for joining and hope to talk to you again soon. Thanks. Thanks for having me.
41:44All right. That is Practical AI for this week. Subscribe now. Now, if you haven't already, head to practicalai.fm for all the ways. And join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our Beat Freakin' residents, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. that's all for now we'll talk to you again next time
From the publisher
We discussed “🥦 Broccoli AI” a couple weeks ago, which is the kind of AI that is actually good/healthy for a real world business. Bengsoon Chuah, a data scientist working in the energy sector, joins us to discuss developing and deploying NLP pipelines in that environment. We talk about good/healthy ways of introducing AI in a company that uses on-prem infrastructure, has few data science professionals, and operates in high risk environments.
Changelog++ members save 5 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Intel Innovation 2024 – Early bird registration is now open for Intel Innovation 2024 in San Jose, CA! Learn more OR register
- Motific – Accelerate your GenAI adoption journey. Rapidly deliver trustworthy GenAI assistants. Learn more at motific.ai
Featuring:
Show Notes:
Something missing or broken? PRs welcome!




