In short
Podcast Notes: AI Today - Episode: Picturing Tomorrow: Exploring the Future of AI in Computer Vision with Sumedh Datar
Episode Overview In this episode of "AI Today," host Jaden engages with expert Sumedh Datar to explore the future of artificial intelligence in computer vision. The discussion covers Datar's journey into the field, key challenges faced, real-world applications, ethical considerations, and emerging trends in AI and computer vision.
---
Key Themes and Discussions
- Sumedh Datar's Journey into AI and Computer Vision
- Background: Began in biomedical engineering, specializing in computer vision and medical imaging.
- Initial Interest: Developed an interest in imaging through a thesis on identifying cancerous regions, highlighting the societal impact of technology.
- Challenges in Computer Vision
- Data Scarcity:
- Notable issue when starting projects with limited data.
- Strategies include sourcing open datasets and creating platforms to collect more data.
- Hardware Limitations:
- Importance of selecting appropriate hardware for specific applications, especially in health care.
- Data Annotation:
- Challenges with obtaining labeled data for training models.
- Discussion on manual versus automated annotation strategies.
- Approaches to Data Quality and Model Development
- Importance of deploying initial solutions quickly to gather real data, even if it's imperfect.
- Iterate on models based on collected data to improve accuracy over time.
- Real-World Applications and Impact
- Healthcare: Developed non-invasive methods for oral cancer detection, encouraging more patients to seek diagnosis.
- Retail: Implemented facial recognition for automated attendance systems, showcasing practical applications of computer vision.
- Future of Computer Vision
- Merging visual data with tabular data to enhance insights and decision-making.
- Potential for significant advancements in analyzing visual content in social media and other platforms.
- Ethical Considerations
- Challenges with model interpretability, particularly in healthcare.
- Necessity for AI solutions to work alongside medical professionals, ensuring explanations for model outputs are clear and comprehensible.
- Emerging Trends
- Self-Supervised Learning: Models that can label their own data, reducing the need for manual annotation.
- Generative AI: The convergence of text and vision models, creating more effective applications (e.g., ChatGPT paired with visual data).
- Challenges in Retail Item Recognition
- Complexity due to the sheer number of products and variations in packaging.
- Need for adaptive models that can recognize changes in product design.
---
Measuring Success of Computer Vision Projects
- Importance of ML Ops and model monitoring.
- Use of metrics tailored to business goals rather than just accuracy.
- Continuous evaluation and adaptation to ensure alignment with user needs.
---
Advice for Aspiring Engineers
- Fundamentals: Emphasize strong understanding of algorithms and data structures.
- Learning Resources: Recommended Stanford's CS231n course to grasp deep learning concepts.
- Software Engineering Skills: Importance of understanding the end-to-end process and customer needs.
---
Conclusion The episode provides valuable insights into the current state and future of AI in computer vision, emphasizing the importance of data quality, ethical considerations, and the potential impact of emerging technologies.
---
Resources Mentioned
- Invest in AI Box: [AI Box Investment Link](https://republic.com/ai-box)
- AI Box Waitlist: [AI Box Waitlist Link](https://aibox.ai/)
- AI Facebook Community: [Facebook Link](https://www.facebook.com/groups/739308654562189)
- AI in Music: [Musical AI Link](https://musicalai.pro/)
- AI Models: [AI Models Pro Link](https://aimodelspro.com/)
---
For further inquiries or to connect with Sumedh Datar, reach out via LinkedIn.
Note: Ensure to engage with this episode for a comprehensive understanding of the impact of computer vision in various sectors, especially healthcare and retail.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28Welcome to the AI Chat Podcast. the show. Hi, thank you so much, Jaden. Just so that to know, all the views are completely my own and talking purely based from my past experience. I do not represent any company or anywhere that I work for. So yeah, thank you so much for having me, Jaden. Super excited to have you on the show. And yes, that sounds fantastic. The first thing I wanted to ask you about is what kind of got you interested in working in this space in general in the beginning? Have you always been interested in tech? What was your kind of journey into this? Yeah, so my journey was somewhat very different compared to others.
1:10So I actually started my undergrad in biomedical engineering. And there, I was fortunate enough to just do my specialization in computer vision and medical imaging. And my final year thesis was actually on identification of cancerous regions. And that's where I got super interested into the imaging world and saw how impactful it is and how useful it is to the society. That's amazing. That's so interesting. And, you know, growing up, was this an area that you had been interested in? I guess, what kind of introduced you to it? Did you have family members or friends or people that were kind of interested in this space?
1:52Oh, no, actually, it was nobody. It was just by myself. I clearly didn't know initially. I thought hardware is where I think I'm more interested in. And then later, I realized that's not my cup of tea. And I think medical imaging was closest to software that I really got exposed to while doing my undergrad, because since my undergrad was not in computer science, and it was not code intensive, The closest to coding was medical imaging. And that's when I was super thrilled that I could actually contribute to a good amount of code. And that got me interested in this space. Very cool. That's amazing.
2:31I'm wondering, can you elaborate a little bit on some of the key challenges you've encountered in applying computer vision to, you know, healthcare and the retail sector? Oh, yeah, sure. So the biggest challenge is scarcity of data, right? Throughout my experience working in the computer vision space, we have always started where there's absolutely no data, right? You still need to have a solution, but you don't have data. How do you solve a problem? So you need to find creative ways to actually identify. And when it comes to data scarcity, okay, you have data scarcity, but you need to overcome and start somewhere.
3:12Right. So the possible ways to start are getting a few actual images that you are going to use later, which is really hard. So, for example, when I'm working on the cancer, working in the cancer space, I would actually go to the patient and actually take an image of the patient and then come back and then train my model and see how that works. OK. There are other techniques as well. There are some open source data sets that's available We can take that initially to just start off and then see how the model performs and then put out in the wire Put out in the wild and then create a platform to collect more data So data scarcity is the first thing then the second thing is hardware limitation, right?
4:01Based on the problem that you are solving you need to exactly get the right hardware as well right at what speed the camera has to run what kind of accuracy you need when it comes to health care you should be extra careful right and it's not health care when it's like in other space sometimes if you skip a few frames it's fine but and also these deep learning models are super heavy what kind of models do you want to use we want like super accurate models which has to detect every single frame or it should be like yeah you're okay for you but it's important to you know have the speed rather than accuracy so I worked in both areas and both actually have equal importance so and it's a trade-off too and last thing is the annotation of data you don't get labeled data you don't get annotated data so when I say annotated data say if you have a car in an image so someone has to manually put a box around the car and And similarly, you need like thousands of images like that.
5:08And that you need to feed, that you feed to the model and then model starts doing the prediction, right? But annotating that data, what strategies can you use apart from doing it manually to annotate? These are some challenges that I've faced. Wow, that's very, very interesting. How do you kind of approach some of those issues? For example, the issue of, you know, data quality when developing computer vision solutions. How, what's your approach on that? The first thing is to have something out there, right? If you sit back in your lab or if you sit back in your office waiting to get the best possible predicted model out there, that's never going to happen.
5:49You need to have a solution first out there because that helps you in getting more data. And once you have more data, you can build better models. And the data that you get is the actual data that you'll be dealing with in future, right? So getting a platform out there and having that platform to collect as much data as possible, it'd be wrong or it'd be like, say, 20 % accurate and 80 % wrong. That's fine. But you're getting a lot of data. You can go back. You can iterate on it, train your model, and then redeploy back rather than waiting forever to have the best model out there. That's one of the strategies that I've learned.
6:30That's really interesting. Interesting. What would you say are some of the, for you and all the things that you are working on, right? You have a lot of knowledge in this space. What are some real world problems you've solved using computer vision and kind of what impacts have they had? so um i walked in the healthcare space especially in the oral cancer space right uh so uh patients at least this is this is an indian scenario in rural areas where uh people are predominantly have smoke and uh they suffer from oral cancer but they're very reluctant to go get themselves tested because the procedure is so painful uh they actually have to go through a biopsy test which hurts a lot and many people are reluctant to do that how do you make it painless right so i basically saw this problem and i was like can we solve this painless right non-invasive way can we really solve so i actually went and took a photo and then once you have the photo you can actually say put a box and be like yeah i think this could be cancerous maybe you have to go for diagnosis right so it may not be accurate but what's happening is you are actually telling the patient that hey look you seriously have a problem so you have to go for a better treatment to just save yourself right so that way it was more convincing to the people and that helped them to get better care so that's when i saw like real value in computer vision wherein you can solve the problem without pain so that was one space and the other space was basically like like doing like face recognition like an automated attendance management system and rather than using a book and a pen you can basically take a snapshot and then you can recognize the faces and then do attendance that was the other space so computer vision what it has done is the algorithms are not too crazy.
8:37You don't need to have like bunch of loops or crazy dynamic programming or anything. It's just about finding that small problem for which you can apply a computer vision model and then the impact is so huge. That is how computer vision has always been. So that way I enjoyed computer vision. Very, very interesting. And I would be curious to hear your opinion on this. You've obviously worked a lot with computer vision and ways that this is, you know, helping making big impacts in some really incredible spaces, for example, like healthcare. I absolutely love your examples there because, you know, you really can see this is something that is helping the patient so much as making such a big positive impact.
9:19I'd be curious from your perspective, as you kind of look to the future with this technology, where do you see, you know, computer vision mixed with, you know, artificial intelligence and everything we're developing in these spaces, where do you kind of see that in the future? What kind of changes do you think will happen? What kind of technology do we maybe not have today, but you think we may have in the future? Yeah. So like what happened is if you take like, say, a decade ago, right, was when data science became extremely popular. Data science, I mean, tabular data, right? Like say, transaction data or say in the insurance space, basically it's just what you're doing is you just have rows and columns and you're just having more and more features, right?
10:05And you could build predictive models with statistical techniques like simple things like linear regression, logistic regression, random forest, things like that, right? But now what's happening is computer vision is almost paired with the tabular data as well and the results on the computer vision side of it is so good that you can actually reason out as to why it is good and then you can find ways of how you can tag the computer vision features along with the tabular features and make bigger sense and have better decisions or maybe have decisions that probably you might not think but the model is actually giving you like that's where the future is kind of moving towards.
10:51Okay. Because what has happened is 10 years ago, you had, for example, like say Facebook, right? But today you have TikTok, Instagram, Reels, lots and lots and lots of visual content, right? So you can get a lot of information from visual content. So when you pair visual content with tabular, you can get way more insights. I think kind of that's where the future is heading towards. Very, very interesting. Something that you mentioned earlier also struck me as fairly interesting, you kind of mentioned the importance of like labeled high quality visual data. What strategies do you recommend for organizations to source and utilize that type of data?
11:30Yeah, so the thing is, as a machine learning engineer or as a data scientist, you are not just doing a simple model training or changing the architecture of the model. But you should also be okay with doing a lot of legwork like the data quality. Why I am coming to this is say for example you have a vendor who is actually doing the data quality for you right Once they do it and when it comes back to you You have to actually check every single data that you have And once you feed it to the model see how the response is like And then if there are issues which there will always be You have to go back to the vendor and tell them Because you have the maximum context and they don't have it right The context that you have they don't have it so the high quality data starts with you and ends with you uh the other guys are just uh supporting you but you should never assume that they are doing your job and you don't have to it's like it's your job in the end so that's where um high quality data comes into picture and the involvement of the stakeholders here very interesting you know earlier you you know you talked about the fact that you've done some really impressive things interesting things in the healthcare space, I'd be curious if you can kind of discuss some of the ethical considerations that come into play when deploying computer vision technologies, especially in sensitive areas like healthcare.
13:05Yeah. So what has happened, and even I struggle on a day-to-day basis, right? When a model gives a certain output, it's very hard to understand why it is giving a certain output right like take for example uh tesla they are purely running on uh visual sensors that should say right and then um they had the they had a image of a pickup truck which had like wood logs or something and that was being shown like a traffic signal right on the display itself on the display itself visually as a human you see it like yeah it's a wood log and it's a truck it's not a signal right right since you but since it's predicting as a signal the decisions are according to that like oh is it green is it red right it's right so deep learning models have this particular problem wherein it's so hard to interpret why the model gave a certain outcome right now coming back to the healthcare angle it's very hard to say someone that hey i think you have cancer right it's it's very hard and you need to have like ample amount of information to back it up right yeah so um a model interpretability is something that there's a lot of research going on in this space wherein um you can clearly visualize why the model actually gave a certain output like you can visualize the layers of the neural network and you can be like hey i'm not sure but this is what the neural network says and this is the reason i feel the output is somewhat like this like the reasoning like this helps and second thing is it should be always backed by subject matter experts especially in the healthcare space right uh three people use three different answers right and the more specialized the doctor is the answers are more different but how do you get that right so it's always good to have an ai solution with the support of a doctor wherein you're actually helping the doctor but you're not really taking over the doctor because it's like way far in the I don't think AI is there yet, but there are some areas where when you have a solution, you should have explainability techniques along with it.
15:41This is why the model told this is the answer. That's when I think you'll be in a better place rather than, yeah, okay, this is the probability 0.9 and this is what it is. Okay. Yeah, that makes a lot of sense. what from your experience do you think are some kind of emerging trends in computer vision that you find particularly exciting or promising today uh the emerging trends are one is like the self-supervised learning wherein uh you just give the data and the model does the labeling for you uh emerging field and of course the gen ai is the next big thing right so you already have the chat gpt wherein uh you give a question and it gives you an answer right uh and there are a lot of vision models wherein you just give an image and it gives the description of it right can you that with chat gpt and build a better model build better models and eventually better applications right of numbers this is the emerging trend very very interesting yeah that that's it's exciting stuff it's exciting to be kind of uh watching is unfold so fast everything's advancing i'm wondering could you share some insights into some of the specific challenges and solutions associated with item recognition in retail i know there's a lot of different kind of challenges and nuances there right right so one thing is in the retail space uh just the sheer number of products right like you have so many products uh it's just so hard to do the recognition so how do you do the recognition for that and uh when it comes to visual recognition when it comes to supervised learning uh what you're doing is uh you're basically doing prediction from the prior data you have now what is happening is when you run some kind of a campaign when the packaging changes right the model results also change automatically interesting oh man that's so complex how do you handle these situations and then the third thing is like the sizes right like you have so many different sizes they all look the same and with vision you can do only so much so that's when you need to think about other signals like how do you handle like the sizes do you want of bring in like point clouds and things like that which helps you in making better decisions these are all the challenges that we face in the retail space very interesting yeah and for the for people listening that are trying to understanding like kind of what we're talking about there's a number of different retail retailers and stores where essentially they have like a little shelf and you can come and put all of your products on the shelf they'll scan them analyze them then they'll just tell you what the price is it's kind of like self-checkout that you know you do a walmart except instead of scanning all the items you can just put all the items on a shelf so there's really cool things happening with computer vision right now but yeah like you're imagining i imagine that it just throws everything for a loop when you know for example you have a can of coca-cola you know what that is but if they do a new campaign and there's a new packaging on there or they change the size of the can or the bottle all of a sudden like you got to refigure out what every single item is as it's constantly changing with thousands of variations oh my goodness i can't even imagine the headache but i'm happy The tackling it is a big, big challenge.
19:06When you do something like this, how do you kind of measure the success or the effectiveness of a computer vision project? Yeah, this is something that's super important, right? Say in your lab or in your environment, it's working really well. Once it goes in the wild, it's not working well, right? so how do you handle the ai adaptability right how many people are ready to accept the technology how many people are using it regularly right so this is when ml ops and model monitoring they come into picture and you need to have metrics you need to have dashboard and continuously keep monitoring and seeing what's going on is it doing really well are the models off are people not happy like what's going on right so data monitoring model monitoring these actually play a big role and the metrics are very different it's not just the regular accuracy metric but it's more on a private level and the metrics should be translated to business right so the metrics are very uh subjective and the definition of the metric is not very straightforward so it's very tailor-made to the use case that thanks to business so tracking these metrics actually help you know how good or how bad the models are doing and yeah AI adaptability is not easy and we need to find different ways of monitoring to make sure we have a successful product.
20:54Very interesting. Really appreciate you coming on the show today. I would love to ask you, it's kind of a question as we're wrapping up here, what advice would you give to aspiring engineers who want to specialize in computer vision and deep learning today? yeah based on my experience coming from a non-computer science background to a computer science background what I really noticed is the algorithms and data structures play a big role because you are coding the equations right so you need to know your basics well having an expertise in one language is enough but being an expert in that language is the most important thing to know the syntaxes that's the most important thing and when it comes to deep learning machine learning i think the stanford cs231n course is the starting point because there they actually teach how to how the algorithms work by actually coding from scratch without using any libraries so So in the process, you learn all the internal components that is hidden when you just use a library.
22:07So that's where your learning is maximum. I think these two are the most important thing. And last thing is the software engineering aspect. So nowadays, what's happening is it's not just on the research level. It's also on the application level, right? Like in the end, it should go to a customer or it should go to someone who has to use it or he, she, whoever it is, they have to use it. Right. So how do you see it from a customer's lens? And I think this is where software engineering comes into picture. Having strong software engineering skill set, which involves like a full stack development, a rough knowledge of the front end, back end kind of gives you an end to end view of how the product looks like.
22:49I think these are good starting points to get into this. Really interesting. Amazing and great advice. So really appreciate you giving that. you know if people want to contact you or ask you questions or you know connect with you what's a good way for people to find you yeah I think they can find me on LinkedIn and ask me any questions that's totally fine thank you so much Sumed for coming on the podcast I really appreciate all of your insights and everything you've shared to the audience thank you so much for listening to the AI chat podcast make sure to rate us wherever you get your podcasts and have an amazing rest of your day.
From the publisher
In this episode, we embark on a journey into the future of AI in computer vision alongside expert Sumedh Datar, envisioning the technological advancements and societal implications that lie ahead.
-
Invest in AI Box: https://Republic.com/ai-box
-
Get on the AI Box Waitlist: https://AIBox.ai/
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
