Sensory Revolution: Meta Unveils Multisensory AI for Thermal, Depth, Visual, Movement, Text, and Audio Processing

27 Feb 2024 · 14 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today Podcast Episode Notes: Sensory Revolution: Meta Unveils Multisensory AI

Episode Overview In this episode of AI Today, the hosts discuss Meta's groundbreaking release of Multisensory AI, which integrates multiple sensory inputs (thermal, depth, visual, movement, text, and audio) to create immersive experiences. This technology marks a significant shift in AI's capability to generate complex environments and scenarios based on simple inputs.

Key Topics Discussed

  1. Meta's Multisensory AI
  2. Overview: Meta announced a new software called Image Bind designed to train AI models using a combination of various data types beyond traditional methods.
  3. Capabilities:
  4. Utilizes inputs from:
  5. Text
  6. Depth: Measures depth in images to create 3D visuals.
  7. Thermal: Detects temperature data.
  8. Audio: Interprets sound clips.
  9. IMU (Inertial Measurement Unit): Tracks motion data over time.
  10. Open Source Release: The software was released on GitHub and gained over 5,000 stars, indicating strong community interest.
  1. Implications for AI Development
  2. Broader Context: Moves away from traditional AI reliant on text and images to a more complex understanding using multiple data types.
  3. Potential Applications:
  4. Virtual Reality (VR): Enhances the realism of VR environments, aiding developers in creating detailed scenes with minimal effort.
  5. Content Creation: Allows for the generation of immersive videos and animations based on combined inputs (text, audio, and images).
  6. Film and Gaming: Significant advancements expected in industries relying on multi-sensory experiences.
  1. Transitioning to Human-like Learning
  2. The episode emphasizes the potential of this technology to mimic human learning processes, where individuals learn from a multitude of sensory inputs.
  3. The goal is to create AI that can process and combine various stimuli to generate realistic environments and scenarios.
  1. Future Predictions and Ethical Considerations
  2. Emerging Sensory Inputs: Meta researchers suggest that future advancements could include inputs such as touch, smell, and even brain signals (fMRI).
  3. Concerns Raised:
  4. The potential for invasive technologies that could read brain activity and thoughts raises ethical issues.
  5. Speculation about the integration of fMRI technology into consumer devices (e.g., VR headsets) and the implications for privacy and data security.
  1. Conspiracy Theory Speculation
  2. The discussion takes a turn into speculative territory regarding the use of fMRI technology by companies like Meta and Apple to enhance AI and marketing initiatives.
  3. Concerns about the ethical ramifications of personal thoughts being interpreted and potentially exploited by corporations.

Key Takeaways

  • Meta's Multisensory AI represents a major shift in AI technology, creating opportunities for more immersive and realistic experiences across various sectors.
  • The ability to combine different sensory data without exhaustive training sets a new trajectory for AI development, moving towards more human-like understanding.
  • The potential integration of fMRI technology into consumer devices raises significant ethical questions about privacy and the boundaries of AI capabilities.
  • The episode concludes with a call to watch the ongoing developments in this field, as they may have profound implications for the future of technology and society.

Links and Resources

  • Invest in AI Box: [AI Box Investment Link](https://republic.com/ai-box)
  • AI Box Waitlist: [Join the Waitlist](https://aibox.ai/)
  • AI Facebook Community: [Join Here](https://www.facebook.com/groups/739308654562189)
  • Learn More about AI in Music: [Musical AI](https://musicalai.pro/)
  • Explore AI Models: [AI Models Pro](https://aimodelspro.com/)

Conclusion This episode of AI Today highlights the exciting advancements in AI technology through Meta's new multisensory approach, while also urging listeners to remain vigilant about the ethical implications of such innovations.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Facebook has just made a major announcement and released I would say a relatively major piece of software and a major update in AI. So today on the podcast, we are going to talk about what that is and what the implications are for this in AI and a lot of other areas in tech. So what it's called is image bind. This is an they actually released this open source. And essentially what this is doing is it's a way for you to train AI models based off of a lot more data than just the traditional image or text image or image to video or image to audio that we're, we're used to seeing. So essentially what this thing can do is based off of a single sound or based off of a single image or based off of an input, it can create an entire environment.

0:44So you might be able to play an audio clip of a street in New York and it will be able to turn that into a picture and then into a video of what that street would look like based off of context clues. And so this thing has a couple different really interesting inputs that it takes as far as what it uses to create these context clues. It works with five different areas. So text, obviously. Depth. It can tell the depth of an image. And it can use that to create different things for videos, making 3D things, etc. It can do heat mapping. So actual temperature of a room. it can detect from an image and it can put that into its calculations audio so the sound like i mentioned before and also imu which essentially imu is what is called um internal as an internal measurement unit and essentially in it what it is is a sensor that provides motion data in a time series format so it can detect if an object has moved and the series of time so it's like tracking that object, I guess, in the most simple terms.

1:53They released this open source on GitHub at the moment, and it has over 5 ,000 stars. And so this is obviously a project that's doing really well. Actually, I believe it's 4.8 thousand stars as of this moment. And it's a project that's doing really well and getting a lot of interest. And so I actually think this is going to make some massive leaps forward for AI. And essentially, this is going to go away from what we're typically used to, right? We have things like these image generators, like mid journey, stable diffusion, dolly, that essentially are just pairing words with images and allowing you to create like a visual scene based off of that text you gave it.

2:30And so now this is going to be a much broader net, it's going to be able to link text, images, video, audio, 3d measurements, which is the depth we were talking about temperature data, like literally thermal, and then all of this motion data from the IMU and it does this essentially without ever having to first train on every possibility it's essentially I think what's the early stages of a framework that could eventually generate really complex environments from an input as simple as text so you could put a text prompt into this you could put an image into this you could put an audio recording into this or some combination of those three right like you could put some audio and a prompt and have it generate something But anyways, what I think this is going to be really powerful for, obviously, I believe this is going to be really big as far as moving machine learning forward in like really bringing it to what human learning is, right?

3:24Because when we are in an environment, when we're in a busy street in New York, and we hear, you know, taxis going down the road and cars rushing by and birds flying and people all around us, like we learn based off of all of this data that our computer, our mind, our self is in like is internalizing. and we know like what situations are safe, what areas are good, what places are familiar. We have all these different context clues and like we have so many different pieces of data that we're inputting to make our decisions. Up until this point, AI is right now just relying on text input. So they're creating this model where this AI can take inputs, like all sorts of inputs and it can create scenarios.

4:03Now, a lot of people are talking about the fact that they believe the reason that Facebook is coming out with this and working on this and open sourcing this, giving this to the public is because they want to create powerful tools for VR, right? This is a perfect application for VR. It's kind of a no brainer. Developers, really, it could take a lot of the legwork off of developers who are trying to develop, you know, really realistic scenes and images and places with sounds and all of these different stimuli inside of the VR to make it seem as realistic as possible. They can now, you know, maybe they put like a background track of a specific location and if an AI can now generate with that location looks like that saves them so much time and energy like can you even imagine the amount of energy it would save them not having to recreate environments and scenes and locations for video games for example or for you know meditation apps or all sorts of things so I really do think this is really powerful really immersive technology for VR but then overall for AI in general this is going to be massive in like the film industry, um, and all sorts of different industries, um, that really have like this multi-sensory kind of, uh, situation that you are looking for.

5:15So I think right now you can obviously use mid journey, um, to give it a prompt, you know, like a dog that's on top of a rock, that's holding a ball, that's, um, wearing a bow tie and it's going to create, you know, the crazy situation for you. Um, and now I think with image buying you're eventually going to be able to create a video of that dog in that situation with corresponding sounds and you'll be able to include, you know, like a detailed suburban living room or the room's temperature, the precise location and anything else that you think fits into that scene or well, anything else it thinks should be should fit into that scene can be created.

5:53So meta researchers, they were commenting on this and they said, this creates distinct opportunities to create animations out of static images by combining them with audio prompts. For example, a creator could couple an image with an alarm clock and a rooster crowing and use a crowing audio prompt to segment the rooster or the sound of the alarm clock to segment the clock and animate both into a video sequence. So really powerful for video. I think for video games, that's obviously going to be big. I think at the same time, content creators are going to be able to make like really immersive videos with realistic soundscapes and movements just based off of text and image or audio input.

6:32I could even see a world where I took this podcast and a picture of myself sitting in front of my desk and had it recreate, you know, me talking in front of my desk and making the video instead of me having to record, you know, theoretically. So a lot of really amazing technology coming out that I think is going to really help enable this. this. Another interesting thing that they did say was, in typical AI systems, there is a specific embedding that is vectors of numbers that can represent data and the relationship to machine learning for each respective modularity. So that's what Meta said, they said image buying shows that it is possible to create a joint embedding space across multiple modularities without needing to train on data with every different combination of module of modularities.

7:18So it's important because it's not feasible for researchers to create data sets with samples that contain, for example, audio data and thermal data from a busy city street, or depth data and a text description of a seaside cliff. So I think, all in all, meta really just kind of use the this technology as eventually expanding into it, like all of your current six senses, so to speak. And And so while they explored six modularities in their current research, they believe that introducing new ones that link as many senses as possible, like touch, speed, smell, brain, FMRI signals will enable richer human centric AI models.

8:02That literally is a quote. Okay, this is, I think, the biggest thing I've learned from all of this and a takeaway that I hope if you listen to my AI mind reading podcast, and in case you didn't and you think I'm like crazy. They created an AI that can read your mind. They put you under an fMRI scanner. And well, essentially, they put a bunch of people under there, had them listen to some things and listen to like podcasts for 16 hours, visualize it. And then they trained it off of the transcript of the podcast and what was happening to their brain while they listened to it. And then they could have people go under it, read something and visualize what they're reading.

8:36And then the AI was able to translate and write down everything they read. So reading their mind, essentially. Okay, the reason this is relevant is because they literally said, while we explored six modularities in our current research, we believe that introducing new modularities that link as many senses as possible, like touch, speech, smell, and brain fMRI signals will enable rich human-centric AI models, blah, blah, blah, fMRI signals. What does this mean? This means my prediction is almost certainly correct my prediction was that they were going to be able to stick fMRI scanners into VR headsets and that those VR headsets are going to be able to read your mind essentially do you think like do would you put that past Facebook first off meta whatever they have an ads business they want to know every little bit of peace personal information about you in fact a lot of people don't know I'm in marketing and Facebook literally they They obviously have a profile on every single person that's on their platform.

9:41They add a ton of data to that profile that you never gave them. They will literally go to credit bureaus and buy data from them about you and add it to there so that they know what your annual income is, they know what your credit score is, and people are able to target ads at you based off of that. They buy third-party data from thousands of sources to create the most complex profiles essentially on you and everything you do and see outside of Facebook. irrespective of Facebook, just so their ads business is better. So would you put it past them to literally like what they just said, that having brain MRI signals is going to make their AIs better?

10:20Would you put it past them to put that technology into a headset? And now a lot of people are like, hey, this is kind of crazy. Like as if, you know, like, because I've talked about like governments using this and companies like Facebook, I specifically called out Facebook. And I'm also calling out Apple because Apple already inputs health data scanning stuff into their Apple Watch. And Apple's going to drop a headset. So Facebook drops this technology. Apple's going to drop it. And they're going to say like, hey, this is to help detect brain aneurysms and a bunch of other things. Sorry, this podcast is turning into the conspiracy theory hour.

10:55I gave you the data. And now this is what I also think will happen from this. And by the way, as far as conspiracy theory, this is not a conspiracy theory. I will bet you money that this will actually happen because they literally put this in their press release that the f MRI sets Data is gonna make it more of a rich experience aka to have to get that data from somewhere And since you're not gonna go into an fMRI scanner They're gonna they're gonna have to get it from you somehow so it's gonna come from VR headsets And anyways people are like oh man, but those machines are so big Well, I was doing a bunch of research and back in 2013 They actually created a technology to get those those devices in a much, much smaller form factor.

11:35And so I believe just based off of the advancements they had back in 2013, I didn't hear a lot about it since, but they're going to get that thing into a headset 100%. They're going to get that into a VR headset, and then they'll be able to read your mind. And when they show you an ad on your VR headset in your metaverse, they're going to be able to detect like, did you react positively or negatively to it? they're also literally going to be able to read the thoughts in your mind so that's kind of alarming like the words you're thinking um and beyond just the words you're thinking there's a lot of data that has come out or not data there's a bunch of research like these ai reports that came out where they are able to you can visualize a face and they can draw the face you can also visualize a scene and they can draw the scene like they can recreate it in an image and it's not as obviously as powerful as like mid-journey it's not like quite literal and some of the pictures like like visualize this face you visualize it and it's like a little bit off but come on this technology is going to get better they literally can pull images out of your mind and literal phrases words text and thoughts out of your mind so in any case uh do with that what you will but uh i for one will probably not use any technology that has an f mri scanner in it that is scanning my brain or has a capability even if they say they won't.

12:54Because imagine like third-party hacks on something like that where like you're using an app and it like accesses your fMRI scanning like capabilities of your device. And it's like, oops, wasn't supposed to do that. Like all like phones have had similar hacks where like, hey, delete this app because it's like a spyware app. Okay, yeah, would you like a spyware app on your VR headset that's reading your literal thoughts, transcribing them and sending them to a foreign government or a company or any organization. In any case, I think this is probably a good place to end off on this podcast episode.

13:32This is going to be a very big place to keep watching. I think this has a lot of implications for AI. This may be one of the biggest ships in the entire industry. We're moving from text to output to any sort of sense to output, and including the fact that Facebook would like to get fMRI information from you. So an area we'll be following. Thanks so much for joining us today.

From the publisher

In this episode, we delve into the groundbreaking release of Meta's Multisensory AI, which promises to revolutionize perception and interaction by integrating thermal, depth, visual, movement, text, and audio processing capabilities into a single platform, ushering in a new era of immersive experiences.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Sensory Revolution: Meta Unveils Multisensory AI for Thermal, Depth, Visual, Movement, Text, and Audio ProcessingAI Today · 14 min
Listen in VO