This Company Mapped the Entire World in 3D. Here's Why.

15 Apr 2026 · 1 h 3 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that frontier AI models lack true physical-world understanding, and that “spatial intelligence” requires a machine-readable, continuously updated “ground truth” 3D world map. It connects maps to multi-agent communication, retrieval for reasoning models, and simulation/forensics where synthetic scenarios must be anchored in real geography.

Guest backgrounds

Peter Wilsinski is Chief Product Officer at Vantor. He previously worked at Palantir (joined 2012), starting as a QA engineer, then software engineer, then product lead. He helped pivot Palantir’s defense work toward operations using mapping capabilities called Gaia, and spent the last five years on Palantir’s ontology system (ontology language/toolchain/runtime). He has worked on mapping “on and off” throughout his career.

Key claims

World models lag language models by ~5–10 years; current “world models” can be hallucination-like, so grounding matters. Vantor’s approach translates pixels into embedding space so text-based LLMs can reason about geospatial entities. Vantor builds a global 3D replica (100M+ sq km, ~3m spherical accuracy, 50cm resolution) and updates it as the world changes.

Notable examples

A story about a new building near Denver not appearing on Google Maps, causing brief doubt in his own eyes. The ProPublica January 6th Parler video alignment story (spatial/temporal arrangement instead of a single timeline). The New York subway map as an example of purpose-driven cartography (context over strict physical scale).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Evolution of Maps and Technology

0:00 to 0:37

Learn about the historical significance of maps in technology and communication.

“The number of bits is increasing and the number of atoms is staying constant.”

Understanding the Gaps in AI

1:00 to 2:17

Explore the shortcomings of AI in understanding the physical world.

“AI can write code, summarize documents, even reason, but when it comes to operating in real environments, it still struggles.”

Spatial Intelligence and Its Impact

2:17 to 4:27

Discuss why spatial intelligence matters in various fields, including finance.

“So the ontology language, the tool chain, the runtime, you know, essentially how to map the digital world.”

Digital Mapping and Its Evolution

4:27 to 6:28

Learn how Vantor creates digital representations of the physical world.

“Intelligence really is about understanding the physical world.”

Making Maps Machine Readable

6:28 to 10:44

Discover how embedding models make maps understandable for AI systems.

“Then the other part of it is, I mean, fundamentally, is spatial intelligence perhaps more important for an AGI-like system?”

The Future of World Models

10:44 to 13:18

Examine the development of world models and their real-world applications.

“And have you, so how much of this have you done?”

Ground Truth World Models

14:00 to 15:10

Learn how world models can enhance accuracy in mapping and simulations.

“But all of the buildings exist, all of the roads are true, all of the sort of fundamental characteristics of the world are accurate.”

The Role of AI in Mapping

15:10 to 19:30

Explore the impact of AI on mapping real-world environments and simulations.

“Would it be accurate to call it a ground truth world model, like what you've built?”

Integrating User-Generated Data

19:30 to 21:40

Discover how users can contribute to and enhance the mapping system.

“So we think of it as really putting every pixel in its place, so that you're starting to get this, you know, patchwork ground truth data based on whatever data you're feeding in.”

Understanding the Tensor Globe Platform

21:40 to 24:10

Dive into the components of the Tensor Globe platform and their functionalities.

“Can you explain Cortex and Nexus for us?”
Show all 31 chapters

AI and Spatial Intelligence

24:10 to 28:00

Examine how AI is transforming spatial intelligence and mapping capabilities.

“which is really focused on looking at the globe, right?”

Exploring High Resolution Data and Spatial Intelligence

28:00 to 29:23

Learn about the advancements in using high resolution data for spatial intelligence applications.

“And so, you know, to your earlier question, like, do we have a compute cluster big enough?”

The Future of Mapping and Semantic Understanding

29:23 to 31:05

Discover how mapping can evolve to include deeper semantic layers and insights.

“on for 15 years, which is really much more image segmentation, sort of traditional computer vision algorithms versus these more advanced sort of reasoning algorithms.”

Developing a Vocabulary for Spatial Data

31:05 to 33:01

Understand the importance of creating a language and grammar for spatial data.

“And maybe you have another stack over here, that's another sentence.”

Predictions and World Models in AI

33:01 to 34:53

Explore how world models can generate predictions and validate hypotheses.

“Okay, I'm going to, I'm going to try and, and I'm going to posit an explanation here and I'm going to see if I got this right.”

The Role of Synthetic Imagery in Intelligence

34:53 to 36:55

Learn about the use of synthetic imagery for intelligence and forensic analysis.

“And then let's diff that image against the image we actually get.”

Digital Forensics and Spatial Reasoning

36:55 to 38:11

Examine how digital forensics can utilize spatial reasoning for better insights.

“You know, you can imagine them being able to do very similar things, right?”

Navigating with Advanced Positioning Technologies

38:11 to 40:16

Discover how new positioning technologies can enhance navigation systems.

“Like if you're like, like, you know, you're in a legal case, you could use this model and say like, OK, this is what we think happened in this in this instance based on the evidence.”

The Future of World Models and Mobile Applications

40:16 to 42:00

Learn about the potential integration of sophisticated world models into mobile apps.

“And I think a lot of the challenge of this, again, going back to sort of what makes it hard to build these systems, you know, there's obviously an AI model development element, which is not my personal area of expertise.”

Data Orchestration in 3D Mapping

42:00 to 44:20

Learn how data orchestration is essential for creating a shared digital representation of the world.

“And, you know, you can really think of in the same way that a Git repository or a version control system would do that for a piece of code.”

Augmented Reality and Global Positioning Systems

44:20 to 46:40

Explore the integration of augmented reality systems with global mapping technologies.

“But what about the vision models and the world models?”

AI Training in 3D Environments

46:40 to 49:40

Understand the importance of 3D environments in training AI models and perception systems.

“I think it's like it's it's always tempting to anthropomorphize things like but we do.”

The Future of AGI and 3D Intelligence

49:40 to 52:20

Discuss the relationship between AGI systems and the necessity of understanding 3D space.

“We also have a lot of publicly available data through our open data program for a lot of sort of wildfire events, things that are of public note.”

The Role of Augmented Reality in Everyday Life

52:20 to 55:40

Learn about the potential applications of augmented reality in practical day-to-day tasks.

“Airplanes moved from military to consumer.”

Revolutionizing Tasks with AR

55:40 to 56:01

Discover innovative ways AR could enhance learning and hands-on activities.

“It's like people sitting on their couch alone, watching slop.”

AI and Augmented Reality for Learning

56:01 to 56:46

Explore how AI and AR can enhance learning experiences such as solving a Rubik's Cube.

“and I was like watching YouTube videos to try to memorize like exactly how to do it.”

The Challenge of Predictive AI Cues

56:47 to 57:51

Discuss the complexities of developing predictive cues in AR systems.

“And, you know, I think it would be an incredible experience to have Ikea in my glasses saying that's, you know, object 12, put object 12 in object 11.”

Onboarding Users for New Technology

57:52 to 58:59

Learn about the difficulties in onboarding users to new voice and vision-based technologies.

“it felt like I had to go to Hogwarts and learn all the spells of like Alexa, timer, Alexa.”

App Discovery Through Physical Interaction

59:00 to 1:00:16

Discover how physical locations can facilitate app discovery and user engagement.

“I think that would be something we'd probably partner with.”

Navigating the Digital World via Physical Spaces

1:00:17 to 1:01:18

Understand how physical environments might guide navigation in the digital realm.

“Like, I think over time, you know, these physical, the physical world is going to be a much more important part of digital product discovery.”

The Future of Agent-to-Agent Communication

1:01:19 to 1:02:10

Explore the potential for direct communication between digital agents in various contexts.

“whether it's QR codes, which I hate, but I think are very effective.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The number of bits is increasing and the number of atoms is staying constant. Maps predated language. Like when you look at cave drawings, there are older maps than structured language. And I think they're a really important part of technology and a really important part of multi-agent communication. I trust Google Maps so much that I actually doubted my own eyes. World models, you know, they're probably five years, 10 years behind language models. And what that means is it's not just that the models are behind or the technology is behind. it means that our words and language and terminology for describing the sort of architecture of the system is very early.

0:37A lot of our vision for what we're building is a world where you never look at a video again. You're only looking at a map. And the video is on the map, but you can zoom out and you can zoom in. And a cure-tax expenditure from the top 5-10 hyperscalers is probably 70-80 % of total U.S. defense spending.

0:57Grant:Welcome, humans, to the Neuron AI Explained podcast. I'm Grant Harvey, writer of the Neuron Daily Newsletter, and today I'm joined by Peter Wilsinski, Chief Product Officer at Vantor, talking about one of the least discussed but most consequential gaps in AI, call it the gaps in frontier AI models right now, the inability to truly understand the physical world. AI can write code, summarize documents, even reason, but when it comes to operating in real environments, it still struggles. And that's where spatial intelligence and Peter comes in. Peter, welcome to the Neuron. Thanks so much, Gra. Happy to be here.

1:29Grant:Yeah, it's awesome. I'm excited to talk about this topic because I think it's something that, you know, we're always talking about language models and how great they're getting and agents and all that stuff. We're leaving this part of the conversation out of it, which is like, how does AI actually play out in the real world? How does it understand the real world? So I guess to start, would you just tell us a little bit about your background and your time at Palantir and what pulled you into this problem space at Vantor? Yeah. Yeah. So I've been working on maps really for my whole career on and off.

1:57And I joined Palantir in 2012, was working as a QA engineer, then a software engineer, and then spent a long time as sort of a product lead there. I kicked off a lot of work on sort of pivoting from intelligence to operations in the defense space with some mapping capabilities called Gaia. And then spent my last five years there really working on the ontology system. So the ontology language, the tool chain, the runtime, you know, essentially how to map the digital world. And, you know, I think about the real convergence of these things as combining a map with a graph, right? So that you have a graph that represents this very abstract digital world and then a map that represents sort of the physical world.

2:36And I think as the digital world gets more and more complicated, the map becomes a way to ground that complicated nature into the unitary thing that is the physical world. And I think, you know, I'm sure all of us feel this way day to day, but whether it's on Twitter or watching the news or we're just overwhelmed by the feeds of digital information, that's really multiplying. Yeah. You can sort of say the number of bits is increasing and the number of atoms is staying constant. And so I think of what I'm excited about with spatial intelligence, there's always a lot of good technology work to do.

3:08But it also is a way to sort of recenter us down on the unitary world that we all live in, that ultimately all of this digital apparatus is meant to make better and improve.

3:20Grant:Yeah, yeah, that's interesting. So let's talk about spatial intelligence then. What does that actually mean? And why should people in finance or marketing care about it? Like, you know, obviously, we all live in the physical world, but but how does it actually impact us? Yeah, I mean, I think it's a really fun time to be talking about this in, you know, February 2026, where, you know, I think the public equity markets are always an interesting signal that they don't determine the future. But I think they do a good job of trying to understand the uncertainty of where we're going. And, you know, we're in the middle of this world where we're really re-underwriting the risk profiles of different businesses, which sort of relates to what does the economy look like in the future?

4:02And I think, you know, after a world where, you know, last 20 years, it's been financialization, digitization, the winners in the economy have been companies that have figured out how to wrangle bits and wrangle the flow of information. A hundred years ago, the companies that were winners were the railroads, were the oil companies, were the sort of aerospace industrial companies. You know, if you look at Germany in the 1800s or the United States in the 1900s, that really was what drove equity values. Intelligence really is about understanding the physical world. And, you know, I think about a lot of what we're doing at Vantor as building a bridge between the physical and the digital world, right?

4:40So that you can sort of take these digitally native AI systems and have them understand the physical world by making it digital, right? By taking pictures of it, by integrating those pictures together, and then allowing those digital systems to start reasoning about the world.

4:57Grant:Does this look like a giant map of the whole world at some level? Do you actually map the entire world with what you do? Yeah, so that's what we do at Vantor. That's sort of the core bread and butter of the business. You know, the company was really founded as Digital Globe, was really focused on this idea of how do you take the physical world, use satellites to make a digital version of it. And at the time the company was really started and went public, the consumer of that Digital Globe was really a person, right? You could think of looking at the globe on your iPhone or looking at the globe through a screen.

5:30And, you know, the exciting thing is we've sort of gone through the transition from Digital Globe to Maxar to Vantor is really centering the role of a machine in the consumption of that Digital Globe, right? So really thinking about how do humans and machines look at the same digital representation, the same digital replica of the world, so that they can be operating and acting and communicating about space, right, about space and time. And I always think about, you know, language and maps. Those are early, very primitive human technologies. And in a lot of cases, maps predated language. Like when you look at cave drawings, there are older maps than structured language.

6:12And I think they're a really important part of technology and a really important part of multi-agent communication, whether that's multiple human agents, multiple digital agents or human computer systems working together.

6:25Grant:Yeah. OK, there's a lot of there's a lot of really cool stuff there. So the first one is like making making the digital maps that you've created actually machine readable, right, by any type of agent in the system. So I'd love to talk more about that. Then the other part of it is, I mean, fundamentally, is spatial intelligence perhaps more important for an AGI-like system? For people who listen to The Neuron, they know what AGI is. But just for people, if this is the first episode you've ever seen, AGI is an artificial general intelligence. So an AI that can reason and generalize across any domain, is spatial intelligence actually central to that if maps became before language to a certain extent?

7:02Grant:Love your take on both of those. Yeah. Well, maybe to start with the first one. You know, I'm sure a lot of your listeners are familiar with embedding models within the text domain and the criticality of sort of how that allowed language models to combine with corpuses of information and really do these sort of retrieval augmented generation workflows and really make the data that existed searchable by the models. One of the big ways in which we've been focused on making that digital globe accessible to AI models is by using embedding systems that embed. If you think about the way our system works today, we take images, we transmit them down in raw, wideband radio frequencies to the ground statement.

7:44We materialize them into sort of a big gigabyte scale file. And then you can open that on a computer and you can look at it. And so that's what we call pixels, right? Sort of that raw representation. If you're using Apple Maps or Google Maps, you know, you're not always looking at the pixels. You're often looking at a map, right? And, you know, I think this is the place where GIS technology and digital technology has really maybe over rotated towards a, quote, true physically true representation of the world, as opposed to a historical map that communicates a lot of context and a lot of substance by warping the way the world actually is.

8:23So if you my favorite map is the New York subway map, right? And the New York subway map Manhattan is the size of Brooklyn, right? Like, obviously, Manhattan is much smaller than Brooklyn from the perspective of how many square kilometers is Manhattan. But from a map making perspective, the process of cartography always used to be not just how do we draw a perfect map of the territory, but how does that map contain subjective context about the purpose, right? Different maps have different purposes. And I think as we've moved to the world of digitization, we've had this sort of one size fits all amazing Google Maps of the world.

8:59But, you know, the historical maps of the world, you know, you had mining maps, you had population density maps, you had maps that showed railroads, you had a lot of different maps that people would switch behind between with different contexts. And I think what we're doing with these embedding models is really starting to, you know, take the raw representation of the Earth in terms of an image and actually project that into hyperdimensional embedding space. You know, the kind of space that you and I, if we look out our window, we say that's a tree, that's a parking lot, that's a car, right? What we're doing there is we're taking a highly dimensional vector of car-like entity and we're matching that to a word, car.

9:37And now we're saying that state space of vectors that are car-like are car, and that state space is forest. And I think we're basically making it so the human language that defines these sort of types of entities, features on a map is computer readable, right? And so now you can sort of translate between the raw pixels from an image, from a sensor, into this sort of almost Rosetta Stone that both people and humans can interact with machines on. So the machine basically knows exactly what the human sees when the human's looking at it because everything has been essentially written in these embeddings.

10:11Exactly. And it's sort of the exact same way that we don't speak embeddings natively. Like if you gave me a hyperdimensional vector and asked me what it meant, I wouldn't really know. But I know what a king is. I know what a queen is. I understand what, you know, a mother and a father are. And so I have the sense, and by sort of mapping between text and embeddings and mapping between images and embeddings, you basically create a translation layer from images to text. And that really acts as a bridge for, you know, taking next generation reasoning models that are fluent in text and allowing them to use that intuition that they've been trained on to apply it to things that are happening in the real world.

10:51That's awesome.

10:52Grant:And have you, so how much of this have you done? Have you done the whole world with this? Yeah, so it's sort of funny. I mean, the word model has so many different meanings. You have a business model, you have a language model, you have all sorts of models, right? Which are meant to represent some sort of compressed version of reality. And, you know, in the geospatial world for 50 years, we've had digital terrain models, digital surface models, digital elevation models, right? This is what I think of as like model as a noun. It's sort of like a toy car model, right? Where you could put it on your desk or a globe, you could put it on your desk.

11:30And then we have these other models that are more like verb models. They take data in and they output data. And so you could think of them as like a noun, raw data in, verb, the sort of world model, output, a new processed form of data. And so what we've been focused on at Vantor, you know, since the very beginning of the company, but especially since we acquired Vrykon in the late 2010s, is taking raw images and producing a super high fidelity 3D model of the world, right? So that you can really take 2D images and really project them into 3D so that you actually, you know, are really representing the world as it is, as opposed to just pictures of the world, right?

12:11And that's something that

12:12Grant:you could actually go inside. Like if you are a human looking at it, like, can you navigate inside this 3D landscape. Exactly. And we've built out the whole world. We have over 100 million square kilometers, you know, at, you know, three meters spherical accuracy with 50 centimeter resolution across the whole world, you know, really, really exquisite tech, you know, they're doing things like how do the seismic, how does how do the waves, the movement of gravity affect the positional accuracy? How do movement of tectonic plates affect that? How does that affect things over a decade. It's really exquisite work.

12:44But the whole purpose is exactly what you're describing so that you could really have a real virtual environment that mirrors the physical environment. And I would distinguish this a little bit from other classes of world models, which are really building virtual environments that are intentionally hallucinogenic, right? They're intentionally not the real world.

13:04Grant:They sometimes feel like a hallucination when you're in there, especially the early ones. It's all moving around and it's like the walls are melting. I think it's like you're on acid or something. That's exactly right. And so, you know, what I think about and world models, you know, they're probably five years, 10 years behind language models. And what that means is it's not just that the models are behind or the technology is behind. It means that our words and language and terminology for describing the sort of architecture of the system is very early, right? It's very early in the way we describe these things.

13:35And so, you know, when I think about world models, we have a physical world model of the physical world. And I think about that in a lot of use cases, especially around, you know, dual use defense and civilian use cases, especially around simulation as almost grounding the world models in the physical world. So that you need that. You could start animating weather, you could start animating, you know, courses of action, you could start moving equipment, and you could have a world model that's generating a world that doesn't exist. But all of the buildings exist, all of the roads are true, all of the sort of fundamental characteristics of the world are accurate.

14:14You know, there have been some great work by some teams at sort of taking mid-resolution data and using a world model to make it, you know, one centimeter accuracy from 10 centimeters accuracy. And that's really just using the imagination of the world model the way a human would use their imagination. If you ask them to, you know, draw the building, they could see that there's a window. They could maybe see some panes of glass. you could draw that window at much higher accuracy, maybe, just by imagining in your knowledge of the world, that's the sort of like CSI enhance button, right? And I think you have to be very careful with these things, because sometimes the synthetic data is accurate, sometimes it's not.

14:56But I really view our role in this whole process as grounding those world models in the real physical world, which isn't the only use case for world models, but I think is going to be the use case that really makes them quite valuable.

15:11Grant:Would it be accurate to call it a ground truth world model, like what you've built? Yeah, I think that's a great, great word for it. Yeah, there was I saw an excellent video on this recently, and I can I can pull it up and perhaps we can show a clip from it. But essentially, it was a guy talking about how all of the current AI tools for video generation are kind of going into this node based direction where it's like node builders. and actually what it needs to be is you should really you really need a viewfinder like as if you're building a video game in order to do any type of video generation but in order to do that you have to give it essentially a ground truth model that it can generate from and you know in the case of you know you're doing something like fantastical or in video generation you're building that yourself right perhaps in like a 3d space and there's some companies that do that but it sounds like if you want to do anything close to like that in the real world this is what you would use.

16:05Grant:You would use a ground truth model like what you've built, and then you could generate, you know, tons of simulations based on that. Exactly. And I really think of it, I think simulation is the right word or scenario is the right word. But I see, I think it's important as a company, you want to know what are we doing and what are other people doing? Like, how are we framing the world for ourselves? And I really view our responsibility as we build that master branch, right? As the world changes, we update that ground truth world model, and we just keep it going. We've been doing it for 20 years.

16:36We've got, you know, 250 petabytes of information about the world changing over time. And we're just going to keep doing that day in, day out, like 6.8 million square kilometers of new data coming in every day. And we'll keep updating the world as it changes. And this is the thing that I think is so magical about what Apple Maps and Google Maps have done is they've built these products that I'll tell a story that, you know, I always think about because it was a crazy experience for me. But I was switching moving apartments and, you know, it was a new apartment building and been built a little bit north of Denver.

17:10And I drove over there and it wasn't on Google Maps. And so I looked because it was new. It had been built, you know, in the last year, last two years. And I looked out the window of my car. I was parking and I saw a building and I looked at my phone and I didn't see a building. And in my brain, I wondered, like, is there a real building there? Like, I trust Google Maps so much that I actually doubted my own eyes. You know, I was like, oh man, if it's not here, it can't be there. And of course that's crazy. And it didn't, you know, I didn't, that was like a couple seconds of worry. But I think there is just this thing, which is we live in such a tiny part of the world, right?

17:47Like when we drive to and from work, when we walk to pick up our kids from school, like we're living in miles, square kilometers, maybe a dozen square kilometers. And so we can really underestimate how much the whole world is changing because most of the change is far away from us. And I think it's underrated how much it's changing and especially how much it's changing now, right? I think one of the things we've been having a lot of conversations with customers across the world is the rate of construction of these AI data centers, of these energy production systems. Those are physical changes in the world that were not happening at that rate 10 years ago, even two years ago.

18:28And so I think, anyway, long-winded way of saying, I totally agree. I think that we're really building that world, trying to keep it as up-to-date as possible, trying to make sure it's the best reflection of ground truth so that you can really bring other models to it that can imagine future scenarios, that can do simulations, that can really let you imagine, like, what would my city look like if we changed our zoning laws? What would my city look like if we deleted elevator requirements, right? Like, you could have a world model sort of hallucinate new types of structures, new types of ways. What if we had all the buildings made of brick, right?

19:00Those are really, I think, valuable use cases that would really let people imagine a future that's really anchored on the present. It's not a total sci-fi hallucination, but allow people to think both in terms of, you know, renovating their own house, building new houses, building new, you know, schools and new systems of infrastructure. I think those are places where world models that are grounded in this ground truth replica, I think can be really, really powerful.

19:27Grant:So what if someone wanted to use this? Because you mentioned that you mentioned, like, potentially even mapping your own, your own, like space around you can can someone add their own images into your system and you can embed it like what what like how how how flexible is it is it just you're updating it every day can people almost like fine tune this uh do you get where i'm going with that yeah i mean this has been really um sort of the key unlock as we've you know i've come on i've been here about 18 months historically our digital globe was really produced by only our sensors and a lot of the work we've been doing on that production system, which we call Forge, which takes raw data from our sensors and produces this 3D world, is starting to make it so that you can sort of pour data from other places, whether it's from an iPhone, an Android phone, an aerial system, you know, really pour it into the production system.

20:21So we think of it as really putting every pixel in its place, so that you're starting to get this, you know, patchwork ground truth data based on whatever data you're feeding in. So we're providing almost a scaffolding, but then all the other data can stitch into it and fit nicely. And again, be very fully contiguous across the whole globe. And as data is coming in, you can sort of start clicking it together so that you have a couple high resolution insets here and a couple mid resolution insets here. And our global layer is sort of providing the accuracy and the grounding, which we have across the whole world.

20:56But yeah, I think that's been a big sort of focus of mine is how do we externalize this and really make this more of a platform that more data can feed into that lets you get sort of the best of both worlds. Because, you know, to your point, you're going to have much higher resolution data that you could collect from the ground than you can collect from space. But what space can do is it can, you know, continuously collect that at a regular cadence with a lot of SLAs. And so I think a lot of our strategy with this Forge piece of the platform has been, how do we bring more sources of raw data into that global grounded world model.

Read the full transcript

21:34Grant:Yeah, I like that. Yeah, let's actually unpack the Tensor Globe at a high level. There's Cortex, Forge, and Nexus. We kind of covered Forge. Can you explain Cortex and Nexus for us? Yeah, for sure. So I think, you know, a great platform always emerges from your own operations. And I like to think about, you know, whether it's the iOS App Store and Apple writing those first couple mail apps and you know, working with Google on YouTube and Google Maps, you know, those initial apps are the thing that becomes the platform. I think the example for us that's a little bit more close to home is the example of AWS, where they didn't set out to build AWS at Amazon, they set out to build amazon.com.

22:15And to build an e commerce system in the early 2000s and late 90s, you know, you had to build storage, you had to build commute, you had to build a whole networking layer, you had to build routing systems and DNS entries and all this infrastructure that they built to power amazon.com. And the really inspired thing they did is then they productized that platform, right? But, you know, I think when I think about Tensor Globe, the platform we're launching and launched last year, it's really productizing the core operations of Vantor. So what we do every day with a team of, you know, 2000 people across our whole operation center doing constellation scheduling, production, quality assurance, distribution, exploitation, the whole sort of what we call in the industry, like the TC-PED intelligence cycle of tasking, collection, production, exploitation, dissemination.

23:09That whole cycle is really what we've done inside of Vantor for 20 years. And what we're doing with Tensor Globe is sort of productizing it so that other organizations can do it. So, you know, with Cortex, that's really a system that's designed for constellation management. managing our constellation, managing our constellation, working with other constellations, and trying to understand if you have a goal, like you want to monitor 150 mines for activity or 250 electricity plants for activity, right? We can schedule that constellation across a bunch of different individual sensors and individual satellites.

23:44So that's really the cortex piece is that constellation management element. Forge we talked about is really that fusion element. So taking data from lots of sources and producing a 3D globe. Nexus is really that API and UI layer for looking at the globe, for accessing the globe, for accessing the raw materials and sort of the evidence behind or, you know, the raw data that is being used to produce that 3D world. And then over the next couple months, we'll be launching another component of that platform, which is really focused on looking at the globe, right? Understanding what's happening in the world.

24:18And this, I think, is something we've been working really closely with the Google team on and working with a lot of their foundation models and agent development kit to really apply reasoning to that globe. You can almost think of it as, you know, that that real time strategy look at the world where you're trying to understand what are the patterns happening across very different parts of the world that are actually connected, whether that's a supply chain or, you know, some sort of order of battle workflow in a defense context. yeah you know that that's really that that fourth piece is is providing understanding

24:52Grant:i'm sure you've seen some of the demos on x like of people trying to do this without your level of ability but like where they'll have like these like booty maps of the globe and it's like here's the map of all the different conflicts happening like currently and and it's it's really cool to see that but it'd be even cooler to see it with your level of sophistication and data yeah i think that's right. And I think I think those are amazing. I've loved looking at them. I, you know, we work with some of the people who've been working on them. And it's been awesome just seeing the creativity of, yeah, you can do with these tools.

25:25I think, you know, one of the things that is is true is that, you know, we've really standardized for the last 25 years on a zoomable map as the sort of fundamental substrate that we think of when we think about digital mapping. And, you know, the thing that's very odd about AI is that it can look at the whole world. And we can't. Like, we have to zoom in and zoom out and zoom in and zoom out. And I think a lot of the exciting opportunities, you know, is really applying AI at that very granular level and then rolling up some of these, you know, observations or alerts to a global view, sort of men in black style or, you know, like situation room style.

26:06but the thing from an uh you know when you zoom out in the world it's very low resolution obviously that's sort of how it works is you're compressing the data as you zoom out so that you just see africa and you just see you know the americas you're not seeing every single one of the hundreds of millions of pixels that makes that up and i think you know as a human we just

26:26Grant:can't do that it's like we even have a compute compute cluster strong enough to do that at the resolution we would want? I mean, I think this is like with many of these agentic systems, product development and technology development are different. And I think a lot of the art of product development is using the technology, seeing where the technology is going, and then sort of cheating to build something that actually works today and is actually useful and tricks the user a little bit. There's a lot of art as you zoom out today in the way that the pixels decompress and recompress and the way the vectors sort of do line simplification.

27:04There's tons of technology there that makes it feel like all the data is down there. Yeah. In terms of compute clusters, this really is part of, you know, why we were so excited to partner with Google is you really, I think, can't overemphasize how much money these companies are spending on compute. And from our perspective, you know, we own a constellation of satellites that works in the public cloud environment. And being able to run AI in a public cloud as opposed to in a classified environment, that's just a real change over the last 10 years in terms of, you know, the amount of compute buildup is, you know, orders of magnitude, roughly the same size is the U.S.

27:50defense budget today. Like the careback expenditure from the top five, 10 hyperscalers is probably 70, 80 % of total U.S. defense spending. And so, you know, one of the things we're excited about from a corporate perspective is just, you know, our sensors, our satellites that we've built and put on orbit over the past 20 years, you know, those are directly, you know, we can plug them right into a Google TPU and running next generation frontier models on the really high resolution data. And so, you know, to your earlier question, like, do we have a compute cluster big enough? Who knows? But I think we definitely have very big compute clusters that are capable of processing really, really high resolution data.

28:31And I think that just is a huge unlock in terms of, you know, what can be done. For so long, you know, I've been in this industry for a while. We've been so focused on computer vision algorithms. And I think that so much of the like exciting shift to spatial intelligence is really, you know, what happens after you identify all the vehicles in a parking lot, right? That's always been the task that felt tractable. And so I think as a technologist, we've put a lot of time into trying to solve that problem, because it was like a well specified task that was close enough to the current capabilities that you could actually go after it.

29:07It was like a little bit crazy to imagine building a talking computer that would look at the pictures and tell you what was happening. And that's not crazy at all anymore, right? Like that is actually totally in scope, but it's a really different ambition than I think what the industry has been focused on for 15 years, which is really much more image segmentation, sort of traditional computer vision algorithms versus these more advanced sort of reasoning algorithms.

29:34Grant:So going back to your earlier point about, well, there's a, there's the cool thing that you're building now where you can use the the i guess the layers of all this information but then to your earlier point about the actual layers of the information you know you can do it seems like you can map a lot more than just the number of cars that are in a built like you can you could almost get to the point where you're mapping like the chemical makeup of everything if you had sensors strong enough where you could say like this car is made yeah you don't want hyperspectral sensors for that yeah which is in our specialty but but yeah i mean at some point though right you could if you're if you're able to embed all this information you can just keep stacking layers and layers of information to the point where you can almost make a realistic 100 complete um yeah there's a famous borges uh short story about sort of these map makers who tried to make the most accurate map in the world eventually i think it's called inexactitude in science but eventually they made a map that was perfect but obviously it was the whole size of the world.

30:35So it was leather that wrapped around the whole world and people were wandering around seeing like pieces of the map falling apart. The nice thing about the digital world is you can truly map the whole world at incredible fidelity and incredible precision to your point. In terms of the use cases, yeah, we really think about a lot of what we're doing on the insights and AI side is, you know, developing a language and a grammar for spatial data. So if you think about, you know, image stack of pictures of the same location over time, you know, you can almost think of those as words that are then forming a sentence, right?

31:07And maybe you have another stack over here, that's another sentence. And so you're creating paragraphs. And I think a lot of what we're focused on are sort of what are those higher level abstractions that we're trying to build towards, obviously mediated by language from a human perspective. But I think for a lot of machines, these are going to be, you know, subspaces within embedding space that have some semantic content.

31:30Grant:There's a great paper from Sandia. This is an old paper, like 2014 or something about like geospatial semantic graphs around like, well, if you see a track next to a building next to a parking lot, that's probably a high school, right? And like, that's a pattern that is like high school, even though it's made up of four or five sub patterns. And I think as you sort of think about mapping the world, going back to that New York subway map analogy, like I'd love for us to be able to have, you know, the dynamism of a zoomable map with the semantic relevance of a lot of these features that, you know, pop.

32:12And I think a lot of the experience that Google and Apple and a lot of these zoomable maps have really nailed is like, as you zoom out, California replaces San Francisco and, you know, the labels change. And that's all based on, you know, hard coded relevance and semantic systems underpinning it like a knowledge graph effectively. But I think, you know, as we think about this vocabulary or grammar or language for events happening over time, you know, you have these ribbons that are fixed, you have ribbons that are moving around, you know, the world is made of entities that move around it. And I think a lot of the language that we're going to develop from a spatial intelligence perspective is going to be, you know, anchored on describing those patterns and formalizing them so that, you know, they become things that both computers and humans can sort of reason about objectively.

33:02Grant:Okay, I'm going to, I'm going to try and, and I'm going to posit an explanation here and I'm going to see if I got this right. So if you do that well, and based on where you're at today, that look like equivalent to zooming in and zooming out, like we think of it on the Google Maps example. But as you zoom out into like a 3D space, it starts to see, it's able to describe things like the San Francisco, California equivalent. But like, let's say you're really zoomed in in this 3D version on like the tire of a car. And it's like tire. And then you zoom out and it's car. And you zoom out and then it's street.

33:35Grant:And then it's zoomed out. And then to the point where you could then, no matter where you are in that level of zoom, you can kind of make predictions and explain what's at any layer of that stack? Is that more or less what you're building towards? Is that what you have today? Is that accurate? Yeah, I think that's right. That would be almost like a semantic hierarchy. And I think one of the things that's always hard when you're designing a taxonomy or an ontology is when is something a difference in degree and when is something a difference in kind? Like when is it a lion and a tiger versus a cat, right?

34:07And we build these taxonomies that help us understand the world, help us communicate about the distinctions. And I think one of the things that's really critical as we think about, you know, the new words and new terminology we're going to need in this future, right? The word molecule didn't exist. The word chemical didn't exist. Like people invented words as science and technology advanced, right? Computer used to be a job, not a machine. It was someone who tabulated numbers. And I think as we think about the next generation of sort of words, we are going to need words for these sort of like spatial abstractions.

34:41I think that's exactly right. In terms of like, maybe what you were getting at, in terms of predictive, you know, this is something that's that's pretty new for us. We haven't done any, you know, deep research on it or deep development on it. But, you know, as we think about, you know, image generation, I think there's a real opportunity with world models to generate synthetic imagery and say, okay, we're planning to take a picture of this location in three hours, let's generate the image that we think we're going to take, right? And then let's diff that image against the image we actually get.

35:14That's really when you think about learning, it's sort of making a prediction and then falsifying it with an observation. That's sort of the classic scientific method is, hey, we think this is going to happen. We're going to run an observation or an experiment, and then we're going to validate that. And I think a lot of what we're going to start seeing from an intelligence tradecraft perspective is a ton of in silico image generation, almost running a synthetic war game where we're generating all the images that we anticipate over the next week. And then we're diffing them against various courses of action to understand as those observations come in, you know, are they corroborating our hypothesis or are they rejecting the hypothesis?

35:53And, you know, we can even use that for benchmarking your models too, right? Like you could use it

35:57Grant:for benchmarking your prediction models. You could use it for benchmarking even your image models where you're saying like, we think that, you know, know, if this world model is accurate and capturing the physics, we think it'll look like this when we go out and take this picture of like, you know, a cannon firing at a, you know, you know, a tree or something. And then you go out, you take the picture, and then you see how accurately it got, you know, the image models. Exactly. And I think, you know, as humans, we can launder information from high resolution into low resolution, we can launder information from low resolution into high resolution, like, we're built with a world model built into our brains, right?

36:30Like, I'm not a neuroscientist, I don't know where it is. But, you know, fundamentally, like, from a very young age, humans are developing a sense of physics, they're developing a sense of causality, they're developing a sense of sort of how things occur, and some basic intuition around like, if this, then that in a physical world, senses of gravity, of sort of permanence of objects, right. And when you think about the projection of sort of in silico versions, or synthetic world models, You know, you can imagine them being able to do very similar things, right? If I take a picture of a car on one day and it's a red car at really high resolution and I know it's license plate and then I take a low resolution picture of a red car, I can infer based on my understanding of patterns, my understanding of what's happening in the world, how things occur, that they might be the same.

37:19Now, I'm not going to set that at 100 % the same, right? I'm going to apply an Asian discount to that. I think that's sort of the thing that's really important about anchoring all of this on a sort of unitary digital globe is ultimately the world is happening, whether you observe it or not. And a lot of the work of intelligence analysis, a lot of the work of spatial reasoning is trying to figure out what's happening when you're not taking a picture, right? Like what's happening between the frames? And the idea of being able to generate synthetic frames, have analysts produce key frames in the middle and say, well, based on this and this, I think something happened here.

37:56Generate me a video of if this happened here, what was the full video of the world? Right. I think that's really interesting, sort of similar to a lot of the generative video techniques, similar to a lot of generative images techniques. But that's really useful for like forensic analysis, right?

38:11Grant:Like if you're like, like, you know, you're in a legal case, you could use this model and say like, OK, this is what we think happened in this in this instance based on the evidence. Exactly. I think, you know, digital forensics is, I think, a really interesting use case for this where, you know, you collect a ton of evidence. You have dozens of video files, camera files, people who saw pieces of the event, people who saw, you know, who got a, you know, body camera view of only a little bit. But there's a really cool, ProPublica did a great story where they took all of the video files off of Parler during January 6th.

38:44And they sort of arranged them spatially and temporally so that rather than, you know, in a traditional video file, it starts at zero and it plays forward to 30 seconds or 45 seconds. What they did is they took all those and they aligned it with like 12.14 p.m., 12.19 p.m., right? So it played from 12.14 to 12.14.30 and 12.19 to 12.19.30. So they arrange them in time, they arrange them a little bit in space. But, you know, I think a lot of our vision for what we're building is a world where you never look at a video again. You're only looking at a map. And the video is on the map, but you can zoom out and you can zoom in.

39:21And you're really thinking of the video more as a viewport of the map versus an actual data file that you're like clicking on and looking at.

39:29Grant:Yeah, love that. So that's so many branching ideas here to talk about. But I want to go back to one that I was thinking of earlier when we were talking about the idea of compressing and how humans, you know, we have sort of like this idea of like we can abstract it on a high level and at a low level and we have a predictive model built inside of us. Could you ever compress the world model that you've built or any of the versions of your world model to the point where it would be something that a robot would run on board? Maybe it's not necessary for it to do that, but like, could this ever be something that's fundamental in a AGI like system, for example?

40:06Yeah. So as part of we launched a product last year called Raptor, which is really designed to be a GPS resilient positioning technology. So, you know, the way GPS works right now is satellites, you know, essentially have atomic clocks on them that are highly accurate. just say this is the time this is the time this is the time you receive times from multiple clocks and you sort of do the trigonometry to figure out where you are in the modern battle space people are spoofing those all the time or jamming them and so the gps can go down often does go down i think if you're in an urban canyon in a sort of normal commercial environment i'm sure you've experienced that thing where your gps is like i'm not quite sure where i am i'm like on this road but i'm actually a road over not the end of the world the way it would be in a battle situation you know annoying and certainly bad raptor is really a product we designed to you know navigate the way humans navigate which is looking around and seeing the 3d world and understanding okay based on my knowledge of london and my observations through my glasses of london this is where i think i am and so as part of that we do compress that world model down we do what's called chipping and shipping where we cut out piece of model at high resolution other pieces of the model at low resolution and then we give that to the device the robot or the machine so that they essentially have a map that they can look at and compare what their camera is seeing in 3d to what our world map sees in 3d and that gets you positional accuracy very close to what you get with gps and with the added advantage of you know where you're looking so with the gps system you know where you are but you don't know where you're looking and yeah and that's a huge deal exactly so i think that's a huge, you know, form of compression that we use today.

41:51And I think a lot of the challenge of this, again, going back to sort of what makes it hard to build these systems, you know, there's obviously an AI model development element, which is not my personal area of expertise. But there's a huge sort of data orchestration element of moving data, updating data, making sure you have the right clocks, the right update semantics, sort of this multi master version control system of, you know, how do you take all the data that's coming in and updating the representation of the world and then make sure everyone understands what version of the world are they looking at?

42:24And, you know, you can really think of in the same way that a Git repository or a version control system would do that for a piece of code. You know, our globe would do that for these physical systems, which are both looking at the world and observing it, but then also ideally working together so that they can tip and queue and sort of operate off of a shared digital replica. That's awesome.

42:44Grant:Would it ever be possible to give that kind of level of feature control to like an iPhone app or something? Totally. So, I mean, this is we, you know, a lot of people, you know, at Vantor have worked for a long time with some of the folks at Niantic. Last year, we had a partnership with them really specifically on, you know, using their augmented reality system against our globe so that you could do augmented reality globally. I think, you know, one of the things about GPS, it's a global positioning system. And so while it's possible to do sort of inset mapping or mapping of particular urban areas, if you want to go off road, literally, you need a global map, a global three.

43:27You need a local positioning system, basically. Exactly. We call it a terrain positioning system, a TPS. Yeah, yeah. That's exactly right. And so taking their VPS and making a global, you know, requires them to look at a global representation of the world. And that may not be particularly, you know, important in certain consumer enterprise use cases. You know, if you're Walt Disney World and you're trying to build awesome augmented reality, you know, experiences for Walt Disney World, you don't need that app to work across the whole world. But, you know, if you're a company trying to build AR experiences that do work globally, you know, I think that's really the genesis of that partnership was trying to bring some of their awesome technology.

44:07We've built technology that helps positioning in the air, but they really are the experts on the ground. And so we've been working with them on sort of connecting those two pieces. That's awesome.

44:17Grant:I want to talk a little bit here before we wrap up about like how businesses can use this. But before we do, it sounds like you have essentially you've been focused on trying to make all of this information and data that this world model that you've created, the ground truth world model, as we've called it, accessible to like today's language models and reasoning models. But what about the vision models and the world models? Like, like, will there will there ever be a version of this data where, like, let's say, like an agent system would actually be going inside of the 3D representations that you've created?

44:49Grant:And how do you think about building for that use case as well? Yeah, I think that's definitely right. There's a bunch of simulation workflows along those lines, you know, whether it's trying to train AI pilots or, you know, robotic systems to maneuver in the digital world. You know, if you if you talk to people at the at the self-driving company, most of the miles driven are synthetic miles, right? They're not actual miles driven on the highway just because the number of surprising events happening in the real world, you know, you can't get enough of a corpus. So you want to create synthetic environments that really push you from a development perspective.

45:23So, yeah, I mean, I think that bringing, you know, let's say models that are trying to understand the world into that digital world, I think is really critical. We currently do most of our perception layer was what we call it is really that sort of segmentation, computer vision. That perception layer today is mostly done in 2D space. A lot of the work we've been doing over the last year is being able to take a single image or a single stereo pair image from our satellite and create an on the fly 3D image. And, you know, have been doing a little bit of work now on the on the AI perception side to see, you know, how much better does the perception model get when you give it a 3D image versus when you give it a 2D image?

46:03And I think that's a particularly interesting area of research where, you know, so much of the perception technology stack has been about perceiving 2D images, segmenting 2D images. But obviously with the third dimension, you're getting a ton more information. You know, certainly when I think about analyzing 2D images, a lot of what I'm doing is if I'm metacognating and really thinking about like, how do I interpret a 2D scene? A lot of it is by constructing a 3D scene under the hood, right? Totally.

46:33Grant:Yeah, I think all living intelligence like lives in a 3D world. Right. So like, I think that that at the end of the day has to be a component of any sort of agentic AGI system is that it has to be able to think in 3D space at a certain point. I think it's like it's it's always tempting to anthropomorphize things like but we do. Yeah. Prior to three years ago, there was one entity in the world that spoke natural language. Right. Like maybe parrots spoke a little bit of it. Like there was one, maybe 10 neural architectures that had like some degree of of language capabilities. And there's a new one.

47:07I think that's very exciting. It's likely there are going to be similarities. It's likely they're going to be big differences. But I think from a sort of pure information theory perspective, providing a perception model with a 3D representation, I think is likely to provide much better like f1 outputs in terms of extraction of vectors

47:25Grant:and semantic understanding of the scene yeah and then I think taking that a step further there's a lot of people who think that for an AGI like system to actually be true AGI has to be an agent that takes actions and reacts that can make decisions and that sort of things then then you know you would you would probably want to train that type of system in a 3d environment yeah yeah Yeah, I mean, I think that's right. I will say, you know, as you zoom out, the world becomes more two-dimensional, right? Like if you take off on a plane, like you look down, it's pretty 2D. You're getting ready to land.

47:57It becomes much more 3D. And so I think it also depends on sort of the vantage point you're operating in. There are a lot of domains in which we, as humans, operate intelligently in effectively a 2D space, right? Whether it's certain types of surgical routines, certain types of workflows, there are certain types of technology that have low-dimensional inputs and low-dimensional outputs that we operate intelligently and interact with in a sophisticated manner. I think our computer has a sort of one-dimensional interface in terms of the keyboard, and we're able to do quite a lot with that one-dimensional system.

48:36That's fair, yeah. I don't know that I buy that, like, this, that you have to be operating in 3D space to be intelligent. But I think that there's a class of very important actions in the world that happen physically that are going to be mediated robotically to some degree over the next five, 10 years. And those actions do require physical intuition. And I suspect that those actions are some of the most important actions that end up getting, getting, getting taken in the world. I just think, for example, like, you know, stock trading, you don't you don't need to walk around Wall Street to do stock trading.

49:09Grant:Yeah, that's true. That's a good that's a good counterpoint. I like that. Well, I do want to touch on before we wrap up here, how can people use the tools that you're created? How can they work with you? What's we talked about a couple of the tools in Tensor Globe? But yeah, just walk us through like how how someone could work with you if they wanted to. Yeah, I mean, I think that the best way, especially for folks in this audience, probably is really, you know, getting access to some of the data as a service products, 3D vector products, really sort of getting access to the mapping data. We host that in a cloud system that's accessible through API and have a good licensing structure.

49:46We also have a lot of publicly available data through our open data program for a lot of sort of wildfire events, things that are of public note. And so I think that's a place where people can play around. We run a number of sort of machine learning exercises and have made pieces of our archive accessible to various university groups and other people who are sort of training models, working on particular tasks. I think we did a SAR model a couple of years ago. We've done some interesting image segmentation, prize opportunities. That's, I think, a natural way. But yeah, I mean, I'm on Twitter. I'm on LinkedIn.

50:21Feel free to reach out to me directly. I'm happy to put you in contact with the right people. Yeah, has been an awesome conversation. So feel free to reach out and I'll be sure to reply. Awesome.

50:33Grant:Thank you very much, Peter. Thanks for joining us. And thanks for walking us through all of this and what it actually takes to build that ground truth foundation model, like ground truth world model. It's really awesome what you're building. Thanks. Yeah, hopefully that language kicks in because I think it'll be, you know, it's always good to lock in the right terms when you're building technology. And I think right now the world model is just it's a little bit too big in terms of it could be it's sort of in the eye of the beholder. And it's also like it can be it can mean a perception model as well.

51:02Grant:Right. Like I think I think of it as like you're generating the world, but that could just be like my model of the world. So you're right. It's totally out there. Right. So we'll see. I think this is going to evolve a lot in the next year or two years. A lot of really good people working in this space. So it's definitely exciting, exciting place to be spending some time. I guess all of this conversation, you know, you brought this up off camera and it makes sense to me, reminds us of the metaverse conversation, right? And I think you have some interesting thoughts there if you want to touch on the metaverse.

51:29I mean, I think we've lived in this weird part of history where the dominant flow of technology has been sort of from the consumer space into the enterprise space, which is really historically anomalous, right? Like if you think about how did technology get developed in the 1800s, you know, a lot of technology was developed in the enterprise space and sort of percolated into the consumer space, right? um and and i think you know one of the key questions is sort of is spatial intelligence an enterprise technology or consumer technology and what's going to be the the diffusion arrow is it going to go from consumer to enterprise the way like an iphone did or a personal computer did or is it going to go sort of from enterprise to consumer the way you know a more traditional military technology did let's say like the internet moved from sort of military to consumer The GPS system moved from military to consumer.

52:22Airplanes moved from military to consumer. Like a traditional pattern of technology development historically has been that, you know, the defense industry has pushed the frontier. And then that technology has diffused, whether you look at the Apollo program or the Manhattan program. Military technologies have been then repurposed in the consumer world to, you know, drive productivity there. And I do think that augmented reality in particular is a technology that's likely to be developed in the military context, or at least in the sort of physical enterprise context, right? It's a very challenging technology to develop.

52:57It requires, you know, people who are really interacting with the world in complicated ways, doing sort of physically difficult tasks for it to really be useful. and I think you know it's really that shift from you know do people really want virtual reality or do they want augmented reality I think augmented reality yeah I mean there's use cases for virtual

53:17Grant:reality but I think we can leave virtual reality for the machines I think what we would like is to all of the information we've been talking about is like accessible to you like if I'm looking around my room you know you and I both wear glasses so we're comfortable with wearing glasses and I think people who aren't maybe are a little less comfortable but if I could just look around the room and see like have the ability to like focus in on something and it give me all the information about that object i think that would be really cool or being able to manipulate you know on my computer in a 3d space instead of a 2d space i think that'd be really cool i totally agree and i think you know ultimately i think one of the hardest things about building products is you have to sort of hold two things in your head at one time one the product is important and other the product is not important.

54:04And so if you're working on Netflix, you're coming into work every day. You're like, Netflix, Netflix, Netflix. It's 40, 60, 80 hours of your day. As a user, it's just Netflix. It's just Uber. It's just a product. And I think especially with augmented reality, which is this really persistent part of your life, I recently got a new oven. And the oven really wants me to use its app and make its app the center of my life. And the app has all these features. And it's like, you could get recipes, we could integrate, it wants to be the center of my life. And I think with augmented reality, it's almost even more dangerous, because it's, it's something where you almost have to be very subtle about the way in which you you create the user experience.

54:45Because, you know, as much as you as the developer, as the developer might want to sort of like, be like, Look, I know all this stuff about the world, you're looking through the glasses this is a dell monitor this is a macbook this is expo what it's like i as a user don't want it showing me all that like i think no that'd be way too busy yeah human computer interaction and i think a lot of this is going to be sort of what's the click effect what are the persistent parts of the augmented reality experience versus the transient parts and i think developing that in a defense in a sort of hardcore industrial setting you know that's where a thing like like Apple Vision Pro, you're like, the market for that is not people sitting on their couch watching Netflix, right?

55:26It's people building stuff, people trying to learn carpentry, people trying to like, fix their car, teach their kid how to play baseball, right? Like, it's people doing stuff in the real world. Like, that's the market. And I think when you look at the marketing of these technologies, it's truly dystopian. It's like people sitting on their couch alone, watching slop. And I just don't think that's the right way to position the technologies. I don't think it's the way the technologies are going to have a great impact on the world. And I also think it's a really depressing vision of the future.

55:55Grant:Well, I'll give you an alternate take, which is imagine AR replacing YouTube. So I was, for example, I was using a Rubik's Cube recently, and I was like watching YouTube videos to try to memorize like exactly how to do it. But imagine if you, instead of watching a YouTube video that you have to go back, click play, you can combine AI with an AR interface. And it's like, just shows you a little arrow, like turn it this way. Now I'll turn it that way. I'll turn it this way. And then I can be like, why am I doing this? And it could talk me through it. And that would be so much more intuitive. And I would never have to watch that video.

56:26Exactly. And I think it's sort of the video is that pixel representation. What you're describing is a vector representation, right? Where you're reprojecting those pixels into your pixel space. And whether it's learning Rubik's Cube or fixing your toilet or, you know, doing a carpentry project, like having lightweight cues. You know, I think about when you're putting together Ikea furniture, you're like constantly looking from pixels to furniture, from pixels to furniture. And, you know, I think it would be an incredible experience to have Ikea in my glasses saying that's, you know, object 12, put object 12 in object 11.

56:59Like, that's actually a really cool experience you could imagine. But it really, again, requires you to almost drop the pixels, use the pixels for the machine and produce vector cues for the end user in the augmented reality space that are sort of reprojecting the content from a YouTube video or from millions of YouTube Rubik's Cube videos into this idea that like Grant is holding a Rubik's Cube and that Rubik's Cube is, you know, rotatable in these ways. and I'm going to produce cues from his glasses that know how far his hands are from his face. Like, that's a really tough problem.

57:34Grant:And then there's a whole other layer as the product developer where you want to have a predictive model that knows exactly when Grant needs those cues, right? Where it's like, he's confused, he's looking at it, like, now I should add them in. I think that's right, and I think that's where, you know, it's going to be really hard. And I think a lot about the Alexa problem of, you know, it felt like I had to go to Hogwarts and learn all the spells of like Alexa, timer, Alexa. You know, I never knew what it could do, what it couldn't do. I think the affordances of a voice-based or vision-based system are really tricky compared to the affordances of either a physical or a screen-based system in terms of teaching the user how to use the technology.

58:15And to your point, the worst version of it is like, hey Grant, like, did you know you can set timers on Alexa? You're like, shut up, Alexa. Like, I don't want that in my room right now. And so I think that's going to be a really challenging Like I suspect for these technologies, there's going to be some sort of almost, you know, in a video game where you start and you learn how to jump and you learn how to there's like a training part of the video game. I think for both augmented reality and just voice technology generally, you need some sort of training environment where you as a human user are like learning the affordances of the system and it's constantly updating you as it changes.

58:49And I think that onboarding piece is really hard. And I actually think people haven't nailed it that well. And it's going to probably be one of the hardest parts of deploying this technology. Is that something you're building or something you want to partner with people to build? I think that would be something we'd probably partner with. You know, that's much more on the sort of like, I think of it as like the first person shooter view of the world versus the real time strategy view of the world. And a lot of what we're trying to do is join those. But I think a lot of that sort of like FPS experience is probably going to be more of like an app store experience.

59:21You're going to have you're going to want a proliferated set of developers. Maybe there's a Rubik's Cube app you install on your glasses. Maybe there's a kind of plugins like Co-Works really pushing plugins. Exactly. Yeah, because I think part of that is, you know, one of the powerful things about iPhone or Android is that, you know, based on the apps you install, you're taking an action to like choose the Marriott app and choose the Chase banking app and like you're sort of configuring it to your world and I think similarly it's going to make a lot of sense to be like install the Rubik's Cube app you're going to know that you did that you're going to know that that affordance is available versus just having like a Rubik's Cube app pop up out of nowhere right but I do think like you know maybe to to close this out a little bit I do think one of the things I'm most bullish about is like really using the physical world for app discovery.

1:00:10You know, you go to Chipotle, your glasses have like a little button to install the Chipotle app. You go to McDonald's, your glasses have a little button to install the McDonald's app. Like, I think over time, you know, these physical, the physical world is going to be a much more important part of digital product discovery. As the digital world gets more and more confusing, I think the physical world, you know, you already see this in sort of channel marketing, where people are marketing in the physical world in order to generate impressions in the digital world, right? The like New York subway takeovers.

1:00:40The purpose of that isn't necessarily 100 % to have people on the subway download the app. It's almost to like take pictures of people on the subway with the posters to post on the Instagram to generate, you know, some sort of salience for the users of Instagram to understand that you cared enough to buy out the subway. And so I think more and more like the exclusivity and rivalrousness of the physical world is going to be a key sort of part of, right? It's like sort of that real estate thing of like location, location, location. I think as the digital world gets less and less exclusive, there's more and more content.

1:01:17You're going to see more navigation experiences in the digital world mediated by the physical world, whether it's QR codes, which I hate, but I think are very effective. Other sort of forms of almost like digital waypoints that are positioned in the physical world.

1:01:33Grant:And a lot of that could be technology that is like agent to agent, right? Like it could even be that like the human isn't even the interface, isn't even the interaction layer there, but it's like it exists so that the agent can prompt it for the human. Exactly. Like I think a great example is if you think about the future of a fast food restaurant, right? You know, will you be ordering, you know, will your agent talk to their agent? How will it know which agent to talk to? You're like, well, probably know based on where you are in the world, which agent to talk to, right? You don't want to talk to McDonald's.

1:02:01You want to talk to the specific McDonald's store that you're in. And probably a lot of that routing is going to be physically mediated, even in information space. I love it. That's awesome.

1:02:11Grant:All right. I don't want to take more of your time, even though I could talk about this forever. Me too. All right. Thank you so much. Yeah, of course. Good to see you. Take care.

1:02:28You

From the publisher

AI can reason about text and images, but it still struggles to understand the physical world.


In this episode, Grant sits down with Peter Wilczynski, Chief Product Officer at Vantor (formerly Maxar Intelligence / Digital Globe), to unpack why spatial intelligence is emerging as critical AI infrastructure. Peter spent years at Palantir building ontology systems and mapping tools for defense operations before joining Vantor, where his team has built a 100M+ square kilometer 3D model of the entire Earth at 50cm resolution.


We dig into how satellite imagery becomes machine-readable through embedding models, why "ground truth world models" are fundamentally different from hallucinated ones, the Raptor GPS-alternative system, simulation and digital forensics, the future of augmented reality, and why the physical world might be the most important thing AI still doesn't understand.


Vantor: https://vantor.com


TensorGlobe Platform: https://vantor.com/product/platform/


Vantor rebrands from Maxar Intelligence (Business Wire): https://www.businesswire.com/news/home/20251001760322/en/Vantor-Rebrands-from-Maxar-Intelligence-Unveils-AI-Powered-Platform


Subscribe to The Neuron newsletter: https://theneuron.ai

More from The Neuron: AI Explained

All 106 episodes
This Company Mapped the Entire World in 3D. Here's Why.The Neuron: AI Explained · 1 h 3 min
Listen in VO