AI Exchanges: The Role of Data

30 Sep 2025 · 17 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown

Podcast Episode Summary

AI Exchanges: The Role of Data

Podcast Title: Exchanges Episode Title: AI Exchanges: The Role of Data Release Date: September 25, 2025 Hosts: Allison Nathan and George Lee Guest: Neema Raphael, Chief Data Officer and Head of Data Engineering at Goldman Sachs

Overview In this episode, Neema Raphael discusses the crucial role of data in the advancement of artificial intelligence (AI) and the challenges that data can present in the AI landscape. The conversation is framed within the context of the evolving AI technology, particularly focusing on generative AI.

Key Themes and Discussions

Introduction to AI's Evolution

  • Generative AI vs Traditional AI: Neema highlights the transition from deterministic computing (where specific rules dictate outcomes) to probabilistic models in AI that learn from examples rather than solely rules.
  • Impact of Data: Data is portrayed as a double-edged sword that can either enable or stall AI advancements. The discussion emphasizes the need for high-quality data to optimize AI functionalities.

Neema Raphael's Background

  • Career Journey: Neema recounts his 20-plus years at Goldman Sachs, emphasizing his transition from software engineering to data leadership, marked by the 2008 financial crisis which underscored the necessity of effective data management.
  • Copter Database: He references Copter, a database developed during the crisis that centralized data across various departments, showcasing how effective data management can foster innovation and efficiency.

Cultural Shift in Understanding AI

  • Mindset Change: Neema discusses the cultural shift required in organizations to embrace probabilistic AI models, especially in sectors like finance where deterministic models have historically dominated.
  • Understanding Non-Determinism: Finance professionals are somewhat accustomed to non-determinism due to the nature of market pricing, yet there remains a need for education about AI outputs.

The Hype Surrounding AI

  • Skepticism of New Technology: Neema expresses a cautious perspective on AI hype, acknowledging its potential while questioning its applicability in enterprise settings.
  • Consumer vs Enterprise Potentials: While generative AI shows promise in consumer applications, its enterprise potential is still being evaluated.

Data Availability and Quality

  • Current State of Data: Neema suggests that many organizations have already "run out of data" to some extent, meaning that they have not fully leveraged the vast amounts of internal data available.
  • Synthetic Data: The conversation also touches on the rise of synthetic data and its implications for AI model training.

Enhancing Business Value with Data

  • Data Engineering: Emphasizes the importance of data engineering to normalize and clean data for better AI integration.
  • Unlocking Proprietary Data: The episode discusses how Goldman Sachs is working to utilize its proprietary data more effectively to enhance decision-making processes.

Key Takeaways

  • Data Quality is Paramount: The effectiveness of AI models is heavily dependent on the quality of the underlying data.
  • Cultural Adaptation: Organizations must adapt their understanding of computing from deterministic to probabilistic frameworks.
  • Future Challenges: While generative AI has opened new avenues, the real challenge lies in harnessing enterprise data to differentiate and drive value.

Conclusion The episode wraps up with a consensus on the critical role that effective data management plays in unleashing the potential of AI technologies. Neema's insights highlight the importance of adapting to new paradigms and leveraging existing data assets to foster innovation in finance and beyond.

--- Disclaimer: The opinions expressed in this podcast are those of the speakers and do not necessarily reflect the views of Goldman Sachs. ```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Neema Raphael:Welcome to Goldman Sachs Exchanges. I'm Allison Nathan and I'm here with George Lee, the co-head of the Goldman Sachs Global Institute. Together we're hosting a series of episodes exploring the rise of AI and everything it could mean for companies, investors and economies.

0:22Neema Raphael:George, good to see you again.

0:24Allison Nathan:Great to see you, Allison. Good to be here.

0:25Neema Raphael:So, George, we've had several conversations about how AI is shaping the economic and business landscape. But today, we want to actually get a little bit more under the hood and talk about the technology itself, and in particular, the role that data will play in enabling or possibly stalling its progress. We have a terrific guest to dive into these issues with us, Nima Raphael, the Chief Data Officer and Head of Data Engineering here at Goldman Sachs. Nima, Welcome to AI Exchanges.

0:54George Lee:Thanks for having me. Excited to be here.

0:56Neema Raphael:So before we dig into the topic at hand, first, just tell us a little bit about how you got here, your career journey, and what your current role here at Goldman Sachs entails.

1:07George Lee:Yeah, this is 20 plus years for me at Goldman. I started right out of college, studied computer science. And during that, I sort of realized that I wanted to apply technology to a domain that I had not known before. And so really the finance was sort of like this black box to me. I came to Goldman and met amazing people, started as an analyst here, software engineering, you know, typing at the keyboard, writing code. And then really the data thing came, I'd say, five years in global financial crisis, 2008, Lehman Brothers collapses. And a group of technologists called CoreStrat at the time was going around the firm saying, hey, we have to figure out what our exposure to Lehman is.

1:54George Lee:We have to figure out our liquidity profile, see what's going on at Goldman. And the way they had structured it was to try to get all of the data from the front office, middle office, and back office together in one place to sort of figure out the end-to-end exposure to Lehman in a very technology and data-heavy way. and that was sort of like the genesis of I'd say my data journey here. That project actually was super interesting because we had heard other banks and other financial institutions actually have to go into their filing cabinets to dig out their ISDAs that were assigned with Lehman to figure out what their contracts were.

2:34George Lee:We luckily had a lot of our data sort of corralled in one place. And actually, that database we built, it was called Copter, ended up becoming this place, not only that people realized the power of data, not just being sort of like an exhaust, but actually an enabler for the business. And then not only did people recognize, okay, this database sort of saved the firm in some interesting way. But then when we gave that same data to traders, salespeople, strats on the desk, quants on the desks, people started coming up with new innovative ways to use that data for helping our clients and just running the firm a lot more efficiently.

3:19George Lee:And so it became this sort of launching pad for people outside of technology to say, hmm, data maybe can be a powerful concept here. That's great.

3:29Allison Nathan:And so in your role as chief data officer now. You oversee all of that. Other things you've done in your career have been involved with what we used to call in the dim, dark past machine learning. The point is, you know, AI has been around. We've used it at the firm. It's been, you know, broadly proliferated. But the rise of generative AI has garnered so much attention. Is it fundamentally different than the journey we've been on? Or is it just an extension of the continuum of good old fashioned AI?

3:56George Lee:Yeah, I think I'd say a little bit of both, because it feels like some sort of step change function from the historical, you know, I always talk about the first 50, 60 years of computer science sitting down and humans had to code rules to tell the computer what to do. And we talk about determinism, like the rules were deterministic, like if you push this button, please do this, or if you type these keys, please do that. And so there was this really fundamental shift, I guess in machine learning in general, which is like learn by example instead of learn by rules. And so in some ways, the generative AI stuff is just a continuation of learn by example.

4:35George Lee:But I don't think people naturally saw it go from, hey, I could learn maybe how to predict some patterns to now the computer could create anything. And so there's a little bit of that continuum, like, hey, if we just feed the machine more and more examples, more and more data, it could start learning things, is probably the path of continuum. But the sort of step change was like, oh, well, can we feed it and create images, create audio, create language? And so I think that's sort of the novel step change in the generative part.

5:10Allison Nathan:And you illustrated something I think is very fundamental in terms of company culture in this shift, which is we're used to deterministic computing. For a given input, the outputs are correct, repeatable, and traceable. We're no longer in that sphere. As you pointed out, these are probabilistic machines. Something emerges from it that you can't trace and is often right, but not always. Talk about the mindset difference inside an organization of getting business users in particular to be comfortable with that.

5:40George Lee:I'd say a little bit in finance, people maybe have understood that because of our pricing models and derivative pricing. I mean, it was always stochastic in that way anyways. So there was always a little bit of, okay, like the world is non-deterministic. And so prices are non-deterministic. The markets are non-deterministic. Economies are non-deterministic. So I think there was maybe a willingness to sort of understand that here in the finance world. But I agree. I think when non-engineers sit at a computer, they sort of want a thing to be repeatable pattern. That's how we build workflows here.

6:12George Lee:That's how we build client insights or anything we do here to help our clients. So I think it's really about teaching people this isn't just some magic crystal ball, right? What it's really doing is taking a lot of examples and giving you an extrapolation from those examples.

6:31Neema Raphael:Let me just have a follow-up to that though, because we've had a lot of conversations on this podcast about the ultimate potential of the technology. There's so much hype around it. We're having another, I think, leg up in the hype in the last month or two here. Given what you know about the technology, do you think it's overhyped or maybe even underhyped?

6:52George Lee:Yeah. So as George knows, I'm always a little bit of a skeptic of new technology. Historically, we've talked a lot about blockchain and things like that, and that was supposed to revolutionize. And is the next thing going to revolutionize? And look, I think from an AI perspective, it's obvious that it's real. It's here to stay. There is absolutely a hype to it. But also, when you go on your phone and you ask, Claude, Gemini, GPT, take a picture and you ask, like, what is this? Or you ask, give me some research on a topic I'm curious about. And you get great answers and you research more. It's definitely, definitely real in the sort of consumer world, I think.

7:32George Lee:I think where the hype, I don't know, I would say it's slightly differently than hype. I'd say the potential, I think, in the enterprise is still to be seen. I think there's some really slam dunk use cases we've seen, right? Like agent coding, for example, is the thing that sort of flipped my brain from this might be vaporware to like, wow, this is like really real. When I sat down at the computer and I was like coding with an agent and it was helping me with problems that I've never been able to solve before, I was like, wow, this is incredibly powerful as a superhuman ability, you know, like amplifying my abilities.

8:09George Lee:So I think there's definitely real there. I think from an enterprise perspective, the thing to be seen is where can people harness their data and their enterprise data and the proprietary data they have to make some differentiation in the enterprise space. That's the to be seen part.

8:26Neema Raphael:But we're only a couple of years into this newer generation of these models. Do you foresee a future where we actually do, though, run out of data? I mean, we're early here, but is that ahead?

8:37George Lee:I would frame it a different way. We've already run out of data. We've already run out of data. When you read about the new models, the undertone of what people say, and you've seen this in like the deep seek moment and things like that, is like everyone wonders how did they do that with less money, less money. And one of the big hypotheses is they trained against another model, And so it already incorporated the previous thing. I think the real interesting thing is going to be how previous models then shape what the next iteration of the world is going to look like in this way.

9:11Neema Raphael:So let me reframe my question, which is more that do you think this is going to restrain the potential of the technology?

9:20George Lee:No, I don't think so. the explosive nature of the synthetic data and the fact that now the computer could generate infinite amount of more data. Again, I think there'll be a sort of a cursor of what people call like AI slop versus maybe more insightful data. But I think, I don't think it's going to be a massive constraint only because a lot of trapped enterprise data still has not been harnessed. And I think you see that in the work that we're doing at Goldman, for example, We want to help our salespeople, our traders, our quants, our PMs to sort of, again, get that superhuman capability, that information synthesis, being able to help with their hypothesis.

10:01George Lee:And there's still a lot of data here at Goldman that can be used for that. So I think from a consumer world model, I think it's interesting. We've definitely in the synthetic sort of explosion of data, but from an enterprise perspective, I think there's still a lot of juice, I'd say, to be squeezed in that.

10:20Allison Nathan:Yeah, I would echo that. I think these machines have come an enormous distance in their quality, and they've done it largely in the back of publicly available and synthetically generated data. The amount of data that lives behind firewalls, trapped inside corporate repositories that's highly salient to garnering business value, that has yet to be unlocked. It's the work that NEMA is doing here. Then there are also other horizons of think about all the video data in the world. Think about spinning up virtual environments where you're creating a platform for virtual robots to generate their own data about understanding the world.

10:55Allison Nathan:I think while we've exhausted one pool of data, there are many others to go attack.

11:02George Lee:I think it'll get a little philosophical out of my realm, but I think what might be interesting is people might think there might be a creative plateau. I mean, if all of the data is synthetically generated, right? Then like how much human data could then be incorporated, new human data, new human intellect, new human creativity. I think that'll be an interesting thing to watch from a philosophical perspective. For sure.

11:26Allison Nathan:You know, one or two of our prior guests have made the observation germane to this discussion that the quality of outputs from these models, particularly in enterprise settings, is highly dependent on the quality of the data that you're sourcing and referencing inside the business. Do you agree with that? Kind of goes to this, what's the value of these behind the firewall data stores? Maybe just illustrate a little bit of that, how we can make models smarter with our own proprietary data stores. Yeah.

11:52George Lee:I think first to get, you got to remember what this thing is doing, what this machine is doing, right? Whatever data patterns you are feeding this machine is what it's going to learn and what it's going to extrapolate from. And so I think from an enterprise value perspective, cleaning your data, normalizing it, having the semantics of that data well understood, how it links to other pieces of data, all of this stuff is what's gonna allow enterprises to level up from what we think the consumers get to what enterprise value could be created.

12:24Neema Raphael:So what are some things that Goldman is doing to unlock that value?

12:28George Lee:Yeah, look, I don't think people had always thought of data as sort of like this thing that could give more insight to the world. I mean, it's always historically been thought of as like business exhaust in some way, right? Like a trader executes a trade. They're sort of like, okay, I'm done. Now I'm just managing the risk. But there's a whole machine behind that about what happens after that, all the workflows that happen after that and before that. And so the real challenges are getting that disparate data into some place where you could organize it in a sane way and then normalize it in ways where the data is correct when When you ask it a question, it's linked to the other facts of the world when you want to navigate from that fact to another fact.

13:12George Lee:And so all of these challenges are really, you know, that's why the role of data engineering was even created. People are like, we need a practice of engineering that's like software for data. And so just like people write code in a specific way and there's specific architectures and engineering practices to that is the same in data. You have to sort of understand what the data actually means. You have to understand, are these two concepts the same? Are they linked differently? And so really the challenge is understanding the data, understanding the business context of the data, and then being able to normalize it in a way that makes sense for the business to consume it.

13:52Neema Raphael:Can you actually use the models to help you organize the data? Is there some synergy happening there?

13:59George Lee:Definitely. Like people have built software agents. People have built engineering agents to do like this cleansing, this normalization, this linking. So absolutely in the same way where we're seeing software being created by these agents, there's also a feedback loop of data cleansing and normalization and wrangling. So it's a good insight.

14:18Neema Raphael:We often close these interviews by asking our guests how they might actually use AI themselves. Either we talked a lot about how you're using it in the office, but even personally, what do you find is the most interesting and helpful usage?

14:31George Lee:Yeah, you're asking a tech nerd. So obviously, I think it was like the tech nerd. The coding answer is like the base case. But I also have a three and a half year old son and he's in his why phase, which is awesome. But I'm just like, I run out of run out of like the turtles of the why. And so like, actually, a lot of the times, he's like, what is that? Why is that? And, and actually bouncing ideas off of him with the agents, I think is really cool. And I think it's been powerful for me. So now he asks me questions. I ask the AI questions. We learn together about questions he's curious about. So I love that.

15:10Neema Raphael:And as he gets older, he's going to be able to ask himself. So when my teenagers ask me questions, I say, look it up. Use AI. Let me Google that for you. Exactly.

15:20Allison Nathan:Though part of Nima's genius is he gets to free ride on the knowledge acquisition of his son and be involved. So I love that. Exactly.

15:28Neema Raphael:Well, thanks very much. Nima, that was a fascinating conversation. Thanks.

Read the full transcript

15:32Allison Nathan:Thanks for having me.

15:33Neema Raphael:I mean, George, Nima had so much insight. What really stood out to you the most about the conversation?

15:38Allison Nathan:Well, as I predicted, grounded, objective, thoughtful. I agree it was a great discussion. People sometimes peer past the data problem. But as Nima, I think, illustrated well, it really lies at the heart of bringing value from these systems in business. So onward and upward.

15:54Neema Raphael:As always, thanks for the conversation, George. Always great talking to you.

15:57Allison Nathan:Thank you. Thank you, Nima.

15:59George Lee:Thank you both. It was awesome to be here.

16:01Neema Raphael:This episode of Exchanges was recorded on September 25, 2025. I'm Alison Nathan.

16:29Allison Nathan:This material may contain forward-looking statements. Past performance is not indicative of future results. Neither Goldman Sachs nor any of its affiliates make any representations or warranties, expressed or implied, as to the accuracy or completeness of the statements or information contained herein, and disclaim any liability whatsoever for reliance on such information for any purpose. Each name of a third-party organization mentioned is the property of the company to which it relates, is used here strictly for informational and identification purposes only, and is not used to imply any ownership or license rights between any such company and Goldman Sachs.

16:57Allison Nathan:A transcript is provided for convenience and may differ from the original video or audio content. Goldman Sachs is not responsible for any errors in the transcript. This material should not be copied, distributed, published, or reproduced in whole or in part, or disclosed by any recipient to any other person without the express written consent of Goldman Sachs. Disclosures applicable to research with respect to issuers, if any, mentioned herein, are available through your Goldman Sachs representative or at www.gs.com slash research slash hedge dot html. Copyright 2025 Goldman Sachs. All rights reserved.

From the publisher

Goldman Sachs’ Neema Raphael, the firm’s chief data officer and head of data engineering, discusses the role of data in enabling or possibly stalling AI’s progress with George Lee, co-head of the Goldman Sachs Global Institute, and Allison Nathan, senior strategist in Goldman Sachs Research.

This episode was recorded on September 25, 2025.

The opinions and views expressed herein are as of the date of publication, subject to change without notice, and may not necessarily reflect the institutional views of Goldman Sachs or its affiliates. The material provided is intended for informational purposes only, and does not constitute investment advice, a recommendation from any Goldman Sachs entity to take any particular action, or an offer or solicitation to purchase or sell any securities or financial products. This material may contain forward-looking statements. Past performance is not indicative of future results. Neither Goldman Sachs nor any of its affiliates make any representations or warranties, express or implied, as to the accuracy or completeness of the statements or information contained herein and disclaim any liability whatsoever for reliance on such information for any purpose. Each name of a third-party organization mentioned is the property of the company to which it relates, is used here strictly for informational and identification purposes only and is not used to imply any ownership or license rights between any such company and Goldman Sachs.

A transcript is provided for convenience and may differ from the original video or audio content. Goldman Sachs is not responsible for any errors in the transcript. This material should not be copied, distributed, published, or reproduced in whole or in part or disclosed by any recipient to any other person without the express written consent of Goldman Sachs.

Disclosures applicable to research with respect to issuers, if any, mentioned herein are available through your Goldman Sachs representative or at http://www.gs.com/research/hedge.html.

© 2025 Goldman Sachs. All rights reserved.
Learn more about your ad choices. Visit megaphone.fm/adchoices

More from Exchanges

All 81 episodes
AI Exchanges: The Role of DataExchanges · 17 min
Listen in VO