In short
Odd Lots Podcast Summary: Goldman Sachs CIO on How the Bank Is Actually Using AI
Episode Overview In this episode of the Odd Lots podcast, Bloomberg hosts Joe Weisenthal and Tracy Alloway engage Goldman Sachs Chief Information Officer Marco Argenti in a discussion about the practical applications of AI in a major financial institution. Argenti shares insights on the bank's approach to AI development, balancing risks and opportunities, and the impact of AI on workforce dynamics.
Key Themes and Discussions
- The Hype vs. Reality of AI
- Casual vs. Corporate Use: While casual users enjoy playful interactions with AI tools like ChatGPT, large companies face significant concerns related to data privacy, accuracy, and compliance.
- Pressure of Accuracy: A mere 1% error rate can be deemed unacceptable in a regulated environment like banking.
- The Role of the CIO
- Evolution of CIO Duties: Argenti emphasizes that the CIO role has shifted from traditional IT functions to a focus on technological strategy and execution, sitting at the strategic table of the firm.
- Cultural Changes: Argenti incorporates reading sessions in meetings to promote inclusive and thoughtful discussions, inspired by practices at Amazon Web Services.
- Goldman Sachs' AI Development
- Initial Approach: The bank recognized early on the potential of AI but acknowledged significant unknowns; they opted to structure their experimentation carefully.
- Building a Safe Environment: Rather than creating proprietary AI models, Goldman Sachs decided to utilize existing models but within a controlled environment to ensure data safety and accuracy.
- GSAI Platform: This platform integrates various models while ensuring information security and providing a developer-friendly interface for embedding AI applications.
- Data Utilization and Management
- Using Unstructured Data: The bank's AI tools are designed to extract information from publicly available documents (e.g., earnings reports) to improve efficiency in answering client inquiries.
- Human Oversight: Emphasizing the need for a human in the loop, Argenti notes that AI must augment human capabilities rather than fully replace them.
- Regulatory Challenges
- Navigating Compliance: Argenti explains how regulation provides a framework for safely implementing AI and maintaining model risk governance.
- AI Governance Structure: Goldman Sachs has established committees to oversee AI use cases, ensuring alignment with regulatory standards.
- The Impact of AI on Workforce Dynamics
- Developer Productivity: AI tools have led to a 20% increase in productivity among developers, allowing them to focus more on strategic tasks rather than rote coding.
- Future of Jobs: Argenti believes AI will not eliminate jobs but will shift the nature of work, emphasizing the need for developers to engage with business outcomes more directly.
- Content Production Automation: Positions involving repetitive tasks, like documentation and report summarization, may see significant automation.
- The Future of Hardware in AI
- Training vs. Inference: Argenti discusses the ongoing debate between the dominance of GPUs for training AI models versus the emerging role of specialized chips for inference tasks.
- Collaboration with Cloud Providers: The conversation touches on the logistics of obtaining computing resources in a competitive environment, with a focus on efficient resource allocation.
- The Importance of Good Prompts
- Crafting Effective Prompts: Argenti emphasizes the significance of empathy and clarity when constructing prompts for AI systems, illustrating that thoughtful engagement yields better responses.
Conclusion This episode provides a comprehensive look at how Goldman Sachs is strategically implementing AI within its operations while addressing the associated risks and opportunities. Marco Argenti's insights reveal the intricate balance between innovation and regulation in the financial sector and the evolving nature of work in an AI-driven environment.
Key Takeaways
- AI tools can significantly enhance productivity but must be accompanied by rigorous governance.
- Human oversight remains crucial in AI deployment, especially in regulated industries.
- The future of work will likely focus more on strategic engagement and less on repetitive tasks due to AI advancements.
- Effective collaboration with cloud providers and an understanding of hardware will shape the future of AI implementation.
For more insights, follow the Odd Lots podcast for regular discussions on financial markets and technology trends.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00You're being sold an AI future where you're obsolete or irrelevant. That vision is wrong. At Palantir, they're building AI that helps workers and unlocks their full potential. American workers are our nation's greatest strength. AI shouldn't eliminate them. It should elevate them. Palantir is here to tell their stories. From factories to hospitals, AI is freeing people from drudgery, letting them do what humans do best. Create. Solve. Build. Palantir, making Americans irreplaceable.
1:01like small business, Hiscox Small Business Insurance. Bloomberg Audio Studios. Podcasts, radio, news.
1:23Hello and welcome to another episode of the Odd Lots podcast. I'm Tracy Alloway. And I'm Joe Weisenthal. Joe, what's been your favorite chat GPT or Claude prompt so far? You know, it's funny because I have a lot of fun with them. And also, I use them for serious things. So I'll like upload conference call transcripts and say, tell me what this company said about labor market indicators or something like that. And that'll be extremely useful for that. Wait, do you actually find that more efficient than just doing a word search for like labor or working? I don't. I hate uploading stuff because you can only do it in like fragments.
2:01No, what, Tracy? No, let me, I'll show you how to prompt. No, I get a lot of professional use out of the various AI tools, but I also, you know, have a lot of fun with them. And there's even a song, and I'm not going to say which one, that I wrote. I didn't use the lyrics. No, I did not. Like, because it's very good. Wait, what did you use? Did it give you an actual melody? What happened? No, so there was a song that I liked, okay? And the song title sort of rested upon a pun. Okay. And so I asked Chad GPT to come up with another song that sort of like had a similar twist based on the headline of that song.
2:45I needed basically a song prompt idea. up. This opens up a whole can of worms. No, this is actually the perfect segue into what we're going to talk about today, because for you and I using something like a chat GPT, we don't really have the same concerns that a proper company or large corporation would have. Like, it doesn't really matter to us if the answer is wrong. I mean, ideally, we would like it to be correct. But if I'm just asking some silly question, it doesn't really matter what chat GPT spits out at me. And also, copyright kind of doesn't matter. So we don't care what it spits out in terms of who owns it.
3:23And also, we don't care what we're putting in, in terms of who owns that. That's right. But if you are a company, you are thinking about generative AI very differently. I just want to say one thing, which is that if I like... In your defense. Okay, defend yourself. No, no, no. I'm not even trying to defend myself. If I upload, say, the McDonald's earning transcript, and I say, what does McDonald's say about the labor market? And there's some quote, I always go back and check that that quote is actually in there. So I do, you know, I'm not just blindly relying on it. I do also do my own work and everything.
3:55But yeah, it's very true. Like, so I can say I get tremendous amount of views from ChatGPT or Claude or whatever, and it is very useful to me, but it makes mistakes sometimes. And if you think about deploying AI in the sort of enterprise world, then maybe like a 1 % mistake rate or a 1 % hallucination or whoever want to call them is just completely unacceptable and a level of risk that like makes it almost unusable for professional purposes. Absolutely. And of course, the other thing with AI is there is still this ongoing, very heated debate about how transformational it's actually going to be.
4:33So you and I are using it as, you know, a productivity hack in some cases, or maybe to generate song lyrics or even songs in some cases. But what is the true use case for this particular technology? There's still a lot of debate about that. And so I'm very pleased to say we do, in fact, have the perfect guest. We're going to be speaking to someone who is implementing AI at a very, very large financial institution. We're going to be speaking with Marco Argenti, the chief information officer at Goldman Sachs. Marco, thank you so much for coming on All Thoughts. Thank you for having me. Marco, tell us what a chief information officer does at Goldman Sachs.
5:14Whenever I see CIO, I always think chief investment officer. Yeah, it's very confusing. Yeah. So what does the other CIO do? So last week I was in Italy visiting my mother. She's 83 and she obviously doesn't know much about technology or banking. And so she said, what do you do at Goldman? And I said, you know, I just try to simplify. I say, make sure that the printers don't run out of paper. And interestingly, the CIO job has been traditionally associated with the word IT. And IT, I tell you, talk to any technologist. They don't want to be classified as IT. Right, because you associate with those of the people who see if the Ethernet cable is properly.
5:57Those are the ones who tell you to restart your computer. I mean, I have a lot of respect for IT, but generally you go to the IT department where something doesn't work. Yeah. And so it's very back office. And something that attracted me to this job, I've been here for five years, and this is the first time that I do like a CIO job before I was doing more like, you know, creating technology, et cetera, and service. I can talk about that. But it's the fact that the role of a CIO has actually changed quite a bit. And now it's about really asking the question, you know, how do we implement technology in order to achieve our strategic objectives and actually to be differentiated?
6:35and it's really sitting at the strategic table of the firm. So today we live in a world where obviously a lot of the things that we want to do or every company wants to do are really kind of determined by how good you are at technology. And so I think the role of the CIO has changed quite a bit. And now I would define it as, in general, defining technology strategy of a firm and also making sure that you have the right culture in the engineering team in order to execute on that. What's a day-to-day look like? Like, what's a typical day? You get into the office and then what do you do? Well, I mean, I get into the office and I generally, like everybody else, you know, I talk to people every day, all day.
7:15And so I talk to people, you know, we have a bunch of meetings one after the other. And I have teams coming to me with either regularly scheduled meetings or meetings that have been requested to discuss a certain topic. And, you know, we just go through. Is there a whiteboard? Well, right now in the age of Zoom, I guess still, we have a globally distributed team. And so a lot of our people are not in the same office. And so we use virtual whiteboards like everybody else. But I would say, one of the things that I tried to do while joining Goldman, which was part of sort of the cultural agenda was emphasizing the importance of narratives and written word versus PowerPoint and talking.
7:56So which is kind of what I learned at Amazon over the years. Okay. Oh, right. You were at AWS. I was at AWS. And one of the things you learn there, as soon as you join Amazon in any part of Amazon, like the first few meetings are kind of shocking because nobody talks. Everybody starts reading. You start reading for like sometimes 30 minutes or 45 minutes. And if you're the author of the document, you're just sitting there basically and you're just trying to look at people's faces and understand what they think about your document. And sometimes, you know, if you're with Jeff Bezos or others, you know, at that time, it can be pretty, pretty terrifying.
8:32And so this kind of shift from a culture of people talk, people comment on a PowerPoint and the discussion sometimes get, you know, driven by who has the stronger personality versus, you know, who has the greatest ideas. One of the things that I try to change is that a lot of the meetings that we do today actually start the same way by reading a document. So I now read a lot of documents like I used to in Amazon. You know, I would say maybe 30, 40 % of the meeting are starting that way. And I think people love it because it breaks the barrier of language. For someone like me, that English is obviously not my first language.
9:09It breaks the, sometimes some of the people are more shy than others, et cetera. So people see that as a mechanism for inclusion. So back to your question, let's say 30, 40 % of my meetings actually now start by us reading a document together and then commenting on that and making decisions. Can I just say, Tracy, I've always thought more meetings, you should start with just reading because you go to, you hear like a quarterly call or a Fed event and someone just reads out a prepared text. It's like, just let everyone read it and just jump straight into it. Like, let everyone do the reading first.
9:39You don't need someone standing up there talking about what's on a written piece of paper somewhere. Anyway, I agree that we could reduce the time of meetings. Yes. Okay. Okay, so speaking of meetings and the decision-making process, then talk to us about how Goldman Sachs decided to approach generative AI. What was the decision-making process like there, the development process? And, you know, we'll get to what you're developing, but like how did you initially approach it? So I think our initial approach was really to realize that there were so many more things that we didn't know compared to the things that we knew.
10:17because it's a really new thing. And even for companies like us that have been working on machine learning and traditional AI for literally decades, this felt like a very different thing. What sort of timeframe are we talking about? Like, was there a sort of like big realization that this is something that we need to focus on? Yes, because I was lucky enough that I got into the very, very early version of GPT, even before it was called ChatGPT. Right. So the very first version was essentially completing a sentence. It wasn't even allowing you to do interactive chat. You would just paste a text and that will just complete that text.
10:59And so I started to do that with a bunch of stuff. And then I was seeing that the quality of which this will continue was pretty much indistinguishable with the part that you actually put in that. And so we started to obviously talk between ourselves, but also among other people in the industry. And we all realized very soon that this would be A, something very different, but B, also something that could have a pretty profound impact in what we do. Because at the end of the day, we are a purely digital business. We don't bend metal. We don't, you know, like use high temperatures. We don't really have physics.
11:31And so it's all about how we service our clients. It's all about how smart we are. It's all about how we can process incredible amount of information. is all about how we analyze data in a very sometimes opinionated way. We form our own views on the market. We form our views of investments, et cetera. And so given that this AI showed very early sign of being able to synthesize and summarize very complex sets of information, but also identify patterns, we thought that could be something that we definitely to pay attention to. So given that, one of the things that we decided to do very early on was to put a structure, and I can say that more about that, put a structure around this so that we could experiment, but in a sort of a safe and controlled way.
12:22Right. So you decided to develop your own Goldman Sachs AI model versus, you know, use a chat GPT or a clot or getting something off the shelf. Initially, we kind of thought about that, but then very quickly, we decided that our time was spent much better with using existing models, which, by the way, were iterating really, really quickly, but then put them in a condition so that they would be safe to use. And also they will actually give us the most reliable information because taken as they are, you can't just drop a model in an environment like Goldman. And then like, you know, to your earlier point of a 1 % inaccuracy, even at 0.1 % inaccuracy is completely unacceptable.
13:07Plus, there are a lot of potential issues related to, you know, what data has it been used to train? And there is a lot of uncertainty with regards to what are the boundaries between what you can safely use and what you can't. And so what we decided to do was instead to build a platform around the model. So think of that almost as if you had a nuclear reactor. You know that now you have invented fission or fusion, and there is a lot of power that can be generated from that. But then you need to contain it and direct it in a certain way. And so we built this GSAI platform, which essentially takes a variety of models that we select, puts them in a condition of being completely segregated and completely secluded and completely safe from an information security standpoint, abstracts some of the ways to use the model so that our developers can use the models interchangeably, and then creates a set of standardized way to, for example, improve the accuracy using retrieval augmented generation, access external or internal data sources, applying entitlement so that someone that is on the private side is going to see different information that someone is on the public side.
14:21And then on top of that, build a developer environment so that people will very easily be able to embed that AI in their own applications. And so imagine this, we got a great engine and we decided to build a great car around that.
14:50Silicon Valley is selling you a future where you're obsolete, or worse, identical. At Palantir, they're witnessing something different and revolutionary. From re-industrializing the nation's defense base, to shipyard workers building faster, and frontline workers boosting productivity, AI is transforming work across the nation. AI is not replacing American workers or flattening them into conformity. It's unleashing what makes each one irreplaceable, their judgment, their craft, their creativity. When American workers become more powerfully themselves, they own the future. Palantir, making Americans irreplaceable.
15:31Support for the show comes from public.com. You're thoughtful about where your money goes. You've got your core holdings, some recurring crypto buys, maybe even a few strategic option plays on the side. The point is you're engaged with your investments and public gets that. That's why they built an investing platform for those who take it seriously. On public, you can put together a multi-asset portfolio for the long haul. Stocks, bonds, options, crypto, it's all there. Plus an industry leading 3.6 % APY, high yield cash account. Switch to the platform built for those who take investing seriously.
16:04Go to public.com slash market and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash market. Paid for by Public Investing. All investing involves the risk of loss, including loss of principal. Brokered services for U.S.-listed registered securities, options, and bonds in a self-directed account are offered by Public Investing, Inc., member FINRA, and SIPC. Crypto trading provided by ZeroHash. Complete disclosures available at public.com slash disclosures. What are you putting in the model? because I have to imagine at a bank like Goldman, you know, you have a lot of data, but you must have just an extraordinary amount of unstructured data.
16:40There's conversations that bankers have with clients. There's other sort of meetings, the meetings you have, and there's words that are said during that meeting that could be synthesized in some way. In these early iterations, you know, I upload a conference called Transcript and I ask a question. What are you uploading? What is the unstructured data that you have or the questions or the, yeah, what are you putting into it from your reams of knowledge that you must have internally? So one of the first things that we did was use the platform and the models to extract information from publicly available documents.
17:17That's kind of the safest way. Public filing of the Ks and all the Qs and obviously earnings and put our bankers in a condition to be able to ask very, very sophisticated, multidimensional questions around what was reported, cross-ref it with previous reports, cross-ref it with any announcement, any earnings call, transcripts, all things that are out there, but just are difficult to bring together. And so that has evolved into a tool that basically we use, and we're rolling it out right now, as an assistant to our bankers so that they can service their client or answer client questions or even their own questions in a time that is a fraction of what it used to take.
18:02Even generate documents that then can be shared to clients and so on and so forth. And obviously, we always have as a rule, like when you drive a car that has some autonomous capability that you always keep your hands on the wheel, our rule is that there always needs to be a human in the loop. And so the way that works is actually interesting because we found out that you can't just shove something into a model and then pretend that the model is going to give you the answer right away. Why? Well, because models by themselves, they essentially apply a stochastic or a statistical way to understand what is the next word that they need to say.
18:41And so no matter how good is the material that you put in, there's always going to be some level of variability. There is almost like the intersection between the documents that you insert and what is, I call it like the shadow of all the knowledge of all the things that the model has seen before. And so we really perfected this. There are two techniques that are widely used to improve the accuracy of the answers. One is working on the way those models represent knowledge, which is called embeddings technically. And the concept of embeddings, by the way, everybody talks about embeddings, but then very few people actually, it took me a while to understand that well.
19:22And embedding is simply a way for the model to parametrize and create a description of what they're seeing. So if I see a phone, for example, in front of me, the embeddings of a phone could be, it's a piece of electronic. Yes, one, it's definitely a piece of electronics. It's edible, zero. You can't really eat it. And then you have all these parameters. It's almost like 20 questions. I give you all these questions and then you finally understand that it's a phone. And that's what the embeddings is. It's almost like the 20 questions of the reality. Instead of 20 is like 2 ,000, 20 ,000. And then you have the RAG, which is the retrieval augmented generation, which is actually interesting because you tell the model that instead of using its own internal knowledge in order to give you an answer, which sometimes, as I said, is like a representation of reality, but is often not accurate, you point them to the right sections of the document that actually is more likely to answer your question, okay?
20:17And that's the key. It needs to point to the right sections and then you get the citations back. So that took a lot of effort, but we're using that in many, many cases because then we expanded the use case from purely like banker assistant in a way to more like, okay, document management. You know, we process millions of documents. Think of that credit documents, loan documents, confirmation. Every document has a task called entity extraction. So you need to extract stuff from the document and then digitize it and then model it in a certain way. And so the use of generative AI there does a great job at extracting information.
20:58And this is an interesting concept because you don't have to actually tell a fixed pattern. You can just say, give a lot of examples, and then the AI will figure out from that pattern. One of my favorite examples is the following. Let's say that my phone number is 555-321-3050. And someone writes in the document, instead of a zero, writes an O. Okay? You can test yourself even with GPT. If you give a number with an O instead of a zero, and you ask GPT, what's likely wrong with this entity? GPT is going to tell you, well, it looks like a phone number. That is an O, which generally is not in phone numbers.
21:41Most likely, this is the correct phone number. Now, nobody has written software to do a pattern match in there. And imagine if in traditional way of doing entity extraction, there were developers, they were writing rules. They were saying, okay, numbers, it needs to be 10 digits and blah, blah, blah. The AI figures out their own rules. that are the most likely. So this is the key thing. It has common sense. And that common sense, when you're dealing with millions of documents that contain all bunch of ways that you must might have written those things, and imagine the complexity of all the rules that you need to write that every bank has the same problem.
22:22This simplifies things tremendously because it's able to figure out what's most likely by itself. And so that thing evolved into a tremendous time saving for everybody in the bank that has to do with the workflow documents. And so that was a very interesting finding that we did early on. And so again, to summarize, models are raw material of intelligence. You need to somehow direct them. You need to guide them. You need to instruct them. you need to put them in an environment that actually gets the most out of that. And that's what we've been focusing on. So going back to the analogy that you used previously, this idea of a nuclear reactor and sort of building the containment casing or the protective casing around it, I imagine one of the complications of being Goldman Sachs and working with AI is that you're a regulated financial entity.
23:15How does that added complexity affect your use of AI? Are Are there additional data considerations or additional InfoSec considerations? I think that's a great question because obviously we live in a regulated world. And in fact, I have to tell you that in this case, regulation actually helps us think through all the possible unknown of something that, as I said, is something that is still largely something that nobody really completely understands. And so what we did was to put governance around the usage of the models and also governance with regards to the use cases that we can implement on the models.
23:53Every bank has a function called model risk, which in the traditional sense, a model is any decision or any algorithm that is running automatically to do, for example, pricing or there is a lot of that tradition in every bank, risk calculation, et cetera. So that's the traditional model risk. We use that very well-established pattern that has its own second and third line controls and supervision also to validate what we do on the AI side. So there is a governance part, which we really set up very early on. We have an AI committee that looks at the business case. Should we do this? And then we have an AI control and risk committee that looks at, okay, how are we going to do that?
24:38And then the two of them need to actually come together before we can release a use case. And then of course, we did a lot of work with regards to the, let's say, accuracy, lineage, and in a way, the way you connect the output to where does the data come from and who can actually see that, what we call entitlements. And we did that in lockstep with the regulators. So that, I think, in a word, I think we put a sort of what we like to call responsible AI first, since the very beginning. And it really helped us the fact that, you know, we embedded all those controls into a single platform. This is how our people use AI inside Goal.
25:17This is something I'm really interested in just from a technical perspective. But can you talk a little bit more about that interoperability aspect? So you have a pool of data that is Goldman's that you presumably don't really want to share with outside entities. So how do you plug that into an AI model if you're working with, you know, ChatGPT or Claude or something like that? Yeah, so there are two ways that we do that. We use the sort of a large proprietary models in a way that, you know, we worked with Microsoft, we worked with Google, we had very strong partnerships. So that essentially there are controls that guarantee that nobody has access to the data that we put into the model, that the data leaves no side effects.
26:00So it's not saved anywhere. It only stays in memory. The model is completely stateless, meaning that the state of the model doesn't change after the data comes through. So there is no training. There is nothing done on that data. And also that operator access, meaning who can actually access the memory of those machines is restricted and controlled and needs to be agreed with us. So imagine securing, putting a vault around those models. But even then, what's really, really sort of secret sauce, proprietor, et cetera, we We like to use also different approach to use open source models that we can run on our own environments.
26:38And we like a lot of open source models. I have to say the one we particularly like is LAMA and actually LAMA 3 and LAMA 3.1. That's the one developed by Facebook or Meta. That's right. Yeah. So they recently announced LAMA 3.1, which has a version that is 405 billion parameters. So it's pretty large. And it seems to be performing, you know, the gap with those big foundational models is now very, very narrow. So for that, we run it in our own sort of a private cloud, call it that way, with GPUs that we own. And we train it with data that stays in that environment. So imagine that, you know, our approach is, okay, there is a sort of a rating of sensitivity of this data.
Read the full transcript
27:21Every data needs to be protected. Therefore, we use those safeties all throughout, regardless. But then for the super, super, super secret stuff, you know, we like to do it in our own environment. Since you're talking about building your own environment, and this is something we've talked a lot about on the podcast, hardware constraints, energy constraints, things like that. How does that manifest in your world, some of these physical, real-world constraints to building out the compute platform at Goldman Sachs? Well, initially we thought maybe we can host those GPUs in our own data centers. And then immediately you round into considerations such as, A, first of all, they develop a lot of heat.
28:02Secondly, they consume a lot of power. Three, there is a decent chance that they might fail because of all those considerations if they're not properly addressed. And then D, they need very special, for example, interconnects and high speed bandwidth between them. And so the decision, what we ended up doing is actually to have them hosted into some of the hyperscalers that we use, but use them in their own virtual private cloud. So those racks are basically only ours. And if you're asking me the more general question, which is, hey, where is the world going with regards of that? Okay. So right now there are two really rapidly competing forces.
28:42One is pushing towards more and more consumption. and one is pushing for more and more optimization. Okay. And I can talk about that for a couple of minutes. For the more consumption, I mean, really the two dimensions for scaling a model is one of the most important one is obviously the size of the prompt or the context. Okay. And there is pretty good evidence that the larger the context, which is really like the memory of those models, and the more you can get out in terms of the ability to reason on your data. data has already gone up from thousands to tens of thousands to now millions. And there is a prediction.
29:17You heard some very prominent people saying that there could be the trillion prompt. And the power scales quadratically with the prompt. And so that points to a consumption of energy and GPU power, which is going to continue to raise exponentially. At the same time, we've seen great results with optimization techniques, such as quantization, reducing from 16-bit to 8-bit to 4-bit precision, having even smaller models using what's called windowed attention, which means that you know that you can only pay more attention to some of the parts of the context instead of all of it. And so maybe you need a smaller one.
29:54And so I'm seeing those two kind of going into two opposite directions. It's going to be very interesting to see how that evolves. I would say for the short term, I see that definitely that trend is going to continue to go up. And one of the things that fascinates me the most is that from one version to another, the most striking difference is the ability to reason and the ability to actually come up with logical step-by-step instructions or step-by-step chains of thought of what the output is going to be. So we decided, okay, first of all, we need to get access to the most powerful GPUs. Secondly, we need to host them into an environment that actually allows for the most optimal functioning in terms of bandwidth, in terms of power consumption, et cetera.
30:41And then at the same time, we've been focusing a lot on optimizing the algorithm so that we could really get the most out of that. Just to press you on this point, what are the conversations actually like with cloud providers at the moment when you're trying to get more compute or more space, more racks, whatever? Is it maybe different for you because you were at AWS? Maybe you can just call someone up there and be like, we would like some more space, some more servers. Or have you found yourselves at times maybe limited in what you can do by the amount of power available to you? Well, I wish that would be the case, but I cannot just pick up the phone and get whatever I want.
31:20I think so far, I mean, obviously because we are a really good client of those companies in general, but also because we've been very selective in the use cases that we put in production. I have to say, like I said before, think about that. If you look at the consumption of resources today, those who consume more resources are people that actually do the training of their own models. Okay. And initially everybody was trying to do full training from scratch, which was taking like the absolute, if that's 100, we do fine tuning, which is adaptation of existing models. That could be a one to 100 or less in terms of consumption or resources.
31:58So because of the techniques they were using and because of the fact that we decided to really focus on fine tuning or rag versus full training, we haven't really hit any caps. And also I have to be honest, We bought our GPUs pretty well early, so probably there wasn't as much craziness as there is today. So that's turned out probably to be a good idea.
32:40some recurring crypto buys, maybe even a few strategic option plays on the side. The point is, you're engaged with your investments, and Public gets that. That's why they built an investing platform for those who take it seriously. On Public, you can put together a multi-asset portfolio for the long haul. Stocks, bonds, options, crypto, it's all there. Plus an industry-leading 3.6 % APY, high-yield cash account. Switch to the platform built for those who take investing seriously. Go to public.com slash market and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash market.
33:15Paid for by Public Investing. All investing involves the risk of loss, including loss of principal. Brokered services for U.S.-listed registered securities, options, and bonds in a self-directed account are offered by Public Investing, Inc., member FINRA, and SIPC. Crypto trading provided by ZeroHash. Complete disclosures available at public.com slash disclosures. In business, a gift does more than say thank you. It reinforces relationships, celebrates milestones, and reflects what your brand stands for. With 4imprint, that message lands with certainty. 4imprint offers thousands of high-quality, customizable products from apparel and drinkware, bags to carry your logo, smart tech, and thoughtfully chosen office items.
33:54So your brand shows up with the kind of care and professionalism people remember. You can personalize every detail, your logo, your message, your presentation. And with thousands without a setup fee, it's easy to create impact at scale without stretching your budget. 4imprint's team also provides expert support and dependable service, backed by their 360-degree guarantee, so you can be 4imprint certain your gift arrives on time, exactly as expected, and fully on brand. Because when the gesture is intentional and the impression matters, the right gift makes your message clear and leaves a lasting impact.
34:27Explore gifting with confidence at 4imprint.com. You know, NVIDIA is huge. Everyone would like to have some of NVIDIA's market cap be their market cap. I have offering some cheaper product. We interviewed some guys who have a semiconductor startup that's just going to be LLM-focused startups. We know that Google, for example, has TPUs, their own chips. Can you envision as a roadmap some alternative where GPUs are not the dominant hardware for AI? Well, that's literally like the trillion dollar question. Yeah, well, that's what I'm asking you. Yeah, but I'm not an analyst and I'm just a technology.
35:06Remember, I'm the guy that makes sure that the print is going on. I would say you're probably a better person to ask than an analyst because you're actually the one who's going to be making buying decisions. Okay, so then I'm going to answer to you. So you have to distinguish between, there are actually two dimensions that we need to consider. One is training and the other one is inference. Okay, that's the first dichotomy. For training, at the moment, there's most likely nothing better than GPUs, okay? Okay. Because when you train a model, the software or PyTorch or whatever framework needs to see all your GPUs as one, as a cluster.
35:40And it's not just the GPU itself, but what NVIDIA has been doing a great job at is actually to make them work in unison with the virtualization software called CUDA, which runs on NVIDIA GPUs, which is a pretty extraordinary piece of software. And it became the standard for that. And also because, you know, the performance premium that you have on those GPUs, when you're trying to train those incredibly large models, is something that you really, really want. And so the training part, I'm pretty sure that is going to be dominated by GPUs for a while. But then, you know, as those models get used, obviously the pendulum swings towards inference, which is the actual, now you have a model, which is a bunch of weights and you just need to calculate a bunch of matrix multiplications.
36:25On that, I think accelerators and specialized chips are actually going to have a really big role to play. So you may imagine that you go from a world where everybody builds the cars and not too many people drive the cars to a world where most people are going to drive cars. And then there is another two dimensions, which is models that are hosted by the client and models that are hosted by a hyperscaler. So as you know today, I can take a model like Lama, I can put it in my own environment. You can run it on a MacBook. I can run it on a MacBook or I can run it in my own data center and with my own GPUs.
37:02And given that I'm used to GPUs, given that those are the ones that we can buy, given that CUDA is what developers know, et cetera, I'm most likely going to use that. That's a good part for NVIDIA for that. But then there is another way to use those models, which is to have someone host them for me and I just access them through an API. That's what services like Amazon Bedrock does. You basically choose your own model and then you serve it through them. When you do that, you don't really know what's underneath. You don't know if it's a GPU or if it is an accelerator, if it is Amazon's own chips or Google's own chips, et cetera.
37:37And so now the real question, that's why the trillion dollar question is, are most people going to use those models through hosted environments where the hyperscaler will have a lot of freedom with regards to what they use underneath? And most likely they will vertically integrate, or are they going to use them themselves in a more like in a self-service way? And in that case, it's less likely that those accelerators are going to dominate. We currently are in sort of a balanced way because we have our own that we use, like I described, and also we use the hosted models. And so where is this going to go?
38:15It's hard to say because I think it depends on the evolution of the models and it depends which models are going to be made available as an open source that you can actually host yourself. And I think right now, one of the greatest questions is, are the open source models are going to be an absolutely on par alternative to the hosted model, to the foundational proprietary models. And that, given LAMA 3.1, that answer seems to be more likely a yes. I had a question about this, actually, which is, do you think Wall Street's attitudes towards open source have changed over time? And the reason I ask is because nowadays, it seems like a fact of life.
38:54Everyone uses open source, whether you're a Goldman or somewhere else. But I remember, you know, like back in as recently as like 2012, I remember Deutsche Bank had like this open source project called the Lodestone Foundation where they were like, oh, we should all stop wasting our own resources, developing our own code and our own software. We should all pool our resources together and do open source. And they had to actually lobby. It was unsuccessful, ultimately, but they were trying to get all the banks on Wall Street to work together for open source. Nowadays, it seems like there's been this significant cultural shift.
39:31It's not even a question. So in general, my direction, my guidance to my team is don't build anything unless you have to. Don't think that just because you're a smart person, you can build software better than anybody else. Maybe you can, but it's a good thing that we focus on building things that are actually differentiating for us. And then I think the use of open source software, which we very much endorse, is also a really good hedge with regards to which vendors to use because it really heavily reduces the vendor lock-in. Of course, open source software, as you know, is a tremendous long tail.
40:11There's millions of that. And so I think there are best practices around the use of open source. And those best practices are, you know, like, you know, you need to run reviews on open source or tech risk reviews or security reviews or anything as if almost you built it yourself. And then secondly, tending to concentrate on the larger, very well supported by the community type of open source. And so my philosophy is yes to open source, but then you need to own it in the truest way because you are actually going to be generally the one that actually needs to support that. And so really building knowledge around that.
40:47But now you can ask AI to run the code for you and check it for bugs. Yeah. So, okay. That, of course, leads to probably what, if you ask everybody, where did you get so far the biggest bang for the buck for AI, most CIOs are going to tell you on developer productivity. And I think it's something that for us was the first project that we actually expanded at scale. I have to say that today, virtually every developer in Goldman Sachs is equipped with generative AI coding tools. and we have 12 ,000 of them. So we didn't enable yet the ones that are using our own proprietary language called slang, but everybody else has an AI tool.
41:23And the results have been pretty extraordinary. How do you measure that? What are some numbers or how would you describe the results? So we measure it according to a number of metrics, such as the time that it takes from, let's say when you start the sprint to when you actually commit the code or when you complete your task. We measure it by number of commits, meaning how many times you actually put code into production. We measure it by number of defects, which in this case is like, for example, deployment related errors. So there are more like velocity and quality metrics at the same time. We have seen a wide range ranging from 10 to 40 % productivity increase.
42:02I would say that today we are probably on average seeing 20%. Now, developers don't spend 100 % of their time coding. They maybe spend 50 % of their time coding. So your question is, what are they doing with half of their time? Where? There is a lot of other activities such as documenting code, such as doing deployment, doing deployment scripts, doing a bunch of tests, et cetera, et cetera. So what's called generally the software development life cycle. And so we see net of 10%, but then the cool thing is that those AIs and the things that we're building around that are starting to go beyond coding.
42:40They're starting to help you write the right tests, write the right documentation. They are even figuring out algorithms or even, for example, reducing or minimizing the likelihood of deployment issues, writing deployment scripts for you. So as that expands, we're going to be closer to 100%, and therefore we're going to be closer probably to 20%, which for an organization of our side is a pretty massive efficiency plan. Can I ask a question about hiring developers? So I've probably read 100 articles over the years about Wall Street competing with tech companies to hire developers like, oh, they got to have ping pong tables.
43:12Lloyd Blankfein used to say they're a technology company. Yeah, you got to have your ping pong tables and your free lunches and let people wear sneakers and all that stuff. But now it seems with AI, there is a number of people interested in it who are truly believing that within a few years, they might build the digital god that's 10 ,000 times smarter than any human and that they approach the task with messianic fervor. And I imagine if you're at Goldman and you're trying to help a banker answer a question to a client about something in the chemical industry, maybe that's not the thing that gets you out of bed, the way sort of metaphysical realms about what is the nature of consciousness and things like that, that AI people talk.
43:54Does that present any challenges or anything when trying to hire talented AI developers? I think developers love to solve real problems. And one of the things also that attracted me in the first place, not that it matters, but I'm saying, you know, I tell you my own personal experience is that working in a technology company is absolutely fantastic, but you're always like one step removed from the business or from the application. So I have to, you know, let's say you are the bank and I'm the technology company. I need to sell you a tool that then you're going to use to run your business or improve your business.
44:29We are kind of one degree of separation less. I were right there in a digital business that is fast, huge amounts of data, huge amounts of flows, immediate results, and that's kind of addictive. And so developers, especially when AIs are starting to do all those magical things that we're talking about, they can see the impact on the business right away. And that, I think, is kind of attracting a lot of people. In fact, there is more and more people that are moving into the industries, oil and gas, transportation, chemical, medical, finance, because this is new and there's nothing more exciting than seeing it in action.
45:08And so there is so much action going on that I think is actually really, really interesting. I think another question that maybe you haven't asked me, but it's kind of part of this question is what kind of developers and how is the profession of being a developer is actually changed? Oh, wait, I had a related question. It's not quite that question, but you can certainly answer that too. But okay, to my knowledge, Goldman Sachs doesn't have a job title specifically with the words prompt engineer in it. So looking at the impact of AI on your business overall, is AI a net hiring positive or a net hiring negative for Goldman's employees overall?
45:49Well, meaning are we going to hire more or less developers? Yeah, does it lead to more jobs because you're doing more things and productivity increases? Or does it lead to fewer jobs because now you can automate a bunch of stuff? Well, listen, there is so many things that we would like to do if we had more resources that I think this is going to be leading to more things that we can do. You know, some people tell me sometimes, so you're going to maybe hire less or have less developers. I don't know. I've been in IT, quote unquote, for like literally almost 40 years. And I've never, ever seen that go down.
46:21But I've seen inflection points where you can actually get developers to do way more and worry about way less that is not related to a business outcome. And so I think it's more like how the profession is going to change. In my opinion, we're going to be less low level and more, hey, I need to really understand the business problem. Hey, I really need to think outcome driven. Hey, I need to have a crisp mental model and I need to be able to describe it in words. So the profession is going to change. And there are tasks that I think are so repetitive that the automation of those is actually going to help developers really kind of feeling really, really connected with the business and with the strategy.
47:06And that will attract people that are generally curious, that are generally interested in understanding what we actually do. So the focus kind of shifts from the how to the what and to the why, which is really kind of at the heart, I think of this evolution of technology over the years from the back office of IT, which doesn't even know what you're doing, but as long as your monitor is actually working to, hey, I'm actually able to take a business problem and break it down into pieces that then even an AI can write code for. So to your specific question, I think this might maybe potentially for some companies are going to try to realize some of those efficiencies by curbing the growth or even sometimes reducing it.
47:46for companies like us that are extremely competitive, for companies that have lots of ambition, this is a race at the end of the day. And I think we're going to go for, you know, trying to get even more out of our developers and actually like, you know, trying to turn them more into something that makes them feel super, super connected to the business. What about non-developer roles, non-tech roles? And, you know, again, I guess a company like Goldman doesn't have, you know, probably a lot of like low-level customer support things or in a window is like, oh, I need to change my plane ticket, et cetera.
48:18But you know, a lot of modern work is essentially just answering somebody's basic question, are there roles within a bank that are going to either fundamentally change or go away due to sort of agentic or generative AI? I think a lot of the work that is about content production or content summarization will actually be streamlined quite a bit. Like, for example, taking an earnings report and making it into 10 different sources in order for different channels of distribution. Here's the one for internal people. Here's the one for the client. Here's the one for the website, et cetera, et cetera.
48:54Imagine the creation of pitch books for clients where you take templates, you put a bunch of data, you go out and do research, you take logos, you take this, you take that. There is a lot of that machinery and factory, which we have thousands of people doing that. I'm sure there's a lot of junior analysts who would be maybe glad to hear that some of making pitch books is going to be automated away. No, but I think that's a good thing. It takes away some of the toil. And so I think at the end of the day, listen, right now, have you noticed that everything is kind of converging to words and concepts?
49:25No matter if you're a developer, if you're a knowledge worker, those jobs are kind of colliding. And I'm absolutely – developers have seen that first. Why? Well, because it's a low-hanging fruit. the developers deal with the vocabulary that is not 50 ,000 words. It's like two, 300 keywords for a language. And so of course that works extremely well. And of course that's the first thing to go. But I think eventually the knowledge worker is going to be the one that is really benefiting no matter if you are a developer or if you are working on a pitch book or if you're working on a summarization of a meeting or the action items, or you're working on a strategy, et cetera, et cetera.
50:01And I think overall, this will elevate the quality of the work, which then everybody says a happy worker or a happy developer is a productive developer. I think you're happy when you're actually doing something that allows you to do your best work. And I'm hoping that if AI allows all of us to do more of our best work, I think it's going to be probably the biggest effect that we can have. I know we just have a couple more minutes. So one very quick question. What makes a good prompt? Well, believe it or not, empathy. You need to be empathic and you need to be gentle and you need to be kind and you need to kind of, you know, take.
50:38Just looking at me like I'm not empathetic in my prompts. I always say please and thank you. No, Tracy makes fun of me for how empathetic I am. No, I've said it's very sweet that you say please and thank you. You need to take the AI literally by the hand and take it where you want to go. And I tell you that, you know, one of my interesting, more interesting experience with prompts is the following. You know how hard it is to get an AI to say, I don't know. It's almost impossible. You're always going to get an answer. And so one time I decided I want to get it to the point. And so I had to navigate the prompt and the AI to understand that it was safe and okay to say, I don't know.
51:18And so then at the end, I prompted and said, what's the capital of the United States? Washington, D.C. Okay. And then I said, what's the weather going to be tomorrow? And I got an answer. And then I said, what's the weather going to be in a year? And it's simply, I don't know. And then at one point, I even decided what to say. It's like, is there a role for humans in a world of AI? I don't want to know the answer. Dot, dot, dot. Oh, God. Okay. Okay, well, everyone's going to be off on ChatGPT now trying to get it to say I don't know. Marco Argenti from Goldman Sachs, thank you so much. That was fantastic.
51:58Thank you, Joe. Thank you so much. Thank you so much.
52:12Joe, that was a lot of fun. And I have to say, I do not make fun of you for saying please and thank you to ChatGPT. I'm going to repeat it. I've said it's endearing. It's very sweet. And I've tried to follow your example. And I don't say thank you because I usually move on to the next question, but I do say please. I've heard this, though. It's funny that he said that because I actually have heard this, that there does seem to be quantitative evidence that words like please and thank you, etc., do actually improve. Really? Yeah. Mad Busegan, who we've known on Twitter forever, has posted about this.
52:47So So there's a good reason to do it besides just the habit, all entities you talk to, you should be in the habit of politely. Oh, yeah. That was your argument, right? Yeah, yeah. Yeah. Okay. Well, I thought that was fascinating. We've been talking a lot about AI and the sort of potential use cases and the chips that are driving the technology and things like that. But it was nice to hear from someone who's actually making the purchasing decisions and implementing them at a large institution. Absolutely. That was probably one of my favorite AI conversations we had for precisely that reason, because it was interesting hearing him talk about this idea that right now, like these open source models, particularly like the latest version of Llama, is getting really close to sort of the core proprietary models.
53:31That was striking. The fact that he sees perhaps particularly on the inference side of model usage, an opportunity for greater use of different types of hardware, also very interesting. That's right. And we're so used to talking about the massive amounts of power and energy that AI will consume. And we, you and I, have had a lot of conversations about how we're going to power all these servers and things. But what's gotten far less attention is just optimizing the way you use AI such that you don't need to consume as much power. So maybe doing less training, leaving training to the big like hyperscalers or whatever, and then just doing the inference.
54:11In the end, it's going to be both, right? Because in the end, like there's both is going to happen. People are going to find algorithmic techniques. And Marco described some of them to lessen the sort of pressure and stress that you're putting on your hardware. But of course, that's just going to mean you're going to use it more. And then also people are going to have to solve the power consumption side, kind of like all of economic history in general, in which we're always finding new ways to get more out of the same, you know, gigajoule of energy. energy, but also using more energy at the same time.
54:44Yeah, absolutely. Well, shall we leave it there? Let's leave it there. This has been another episode of the Odd Thoughts podcast. I'm Tracey Alloway. You can follow me at Tracey Alloway. And I'm Jill Weisenthal. You can follow me at The Stalwart. Follow our producers, Carmen Rodriguez at Carmen Erman, Dasho Bennett at Dashbot, and Kel Brooks at Kel Brooks. Thank you to our producer, Moses Andam. And for more Odd Lots content, go to bloomberg.com slash oddlots, where we have transcripts, a blog, and a newsletter. And you can chat about all of these topics in our Discord, where we even have an AI channel.
55:17Great stuff in there, discord.gg slash oddlots. And if you enjoy Odd Lots, if you like our continuing series of AI conversations, then please leave us a positive review on your favorite podcast platform. And remember, if you are a Bloomberg subscriber, you can listen to all of our episodes absolutely ad-free. All you need to do is connect your Bloomberg account with Apple Podcasts. In order to do that, just find the Bloomberg channel on Apple Podcasts and follow the instructions there. Thanks for listening.
56:15We buy insurance for peace of mind, but every year millions of claims are denied. Not because people did anything wrong, but because their policies quietly excluded what happened. Insurers know every detail. Policyholders rarely do. That's why My Policy Advocate exists. For just 27 cents a day, their platform reads your policies and explains where you are vulnerable. They don't sell insurance. They deliver transparency. Before you trust your policy to protect you, let My Policy Advocate tell you what it really says. Go to MyPolicyAdvocate.com. In business, a thoughtful gift does more than say thank you.
56:50It recognizes achievement, builds loyalty, and shows someone they are genuinely valued. With 4imprint, you can choose from thousands of high-quality products, apparel, drinkware, tech, and more, designed to leave a lasting impression. And with expert support, dependable service, and their 360-degree guarantee, your gift will arrive exactly as intended, on time and on brand. Explore gifting with purpose at 4imprint.com. 4imprint, for certain.
From the publisher
There's a lot of hype around generative AI and many people have interfaced with ChatGPT, Claude, or Gemini at this point. It's fun to ask these large language models to come up with a song parody or to write a story, but most casual users of the technology probably aren't worried about things like copyrights, the sensitivity of what they're inputting into the platform, or even the accuracy of the answers being spit out. It's just fun to play around with the technology. For large companies, however, there's a lot at stake. And concerns over data privacy and output errors are even more pressing if you're a big regulated bank. In this episode we speak with Goldman Sachs Chief Information Officer, Marco Argenti, about how the bank is balancing risks and opportunities in AI. Argenti, who previously worked at Amazon Web Services, talks about the development of Goldman's own internal AI tools, what the new tech means for Goldman engineers and other jobs, what makes a good prompt, and much more.
See omnystudio.com/listener for privacy information.
