In short
Eye On A.I. Podcast Episode #123 Summary: Aidan Gomez - How AI Language Models Will Shape The Future
Episode Overview In this episode of Eye On A.I., host Craig S. Smith engages in a conversation with Aidan Gomez, co-founder of Cohere and a notable AI developer. They discuss the advancements in AI language models, particularly the transformer algorithm, Aidan's journey in AI development, and the implications of these technologies on the future.
---
Key Topics Discussed
Introduction
- Host: Craig S. Smith
- Guest: Aidan Gomez, co-founder of Cohere
- Focus: The evolution and impact of AI language models, particularly through the lens of the transformer algorithm.
Aidan's Background
- Started his career at Google Brain, working under Geoffrey Hinton.
- Contributed to the development of the transformer algorithm which underpins many modern AI applications.
The Transformer Algorithm
- Initial Concept: Aimed to create a multimodal model that could handle diverse data formats (text, images, audio).
- Architecture Breakdown:
- Simple structure with Multi-Layer Perceptrons (MLPs) and attention layers.
- Simplifies the complexity of previous models (e.g. LSTMs) while enhancing performance.
Attention Mechanism
- Purpose: Helps the model understand relationships between components in a sequence (e.g., adjectives and nouns).
- Functionality: It enables the model to learn contextual relationships, enhancing language understanding and generation.
Scaling AI Models
- Early scaling involved increasing model size and dataset size, leading to unexpected improvements in language comprehension.
- The simplicity of the design has allowed for rapid advancements in capabilities.
Cohere's Vision
- Foundation: Aimed to democratize access to language models, enabling developers to create applications powered by AI.
- Product Focus: Offers flexible deployment across various cloud platforms, addressing privacy and compliance issues.
Current Trends in Language Models
- Applications: Language models are evolving to handle more complex tasks and interfaces, including conversational agents.
- Challenges: Addressing hallucinations in AI responses, with mechanisms to cite sources improving reliability (e.g., Retrieval-Augmented Generation).
Future Implications of AI
- Discussion on how AI language models are poised to redefine user interactions across different industries.
- Emphasis on the expectation for conversational interfaces to become standard in digital products and services.
Concerns About AI
- Aidan and Craig touched on the ethical implications and potential dangers of AI, including concerns raised by Geoffrey Hinton regarding unforeseen consequences.
---
Key Takeaways
- Aidan Gomez’s journey illustrates the rapid evolution of AI language models, particularly through his work on the transformer algorithm.
- The transformer architecture, while deceptively simple, has significant implications for scalability and performance in language processing tasks.
- Cohere aims to empower developers by providing flexible, secure AI solutions that enhance user interaction with technology.
- As AI models advance, it is crucial to address challenges such as data privacy, hallucinations, and ethical concerns surrounding AI deployment.
---
Conclusion This episode of Eye On A.I. presents a deep dive into the transformative potential of AI language models through the experiences of Aidan Gomez. As AI continues to evolve, understanding its architecture, application, and implications becomes increasingly vital for developers and users alike.
---
Additional Resources
- Cohere Website: [Cohere](https://cohere.ai)
- Twitter Handles:
- Craig Smith: [@craigss](https://twitter.com/craigss)
- Eye on A.I.: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
For more insights into the world of AI, tune in to future episodes of Eye On A.I.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00It was quite extraordinary, quite extraordinarily convenient that simply by scraping more data off the web, not necessarily clean data, messy data, just web data, taking in everything, there's tons of everything. out there, but taking in a very noisy, messy, massive data set and just making the model bigger, throwing some more chips at it. And what came out the other side was something that understood language in a way I personally thought we were decades from. We're talking this week to Aidan Gomez, who helped develop the transformer algorithm, which lies at the heart of generative AI empowers large language models such as GPT-4.
0:45AIDAN now leads a startup, Coher, a platform that offers users access to pre-built LLMs as well as allowing users to create their own LLMs. First though, I want to give a shout out to our sponsor and encourage anyone with a business to take advantage of a deal from Oracle, which is offering a full NetSuite implementation with no down payment and no interest for six months. NetSuite is a cloud-based business management software for enterprise resource planning, financial management, customer relationship management, and e-commerce. To take advantage of the offer, go to netsuite.com backslash I on AI.
1:40That's netsuite, N-E-T-S-U-I-T-E dot com backslash E-Y-E-O-N-A-I. Now, let's get back to Aiden. I'm Aiden. I'm the CEO and co-founder of Cohere. I started the company with Nick and I in about three and a half, four years ago. Before that, I was kind of the perpetual intern at Google Brain during my undergrad and then later my PhD. I started down in the Bay Area and Mountain View. And I was part of the team that created the Transformer. And it was incredibly exciting. It took the world by storm, I think, certainly to my surprise. And I think everyone on the team was quite taken aback by its popularity.
2:38But before Google, I was an undergrad, or also during Google. I was an undergrad at U of T. I grew up in rural Ontario, Canada in a maple forest. And so I'm the world's most Canadian man. Yeah. But yeah, that's me. And so you were at U of T, were you studying with Jeff Hinton? I guess he was probably kind of retired from teaching by then? He was definitely not teaching, but he was still at the university. This is before the Vector Institute was created. And so, yeah, he was, you know, like I I didn't really get into deep learning until after second year. And then when I started looking into it, I became obsessed and I was just reading papers night and day.
3:39I would fall asleep with a research paper sitting on my bedside. I would, in between sets at the gym, you know, have a stack of papers that I was reading through. And I kept seeing this name and his affiliation was U of T, which was where I was. And so I reached out to Jeff. This is before Google. And I just, you know, I'd been reading his papers. At that point, I was studying Reloos and MLPs and just the most simple piece of the AI deep learning stack. And I was like, why do you have these functions that are just flat and then up? I think that they should be periodic. And so I emailed him with an idea of being like, hey, why did you make this decision?
4:28I think they should be periodic. There should be some regularity. and it should be bounded so that it doesn't go to infinity if we get a large input. And to my surprise, he responded. And he actually explained the decision. And so that was pretty amazing. That was my first interaction with Jeff. And then when I came back from Google in Mountain View to Toronto, Jeff said, hey, come work with me in the Toronto Brain office. And that was where I met my co-founder, Nick, was there. And just on, so you worked on the Transformer algorithm with a team in Mountain View at Google. Google Brain, was it?
5:16Yeah, Google Brain, yeah. So can you explain that periodic versus stable or which algorithm were you talking about? yeah i mean it's um it's not very important because i was wrong uh so in some sense it doesn't it doesn't really um matter and i think it's more just to jeff's credit the fact that he responded to a second year undergrad with uh you know wacky idea earnestly um and this guy was literally the top of his field. He had time. He had time for me. And so I think that particular piece, I mean, maybe it's interesting. So for instance, in deep learning and neural networks, we have these neurons.
6:09These neurons fire. There's some function that determines they're firing. There's generally some threshold at which they don't fire, they stay dormant and then above that they fire. And so when they're firing, they basically, they fire linearly proportional to the input intensity that they're getting. So if the input intensity is high, the output intensity is high when they're firing. But that leads to potentially unstable behaviors. If you have, for whatever reason, some sort of blow up or some sort of like burst of signal coming in, then you'll get a huge burst out and that'll propagate and make things more and more noisy.
6:58And that leads to instability. It makes things complicated in training. And so my proposal was instead of just firing linearly proportional to your inputs, instead have some sort of predictable regular periodic pattern, like a sine wave or something, so that you always know your output is bounded between some values. But that has not taken off, and we've since solved the training instability and the blow-ups and that type of thing. So, yeah, that was just my first email to Jeff, I think six months into my study of deep learning. Wow. That's impressive. And from Maple Forest. Sort of. I love that, but I go back often.
7:50Yeah. Then at Google Brain, what was the project that you were working on? What was the initial idea that led to Transformers? So I was on the infrastructure side. Like the original idea I joined Google for, I was working with Lukasz Kaiser. And what we wanted to do, I think Lukasz operates half a decade to a decade ahead of his time constantly. And so the project that I joined for was actually this paper called One Model to Learn Them All. And the idea was, we're going to take every single data set that machine learning researchers have compiled, and we're going to put it into one model. And that means it needs to be multimodal, because we have data sets for images, we have data sets for audio, video, you know, text, everything.
8:46And so what we wanted to do was throw all the modalities in as well as out. So you can consume video and let's say describe the video or you can consume audio and transcribe it. But you can also take in some text and then produce audio. You can also describe the video that you want and video comes out the other side. So it's just like fully multimodal on both input and output side. And we just train on everything, like truly everything we've come across. This now sounds kind of familiar, right? Because this is sort of the project roadmap that we're on right now with these big, large language models that we're throwing everything we have, the entire internet.
9:30And now we're starting to add in every modality that we can. and so that was what I joined for that was a different project altogether to support that project we built this we built this piece of software this piece of infrastructure because that model was going to be huge and the data pipelines were going to have to be extraordinarily complex and so we needed something to to suit that and so what we did was created this program called Tensor to Tensor. It could distribute across arbitrary numbers of GPUs, like thousands and thousands and thousands. And it was very focused on autoregressive modeling, which is the type of modeling that the transformer is.
10:15And so at that time, I was sitting next to Noam, who was fiddling with autoregressive models and in particular attention-based models. He was really interested in attention. And then we heard about a team over in Translate, which was being led by Jakob, which was also interested in attention-based autoregressive models. And so Lukash convinced Noam and Jakob, come over, build it on our stack, build it on tensor to tensor. And they did. And so over the next, I think, 10 weeks, it was just a sprint to build this model. And the intensity just ramped up and ramped up because the results we were getting were extraordinary.
11:06So I think this was like, it wasn't the first, but it was one of the very early, extremely successful scaling projects, like hyper-scalable architectures, massive data, massive model sizes. and massive GPU clusters just led to extremely high performance. And first of all, the tensor-to-tensor, that's a framework or an orchestration layer? Yeah, yeah. So it was built on top of TensorFlow at the time. but it was basically just a library to support large distributed model training and it had all the latest kind of tricks and hacks with learning rate schedules and initialization techniques and it had all this stuff built in and so it led us experiment really rapidly.
12:16I think, if I'm being honest, Tensor2Tensor was a mess. It was crazy. It was just all over the place. It supported everything. We were just throwing every new paper that was coming out into it. It was a little bit chaotic. And there exists far, far better systems nowadays. But back then, it did the job. it did the job we were able to move insanely fast um and so i'm quite proud of it and you were attention was already something that was being talked about uh a couple of questions in that process what was your role uh i mean i'm i'm not i'm a journalist I imagine you guys sitting next to each other furiously coding.
13:12I mean, were you coding or is it more that you're in a room with a whiteboard trying to figure out the architecture or is it something else? There's a lot of like whiteboarding and diagrams and just conceptual structuring these building blocks and putting them together. thinking about the architecture itself. There was a lot of that. And that was mainly done by Noam, Shishniki and Jakob. I think for me, I wasn't sleeping. I was working like 14 hour days coding, building up this infrastructure, making it more robust, running experiments. And so it was very much hands-on. um coding no one was sleeping everyone was just hacking experimenting running little tweaks little ablations to see if i add this in what changes if i if i remove it if i tweak it um every single one of us was just messing with everything and trying to figure out what was the optimal configuration um and so that's how we got to that that finished product yeah and and certainly uh The result now is leading to auto code generation.
14:41Were you using any tools to speed up the writing of the code? At that time, nothing existed. Truly nothing existed. It was all, you wrote it yourself. Yeah. Yeah, that came later. And that was powered by Transformers. They kind of enabled that. Yeah. And for myself, I mean, I've read the paper and certainly talked to a lot of people about Transformers and their progeny. But can you explain, in as simple terms as you can muster, what the transformer algorithm is and what it does. And I'm just curious too, if you were to send me the transformer algorithm, sort of the basic algorithm, is it a million lines of code?
15:50Is it 20 lines of code? I'm just curious what it looks like. Yeah, nowadays it's probably closer to 20. um 20 lines of code uh extremely extremely simple um i think a big part of the beauty of the model the architecture was the fact that it was just so simple um like it it is among the simplest architectures that were going around at the time it was just some, like the most basic layer, the layer that has existed for like, I don't know how many years now, maybe over half a century. Like the basic layer is called like a MLP. That's just what it's called, MLP. And really, the transformer is like a simplification, but it's just some MLPs stacked on top of each other plus an attention layer and nlp you're saying like natural language process m no okay yeah yeah um so this is just the name doesn't matter multi-layer perceptron is the actual acronym but multi-layer perceptron sounds like a neural deep not but totally yeah That's the fundamental unit.
17:19And before Transformers, there were these very complicated LSTM architectures with gates and all of these confusing bits and bobs that just made it work. With the Transformer, all of that was torn away. And the layer became MLPs plus one attention. That was it. And so that was super, I don't know, it was beautiful that you could just carve away so much stuff and just leave something so simple that performed so well, that was so scalable. So the architecture is not this hyper complex beast. it's actually just a very simple scalable compute saturating you know thing explain what it does so you have the multilayer perceptron as the base how do you create attention how do you create attention yeah so attention is like this idea that you want to relate parts of a sequence to other parts.
18:38It's a fundamental property that there are relations. If you have a sequence of things, a thing in a list, in an order, there are going to be relationships between those things. Obviously, that appears in language very, very strongly. You have adjectives, which are tied to nouns and tons and tons of structures like this. And so since we were developing this explicitly for language, we wanted the model to be able to represent those relationships quite easily. And so that's what attention does. Attention says for this word in this sentence, I'm going to learn which other words or which other word in the sequence it's related to.
19:21And so for the sentence, the brown dog, you're going to want to learn that brown refers to dog.
19:34And maybe the refers to dog. So you're going to want to model those relationships. And attention enables you to do just that. And it's not that simple. It's not just like the model is learning adjective noun relationships. It's learning far more complex. stuff that we probably don't even have language to describe, but we just do it intuitively in our heads. So that attention layer is the fundamental unit of learning relationships in sequences. And it turns out to be extraordinarily powerful. And how then does that, because I've spoken to Ilio on the podcast And he talks about seeing the paper and like the next day implementing it in what they were doing that led to the GPT models.
20:37How does that scale then into the large language models that we see today? in their earliest form it was like a very naive scaling it was just take it take the model and make it bigger and the way that you do that is you add more neurons to the network you add more layers so it becomes you know a much taller model much more deeply stacked and you just take a much larger data set than the one that we were considering and a much much larger model than the one we were considering um and a much larger pool of compute you plug those all together and what comes out the other side i think it shocked virtually everyone um it was quite extraordinary uh quite extraordinarily convenient that simply by scraping more data off the web not necessarily clean data right like messy data is just web data it you're just taking in everything there's tons of junk out there but taking in a very noisy messy massive data set and just making the model bigger you know throwing some more chips at it and what came out the other side was something that understood language in a way I personally thought we were we were decades from um yeah it was quite a extraordinarily convenient and um exciting reality so uh and and that led to uh bert is that right uh that that in particular like bert predated or maybe i have them in the wrong order there's some order there's there's gpt1 which was the first of these scale-up large language model papers.
22:40I think BERT predated GPT-1, I think. But BERT is a different thing. BERT is kind of like a different beast. Instead of learning to generate language, it learns to represent it. And it's a subtle distinction. now like we're all paying attention to the generate side because it's so it's visceral right it's like you can talk to these things they can write back to you it feels there's a very visceral human reaction to something that can speak to you but there's another side to this whole thing which is
23:22representing language in a numerical form and that's extremely important it's hard to overstate how significant that is. And that was like the first killer application of Transformers. It was integrated into Google search and Google themselves describe it as the most significant advance in search quality in I think it was two decades, 20 years, like basically Google's entire lifespan.
23:55so that was that was amazing we got something, we got a model we got a program that was capable of representing language to be used downstream for applications like search and classification etc
24:13extremely, extremely faithfully in a very, very high utility way in a way that just boosted performance in a way we really didn't expect across pretty much any tasks you threw at it. And anytime you wanted to use language for some downstream thing, putting a BERT model there and taking the representations from that and running with those representations, you beat state of the art. You outperformed everyone else. so maybe BERT was like the first seed of this idea we can take a transformer we can set it against a very simple task on a very diverse set of data and what comes out is something that seems to get language it seems to just get it if I'm right that BERT predated GPT-1 I'm not sure that's true you'll forgive me i want to get to cohere but i i i'm a layman my audience is somewhere in between me and and you i mean they're they're fairly sophisticated but so you've got 20 lines of code you feed it some data, let's say a sentence.
25:36How is it, and it's relating within the neurons or the perceptrons of the multilayer perceptron, it's relating one piece of data, one word to another word. How is it doing that? Is it by feeding it huge volumes of data that it begins to see patterns? Or within that 20 lines of code, something incredible is happening? Is it possible to explain that? I think it's maybe one line of code. that leads to that behavior. The other 19 are support. I would say the one line is the objective. It's like what you're asking the model to do with the data. You're feeding through this like hyper complex pool of data.
26:47And what does it mean to feed it through? Well, what you're actually doing is you're saying in the generative case, this is like the GPT style case, you're asking it to, given all the words up to a point in a sentence, predict the next one. And that sounds simple. It sounds like stuff we've had for a while, which is like autocomplete, tab autocomplete, or no, it's like that objective is horrendously complex. Because if I give you, on the internet, there's examples of translation, right? Like there's forums online where people teach each other how to speak different languages. And someone asks, Hey, how do I say the brown dog in Spanish?
27:31And then stop. And then the person responds, Oh, you say it by, I don't know how to speak Spanish, but whatever it is. Right. And so if you ask your model to model this, the only way for it to accurately model this, it has to know how to speak Spanish because it's seeing the English part saying, hey, how do I translate the brown dog into Spanish? Stop. And now I need to produce the Spanish translation. And so you can see, like, just organically, by learning to generate sequences in order, you're forced to learn extremely complex behaviors like translation, like classification, like writing code.
28:13at the top of a piece of code, you'll have a function signature, you'll have a comment, a doc string saying this function does X, Y, and Z, it takes these inputs of this structure and outputs the following. And then if you're going to model that code, you have to learn to program because you're just given a function signature and then a doc string that humans wrote for other humans to read. And so I think one of the most beautiful things that falls out of this is using this very, very simple structure, which is just here's a ton of data. Learn to generate it. Learn to predict the next token. You think you're asking the model to do something quite simple and minimal.
29:00The reality is you're asking it to do an extraordinarily complex set of tasks. You're asking it to understand our culture, our language, the interactions between us. You're asking it to understand that data at the deepest level. And so what you get out the other side is a model that, you know, roughly does understand and does have the capacity to do all that stuff, does understand our culture. I think that's another one of these like beautiful simplicities. Such a simple objective. Such a simple objective. Pick the next word. And what falls out of that, what you're actually asking it to do, it's so extraordinary.
29:50and when you're so there's what five of you working side by side uh how many people were working on the project i think weren't there five or six names on the paper i think there were uh you know eight or seven eight yeah but in any case you're it was there a moment or did you know going in just from whiteboarding that wow this this could work or or was there a moment when you were you know running tests that you began to see these extraordinary results and knew you were onto something amazing yeah there are definitely moments where like someone would come running over from their desk and be like, yo, come look.
30:44And they had just run the eval and it was like, it was state of the art. It beat everything that came before. And then we would all be like, nice. Okay, let's keep pushing. And the funny thing is it came together so quickly. It was really like over the span of three months. This wasn't like a year long effort or anything like that. It was this super fast iteration pace.
31:13I don't know if there was a moment. I really don't think anyone fully grasped the significance. And that's mostly because the significance wasn't there at the time.
31:31The significance came from the fact that people adopted it. They could have adopted something else. They could have leaned into something entirely different. They chose a transformer for whatever sort of memetic effect led to that. But they chose a transformer. They started investing, the community started investing tons of time in building infrastructure and support all the way down to the hardware level for this particular architecture. and they enabled us to, us being the entire AI community, to consolidate on one architecture. And so I've said this before, and I feel quite confident almost everyone on the paper would agree.
32:22The Transformer could have, it could have been another model. Frankly, it could have been another model. The transformer was just the simplest. It had the best support. And then the community reinforced that. And the community made some sort of decision to consolidate on this architecture and really invest in it. And they made it a success. It could quite easily have been another architecture that similarly scaled up well, saturated compute well. You think there are other architectures out there that just haven't been discovered or explored that could lead to such dramatic results? Absolutely.
33:11Like unequivocally. I think definitely they exist. They're out there. And with enough work and effort, maybe we could flip to another architecture. but we've already done half a decade of infrastructure development and software support and writing highly optimized kernels for the hardware for transformers. And so there's this resistance to moving. It would take a lot of community willpower to move away from the transformer. And the only thing that would motivate that is some new substantial breakthrough. at the architecture level.
33:58Yeah, so I don't see that happening, but I also don't make the claim that the transformer architecture is some divine, entirely unique piece of software. I mean, presumably these large language models themselves could at some point suggest other architectures. yeah people have wanted to use models uh in that sort of like feedback loop um yeah yeah i i think that's definitely we're already starting to see um chip architectures being uh decided by by models um oh is that right yeah and so the chips uh train the model and the model change you know decides the next generation of the chip and there's um this feedback loop that who's doing that uh google mostly their v four or five tpu chips were model uh placed designed um yeah so i think that's that's exciting that happens on a super slow time scale because it just takes so long to actually fabricate chips push them out verify them um so that happens at too slow a time scale the stuff that you're describing like the architecture search uh projects i i would say those have actually surprisingly been quite low yield um and that's probably because humans have spent so much time on neural net architectures they've explored that space so thoroughly and done a pretty like pretty compelling job at it and so when we threw models at it like the gains were marginal always or or they like rediscovered stuff that we had discovered previously and kind of missed and they just brought it to light they surfaced it again
36:11so people have kind of tried that but it seems like in architecture space it's actually it seems to have been saturated or perhaps the methods used this was also at google perhaps the methods used weren't the right ones it's hard to say but there was an effort to try to get models to produce new model architectures and have this self-improving feedback loop and i I would say that it largely fell flat. So you went then from Google. Tell me about how you started Cohare. Yeah, so I spent the better part of three years bouncing around. So I was in Mountain View for the Transformer, and then I went to Toronto, and Jeff said, hey, come and hang out at Google in Toronto.
Read the full transcript
37:05And then I graduated from undergrad. I went to Oxford for my PhD. Jakob, from the Transformer paper, he had actually decided to leave Mountain View and go back home to Berlin. And he was like, you know, I'm going to set up a brain office in Berlin. And so I was like, hey, that's pretty close, like a 40-minute flight from London. Let's work together. And so then I was on a plane every two weeks to Berlin to see Jakob and work there. and eventually
37:42eventually I just realized like there was a revolution kind of promised back when I was in Mountain View just after we had released the Transformer paper publicly Nome immediately started working on language modeling and scaling the models up. And he was actually deeply involved in the GPT-1 paper. He was helping opening AI with it. And then I went back to Toronto and I got an email from Lukash. And he's like, hey, have a look at this. And in that email, there was a Wikipedia article. And the title was The Transformer. um and then i so i was like oh hey there's a wikipedia article on this uh i kept reading down and then with a japanese punk band and it consisted of these members and this member had left and i i was just like what the fuck like lukash what is this and he was like a transformer wrote this i just put in the transformer as the title everything else
38:55and that was just like you're kidding like it was like surreal it was just like you you know you went to bed one night and models could barely spell and then you woke up the next morning and they were writing as fluently as like such a plausible story about a Japanese punk band called the transformer um and i uh i think that was like the moment that i was like okay this unlocks in product space this unlocks something categorically different like just uh something extraordinary um and i thought it was going to happen and i waited and i waited and i you know i was doing my phd and I was putting out new research, improving fundamental methods.
39:39And after three years there, nothing had changed. The world was the same. And Nick and Ivan, my co-founders, like, I think we all felt the same disappointment. Nothing had changed. We saw something magical three years ago and nothing had changed. No one's talking about it. And so eventually that disappointment turned into resolve to do it ourselves. And so we decided, okay, let's leave and let's go build Cohere to bring this to the world. This is before GPT-3, just after GPT-2 in 2019. and back then the mission was really just A, this is the most amazing technology that humans have ever created let's model the web let's build a model of the entire internet and B, let's put it into the hands of every single developer on earth and let's inject it into every single product and just create a new generation of magical product experiences.
40:59So that was really the seed. And then, so Cohare is, at its core, a large language model or a suite of models for different vertical tasks or describe what it is and how people use it. So at its core, yeah, it's like we're an intelligence factory building these big models, making them as usable, as useful as possible. There are like a suite of models. We have both sides of that coin that I was describing before where there's the generative and then representation. So BERT style is representation and GPT style is the generative side. So we have both of those and we build them in-house. The way that we bring them to the world is that we partner with enterprises and we solve really what are some of the today's largest blockers for adoption, which are privacy, privacy blockers, data compliance blockers.
42:16If you're really going to put these large language models into useful applications at the forefront of your product, they're going to be touching data that's the most sensitive, like user data, right? Like people's private data. And so that puts a very, very high security bar. so for us one of the benefits of being independent our competitors mostly are sort of bound to one cloud provider there's exclusivity there for us being independent means we can play with everyone and with the enterprises that use us they don't get vendor lock-in so they're not trapped into one cloud provider they can bounce between and we can deploy wherever they go So for Cohere, one of our core efforts right now is making it so that these models can be deployed on any cloud provider in situations where the data is the most sensitive, because that enables the most interesting and impactful applications.
43:24otherwise you kind of get what I've been seeing a lot of recently which is superficial deployments of these models they're not real not not product changing not like fundamental shifts uh in infrastructure but more like here's my product and I'm just tagging it onto the side and here's like an auxiliary experience um I think that makes a lot of sense given the fact that this year everyone just kind of like woke up and so um it's going to take a while to actually replace this with the the thing that we want um so it makes sense but really the piece that's blocking this is the fact that there's not a lot of trust uh in some of our competitors due to the fact that in the past they've trained on their user data and they've disintermediated people.
44:19And so for us, we want to regain that trust and be the trusted partner for enterprise to actually bring large language models into like a truly transformative way. So I think there's like right now, there's a product transformation that's kind of simmering under the water, because the whole world just woke up every single company now is trying to figure out what does this mean? What does this technology mean for my product, my experience? What are my users, the consumers, what are they going to expect from me? How do I not get left in the dust by my competitors who are going to reinvent their product on the back of this technology?
44:58And so they're starting to do the work. In 18 months, the product space is going to look completely different because right now everything is shifting behind the scenes. and so for Cohere we really want to power that transformation and be a trusted partner to the largest enterprises and the best developers on earth and enterprises spans the gamut of industrial verticals or are you focused on one industry it's totally totally horizontal so it impacts everything um like i think you're going to be doing your banking with a conversational agent you're going to be doing your shopping with a conversational agent i think it's really hard to think of a particular vertical or industry that's that doesn't need to be changed by this because consumer expectations are going to be there's going to be this inter when i show up to this new product there's going to be this interface that i expect, which is language.
46:08So with these interface level changes, and in the same way that if you're a product or a service, you have to have a mobile app because everyone's on their phones and that's how they want to interact with products and services. In that same way that that mobile transition led to everyone having to support this interface that the consumer expected, everyone is going to have to support conversation and dialogue with an intelligent agent. as an interface onto your products and services so there's like this resurfacing of product space that is literally happening right now is there an example uh without naming names that you can give that you think is going to blow everybody away i mean i it's no secret that we're starting to see some very compelling assistant-like offerings.
47:10There were the promises with Siri and Google Assistant and Alexa that came 10 years ago or whatever it was. And those fell flat. And I think the technology truly just was not there to support it. There is now the possibility of a truly general assistant. We actually have the technological bedrock to support that. I'm sorry? It's emerged recently. It's a fairly recent development that that has been unlocked as a thing you could possibly build. Yeah. You know, I talked to Ilya about RLHF, reinforcement learning with human feedback, as a way of kind of guiding the model toward more grounded responses.
48:10But I've talked to other people who say that's still speculative and takes a lot of time. And they're using vector databases and loading vector databases with authoritative data. And then the language model, in effect, is just the mouthpiece. It's not calling up the answers from its accumulated knowledge. It's referring to this vector database. how do you guys deal with hallucinations um yeah i there's like there's someone like sarah hooker at cohere she said this before and i i really i really like it um you have to distinguish between the hallucinations that you want which are like creativity and the hallucinations that you don't want it like it's great when it hallucinates a story or a new joke or you know um you want that and so you don't want to like beat that ability that capability out of the model um at the same time you need ways to control it so for instance um if you're doing knowledge gathering or research you definitely don't want anything made up there's like almost zero tolerance for um hallucination and so you kind of want a gradient or a parameter that you want to set uh which might be the creativity parameter um and i think that's becoming increasingly possible another another really good way to get models to be more truthful is to actually force them to cite their work so there's Patrick Lewis he was the first author at Meta on creating RAG, it's called Retrieval Augmented Generation and so that's this idea that you have a model and you have an external knowledge base or maybe multiple external knowledge bases maybe one's Google, one's your private emails, one's your blah blah blah and what the model can do is it can go out and it can query these sources so it can say hey the user just asked me about this.
50:44I think I should query Google. And then it gets back from Google some documents or it gets back from your email, whatever emails you're looking for. And then now that it can read those, it can generate a response and it can cite back to them. It can say, hey, you asked me this. I think this is the answer because of this sentence inside of this document or this web page. by forcing the model to learn to cite its sources you get two things one is the fact that you can actually check it's um you can verify it right you can check that it's telling the truth you click into that link you can read the thing and you can say it lied or you can say oh no it's right yeah you know that checks out um so one is you get it to cite its sources uh the other thing is that you force a behavior, you reinforce a behavior into the model of not making claims without grounds to those claims.
51:46And so it starts to learn these scenarios where, you know, when I'm writing stories, I don't really need to cite sources. I just need to write and the user is happy and content and, you know, I get a good reward. And in the scenario of, hey, I'm doing research on a topic, can you tell me about X? It starts to learn, okay, shit, in this case, I need to have a very rigorous bibliography. I need to be able to tie that back. And if I mess up, if the user clicks through and sees an error or a hallucination, I'm going to get a super strong negative feedback. And so it learns to differentiate between these scenarios.
52:25So I really do believe retrieval augmentation is going to be one of the key pieces of along with human feedback. It's going to be one of the key pieces of making these models more reliable, more grounded. That's fascinating. I'm coming up to an hour. Can I ask a few more questions? Yeah, go ahead. Yeah. I've got to ask this. You know, this has set off sort of the public release of ChatGPT has set off this debate about how dangerous these models can be. To everyone's surprise, Jeff has gone public saying some really dire things, which I don't know him like you do, but I've known him for a while, and it surprises me.
53:26I've never heard him speak that darkly about something. Do you have a view on that? That's one question. And then the other is this debate about sentience or self-awareness. I mean, you had your fingers in the brain of these things. do you think that sentience or self-awareness could really emerge? Or do you think that these are bits of code and it's all an illusion? There's a lot to say. that we need we need another hour or two together to properly represent my beliefs around that question I think the first part for Jeff Jeff is like Jeff went through the same thing that I think many of us in the field went through where our timelines got pulled forward massively And so we thought we'd have models that could write compelling English in a few decades, and then suddenly it shows up a year later.
54:52And so that throws you into this state of shock and uncertainty, and you're quite caught off guard. And he's spoken about this, I think, publicly. The sentiment of surprise at progress and rate of change. I remember having conversations with him myself where both of us were kind of like, these people who talk about AGI, you know, what nonsense, ha ha ha. And this was back when models could barely spell.
55:30But then you kind of get surprised and shocked and your uncertainty blows up. And sometimes that can have the effect that, okay, anything's possible. Oh, my God. I was so far off on that that now I am shooting up my uncertainty across all. Anything could be possible. Super intelligent God. Okay, maybe that's even possible. So I think a lot of folks were all reckoning with that and recalibrating and adjusting our own timelines and understandings of progress and pace. Jeff is extraordinarily thoughtful, and he's been thinking about this since at least the beginning of Cohere, so at least the last three and a half years.
56:19He's been thinking about this very, very deeply.
56:25And so I think people should take him very seriously. I think there will be a lot of sensationalism and a lot of extrapolation from what he's saying. But if you actually listen to what he's actually saying, it's quite a measured, he's like, I'm highly uncertain about what can happen. And that means we should take this stuff seriously. Because we just don't have certain bounds. We don't have certainty around the future. And so we should be taking all the different possibilities quite seriously. Not saying that they're likely to happen, but just saying that they can't be ruled out yet. And so let's take them very seriously.
57:08I think there's a lot of journalistic takes and headlines and clickbait and nonsense. but if you actually listen to Jeff I think his take is quite measured and reasonable
57:24and actually I'd love to have you back on to talk at length about these things but on the idea of sentience or the illusion of sentience I mean you know more than almost anybody having built these models, both what they're capable of and what's behind their expressions. Do you think that, I mean, it's a philosophical question about what sentience or consciousness is, whether it's, you know, whether our consciousness is just an emergent property from the neural activities of our brain and it's largely illusion. I mean, just what would you say to all of that? Yeah, I would say I don't place like a divinity on humanity.
58:41I think that consciousness is in the brain and it is like a physical process. And it's maybe like maybe consciousness is what computing feels like, like what processing feels like. And if that's the case, it's really hard to argue that that same phenomena couldn't be present in silicon. I think you really have to, I think there has to be a leap, right, to say that the circuits in our brain, because they're human or because they're biological, have some sort of fundamental distinction. I think you really have to take a leap of faith there. And so if I'm just being pragmatic and reductive, again, we would need two hours to discuss this, I think, more completely.
59:44But I think it'd be really, just as a scientist, I think it'd be really hard for me to say there's no way these machines could become sentient. I just, I can't construct an argument that ends at that. Yeah. Yeah. Yeah. Yeah, well, let's leave it there. But can I get a promise that you'll come back in a few months and we can go deep on that subject? Yeah, I'd love to. Yeah. Okay. Yeah, Aidan, this has been really fascinating. I'm delighted. And I'm sure you heard at the MIT Tech Review Conference, somebody asked Jeff, he was on virtually from the UK, but somebody asked him whether he would divest himself of Coher.
1:00:48and he said no no he's he's gonna stay invested so yeah yeah that's a that's a funny question um yeah yeah okay great well i really appreciate your time and uh and we'll talk again that's it for this episode i want to thank aiden for his time i also want to remind you to check out net Suite, Oracle's business management software for enterprise resource planning, financial management, customer relationship management, and e-commerce, among other things. Go to netsuite.com, that's N-E-T-S-U-I-T-E dot C-O-M backslash I on A-I, E-Y-E-O-N-A-I, all run together to take advantage of this offer. And remember, the singularity may not be near, but AI is about to change your world, so pay attention.
From the publisher
Welcome to Eye on AI, the podcast that keeps you informed about the latest trends, obstacles, and possibilities in the realm of artificial intelligence. In this episode, we have the privilege of engaging in a thought-provoking discussion with Aidan Gomez, an exceptional AI developer and co-founder of Cohere.
Aidan's passion lies in enhancing the efficiency of massive neural networks and effectively deploying them in the real world. Drawing from his vast experience, which includes leading a team of researchers at For.ai and conducting groundbreaking research at Google Brain, Aidan provides us with unique insights and anecdotes that shed light on the AI landscape.
During our conversation, Aidan explains his collaboration with the legendary Geoffrey Hinton and their remarkable project at Google Brain. We delve into the intricate architecture of AI systems, demystifying the construction of the transformative transformer algorithm. Aidan generously shares his knowledge on the creation of attention within these models and the complexities of scaling such systems.
As we explore the fascinating domain of language models, Aidan discusses their learning process, bridging the gap between code and data. We uncover the immense potential of these models to suggest other large-scale counterparts. We gain invaluable insights into Aidan's journey as a co-founder of Cohere, an innovative platform revolutionizing the utilization of language technology.
Tune in to Eye on AI now to immerse yourself in a captivating conversation that will expand your understanding of this ever-develop field.
(00:00) Preview
(00:33) Introduction & sponsorship
(02:00) Aidan's background with machine learning & AI
(05:10) Geoffrey Hinton & Aidan Gomez working together
(07:55) Aidan Gomez & Google Brain's project
(12:53) Aidan's role in building AI architecture
(15:25) How the transformer algorithm is built
(18:25) How do you create attention?
(20:40) How do you scale the model?
(25:10) How language models learn from code and data
(29:55) Did you know the potential of the project?
(34:15) Can LLMs suggest other large models?
(36:45) How Aidan Gomez started Cohere
(41:10) How do people use Cohere?
(46:50) Examples of language technology models
(48:40) How Cohere handles hallucinations
(52:53) The dangers of AI
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI




