Towards Improved Transfer Learning with Hugo Larochelle - #631

29 May 2023 · 39 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes - The TWIML AI Podcast #631: Towards Improved Transfer Learning with Hugo Larochelle

Episode Overview

  • Host: Sam Charrington
  • Guest: Hugo Larochelle, Research Scientist at Google DeepMind
  • Focus: Transfer learning, deep learning models, and creation of the Transactions on Machine Learning Research journal.

Key Topics Discussed

Introduction to Hugo Larochelle

  • Education: Undergrad and PhD at the University of Montreal, worked in Yoshua Bengio's lab.
  • Career Path:
  • Early work on neural networks during the rise of deep learning.
  • Postdoc with Geoffrey Hinton.
  • Faculty position at the University of Sherbrooke, startup acquired by Twitter, and then transitioned to Google Brain.

Transfer Learning

  • Definition:
  • Broadly defined as adapting a pre-trained model to a different task or data distribution.
  • Distinction between pre-training (knowledge acquisition) and fine-tuning (knowledge mobilization).
  • Phases of Transfer Learning:
  • Neural Knowledge Acquisition: Pre-training models on extensive datasets to embed knowledge.
  • Neural Knowledge Mobilization: Extracting and leveraging knowledge from pre-trained models for specific tasks.

Fine-Tuning vs. Other Approaches

  • Fine-Tuning:
  • Traditional method where a model is fine-tuned for specific tasks but can be expensive in terms of compute and memory.
  • Alternative Methods:
  • Linear Probing: Using only the final classification layer without modifying the rest of the model.
  • Task-Specific Parameters: Adapting only certain parameters (e.g., batch normalization) to save on computation.
  • Head-to-Toe Method: Using a sparse linear classifier that can dynamically access features from various layers without retraining the entire model.

Current Research Directions

  • Neural Knowledge Mobilization:
  • Focus on efficient ways to adapt models without full fine-tuning.
  • Addressing issues with downstream tasks that may have limited labeled data.
  • NN Probes:
  • Probes used to investigate how features are used across different layers of a model.
  • Findings indicate that lower layers often help with tasks very different from those in the pre-training phase.

Application to Code Completion

  • Work on Code Using Language Models:
  • Research on using large language models (like Codex) for code completion.
  • Exploration of fetching contextual code from various files within a repository to improve completion accuracy.
  • Investigating how the language model handles out-of-distribution context.

Future Directions

  • Interest in applying transfer learning techniques to environmental problems, particularly using remote sensing and satellite imagery.
  • Leveraging AI for species recognition and environmental monitoring.

The Transactions on Machine Learning Research (TMLR)

  • Overview: A journal created to address shortcomings in current publishing in machine learning.
  • Goals:
  • Open review system and flexible submission timelines.
  • Separate the scientific validity of work from its excitement level.
  • Collaborating with conferences to feature TMLR papers.
  • Current Status:
  • Growing steadily with around 1,000 submissions per year.
  • Positive feedback from the research community.

Key Takeaways

  • Transfer learning is evolving beyond simple pre-training and fine-tuning into more sophisticated methods that consider the structure and features of deep learning models.
  • Recent advancements in NLP and applications such as code completion highlight the versatility and robustness of large language models.
  • Continued research in transfer learning can lead to significant advancements in various fields, including environmental science.

Closing Remarks

  • The conversation provided deep insights into the current state and future potential of transfer learning and its applications in various domains, underscoring the importance of adapting deep learning models in innovative ways.

---

For more details about this episode, visit [TWIML AI Podcast Episode 631](https://twimlai.com/go/631).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:07All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm excited to be joined by Hugo LaRochelle. Hugo is a research scientist at the recently rebranded, renamed Google DeepMind. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Hugo, welcome to the podcast. Thanks for having me. Really glad to be here with you. Same. We will be digging into, I imagine, a pretty broad range of topics, but we'll be focusing in on your particular area of research interest, which is transfer learning.

0:47But before we do that, why don't we take a moment to have you introduce yourself to our audience and tell us kind of how you came to work in ML. Sure. So I did an undergrad and a PhD at the University of Montreal where early on I had the chance to work in Yoshua Bengio's lab. And I had a vague interest in AI when I started my undergrad and started looking up some profs that were working in this area. And this is how I met Yoshua the first summer of my undergrad. I wanted to do an internship at his lab. He said, not quite yet. You don't have enough stats and math courses. So I waited one more year.

1:32And then the other year he took me in and I started learning about neural nets and machine learning. Since then, I've not looked back. And that sort of led me to kind of fortunately, I've been very privileged to kind of happen to be in a lab that was doing neural nets at a time that where these ideas were not popular in machine learning. And so to be working early for that period or be in a small set of people doing that kind of research. And then we started seeing the kind of deep learning explosion. So the term deep learning was more or less coined during my PhD. And I was lucky to do early contributions to this sort of resurgence of ideas in neural nets, like denoising autoencoders.

2:16And at the time, we're doing layer-wise pre-training using unsupervised learning. Also did some work on very early ideas related to zero-shot learning, going from task descriptions to a classifier. And then I had a chance to continue in that direction, go do a postdoc with Jeffrey Hinton, then got a faculty position at University of Sherbrooke. At the same time in Toronto, I met Jasper Snook, Ryan Adams, and others. And we did some work on Bayesian optimization, apply to machine learning to do hyperparameter tuning. And we turned that idea into a startup, which was acquired by Twitter. So then I left.

2:58This sort of happened while I was a professor at the University of Sherbrooke. So I left for Twitter, worked in Boston for a little bit, and eventually got the chance to start a Google Brain group in Montreal. I was looking forward to kind of come back to Canada and Montreal in particular, give opportunities to students and professionals in AI in Montreal to actually be in a big corporate lab doing state-of-the-art deep learning research. And then I've been here since then. It must have been summer when you had that opportunity to reconsider returning to Montreal. Well, as it turns out, I do enjoy winter.

3:39So that's something, though, that I think is an acquired taste, perhaps. But I do enjoy the changes of seasons when usually there's a first. Actually, for me, like the NeurIPS conference in particular, oftentimes when we would come back from NeurIPS, we would get the first snow. And it would also be like around Christmas. So it would always be kind of like a really nice moment where I just heard a bunch of exciting things at NeurIPS. and I know I'm going to get a lot of free time with the holidays to kind of think about these things, but also relax. It has its benefits, I will say. I will defend winter a little bit.

4:17Not AI winters, but winters in general. Nice. So your research interest has kind of centered around transfer learning of late. Do you feel like the accepted definition of transfer learning is kind of the obvious place to start? Or are there nuances that are underappreciated in the community? I mean, I think of transfer learning very broadly as like you have a model that's been trained on some set of data, maybe many tasks. And then from that, you do a follow-up training or adaptation to another task that's different in some ways. Maybe it's different because the input distribution, the type of data that it's seeing as input is different.

5:01Or maybe the task is completely different. If you're doing image classification, maybe you're classifying different properties about these images. And I think of it like as very, very broadly. I've seen sometimes in the literature, sometimes, you know, essentially when they say transfer learning, they really only mean pre-training and fine-tuning. which is like the simplest approach to transfer learning for sure, but it's not the only one. And I think there's some research to do in that respect, but I think of it very broadly. And I like to think about kind of there being two phases. And some of the talks that I've given recently, I sort of call the neural knowledge acquisition phase and the neural knowledge mobilization phase, thinking of it as the pre-training is sort of like us trying to put as much knowledge as we can into a base machine learning or neural net model, because neural nets are what we're pretty effective these days at solving tasks with.

6:04And I say neural knowledge because it's not explicit knowledge. It's not written facts. It's just there's some information in that pre-trained neural network that we can hopefully mobilize and use to solve downstream tasks. And then there's the mobilization or neural knowledge mobilization phase, which is about, okay, how do I use this large neural net that has some implicit knowledge to extract the right things that I need to solve a particular task as one user? And there might be many downstream users or many downstream tasks we want to solve. And I've been particularly interested in the second phase of this, largely because the first phase of knowledge acquisition, now we would pre-train very large models on lots and lots of data.

6:46This is not accessible to everyone anyways. They're compute intensive, so there should be a limited number of initiatives trying to do that. And in some ways, I feel I've contributed some ideas for this. And in my PhD days, for instance, we see denoising autoencoders, versions of this being used behind, you know, there was the BERT model and there are masked autoencoders now vision that are becoming popular and being used frequently. And so I was fortunate to be able to contribute ideas that I think now are used for that pre-training phase. And I've been more recently been interested in trying to see on the knowledge mobilization phase.

7:26So how do we take this pre-trained model and extract the right information from it to solve well some downstream tasks, in particular downstream tasks that perhaps don't have a lot of labeled data, so that we get essentially a procedure and technology that can be used by many people for their own type of task that maybe doesn't allow for large collection of data. And my more recent research has been more interested in that sort of second phase and trying to see whether we can do something better than fine-tuning. Logically, because fine-tuning has some downsides. If you fine-tune a model for any given downstream task, then for each of these tasks, you need a full copy of the weights of the model you fine-tuned.

8:07And that can be expensive if you want to support a lot of different downstream usages. It's also expensive from the point of view of compute. You need to do full forward and backward passes over all of the network. And as they're becoming larger, that's more and more expensive. And that's something you might want to avoid, at least in terms of training a downstream model. And in some ways, there's no particularly good reason why fine-tuning should be the best thing to do. To me, it's always been kind of a hack in some ways. I think the first time I remember the idea of pre-training and fine-tuning being used was maybe around 2006.

8:45I think it was just a tech report by Jeffrey Hinton, where at the time he had proposed deep belief networks. It was going to train a big multi-layer generative model on digits, a joint model of the digit image and the class ID. But because it was a deep model, it was hard to train at the time. You would have the compute, but also perhaps the ideas for training a large generative model like that. And so he proposed pre-training individually each layer using the learning algorithm of a restricted Boltzmann machine, and a type of generative model and doing it one layer after another. And he had this whole theoretical sort of reason why that was a good idea.

9:26It was optimizing some bound over the objective of the final deep belief network you wanted to train. But then in this parallel tech report, it was just like, well, you can just do that. This form of unsupervised learning for a normal feedforward neural net, add a soft max layer at the end and then just fine tune from that. And that's also just about just as good. in terms of classification performance at the time on MNIST. So it was originally thought of as more of a kind of an initialization approach as opposed to actually trying to achieve transfer learning? Well, I think there's an idea that doing unsupervised learning is going to capture something more informative about the structure of the data.

10:08And those features were essentially basically good, and they just needed to be adjusted a little bit so that on the task you were interested in. And in this case, it wasn't even transfer learning. It was actually, you were pre-training on MNIST and you were fine-tuning on MNIST. I think the idea of then fine-tuning on a different distribution after that, I wouldn't be able to point out would be the first paper. And maybe before Jeff Hinton's tech report, the idea of pre-training and fine-tuning had come before. But to me, that's like the thing that I have in mind as like the first idea of separating into these two phases.

10:40But even in that tech report, it felt like it was presented as an intuitively good idea. but not really a theoretically grounded idea. And so I think to me, that means that, you know, there are other ways of taking this neural net that contains basic features and leveraging that information in neural form to solve a downstream task. And so I've been interested in trying to explore that in various ways. And how do you characterize the research landscape around this idea that there's more than just fine tuning to transfer learning? Where are we in terms of the way we think about it? Yeah, so I think for a long time, you were either fine-tuning or you were doing this thing called linear probe, which is essentially just string the softmax layer at the top and not fine-tune the rest of the layers, keep those constant.

11:25So think of just these features that are already being correct. And then slowly people started thinking, okay, well, maybe there's something in between. Maybe you want to do some fine-tuning of the network, but maybe just a subset of these parameters. Maybe you only want to fine-tune the biases of the units, where the intuition there would be, maybe you sort of identify a sub-network within the pre-trained model that is essentially effective because you can turn off units essentially by having very large negative biases, such that no matter what their input is, they're going to stay, if you're using ReLU, say they're going to stay zero.

12:05So some of my work with people here in my group, LNE, Triantafilu in particular, we've kind of explored this idea of having kind of a, we'd call that a template, which would be the pre-trained model. And then we would fill the template, essentially the batch norm parameters that it would be using. So the bias and the scaling parameters in each layer. And think of solving different tasks as just swapping different batch norm parameters. And that has the advantage that when you're, say, adapting to a downstream task, you only need in terms of memory to memorize for that task, what are the new batch known parameters and what are the output weights?

12:41You don't need all of the other parameters because they're the same from the pre-trained model. So that allows you to have a lot of potentially thousand downstream tasks where you essentially think of the pre-trained model as a template and each downstream task fills in the template with these task-specific batch known parameters. But that also has a disadvantage, which is that when you fine tune these batch norm parameters or whatever subset of parameters in the pre-trained backbone that you end up training for the task, you still need to do a full forward, a backward pass whenever you're training.

13:15Again, if the pre-trained model is very large, that can be expensive. And that led to some more research of trying to avoid that, where the idea there with a method that we call head to toe, and that was led by Utku Efchi in my group as well. was to also think of a pre-trained model as essentially having all the right features, but that essentially the features, some of them are just not at the very top of the model. They might be hidden somewhere within the model. And maybe all that fine tuning is doing is kind of emerging, making sure that this information surfaces all the way to the top. But then another way of accessing that information might just be, I'm going to have a linear probe, but that has access to all the layers.

13:58So it has these kind of skip connections into intermediate layers of the pre-trained model. But that's a lot of features. So to try to address that, we would also enforce some sparsity in the connectivity of that linear probe. So it could maybe have no connections with the second layer, a few in the third layer, no connections in the fourth and so on, and kind of essentially infer which connections it wants to keep and which features it wants to use for that downstream task itself by essentially training a sparse linear classifier on all of these features, which was kind of the next iteration in this research agenda of trying to figure out other ways of doing this neural knowledge mobilization, but by avoiding full fine tuning.

14:39And we've shown that in some ways, indeed, we're able to match the performance of fine tuning by doing that. But presumably much more computationally efficient? That's right. Yeah, I think we're taking less than 10 % of flops in terms of training time because of the sparsity of the classifier. And again, also because the probe is sparse, the number of parameters it would have would also be below 10%. It would take much less memory to store for this particular downstream task you'd be solving. And when you talk about these probes, the problem is still classification in all cases, or are there other transfer applications where you can utilize these skip connections and some of this architecture that you propose?

15:21Yeah, we focus on classification, a lot of it out of convenience. It's a task that a lot of people are interested in solving. There's nothing for head to toe that would preclude us from tackling other tasks, but partly out of convenience and also because it is a very popular task. We've mostly focused on that, but I don't see a reason why it wouldn't be applicable to regression. It would be interesting to see if it could be applicable to things like semantic segmentation or other things like that. So this points to potentially interesting follow-up work for sure. And here, you can help me kind of put together this question, but you've trained the base model, you've kind of learned these kind of linear pieces.

16:05Are there other things that you can do with them beyond the tasks that you initially trained on or the tasks that you tuned on? Are they useful like embeddings in a general way? I guess that's the picture that came to mind for me. Potentially. I think one thing that I thought that I really liked about this project that I thought just doing that research actually helped us bring some insights into what is being learned by these models was to kind of look at for different types of tasks, where is it that it's pulling features? And I wouldn't go as far as saying like it would help with interpretability or anything like that.

16:40But I think to me, doing this project was also about testing a particular theory as to how is information captured by these models? Is it completely black box? Or are there ways in which these layers are separating out information in a more semantic or somewhat conceptual level? And I think the answer there was a bit mixed in that, interestingly, I think we found often that the first hidden layer, the one that's closest to the input, was fairly frequently useful. And it was particularly useful, if I remember correctly, when the downstream task that we were solving was very different from the pre-training tasks.

17:26So to give more details on the experimental setup, there we would use a pre-trained model in ImageNet and then we'd look at the number of other classification tasks. In particular, we looked at the VTAP benchmark, which comes from colleagues at Google in Europe. And I think it's called the VTAB stands for Visual Task Adaptation Benchmark and includes a number of different tasks. Some of them are synthetic. Some of them are natural images. Some of them are sort of very different from natural images, but still real tasks like doing some form of predictions based on satellite imagery or medical imagery.

18:02And what we found is that these intermediate layers are particularly useful when the downstream task is different. So essentially, intuitively, if you're going from ImageNet to, I'm trying to remember which task was considered very similar, but I think something like classifying images of pets. There are a lot of animals in ImageNet, but there's this other data set, this pets data set. And indeed, we found that you didn't need as much internal features. The top embedding was pretty good already for solving this pets downstream task. But if you look at, say, medical imagery or other synthetic data set are very different, then accessing these very early layers was useful.

18:41And we couldn't otherwise see better organized structure, except for essentially the first hidden layer, which were probably going to be more or less edge detectors, because we know we learned that from training on ImageNet. That was often very useful in adapting to other downstream tasks. but then the rest of it was a little bit more diffuse or it was hard to sort of determine from one task to another what kind of structure there was. There were definitely more useful often or complementary to the topmost layer. And so irrespective of the success of the method, I think probes in the past have been used and they're called probes partly because the idea was to try to probe the information from these deep networks and access it more directly within the layer.

19:26I think that's roughly where the term probe comes from. And in addition to being a successful method, if we start exposing all layers, for us, it's been interesting because it also was allowing us to kind of probe, okay, how is information distributed across an ImageNet pre-trained model? You can imagine doing that, actually, sort of taking a pre-trained model, taking a number of downstream tasks you're interested in, and using this head-to-toe method to try to maybe infer perhaps how high-level these features are for different downstream tasks. and maybe characterizing a little bit more, like where is PETS information distributed within this model versus other types of semantic categories or other type of concepts that are not from natural images.

20:09And so I think that's a nice aspect of this project that I've enjoyed. Got it. So we've talked about a couple of broad categories or approaches to transfer learning in the research. One is kind of the fine tuning, chop off layers, starting from the head and work your way back and fine-tune. The other is kind of keep things relatively static and access internals via probes are there. Is the research dominated by approaches that fall into one of those two categories or is it broader than that? Yeah. So I think for computer vision right now, this is a fairly common approach to transfer learning. And the area of NLP and notably with the rise of large language models as a way of capturing a lot of background information about language, about code also, and many more things.

21:02Another approach that we've seen has been very popular has been prompting, essentially including information in the context of the large language model that would describe the task that you're interested in solving now, leveraging all of the implicit knowledge and information captured by the large language model. People have looked at doing literal prompt where you're actually putting text. The text might describe different pairs of input-output examples for the downstream task you're interested in. There's the notion of soft prompt where you're actually doing backprop into a free vector or a couple of free vectors that are going to be fine-tuned on your small training set in your downstream task.

21:42And that has been also quite interesting to see evolve as sort of an approach to doing forms of transfer learning. I mentioned I did some early work during my PhD on zero-shot learning. The method we call zero-data learning was the term that we used. Because at the time, as far as I knew, zero-shot learning was not a term that people had coined. And I think Yoshua Bengio at the time suggested that. And to this day, I regret using this term because zero-shot learning was clearly the better term. people kept asking us, what does that mean, zero data learning? Because it makes no sense. You're learning from no data.

22:17I could see why people were confused. And not so long after our paper, in fact, there was a zero shot learning paper. And that was essentially trying to use the same idea, which is in zero shot learning, you're trying to essentially use, let's call it metadata or data about the task that might be a worded description of the task and have a model that can take that as input and then give you a predictor to solve that particular task. We've seen that being enabled by these large language models. So for me, it's been really exciting to see kind of this idea sort of being now used in a much larger scale and much more successfully.

22:57And one ways in which I've been interested in trying to explore that in a more modern context. So doing a form of this neural knowledge mobilization that I've been talking about, but using prompting has been in the context of machine learning models for code. So essentially completing a line of code based on code that comes before it. It's been a huge application for these kinds of models. Me and Danny Tarlow in my group, I've been quite interested in that. And we've been advising students, particularly in this area. The upcoming ICML, there's a paper we'll be presenting on trying to do this form of prompting, but trying to see whether we can use, and this is work led by Disha Srivastava, where what we wanted to see is that, so large language models typically are trained, and this is a project we started some time ago, and they will have a fairly limited context.

23:50And the context will just be the previous tokens before, say, where your cursor is, where you want to start doing some completion of the code. It would typically only be the preceding code because it's kind of a quote-unquote causal model that sort of predicts one tokens based on the previous ones. But really in practice, when you're writing code, often you're maybe in the middle of the file or your file of your project is part of a repository. So there are other files and other directories in the repository, which might include functions you want to call, classes you want to instantiate. And so our question was, well, even though if we have a pre-trained model that's only been trained in a causal way within single text, just predicting the next token, we know that there's these kind of emerging properties, as we've called them, where you can still provide sort of general task descriptions, input-output pairs in the context.

24:47And somehow it's learned to do this kind of zero shot or few shot generalization. And so the question we wanted to see is that, would it actually be able to take snippets of code that comes from other files that is not the current file and still kind of figure out how to leverage that information to do better code completion? And in a way, I wasn't very optimistic going in because I thought, well, we're clearly going to be constructing contexts that are nothing like what the large language model, in our case, we use Codex for the project. Can I jump in and have you clarify the novelty here? Codex, for example, part of what it's doing is it's pulling context from different places.

25:29And as they've evolved Codex, sorry, Copilot in particular, like they have talked about how one of the big challenges that they've taken on is like where they get the context from exactly. And what you've done in this paper? I think at the time when we started this project, I don't think Copilot was doing that or it wasn't known if it was doing that. And so I think we sort of did that about the same time. And so that's why we're using Codex and Codex. But it's the same general idea here. That's right. Yeah. Okay. And this was like when we've done this project for some time now and only now managed to get it published.

26:07It's been on archive for a bit longer. That's how quickly the field evolves. But for me, what was really interesting also about this project is trying to see what kind of capabilities these large language models have, that it goes outside the task that it's been trained to solve. Therefore, when we were putting in context that came from outside, actually the context that the language model would see would not even be code that actually compiles necessarily, because we're kind of putting in code that comes from elsewhere. And we're just trying to see, is it still capable of leveraging that? And also one of the questions was, well, how do we decide what code from where to put in?

26:45And so in this work that we did, it's going to appear at ICML. The way we address that is that we would define a number of different templates for where we could take code. It could be from the same file, but later in the file, it could be the function definitions of sibling files in the same directory. It could be the variables and the file that we were importing in this current file. And essentially trying all potential options on the small training set and sort of seeing which ones and what context tended to work and allow codecs to come up with the right answer in terms of completion. And then trying another neural net that would essentially predict the success of each of these different templates.

27:31Yeah, I think early on with Copilot, it was more rules-based or heuristics-based where they pulled code from to do prompt generation. But what you're saying here is this is learning which code is going to be most helpful to the specific generation. Well, you say repository level. Are you learning this at the repo level or at the query level or in the context of a particular query? So it's repository level because we'll, for a given context, pull in information from anywhere in the repository, but which template you invoke, so which prompt you construct, will depend on the context. So if you were doing completion multiple places in a given file, you might be pulling different contexts for different lines in the file, for instance.

Read the full transcript

28:18So it would be sort of adaptive to the particular situation you're trying to extract information. So it'd be more adaptive. But yeah, I think more or less we're exploring this exactly. I mean, when we came up with the idea, part of me was excited because I thought that seems like it has a good potential. At the same time, I was like, I'm kind of not sure if it's going to work because we're really going to be creating context that will be completely out of distribution with respect to what the language model would have ever seen. Because the context we're constructing, they don't run the way we sort of put them together.

28:51We were thinking maybe we'll need to be clever and be like, add the pulled in context as comments in the code as opposed to just putting the code there. But we didn't even have to be clever in that way. We could just put in the code and the language model was sort of oftentimes be able to leverage it better than if it didn't have any context. And just to make sure I understand that fear, the thought was that you'd be pulling in these snippets that were taken out of context and not complete like the kind of data the LLM was trained on, like entire functions that could run somewhere. And the thought was maybe it would get confused and start generating gibberish as a result.

29:32Well, yeah, that's right. Essentially that context, the way it'd be, because it'd be kind of like glued in snippets of different parts of different files would look nothing like code that someone would actually write. And so would that confuse the large language model or would it still kind of be able to overcome that out of distribution-ness of what it's seeing? And so again, like to me, a lot of these projects are also about trying to uncover interesting properties of these various ways we can pre-train large models from lots of data. And here it was kind of interesting to see that it had some amount of robustness.

30:10It's more than robustness. It actually can, if you put in better information, it will actually leverage it. An interesting other dimension that I sort of put into my general interest in transfer learning. And did you need to tell it that, hey, these are going to be different than what you might be expecting? You know, these are just snippets, that kind of thing. Like, did you explore whether that was an important part of the prompt construction? So yeah, we didn't have to do anything special, which was surprising to me. So this is why I was saying, I thought maybe these snippets, we'll need to put them as comments in the actual code so that it knows you're not running these.

30:44They're just there and you can take from it. But yeah, no, we didn't even have to do that. So surprising property that these models have. And kind of looking forward, you mentioned the rise of LLMs and generative AIs, kind of all the buzz now. How do you see that broadly impacting the direction of your research? And are there other interesting questions that arise for you in kind of the intersection of transfer and LLMs? Yeah, I would say, I mean, I find the way that this area is evolving very fascinating. I will say I naturally have a tendency to try to avoid problems that a lot of people are working on.

31:27I just find that less motivating. I mean, even for this project, initially we thought no one's going to be crazy enough to try this. And then later on, we learned our co-pilot is kind of doing apparently a version of this. And we were in some way scooped. And so we're really glad actually that it got accepted at ICML this time, because I think otherwise we're like, people might just think it's old news at this point. And so moving forward, instead of going in NLP, one thing that I've been interested in is focusing more on vision and in particular, focusing on remote sensing domains. and I've been really interested in trying to tackle environmental problems.

32:08So doing things from satellite imagery, making some predictions about what might be species that are present and potentially using that as an analysis tool for analyzing, getting a sense of how the state of the environment is evolving, maybe using that to drive some policy if we have good tools. And so my student, Mirizan Tang, just presented an early workshop work at iClear, where we presented a data set for doing that from extraction from eBird, where you take some satellite imagery and you predict the probability of observing various types of species. It was presented at the AI and Climate Change workshop at iClear, and we're lucky enough that it was given a paper award at the workshop.

32:51So we're really proud of that, which encourages us to continue forward in this area. I think this is an area that's a bit overlooked. I think we can still study very interesting transfer learning problems because we need to generalize to the future or we need to generalize from locations where we don't have labeled data because labeled data about various problems are not necessarily collected equally well all around the world. So sort of looking at a problem, people don't look into practically all that much, but I sort of believe as we need people doing these kinds of problems. And I think at the same time, we'll be able to do some cool research.

33:25So I think this is one thing in particular I'm excited about. Awesome. Maybe switching gears a little bit, we're about a year into the TMLR experiment. Did you think of it as an experiment or the TMLR being the machine learning research journal that you launched with others? Yeah, Transactional Machine Learning Research. Yep. In many ways, it's an experiment. Obviously, I think we're all obeying Kyungyeon Cho and Raya Hatzel, the other editors-in-chief. we had as initial managing editor, Fabian Pedregosa, it was obvious we're doing something new, certainly for machine learning at a minimum. So for context, the idea here was to try to address a lot of things we didn't like about or a whole that we were seeing in the publishing ecosystem in machine learning.

34:11So people were either publishing at conferences, shorter versions of their papers with only very few deadlines over the year. And the role of a conference is, and I think that this is something that people have some difficulty with, the role of a conference, some people think it's about identifying what are the papers that are correct science. But actually, a conference does more than this. It's certainly wanting to only accept correct science, but it's also trying to identify what does the community think is exciting right now and most worthy of attention. And that second element, that's very much an editorial or subjective assessment.

34:52And I think a lot of the noise in the accept and reject decisions at conferences partly come because of that, because of that nature of what conferences are trying to achieve. So that few deadlines, which I think also makes working as a researcher in machine learning sort of pretty stressful, it seems cramming for the next deadline. Everyone runs out of compute at the last minute. That's right. Also, yes. So we thought having a journal that would be appealing for these kinds of short form publication might be good. Also, we wanted to use the open review system. So doing open reviewing, there were no journals in machine learning that were doing that.

35:28So JMLR is a very well-established journal and TMLR is part of the JMLR family. It's closed reviewing, so it's not open reviewing. It's also usually longer form papers. So we kind of felt that was like an opportunity to contribute something unique. And our approach to reviewing would be that we would only assess whether the claims made in the paper are matched by convincing evidence. And then we would treat any signal related to how exciting the work is as something separate. And for that, in the journal, we think of this as certifications that people would give as like, oh, this should be a featured certification, meaning that it's probably like spotlight or oral sort of excitement level at a conference.

36:10We also now that we've been through this for a year, we're starting to talk with conferences who might want to feature some TMLR papers at their conference because they find the work exciting and they only have to do this assessment. They can sort of be fairly confident that scientifically speaking, the work is solid and that we could give certifications associated with any given conference. So we have our first agreement or experiment with AutoML and also with COLAS now, two conferences, smaller conferences we'll work with to have these event certifications. And for me, the dream is to kind of reach this point where TMLR is where people submit their work at any time, only when it's ready.

36:50And then after that, you start shopping around different kind of venues that might want to feature this work and give it some spotlight. And then we'd have an ecosystem of different certifications coming from different communities. And this is very close to kind of a vision that Jan Lecker had for reinventing, reviewing many years ago that has been a big inspiration for this. And so far it's been going well. We're getting about, I think at this point, our rate is about a thousand submission a year, I think. It's probably going to get bigger with time, but it's a good thing that's somewhat slowly increasing.

37:23And I found that to be pretty rewarding. It's a lot of work, but oftentimes we get really good feedback from people so far about the quality of the experience. And so, yeah, I think maybe in some ways it's still an experiment, but I would say it's less an experiment than it was a year ago and we feel on more solid grounds. That's awesome. Hugo, it was great catching up with you and getting an update on your research. Thanks for taking the time and joining us. Yeah, thank you so much for having me. I really enjoyed this. Thank you. Thank you. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com.

38:05Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.

From the publisher

Today we’re joined by Hugo Larochelle, a research scientist at Google Deepmind. In our conversation with Hugo, we discuss his work on transfer learning, understanding the capabilities of deep learning models, and creating the Transactions on Machine Learning Research journal. We explore the use of large language models in NLP, prompting, and zero-shot learning. Hugo also shares insights from his research on neural knowledge mobilization for code completion and discusses the adaptive prompts used in their system. 

The complete show notes for this episode can be found at twimlai.com/go/631.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Towards Improved Transfer Learning with Hugo Larochelle - #631The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 39 min
Listen in VO