Why Models Are AI’s Next Training Dataset with Damian Borth - #772

27 Jul 2026 · 47 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Damian Borth explains “weight-based learning,” treating trained neural network weights as an input modality to learn a latent “manifold” of models. Instead of training only from data, the method trains an autoencoder on many existing models’ weights to (1) predict properties like accuracy/generalization gap without test data and (2) generate new weights for new models/tasks, potentially speeding training and reducing reliance on scarce high-quality data.

Guest

Damian Borth is a professor of AI and machine learning at the University of St. Gallen. He works on weight-based learning, plus remote sensing and representation learning on tabular data.

Key claims

weights can be compressed into embeddings; those embeddings can predict performance; decoding can generate weights but initially produced “blurry” (low-frequency) details, requiring fixes like windowed reconstruction and scaling.

Notable examples

early results on small “model zoos” (e.g., predicting accuracy on Fashion-MNIST); later scaling to Hugging Face model repositories; a remote-sensing paper aiming for knowledge transfer and large compute savings (claimed ~12,000 GPU hours vs ~350).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring Weights-Based Learning

0:45 to 1:46

Discussion on the concept of using model weights for further learning.

“So we basically thought about this very simple idea.”

Understanding Neural Network Weights

1:46 to 4:39

Detailed explanation on treating weights as an input modality for neural networks.

“And weight-based learning is a quite interesting way of looking at machine learning in general.”

Fingerprinting Neural Networks

4:39 to 6:16

Introduction to the idea of versioning and analyzing neural networks' weights.

“but thinking about that, that you can treat the weights as an input modality gives you suddenly this opportunity of thinking about, okay, what would be language translation with more like neural network models, right?”

Predicting Neural Network Performance

6:16 to 7:39

The process and results of predicting the accuracy of neural networks using weight analysis.

“if you do one update of weights during training, every weight is a little bit different.”

Scaling Up Weights-Based Learning

7:39 to 10:40

Discussion on the challenges and scaling of weights-based learning models.

“And maybe this manifold encodes information about accuracy, what training data was used, what training fraction, learning rate, and all these latent generating factors.”

Collaboration and Community in AI Research

10:40 to 14:00

Insights into the collaborative efforts and community engagement in research.

“They were kind of telling us it's small, but it's interesting.”

Collaborative Research and Community Building

14:00 to 15:28

Learn how collaboration and community engagement shaped early research ideas.

“So, and then, you know, because he said that, I asked him that you have to help to scale that up, right?”

Weight Space and Neural Network Functionality

15:28 to 17:11

Discover how weight space influences neural network functionality and generalization.

“And 2025, we had then the first workshop.”

Exploring Loss Landscapes and Learning

17:11 to 19:25

Gain insights into how loss landscapes affect model learning and performance.

“So it was, you can borrow ideas from other fields and it was an empty field to fill with content.”

Scaling Models and Training Diversity

19:25 to 21:52

Understand the importance of scaling models and incorporating diverse datasets in training.

“So Charles and Charles co-auto is Michael Mahoney, who was then co-auto in our paper.”
Show all 24 chapters

Challenges of Model Adaptation and Training

21:52 to 23:38

Explore the challenges faced in adapting and training diverse neural network models.

“can we train on models that are out there?”

Innovative Tokenization and Normalization Techniques

23:38 to 27:36

Learn about the innovative techniques used for tokenization and normalization in model training.

“So we want to have, you know, different data sets.”

The Vision of Foundation Models in AI

27:36 to 28:00

Discover the future vision of foundation models and their potential impact on AI training.

Foundation Models and Knowledge Transfer

28:00 to 29:38

Learn how foundation models improve the training of remote sensing models using knowledge transfer.

“You just sample the model that you need.”

Sampling Models vs. Data

29:38 to 31:04

Explore why leveraging existing models might be more efficient than training on large datasets.

“That's the reason why the scaling laws are a little bit, you know, considered differently and everybody is moving into test time, adaptation test time training.”

Data Set Prompts and Privacy Preservation

31:04 to 32:38

Understand the concept of dataset prompts and their potential for privacy-preserving model training.

“While with our machinery, with our waste-based learning approach, you could sample a big model, you could sample a rest net, you could sample an efficiency net, depending on what you need.”

Embedding for Efficient Model Training

32:38 to 34:24

Discover how dataset embeddings could facilitate faster and more efficient model training without exposing data.

“without revealing individual members of this data set.”

Potential of On-Demand Neural Networks

34:24 to 36:06

Learn about the future of generating on-demand neural networks and hyper-personalization in AI.

“And you need to aggregate this knowledge.”

Weights and Architecture Mapping

36:06 to 37:37

Gain insights into how generated weights fit into various architectures and the significance of architecture conditioning.

“where you could have on-demand neural networks, on-demand hyper-personalization in the forward path.”

Task Vectors and Fine-Tuning

37:37 to 39:44

Understand how task vectors are used to simplify model fine-tuning and their implications in AI.

“There are definitely more clever ways of building a tokenizer that kind of understands what type of architecture elements are processed.”

High and Low Frequency Noise in Weights

39:44 to 41:17

Explore the concepts of high and low frequency noise in relation to weights and potential research avenues.

“And that's probably also one of the reasons why we still need a couple of fine-tuning steps to get the model very quickly up to performance.”

Exploring Weight Processing in AI

42:01 to 43:38

Learn about the potential for processing neural network weights in different domains.

“Empirically, we observed it, but we cannot explain that way.”

Community and Collaborative Development in AI

43:38 to 45:32

Discover the importance of community and shared language in advancing AI research.

“There's another great workshop that is organized by some of our colleagues at ICML.”

Innovative Approaches to Neural Network Analysis

45:32 to 46:10

Understand new methods for analyzing and manipulating neural networks through gradients and activations.

“Anthropik's been doing a lot of work along those lines as well.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Sam Charrington:One of the biggest questions facing AI today is how foundation models keep improving as high-quality training data becomes harder to find. Some researchers are betting on synthetic data, others on inference time reasoning. Today's guest has his chips on something very different. Every trained model represents thousands or even millions of GPU hours spent discovering what works. Instead of treating those weights just as the end of the training process, what if they're also the beginning of the next one? Damian Borth, professor of AI and machine learning at the University of St. Gallen, sees trained models themselves as data, data that can be learned from, analyzed, and even used to generate entirely new models.

0:42Sam Charrington:When I asked him to explain the idea behind weights-based learning, here's where he started. So we basically thought about this very simple idea. What happens actually if we take the weights of trained neural networks as the input to train a neural network to understand these weights that we have out there much, much better. Thinking about that, that you can treat the weights as an input modality gives you suddenly this opportunity of, can we be much, much faster in creating new weights for particular tasks? Or can we be much more precise in analyzing weights when somebody gives me a neural network that I'm not knowledgeable about and I never saw before?

1:23Sam Charrington:I'm Sam Charrington, and this is the TwiML AI podcast. For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.

1:45We started, like in 2000, 2021, the work on what we call weight-based learning. And weight-based learning is a quite interesting way of looking at machine learning in general. That's currently the major topic. We also do a little bit of work in remote sensing and then representation learning on tabular data. We're now walking and combining everything together to focus more on weight-based learning, which I think is a really interesting way forward, solves a couple of problems that the community currently encounters. and started from a very esoteric idea to something that works surprisingly well.

2:30Sam Charrington:You know, we think about weights as the product of training a model and, you know, we get some utility out of them. We may use them for things like explainability or manipulate them when we're quantizing or something like that. But the idea seems to be that, you know, there's so much more that we can learn from these weights. Exactly. So if you think about machine learning, machine learning has this idea of you have data and some output in classical supervised machine learning data and some predictions. And you train neural network in between to mimic the data set, mimic the distribution of the data set.

3:09And the outcome during this very expensive training procedure is a set of weights, a configuration of parameters that define the neural network, like the DNA of the neural network. This is classical machine learning, supervised, unsupervised, self-supervised, that fuels a lot of innovation over the last 10 years and with Gen.AI, you know, moved to the next stage. If you look at what happened over the last couple of years, more and more of those models have been published publicly, are online accessible at repositories like Hugging Face or GitHub. So we basically thought about this very simple idea.

3:52What happens actually if we take the weights of trained neural networks as the input to train a neural network to understand these weights that we have out there much, much better? So to take another analogy, you know, language models, you take a big model, you train this on every single sentence on the internet. At the end, you have a language model able to analyze language and to generate language. You can take the same idea for pixels. You take a big model, you train on all the pixels on the internet, and you can analyze pixels and you can generate pixels. We do the same idea on all the weights of trained neural networks, so we can analyze weights of neural networks and we can generate weights of neural networks.

4:35As straightforward as it is, obviously there's a little bit more into the details, but thinking about that, that you can treat the weights as an input modality gives you suddenly this opportunity of thinking about, okay, what would be language translation with more like neural network models, right? What would be generation of words and tokens that are words and generation of tokens that are weights? And can we be much, much faster in creating new weights for particular tasks? Or can we be much more precise in analyzing weights when somebody gives me a neural network that I'm not knowledgeable about and I never saw before?

5:13And then, you know, you have this new entire world, this, you know, empty space of things you can do with weights that, you know, you kind of carry into the community and hope that there's somebody listening and continuing and, you know, building up a community which happened over the last two, three years, which is very exciting because there are more people about that. And yeah, weights are exciting, not only as the output of learning, but as the input for learning.

5:41Sam Charrington:You mentioned you started this effort in 2021. Where did you start from? And then we'll work our way towards like where we are now with weights-based learning. Originally, this idea came 2020 and we got the first paper published in 2021. And the idea was very simple. Can we fingerprint a neural network or version a neural network like we can do with software, right? In software, you can do a diff, right? You have a 1 million lines of code. Somebody changes something and then, you know, you do a diff. you know exactly where the difference is. So can we do this with neural networks? Problem with neural networks is if you do one update of weights during training, every weight is a little bit different.

6:21So there's not much you can extract from this. Right, they're fairly unstable locally. If everything is different, nothing is different, right? So we were thinking about, can we find a space where these neural networks, the weights, the models of those neural networks, are a little bit compressed and more understandable. And we started to think about that. And in parallel, there was this really amazing work, the first work from Thomas Untertinger and Daniel Kaisers and colleagues from Google Zurich, and they developed a paper that used weights as input, did some statistical features and crafted features, and then predicted the accuracy of those weights.

7:07And another paper by Ileson that predicted the generalization gap. So people started to use weights to extract information and they were all handcrafted features. So I was thinking about handcrafted features that sold traditional machine learning. So why not end-to-end learning? So then we developed our idea of auto-encoding sequences of weights into a lower dimensional space and then reconstructing it. And if we're able to do this from a population of neural networks, then we can maybe learn a lower dimension manifold that populates actually, where the neural networks populate that manifold. And maybe this manifold encodes information about accuracy, what training data was used, what training fraction, learning rate, and all these latent generating factors.

7:55And we started working on this, and the first paper was really like, well, it works. We can compress neural networks, like very small, tiny, toy examples, right? Embarrassingly small, like thousands, 10 ,000s of parameters. But it was working and we could predict the accuracies and it was very nice to see that. And then, you know, we got published. What exactly were you able to predict? So we took, you know, an autoencoder. We have an encoder and decoder. We learned the autoencoder with, or trained the autoencoder reconstruction loss, a contrastive loss in the middle. And then we took only the encoder and unknown neural networks that we encoded into the latent space.

8:40And these embeddings we put into a simple regression, like a linear regression head to predict the accuracy. So you give me a neural network. I never saw this neural network. And the idea was, can I predict the accuracy of this neural network? Can I test the neural network without the use of test data? Obviously, this worked on this really small neural networks, only in a homogeneous, we call this a model zoo, a population of models. So they were all trained on the same data, same architecture. First, you know, small steps towards this goal. But we could extract this information about the accuracy, the performance, the generalization gap, what kind of activation function was used.

9:22And if you would plot these latent spaces, You would see different initializations and how they evolve because we had 50 EPOS per model training and 1 ,000 models. So little trajectories were visible. So there was structure and they all organized around a latent space.

9:39Sam Charrington:The metadata around performance and accuracy, these are things that you had from the base models and so supervised in that sense. Exactly. So we needed to be able to train this out on Coda. some population of neural networks, Model Zoo. And at that time, you know, laboratory trained under known condition, fully transparent. We knew which model in which epoch had which accuracy. We had obviously trained, test and validation splits. So the Alto Encoder was trained on 600 neural networks and tested on 300 others. And then we could compare, you know, we predicted 90 % accuracy on Fashion Amnesty and the model had 92.

10:23And then we were R-squared and we outperformed the work from Thomas Untertinger and Daniel Kaiser. So we were happy. Outperformed the original weight space. That story, it worked. We got the paper and people, I'm very thankful to the reviewers. They were kind of telling us it's small, but it's interesting. And we were able to publish because that was kind of the ignition of this amazing journey that we had over the last four years.

10:50Sam Charrington:So I can imagine, you know, lots of different directions, including scaling up the models, trying to get more insights out of the space. Like what was next? The thing is, for this type of research, everything was very obvious. It was like lying out and you just needed to do it. I mean, I never had this before, right? So you have an auto encoder, you take the encoder so you can predict discriminative downstream tasks, like what's accuracy, what's, you know, generalization gap. But we have the other thing called the decoder. So can we sample from the space to generate neural networks? It's obvious, right?

11:25We didn't have space in the first paper. So we needed one more year to have a 22 paper published on generating neural networks. Hopefully those neural networks were then better than standard initializations. They were not as good as final or fully trained neural networks. So there was some trouble that we had, which was really, really interesting. thing. We trained this out on Coda, the mean split error was super low. We took the neural networks, we moved them in the forward path, we reconstructed a set, the loss is very low. We plucked the weights back to the neural network, totally screwed up the entire neural network.

12:02We're like, yeah, it was really like, the mean split error is low. And obviously, now obviously, right, a mean split error is an average. So we're very good at reconstructing the average weights, but the little things that make the difference of having this function working or not, they were so important. And there is an analogy to pixels and images. Like when you had generative models for images, the images were always blurry. So people kind of tried hard and, you know, taming transformers for high resolution images, changed the Minsk-Squid error to perception laws, and, you know, did some additional things on quantizing and the gun.

12:41So we knew that reconstruction made the wrong laws and we tried to normalize and play around with the losses and thought about the behavior loss. And then we were able to make those models a little bit better. But there's still a little bit of delta that is missing. We kind of, we generate at that time, blurry weights, right? Low frequency information, high frequency information missing. And we were happy because again, people were kind to us and said like, it's toy examples, but you know, it's interesting. We generate that. The numbers are good. But then I said, okay, we cannot be three times lucky.

13:12So we have to work really hard to scale that up, right? I mean, you're lucky twice, but three times, your karma is gone for the next years. So we then really, and I'm very thankful to Konstantin Schurhold, who was part of that initial phase, and he really worked hard, and we had this idea of instead of reconstructing the entire model, think about the model parameters as a sequence that you window and then reconstruct the windows. And therefore, we would kind of deattach the sequence left of the original model to the autoencoder one. And this was interesting because suddenly we could go to Resnets and beyond.

13:47And this led then to, you know, the work 2024 in collaboration with Michael Mahoney from UC Berkeley. He actually said in one of the discussions, like, Damien, what you're doing is really interesting, but useless. And so, yeah, sure. A lot of research starts like that. So, and then, you know, because he said that, I asked him that you have to help to scale that up, right? So I called him and he was then one of the co-authors on the paper. Also the previous work was with a lot of collaboration, you know, Shavineo and Boris Knassny, they were part of that because in the beginning, the idea was a little bit, as I mentioned, esoteric.

14:22So we were wondering like what other people thinking about. So we very early involved a lot of people from the community to double check if we are the crazy ones or if this idea is, you know, to at least a particular limit, no, meaningful. So we then in 24, we scaled up to larger networks and other people got interested. And we were at the conference and we met, oh, there exist other people that are doing similar things. A lot of work from Technion, you know, Haggai, Maron, Geltschidnik, Yedid. And then people like, you know, Eliahu. It was actually a funny thing. There was one student that came to our poster and he has this kind of, you know, badge.

15:01And at the badge at the bottom, you have always the university written. And instead of the university, he had like, weights are the new modality, Eliahu Horowitz. And I was like, oh, that's exactly what I'm thinking. And then a collaboration started, like a community. And, you know, we recognized the other people. And we said, like, why not doing a workshop? And then, you know, one led to another. And then, you know. That was last year? It was 2024. That was the first one for 2024. And 2025, we had then the first workshop. And then, you know, other people recognized. There are other areas that are very important.

15:39And one of the tricky things with weight spaces is when you have a neural network and you have two layers, let's say for simplicity, fully connected, you can change the position of the neurons. And it actually changes the sequence of weights because the order of weights, but the function is the same. So there's a lot of permutation symmetries and other things in the weight space that do not change the underlying function. So we had it already in the first pair with some augmentation. But there are a lot of people that are very, you know, very specialized on that and much more experienced, much, you know, more theoretical on that.

16:18Sam Charrington:So augmentation in the sense of like applying these identity transformations to your training, your input models and using them to increase generalization. Exactly. Yeah, very simply in our work in 21, we had a contrastive loss and to build a contrast, you need to augment. So, I mean, it's simple to flip an image or rotate an image, but, you know, what's the counterpart in weight spaces, right? I mean, you can do the permutations and you can do other people that invented, you know, scale augmentations and other things. So the field was exciting because you saw things happening in NLP, you saw things happening in computer vision, and you had to translate it into weight spaces.

17:01And, you know, it worked, right? Like augmentations, how can we translate it? there's this perception loss. How can we translate into behavior loss, right? And all those things. So it was, you can borrow ideas from other fields and it was an empty field to fill with content.

17:17Sam Charrington:Along those lines, hearing you talk about these identity transformations makes me think about like other kinds of geometric transformations in the weight space or like Cartesian, the polar transformations or things like that. we saw a Google quantization, I forget the name of the quantization paper that just came out or actually came out a year ago, but it was became, it was revisited a week or so ago. And they did some Cartesian, the polar transformation, like all kinds of stuff that you can do in the weight space that you might, I'm curious how much of that is being explored. So that's, that's exact.

18:03So we came from this one area of motivation. We met the other people, as I mentioned, that are on all the symmetries or the group operations you can do. There's another world of, you know, there's this mode connectivity that, you know, like Gitri Basin where you can align models along some, you know, reference models. And if you just think about this permutation, you know, symmetries, there is so like a loss landscape that is connected. There's so many points that are the same, but they are different with respect to the weight space. And just understanding or trying to understand how lost surfaces look like, how models can evolve along orbits of great work from Bo and then Rose from UC San Diego on all this work, kind of helped me to understand better what actually happens during learning.

18:57And then hopefully, you know, we could move this in our backbone, learning backbone, because, you know, some tokenization, some position encoding in the outer encoder is a transformer outer encoder. You know, we took some of the idea. It helped us. And then, you know, people are also discovering a grokking and phase transitions and how models suddenly kind of converge or suddenly don't work and then suddenly work. So this also is also connected. And I would love to explore this direction more. to kind of understand how can we make training much, much faster, better models or give guarantees is a strong word, but kind of, you know, bands of where models are operational with respect to their, you know, be like their performance occurrences, etc.

19:43Sam Charrington:Yeah, I think the last time we covered kind of weight space, I don't think we talked about it as weight space learning, but kind of this idea of like introspecting weights was with Charles Martin, who does a lot of work on Weight Watcher and crocking and motor collapse and the like. So Charles and Charles co-auto is Michael Mahoney, who was then co-auto in our paper. So yeah, yeah. And it was really funny because Michael was doing with Charles the work of analyzing weights and looking how their shapes are changing so you can make a statement about if they converged or not and they extended, obviously.

20:26Really amazing work. So it also happened in parallel. I'm looking at this like an analytical piece of work. and where ours is like a learning piece of work. So we hope that we can at some point scale up our backbone and process more diverse model zoos, different types of architectures. We have a work where we are now able to train. And that's the vision, right? Can you train different sized models with different architectures, tasks, modalities from open-wage repositories at Hugging Face. Could we download everything from Hugging Face and train a foundation model of neural networks?

21:11Sam Charrington:And you've started down that path. I think the... That rabbit hole, yes. The poster that I originally saw at GTC that led me to you was something about kind of training on the Hugging Space model zoo, right? Exactly. So after we were able to scale up, where then we're thinking like, can we kind of, because we're still limited to the models, right? So I can tell you, I can generate now a new neural network, but I need this 1 ,000 neural networks to have trained before I can generate that one. So you tell me, that's great, but now we have 1 ,001 neural networks. So what are we gaining, right? At the end of it, what are we winning?

21:51So the next step was, can we train on models that are out there? And there's amazing work on analyzing how, you know, hugging face looks like the model atlas from Eliahu Horowitz and Yedit. And there is a lot of, there are a lot of models. So can we download those models? And then independently of what kind of architecture they have or what data set they trained, use our machinery. And it's a little bit tricky because, you know, different sequence lengths, different types of neural network layers. So we have to put some information into it. But we were able to train like the first, you know, weight space learning model that can do generation and analysis of weights, so discriminative and generative downstream tasks on hugging face models.

22:37So this was work by Daniel Falk, which is really amazing.

22:43Again, we thought it's much more challenging to do this, but, you know, you have to scale, you need the machinery, and you need those little tricks how to handle those different, you know, the tokenizer needs to be adapted to arbitrary architecture. That's the thing, yeah.

22:59Sam Charrington:And what's the filter that needs to be applied on the, you know, this Hugging Face's vast library of models that normalizes them to something you can deal with? First of all, and this is work also that, you know, other figures like Eliyahu, there's a lot of content on Hugging Face that is not documented. Around 30 % of the models, they don't have any meaningful metadata. data. So we don't know what, what, yeah, I mean, hugging face is amazing. And so is the first filter, just get rid of all of those? Yeah. Knowing which models are helpful. So we need, we did a little bit of experiments. If we scale, scaling alone doesn't help.

23:37You need to increase the diversity of the models. So we want to have diverse models. So we want to have, you know, different data sets. We mostly focused on computer vision, language is the next. And we kind of developed a scoring function on, you know, how popular is the model? How is it apparent? Or is it some derived work? There's a lot of trees in there to download a set of two, we have 20 ,000 models and from them 2000 models that, you know, are passing some quality checks. And from these models, the billions of parameters, we trained the backbone on open-weight models that then can sample all the different architecture.

24:20We can sample VITs, resnets. This entire thing was focused on computer vision, but we were able actually to sample a GTP2 model. It's still a small model, but there is a domain change. And this model that we sampled, we used as an initialization so it can train faster as compared to training it on regular data sets in our language. So there is some knowledge transfer happening from computer vision models to language models. That's also interesting because now we're still trying to figure out like what are those ways encapsulating, right? And encoding. So that's kind of, you know, the next step would be then to scale it up and to train on different modalities, tasks and architectures.

Read the full transcript

25:04Sam Charrington:You mentioned that there were some tricks that you had to employ to be able to use different types of models. And you mentioned specifically Tokenizer. Dig into that a little bit more and also talk about some of the other tricks that you had to employ to do this. So good models are important. Diversity is important in the model weights you use for training. and the tokenization and the processing of tokens changed strongly inspired by work from Kewang and in Singapore. So that helped us to identify or to process weights in an agnostic way so that we are not bounded by, you know, this is a layer starting, this is a layer ending.

25:52Sam Charrington:And so to be clear, are we talking about your thing is a model, right? And so it has its own tokenizer, or are we talking about like normalizing the tokenizer of the models that you're ingesting or both? So, okay, that's a good question. We take the models that we download from Hugging Face. We strip away the weights in reading order very stupidly. There are probably better ways of doing it. So we kind of destroy the metric structure and then we just sequentialize or vectorize that. And then it's a long sequence of millions of parameters. Got it. So it's just numbers and then you've got a tokenizer as part of your ingestion process.

26:35We also normalize. So normalization plays also an important role. There are different positions where you can normalize on the weights as a pre-processing, during the tokenization, or at the loss function. We try different things and currently it depends a little bit what kind of, on the hugging phase data, we normalize inside of the batches during tokenization. So that's, and other setups, I don't know, but this is important. The tokenization is important. We still playing around with the losses because we want to get some of the high fidelity, high frequency information. So we still kind of suffering a little bit with that.

27:15So just as a simple example, when we generate a model, the model is a little bit damaged, so we need some fine-tuning steps to recover that. Very often, these fine-tuning steps are very quickly able to recover, but given that our decoder has this problem of the blurry weights, we're kind of not getting perfect weights, right? So that's something that we are working towards because if you would think about, So let's think about a world where we could train on all this data, a foundation model of neural networks, and sample on demand your favorite model, whatever you need, then we would be actually replacing and totally replacing pre-training, right?

27:59So why do I have pre-trained model? You just sample the model that you need. So that's the vision of this foundation model of neural networks. We have one additional little thing that kind of causes Oslo a bit of trouble. we need to sample a model and anchor where we sample. And this anchor needs to be a model that we put through the encoder. So the better this model, the better the sampled models, which is leading to the situation that we need a well-trained model to generate another well-trained model, which doesn't make sense if it's in the same domain because, well, I have a model, why should I generate one?

28:36Chicken and an egg problem, right? Exactly. So therefore, we had this paper that we're going to publish soon in CVPR about remote sensing, where we take an ImageNet VIT, use our machinery to generate remote sensing models or remote sensing foundation models. Then we have the domain change, where our encoder, decoder is providing knowledge transfer to generate models that are better than the ones that ImageNet fine-tuning would be able to reach. So here we have a true knowledge transfer, which is really, really great, where we are able actually to outperform or be equally in performance with current models like Tel.fm published at ICLR.

29:17I think the autos claimed they trained for 12 ,000 GPU hours and we are able to do this on 350 GPU hours. That's a factor of 20, 25, 30, depending on how you count. And this is suddenly interesting because you train from models and not from data. So if you think about, we're running out of data. That's the reason why the scaling laws are a little bit, you know, considered differently and everybody is moving into test time, adaptation test time training. We're running out of data to train the large models, but we're not using the weights of older models. So why not using the weights, all the knowledge that, all the compute that people invested, right?

29:59Sam Charrington:And it's also an interesting context. Like if the, if all the data, if, you know, If we're in fact running out of data, at least in specific domains, visual text, etc., then all that data is already in a bunch of models. Why replicate that? And why not just find ways to, different ways to slurp it out of the existing models? That's essentially the premise, right? Totally makes sense. In the remote sensing community, we have 70 foundation models, according to some surveys. And there are still people training the 71st, 72nd one, right? So why not taking this knowledge, compress it all in a weight space learning representation, and then sample models on demand?

30:40Because if you have a foundation model that is a VIT with 800, 900 parameters, and somebody fine-tunes it to a task, this person, the partitioner, needs to use all the 900 parameters, and that's demanding compute, and maybe the performance is a little bit better than a REST with 50 million, 40 million. So, you know, you suddenly bounded in this foundation model world to large models. While with our machinery, with our waste-based learning approach, you could sample a big model, you could sample a rest net, you could sample an efficiency net, depending on what you need. But you can give the architecture as a desired output, and then we sample the parameters for whatever architecture, which is helping for edge devices, helping for foundation model, or helping for other use cases.

31:25Sam Charrington:So you can also think of it as kind of a compression technique in a sense. Exactly. The interesting question is what are we compressing, right? And how much redundancy is there? And can we do it in pruning and also distillation and train with that? Until now, we train with the raw data, but we can generate smaller models. And these smaller models are better than if you would take the original model and do distillation with a teacher's student. So that's also what we show in this remote sensing scenario. So the goal is really like having one big foundational of neural networks and then able to generate on-demand models.

32:08There's one missing piece though, and we have a paper cranking review that might solve that. So if you think about that, we need a model as a prompt to get an anchor to sample other models. It can be a domain change. So what would be really, really nice if you would not need this model as a prompt, but you could prompt with your data set. So give me a model that works well on this data. Exactly. And if this could be done in a privacy-preserving way, it would be even better, right? Because then suddenly people could use this without revealing individual members of this data set. So there's work that does something similar already and Keist, you know, Soro and Andres and others who are working on that and we kind of moved this forward because we have model zoos, right?

32:55We have data sets and we have models. So we have kind of images and we have models. And why not train an aligned space like Clip did with text and images? Then we could use actually a kind of a data set encoder as a model prompt that could move into this space and point into this space. Instead of a model prompt, we use a data set prompt. So you're a bank, you're a financial institution, healthcare provider, whatever. You don't reveal your data. You have your data set of, I don't know, 100 samples, 1 ,000 samples. You create one data set embedding. So you cannot infer the individual members or samples.

33:35You give this embedding to us. We provide you the weights. We give you the weights. You're much faster in continuing training. So this would be really interesting, right? because that opens up to all the potential data sets that are not on Hugging Face.

33:52Sam Charrington:And is the privacy preserving angle there because your process would start with that embeddings anyway, so it doesn't matter who produces it? Or is that a compromise that you could do it with the processes, but you could probably get more out of it if you had the actual data? Good question. So it comes by the method that we use, because imagine you have a data set with 10 ,000, 1 million images. You need some kind, you cannot prompt with all individual images. You need some aggregated. So you need to prompt with a thing. One thing or one vector, right? And you need to aggregate this knowledge.

34:32So like, you know, we do a sentence and you have a text embedding to prompt your image that you generate. So it comes with that. Obviously, you have to check for, you know, trust membership attacks. you have some noise, etc. But the idea is really like, once we have this one embedding, can we generate from this one dataset embedding now the tokens or the embeddings that generate the tokens for the weights? And this would open up, and then the question is, this would open up to datasets that people are not willing to share or not allowed to share because of regulation. A financial institution that would maybe love to share, but they're not allowed.

35:09I have an external PSG student with the Deutsche Bundesbank, the National Bank of Germany. They cannot share, but they might use those embeddings for that reason. And it would also allow us to kind of understand how is the space that we're learning, this latent space, because the most interesting thing would be that we could interpolate the models that we see on Hugging Face. So can a data set that per definition is never on Hugging Face and there's no model that's trained on that, can this data set and the data set prompt generate some meaningful neural networks that live between the space of known neural networks that are embedded?

35:49So do we have a well-behaved latent space? Because suddenly, if we could show that we have one, this machinery could be used for generating neural networks with minimum pre-training or training at all. So think about the world where you could have on-demand neural networks, on-demand hyper-personalization in the forward path.

36:12Sam Charrington:I mean, it sounds a lot simpler than the way we think about like neural architecture search today, which is a lot of very complex machinery. Exactly. And to the end, I mean, the external, you have to give the external signal, you want to understand what kind of architecture, right? And we just generate the way. So we are, I consider this work complementary to the neural architecture search. Ah, so that might define the structure and you might provide the weights to fit into that structure. Exactly, yeah. Or you know, I mean, you know, I want to have this type of architecture, right? Give me a transformer, give me whatever, right?

36:48Because this is, you know, per definition, per design already fixed. And we provide the best weights. To what degree does it produce weights

37:01Sam Charrington:kind of with knowledge of the architecture in which those weights will be used? Or is it just like you've got a parameter, you know, the number of weights and it spits out the number of weights and then you have to, as a post-process, apply that to an architecture in a given way? Exactly. So until now, with the model prompt, you give an architecture. With the dataset prompt, you don't. So you need to give an external signal, like give me a ResNet 18 or a BIT. But with the model prompt, you give an architecture that we fill with weights. There are definitely more clever ways of building a tokenizer that kind of understands what type of architecture elements are processed.

37:47So this would be the next step to be more knowledgeable about that and to also allow the decoder to generate conditioned architectures. architectures. We don't do this until now. We generate the weights. We hope that during learning and training, the backbone saw enough instances of the same architecture. So we probably only can no, we definitely can only reconstruct architectures that we saw in our hugging face collection. But it would be probably much, much better if we could take this conditional signal on top. And some people are doing that, like people in Singapore, Kha 'wang, and also Kaist, they do diffusion on top of that.

38:28And then you have a conditional data set signal and conditional architecture signal. So this would be also the next step to go. But also under the promise, we can do it on open weights model that are diverse, all the chaos, right? Because that's where the most use can be generated from this idea of weights-based learning.

38:46Sam Charrington:Does the model produce a sequence of weights that you then have to map to a position in an architecture? Or does the model also produce, like does it produce a weight and a position in an architecture? So it generates a sequence of weights that you have to then fit to the architecture. and every token that is then translated into a weight has a position encoding which position within the layer, which layer or which block and then absolute position. But that's on the input side. On the input and on the output. It was an autoencoder, yeah. Ah, okay. Got it, got it, got it. So then, you know, you get a sequence of weights and you have to fit it to the right architecture.

39:33You could fit it to a different architecture and then cut it or, you know, slice and dice it. that would be an interesting experiment, what happens then, you know. How many fine-tuning steps do you need to repair that? And that's probably also one of the reasons why we still need a couple of fine-tuning steps to get the model very quickly up to performance. And therefore, I think that conditioning would kind of help us to generate better weights quicker that are diverse enough. And so this would be potential next steps, right, to move. and that condition decoder and look into the decoder.

40:10Sam Charrington:When you produce these weights, are you like overriding an initialized model, like random initialization or zero initialization or something so that the model is always valid? Or do you have to think about, well, you know, the model didn't actually hallucinate it and it generated two weights for this position and no weights for that position, that kind of thing? We entirely override. So we replace whatever is in the model, randomly initialized to have that, and we entirely override and load the checkpoint into that architecture. Since we have the sequence and we know the position coding, this is technically straightforward.

40:51Whatever the model does, it could have hiccups, have repetition. There's no visible pattern that would be like an artifact that we could observe that is kind of repeating again and again and again and again. But you could also think about not only generating one model, but because it's a forward pass, it's very cheap, you can generate ensembles of models, right? And then fit them and you suddenly have this new degree of freedom that you can generate. But yeah, we believe our ways that we can entirely overwrite whatever is there. There's some interesting work that currently is happening on task vectors.

41:26And what is a task vector? So if you have a base model, for example, and you fine-tune to one task, another task, and the third task, you could take the difference of the fine-tuned model to the base model, and this would be a task vector. And you could do a task arithmetic with that. And then in weight space, people are using that in particular in language models, sometimes in computer vision. And when you work on those task vectors, we observe in some cases that it's easier to learn and easier to generate than in other cases. In other cases, full weights are easier. So Empirically, we observed it, but we cannot explain that way.

42:04Why for some setups, the one is easier or the other.

42:07Sam Charrington:You mentioned a couple of times the distinction between high frequency noise, low frequency noise, and the weights, which calls to mind, like doing things in a frequency domain, applying an FFT to weights or something like that. Is that something that people are working on? To the best of my knowledge, not. But I think what would be interesting is, because what we generate is kind of this blurry base, and then you could add some high frequency information on top. And I think in the image domain, this is happening through, you have a kind of a variation autoencoder reconstructing, and then you have a generative adversarial network enforcing high frequency information.

42:50And then you combine lots of both of that. So you could do things like that in our domain on weights. I would not know if somebody really did this. But moving the weights into some other domain and then processing is an interesting idea. Natural frequency would be the right way of doing it. But you could think also about a lot of pre-processing steps until you would go and do the weight-based learning compression or learning the latent representation, the lower dimension manifold on that. There are people who are also moving into that direction. But yeah, let's see. We're not going along that path currently.

43:37Sam Charrington:And will the workshop be continuing this year? There's another great workshop that is organized by some of our colleagues at ICML. It's a workshop on weight symmetries. and we're thinking and continuing it for the next opportunity, which would be then Neuris. But let's see how this looks like. Currently, we have a couple of people that were, you know, PSGs are finishing, new PSGs are coming. So there's a little bit of a gap in the community. With the inflow, a lot of great people are currently finished with the PSGs on the market already in post-op position. So there will be definitely a continuation on that.

44:16and like one of the most important experiences was recognizing that there is a community and once the community is there trying to develop a common language trying to develop common benchmarking trying to develop some ideas that are coherent as a community and moving along that way I think this is now the time to try to move in that direction and bring those ideas together with the liberty of also going to the one or the other idea and direction. There's interesting work from Hager Maroon, for example, and he's doing, taking the ideas of weight space learning, applying it to gradients or applying it to activation spaces, doing representation engineering, which you could also think about a neural artifact that is not the weights, but the gradients during training or the activations.

45:15There's interesting work from Yedit coming around probing neural networks. So not looking at the weights, but having controlled input-output relationships and therefore looking what happens given an unknown neural network when they control the one thing.

45:32Sam Charrington:Anthropik's been doing a lot of work along those lines as well. Exactly. So there's quite interesting work that could be considered complementary to that. and it's always this one thing do we have these artifacts as a collection? Can we learn some representation and then use it to predict properties of the neural network or to manipulate and edit the behavior change and modify it. Very interesting work. Thank you so much for jumping on and sharing a bit about it with our audience. Thanks for having me and yeah, let's see where we will be in three years. Absolutely. Thank you.

From the publisher

For more than a decade, AI has advanced by training ever-larger models on ever-larger datasets. But as high-quality training data becomes harder to find and pretraining grows increasingly expensive, researchers are looking for new ways to keep foundation models improving.

In this episode, Damian Borth, professor of AI and machine learning at the University of St. Gallen, argues we’ve been overlooking an important source of knowledge: the models we’ve already trained. His group’s work on weight space learning treats trained neural networks themselves as data, learning from the distilled results of millions of GPU hours of optimization rather than starting from raw data each time.

We explore what it means to build foundation models of neural networks, how knowledge can be transferred across architectures and domains, why this approach could dramatically reduce the cost of developing specialized models, and whether future AI systems may be trained on collections of existing models instead of ever-growing datasets.

🗒️  Full show notes: https://twimlai.com/go/772.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Why Models Are AI’s Next Training Dataset with Damian Borth - #772The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 47 min
Listen in VO