Performance Prediction for Large Systems via Text-to-Text Regression

16 Aug 2025 · 19 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

A Google/Cornell research paper on “text-to-text regression” using regression language models (RLMs) to predict performance of large, complex systems (e.g., Google’s Borg compute scheduler) by treating system logs/configs as text instead of forcing them into fixed tabular tensors.

Guest backgrounds

No guest names or bios are provided in the transcript; it’s a host-led discussion.

Key claims

Text representation avoids lossy feature engineering for variable-length, nested, and evolving system data; RLMs can do accurate numeric regression with token-based number prediction (cross-entropy), encoder-decoder design, and even training from scratch (no language pretraining). Model size plateaus around ~100M parameters; few-shot adaptation works with ~1–512 examples.

Notable examples

Borg prediction of MIPS per GCU; a ~60M-parameter RLM achieved ~0.99 rank correlation across the fleet and ~100x lower MSE than tabular approaches; digital-twin simulations take 1–18 hours per run, while RLM inference is negligible.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Importance of Prediction in Large Systems

0:45 to 2:42

Understanding why accurate predictions are crucial for systems like compute clusters.

“is tackling problems that, frankly, the traditional methods just couldn't handle.”

Challenges of Traditional Machine Learning Methods

2:42 to 5:47

Discussing the limitations of traditional machine learning techniques on complex data.

“But the problem is they're very restrictive in their applicability.”

Introduction to Text-to-Text Regression

5:47 to 8:40

Diving into how text-to-text regression works and its advantages over traditional methods.

“These machines are in different data centers or cells, as Google calls them.”

Understanding Rank Correlation and MSE

8:40 to 10:35

Explaining rank correlation and its significance in evaluating predictive models.

“engineering setup, invalidating your old data.”

Google's Borg: The Testbed for Predictions

10:35 to 12:34

A brief overview of Borg, Google's compute management system and its metrics.

“They found this leads to more stable training and stops the model over-focusing on tasks which naturally have larger errors, like outliers.”

The Power of Text Representation in Predictions

12:34 to 14:00

How treating system data as text can overcome traditional modeling challenges.

“It's about deeply processing one complex state description at a time to make the best possible prediction for that state.”

Understanding Few-Shot Adaptation in RLMs

14:00 to 14:30

Learn how few-shot adaptation in RLMs allows rapid specialization on new tasks.

“And what about adapting it to new tasks, that few-shot adaptation?”

The Impact of Pre-Training on Performance

14:30 to 15:10

Discover how diverse pre-training tasks enhance performance on new datasets.

“They showed the RLM hitting a strong 0.86 rank correlation with just 512 few shot examples from a totally new task.”

Exploring Model Uncertainty in Bayesian Optimization

15:10 to 15:40

Uncover how model uncertainty guides better optimization strategies.

“It's getting the numbers much closer to.”

Key Design Elements for Effective Regression Models

15:40 to 17:10

Analyze key design choices that enhance the performance of regression models.

“If the RLM is very certain about its prediction for one setting but very uncertain about another, the optimizer knows it should probably explore the uncertain one more.”
Show all 12 chapters

Practical Tips for Fine-Tuning Regression Models

17:10 to 18:22

Get insights on how to optimize fine-tuning processes for better results.

“Specifically, finding the optimal learning rate matters significantly for minimizing errors.”

Future Implications of Text-to-Text Regression Models

18:22 to 19:09

Explore the broader impacts of text-to-text regression on AI and complex systems.

“And looking even further ahead, the paper suggests these RLMs, by accurately modeling numeric feedback, can serve as a foundational aspect for developing sophisticated reward models.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Have you ever just looked at the sheer amount of data swirling around us? All those system logs, complex configurations, and just wish for a shortcut. You know, a way to cut through all that noise, get the real insights, and just know what's going on without drowning in details. Well, today we're diving into something that feels, well, it feels a bit like that breakthrough. We're exploring a really innovative way to predict how massive real-world systems perform. Think Google's huge compute clusters. And the really clever part, it treats all that super complex system data. basically as if it were just simple text.

0:36That's exactly it. Our mission today is to unpack this new research paper. It's from Google and Cornell University, and it shows how this technique, they call it text-to-text regression, is tackling problems that, frankly, the traditional methods just couldn't handle. We're talking about unprecedented accuracy, adaptability, really understanding how these big industrial systems are going to behave. It's genuinely a game changer for anyone dealing with that kind of messy real-world data. Okay, so let's set the scene a bit. Why is prediction so vital for these huge systems like a compute cluster or maybe a logistics network?

1:10Why is getting that prediction right so critical? We're talking latency, execution times, avoiding scheduling clashes, transaction speed. If you can't predict that stuff, you're kind of flying blind, right? Completely blind. And that costs organizations big time. In time, money, efficiency, you name it. And the real challenge with the traditional machine learning approaches, you know, things like random forests, multi-layer perceptrons, they're great for neat structured data, what we call tabular regression. But they really, really struggle with what the paper calls complex systems data in the wild.

1:42Things like config files, system logs. Exactly. Try feeding that into a standard model. The core problem is these methods demand data formatted as flat fixed length tensors. Okay, so basically like a rigid spreadsheet. Pretty much. Everything boiled down, neat columns, fixed size. And getting complex data into that structure, it often takes a massive effort, feature engineering, or sometimes it's just, well, infeasible. And when you try and force it, when you squash all that rich, messy, nested data into simple numbers, what do you actually lose? Well, it becomes what the paper calls a lossy process.

2:19You're compressing so much nuance. You end up treating the whole system like a black box. You lose the context, the relationships within the data. And as this research really hammers home, that just tanks your predictive performance. You lose too much vital information. Yeah, I can imagine. Were there attempts to kind of bridge that gap before like hybrid approaches? There were, yeah. People tried what they called gray box techniques, sort of mixing domain knowledge with machine learning. But the problem is they're very restrictive in their applicability. They need a ton of prior knowledge about how inputs and outputs are linked.

2:53They just don't scale well to the kind of dynamic, constantly changing complexity you see in these huge modern systems. Which brings us neatly to the breakthrough, text-to-text regression, the game changer, powered by all these advances in large language models, they call them regression language models, or RLMs. Exactly. And what's really striking is how these RLMs just nail the core capabilities. The research validates this clearly. First, they can predict numbers like efficiency metrics with really high precision. And they handle complex feature representations and even multimodal outcome distributions.

3:29Meaning, messy inputs are fine, and they can even predict situations where there might be a couple of different common outcomes. Okay. And specifically, on Google's big compute scheduler, Borg, a relatively small 60 million parameter RLM hit a near-perfect 0.99 rank correlation across the entire fleet. Though, 0.9 on average, still amazing. Whoa, hold on. Rank correlation, what does that actually mean for someone listening? Like, in simple terms. Ah, good question. So rank correlation isn't about hitting the exact number every time. It's about how well the model predicts the order or the ranking.

4:02So a 0.99 means if scenario A is genuinely more efficient than scenario B, the RLM gets that order right almost perfectly. It nails the relative performance. Got it. Relative performance. And on top of that, it also achieved 100x lower MSE than tabular approaches. MSE means squared error, right? What's the impact of it being 100 times lower? That sounds huge. It is huge. MSE measures how far off your predictions are on average. So 100 times lower means you're going from maybe broad ballpark guesses to incredibly precise predictions. Think about resource allocation, avoiding slowdowns. It means you can run things much closer to optimal, potentially saving enormous costs.

4:43Less error equals more efficiency. Makes sense. And there's more. This RLM, it easily adapts to new tasks in only 500 few-shot examples, Sometimes even fewer, like maybe 1 to 512 examples. Two shot. So you just show it a 20 handful of new situations, and it just gets it. Pretty much. It shows this enormous transfer learning ability. It's a kind of meta-learning. The model learns such general principles during its main training that it can adapt to a totally new problem with very little new data, like us learning a related skill quickly. It also captures the densities of complex outcome distributions, so it understands the whole range of possibilities and their likelihoods, not just one number.

5:21And ultimately, the paper argues this paves the way for universal simulators of real-world outcomes and can drastically reduce production costs. That's really something. Okay, let's zoom in on Borg for a sec. It was the testbed here. For listeners, maybe not deep in Google infrastructure, what is Borg, briefly? Sure. Think of Borg as Google's giant compute management system. It's like a central brain scheduling millions of jobs across thousands and thousands of machines. These machines are in different data centers or cells, as Google calls them. Users submit task requests, what resources they need, how long they think it'll run, that kind of thing.

5:58And Borg's job is to juggle all that, put the jobs on the right machines to maximize efficiency and keep everything running smoothly. It's a massive orchestration task. And the specific metric they focused on predicting was MIPS per GCU. What's that stand for and what's the intuition behind it? Right. MIPS per GCU is millions of instructions per second per Google computing unit. Intuitively, you can just think of it as, well, the amount of work that gets done per unit of time by using per unit of computing resource. Okay, so a direct measure of how efficiently the computers are actually working.

6:33Exactly. Higher is better. And I bet getting that number isn't simple. What makes predicting MIPS per GCU so tricky? What factors are involved? Oh, a ton of factors. You've got current memory and CPU load on machines, the specific hardware types involved, the complex mix of different kinds of jobs running together. Even just variations in time have user patterns change throughout the day, the week. It's all interconnected and constantly shifting. Right. Now, Borg actually has a kind of digital twin, doesn't it? A backtesting system. It does, yes. A sophisticated framework for replicating cluster states.

7:09They use it for testing changes, running what-if scenarios with scheduling algorithms. But you mentioned it has a big drawback. A major drawback. Speed. Generating just one outcome, one simulation result can take anywhere between 1 to 18 hours of computation. 18 hours. Wow. Yeah. So compare that to the RLM, where the inference time getting a prediction is negligible. The savings are potentially enormous. Huge difference. Plus, other tools they use, like Google Vizier for tuning the scheduler, are stuck with those tabular data formats we talked about. The RLM just bypasses that limitation entirely.

7:43Which really brings us back to the core question. Why text? What makes treating all the system data as text so powerful? Because trying to represent all the input features, cluster names, time windows, hyperparameters, hardware details, those super granular job profiles, trying to cram all that into one fixed length tensor for the old regressors, is, as the paper puts it, notoriously difficult. Yeah, the paper gives some really concrete examples. It mentions features like a list of jobs and platforms can contain arbitrary cardinalities. What does arbitrary cardinalities mean in this context? It just means the list can be any length.

8:20You might have five jobs listed or 500 or 5 ,000. Preditional models hate that variability. You'd have to pick a maximum size, then either cut off data or have loads of empty padding. It's inefficient. And what about new types of machines or jobs appearing? That's another huge headache for the old way. If a new machine type comes online, you'd potentially have to redo your entire feature engineering setup, invalidating your old data. And then there's the deep-nested nature of the features. Think about it. You have job profiles, which contain info about the job and the machines it uses, which also contain info about hardware profiles.

8:57It's like Russian dolls. Data within data within data. Exactly. All nested and interconnected. And squeezing that flat is where you lose so much. Just to give people a sense of scale here, the paper mentions the average character count for just the Chaban machine performance part can hit 268 ,157 characters. Right. That's a lot of text. It's complex, varied way beyond a simple spreadsheet column. So how does text solve this? It's the flexibility. A flexible text-based input representation simply resolves these issues. It handles variable sequence length inputs naturally. Doesn't matter if it's 100 characters or 268 ,000, it just reads the string.

9:34No need to manually enumerate categories or normalize numbers beforehand. The model just observes all the available features in their raw text form. And this minimizes what researchers call epistemic uncertainty. That's the uncertainty you get from not seeing the full picture. By feeding it everything as text, the model gets the complete picture, or as close as possible. Making its predictions much more informed. Exactly. Okay, this is where it gets really interesting, especially for the technically-minded folks listening. The paper mentions some design choices for our RLM that can be considered contrary to popular belief.

10:08What do they do differently? Yeah, this is cool. First, how it handles the output, the prediction. Instead of using a standard error-based loss function, like mean squared error, which tries to minimize the numeric difference directly, They use cross-entropy loss over tokens. Essentially, they teach the model to spell out the number using special tokens. Spell out the number. Sort of, yeah. Imagine tokens for digits assign an exponent. So 72.5 might become plus 725E1. They found this leads to more stable training and stops the model over-focusing on tasks which naturally have larger errors, like outliers.

10:43It learns the structure of numbers better this way. Ah, okay, interesting. What else? The use of encoder. A lot of the big LLMs you hear about now are decoder only, mostly built for generating text. But this research found that for RLMs processing this complex input data, the X, having separate encoder layers are necessary. The encoder's job is to really digest that complex input string. They found that decoder layers alone are suboptimal for this task. You need that dedicated encoder to understand the input properly before predicting the output. So encoder-decoder beats decoder only for this kind of regression.

11:16For this specific complex input scenario, yes, that's what their results show. And here's maybe the biggest surprise. No language pre-training. Really? You don't need to start with a huge pre-trained model like GBT-3 or something? Apparently not. The paper states it's not necessary nor guaranteed beneficial to use a pre-trained LLM checkpoint. They found RLMs can even be trained effectively tabula rasa from scratch. Random initialization. Why is that? Because, as they put it, regression only requires learning the correlations between different structured tokens and does not necessarily benefit from the semantic meaning behind words.

11:52It's not about understanding language like a human. It's about finding patterns in the structured text representation of the system data. Fascinating. So it's pattern matching, not language comprehension. In essence, yes. Yeah. Then there's the Y tokenization. We touched on that clever way of representing numbers like 72.5 as plus 725E1. They call it P10 tokenization. It keeps the model's vocabulary small and avoids needing to normalize the output numbers first, which makes it more robust. It's normalization-free. And finally, the models are context-free in a specific sense. It means the model looks at one complete input snapshot X and predicts one output Y.

12:31This maximizes how much of the model's input capacity, the sequence length, is used just for observing that single X, which can be thousands of tokens long. It's about deeply processing one complex state description at a time to make the best possible prediction for that state. Okay, that makes sense. You're focusing all the attention on the current input. Let's move on to the results then. You mentioned the RLM maximizes scaling on multiple axes. What does that mean in practice? It means performance improves along several different dimensions as you scale things up. But interestingly, not always the dimensions you might expect.

13:04The research pinpoints diverse training data and feature observability as the two most important scaling factors for regression. So seeing more different situations and seeing more detail in each situation are key. Exactly. Better feature observability, seeing everything as text, lets it learn better continuous representations. More diverse training data gives it better coverage of all the possible scenarios it might encounter. But here's a surprise. Model size itself. Not necessarily very important. Really? Not like the giant LLMs we always hear about? Nope. Performance quickly plateaus within the O100M range.

13:41That's 100 million parameters, which is orders of magnitudes lower than state-of-the-art general LLM models within the O1B range. Billions of parameters. Wow. So you don't need a monster model. Right. Which means it requires relatively low amounts of compute, often at most one GPU. That's huge for practicality, makes it much more accessible, cheaper to run. Definitely. And what about adapting it to new tasks, that few-shot adaptation? Yeah, that's incredibly powerful. The process is basically fine-tuning. You take the pre-trained model, restore its weights, and then continue training on a tiny amount of new data, like 1, 512 examples, using a lower learning rate.

14:18And it adapts really fast, only requires a few minutes on a single GPU. It's a form of meta-learning where it leverages its broad knowledge to quickly specialize. And the results back this up. Absolutely. They showed the RLM hitting a strong 0.86 rank correlation with just 512 few shot examples from a totally new task. Compare that to training a model from scratch on that small data set. The randomly initialized model is unable to achieve the same level of performance. Pre-training matters. Okay. They also show that pre-training on more diverse tasks, a higher number of pre-trained tasks, leads to substantially better results on unseen cells.

14:56Diversity helps generalization. And crucially, back to the error rate, the RLM produces much lower residuals overall. Remember, 100x lower MSE than possible with tabular features. This just hammers home the value of seeing all the features, the maximum feature observability you get with text. It's not just getting the order right. It's getting the numbers much closer to. Precisely. And it captures those densities of complex outcome distributions, even showing multiple modes if the situation warrants it, giving you that probabilistic view. And importantly for downstream uses, there's a high correlation between the variance of its predictions and the actual prediction error.

15:33Right. You mentioned Bayesian optimization. How does knowing the model's uncertainty help there? It helps guide the search. Bayesian optimization tries to find the best settings efficiently. If the RLM is very certain about its prediction for one setting but very uncertain about another, the optimizer knows it should probably explore the uncertain one more. It makes the search much smarter. I see. So overall, the takeaway is that the RLM can achieve very precise, point-wise regression across lots of tasks. The paper says the majority of cases reaching above Bureau 0.93 Spearman rank correlation.

16:06That's really, really good. Truly impressive numbers. But let's peek behind the curtain one last time. What were the key design elements that made this work so well? Any specific insights from their emulation studies? Yeah, the emulations where they test removing parts of the model are revealing. We mentioned it, but it's worth stressing. The encoder-decoder architectures perform substantially better than decoder-only for these complex inputs. Even though decoder-only is trendy. Right. For this task, that encoder doing the heavy lifting on the input is critical. Also, feature importance. They confirmed that giving the model longer input sequences, allowing it to observe more features from X, directly improves performance, makes sense, more info, better predictions.

16:50And they showed specific features matter. Adding the time window feature, like the exact start and end time, significantly boosted results, which validates domain knowledge about temporal cycles in system workloads. So the model learned things experts already knew were important. Exactly. It validates the approach. And finally, some practical tips came out for future practitioners. Getting the fine-tuning right is key. Specifically, finding the optimal learning rate matters significantly for minimizing errors. And maybe counterintuitively, sometimes earlier checkpoints from pre-training also can lead to better results when fine-tuning on out-of-distribution tasks.

17:28Earlier checkpoints, why? Because, they suggest, an overly pre-trained model may meta-overfit to the pre-training tasks. It might become too specialized on the initial data mix and less flexible for truly novel scenarios. Sometimes slightly less pre-training retains more adaptability. Huh. Interesting balance there. Definitely. So summing up, the core message is clear. Text-to-text regression is this powerful, general, and scalable tool. It handles raw text from complex systems to predict outcomes. And crucially, it alleviates the burdens of manual feature engineering. That alone is a massive win for anyone working with this kind of data.

18:04Yeah, taking away that bottleneck is huge. And the broader implications. You said it opens new avenues for creating universal simulators for complex systems. Just imagine that across different fields, logistics, healthcare, manufacturing, maybe even climate science, anywhere you have complex systems and messy data needing precise predictions. The potential is vast. And looking even further ahead, the paper suggests these RLMs, by accurately modeling numeric feedback, can serve as a foundational aspect for developing sophisticated reward models. These reward models could quickly provide real-world feedback or experience to other AI systems.

18:41It could catalyze future research on reinforcement learning for language models, helping them learn much faster in complex environments. So it feeds back into improving AI itself, which leaves us with a big question to ponder. right? What does this all mean for the future of predictive AI? Could these universal simulators built on text truly transform how we understand and manage basically any complex system, no matter how messy or textual its data might be? It's definitely something to think about.

From the publisher

This paper introduces **text-to-text regression (RLM)** as a novel approach for **predicting system performance metrics**, particularly in complex industrial environments like **Google's Borg compute cluster**. Unlike traditional methods that struggle with non-tabular data, RLMs **directly process raw text inputs** from system logs and configuration files to deliver highly accurate floating-point predictions. The research highlights the **importance of maximizing feature observability** and **large-scale pretraining** for superior performance and **efficient adaptation to new tasks** with minimal additional data. Ultimately, this work positions RLMs as **versatile and scalable tools** for creating **universal simulators of real-world outcomes**.

More from Best AI papers explained

All 475 episodes
Performance Prediction for Large Systems via Text-to-Text RegressionBest AI papers explained · 19 min
Listen in VO