Performance Prediction for Large Systems via Text-to-Text Regression

30 Aug 2025 · 16 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

A Google Research approach for predicting performance of large, complex industrial systems using text-to-text regression with regression language models (RLMs), avoiding lossy tabular feature engineering and enabling fast, uncertainty-aware predictions.

Guests

No guest names or backgrounds are provided in the transcript; it’s presented as a host “Deep Dive” discussion.

Key claims

Flattening nested, variable-length system data into fixed tensors is “highly lossy” and creates black-box loss of context. RLMs treat both inputs (e.g., YAML configs, logs) and outputs (numeric targets tokenized as strings) as text, reducing epistemic uncertainty. Encoder-decoder architecture and sufficient feature observability improve results.

Notable examples

Google Borg compute cluster scheduling; predicting MIPS per GCU. Reported up to ~100x lower MSE vs tabular methods, ~0.99 rank correlation, near-instant inference vs 1–18 hours per simulation, and multimodal uncertainty via density estimates.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Complexity of Predictions

0:46 to 1:54

Exploration of the challenges in predicting complex systems like city traffic.

“Today, we're plunging into a really innovative approach.”

The Limitations of Traditional Methods

1:55 to 4:17

Discussing the shortcomings of traditional prediction methods like random forests.

“It has real consequences when you lose that context.”

Introduction to Text-to-Text Regression

4:18 to 6:14

An overview of the innovative text-to-text regression technique developed by Google.

“specifically to replicate these cluster states and figure out this MIPS per GCU metric.”

Case Study: Google's Borg System

6:15 to 10:25

A detailed look at how Google's Borg system functions and the metrics involved.

“You just fine tune the existing model with the new data.”

Advancements in Regression Language Models

10:26 to 12:18

Explanation of regression language models and their advantages in handling complex data.

“You can literally see it in the plots in the paper.”

Performance Improvements and Implications

12:19 to 14:00

Discussing the significant improvements in performance prediction using RLMs.

“That really speaks to the power of transfer learning.”

The Power of RLMs in Predictions

14:00 to 14:33

Learn how RLMs drastically reduce computational costs and improve prediction accuracy.

“what you might call universal simulators for real-world outcomes.”

Simplicity Over Size: The Right Architecture

14:33 to 15:46

Discover why smaller models can outperform larger ones in specific regression tasks.

“And, you know, the simplicity and effectiveness of these relatively small models, around 100 million parameters, you said, that's eye-opening.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive, where we sift through the latest research and deliver the most important insights directly to you. So imagine trying to predict exactly how a vast, intricate system, let's say, I don't know, a sprawling city's traffic network will behave under countless different conditions. It's a mind boggling task, right? Often bogged down by just immense complexity. Totally. And the traditional ways of doing this, they usually demand you take all that messy real world information and somehow squeeze it into neat numerical tables. Right. But the problem there, as our sources point out, is when you do that, you inevitably lose crucial context.

0:39And the system you're trying to understand becomes this kind of black box. You get a prediction, sure, but you can't truly understand why it's acting that way. That loss of context is key. Today, we're plunging into a really innovative approach. It's coming straight out of Google Research, and it tackles this head on. We're talking about using something called text-to-text regression, powered by regression language models, to predict the performance of these incredibly complex industrial systems. But before we get to the, you know, the cool solution, let's really nail down the problem they were trying to solve.

1:12Yeah, absolutely. What's truly at the heart of the challenge of these big, complex systems is just how to handle the sheer volume of, well, messy, non-tabular data. Think about configuration files, maybe sprawling system logs, deeply nested operational parameters, stuff like that. It's just not neat spreadsheets. Exactly. And traditional methods like, say, random forests, they really struggle because they demand everything fit into this rigid numerical format, a flat fixed length tensor, as the research calls it. Right. Trying to force all that rich contextual data into simple numbers, like you said, makes it a black box.

1:44Our sources explicitly call this flattening process highly lossy. Meaning vital information just vanishes. Poof. On in translation. basically. And that's not just, you know, an academic quibble. It has real consequences when you lose that context. So our mission today is to unpack how this new text-to-text regression technique tackles these challenges and not just tables them, but achieves pretty astonishing accuracy, adapts quickly, and even understands the uncertainty in its predictions, offering you, the listener, a direct shortcut to getting informed about where system simulation might be headed.

2:18Okay, so let's unpack this with a real-world example. Makes it easier to grasp. Many industries rely heavily on predicting system performance latency, execution times, maybe key productivity metrics. And a perfect illustration of such a complex, sprawling system is Google's Borg, their massive compute cluster scheduling system. Oh, yeah. Borg is truly a behemoth. Imagine a central manager juggling tasks across thousands and thousands of machines. These machines are located in various sales think data centers. And every single user request with its own specific resource needs, its own constraints, contributes to this constantly shifting, really dynamic landscape.

2:54And the system's main goal, it's to maximize a critical productivity metric called MIPS per GCU. That stands for millions of instructions per second per Google computing unit. Intuitively, it just represents the useful work done by the system per unit of time and computing resources. It's a really vital measure of its efficiency. So you want to predict this MIPS per GCU, but the things influencing it are just incredibly diverse, right, and dynamic, like current memory and CPU usage, the specific hardware types involved, the complex mix of workloads running. Even the time of day matters. Right, and all the various hyperparameters for the scheduling algorithm itself.

3:35And here's the real kicker. According to the sources, these features are often deeply nested. Like you have job profiles saying which machines are used, and then those profiles detail the hardware specs of those machines. It goes layers deep. Exactly. Layers upon layers. So given all that intricate nested information, what's the single biggest headache in trying to model that with traditional numbers? Honestly, the biggest headache is the feature engineering itself. It's just a colossal, time-consuming effort. Trying to squash all that multilayered variable length info into flat numerical inputs for, you know, standard models.

4:09It's like trying to describe a really complex recipe on a single line of a spreadsheet. You're just going to miss a lot of crucial details. Good analogy. And our sources reveal that Google actually developed a digital twin of Borg, a really sophisticated backtesting framework, specifically to replicate these cluster states and figure out this MIPS per GCU metric. Okay, that makes sense. Build a simulator. But here's the catch. generating even one outcome from this fancy simulator can take anywhere from one to, get this, 18 hours of computation. 18 hours for one prediction. Yep. Which raises a pretty important question, doesn't it?

4:46How can you efficiently optimize such a critical system or explore different scenarios if each single prediction takes that long? It just becomes this enormous bottleneck for innovation, for tuning, for everything. Okay, yeah. That's a major problem. Which brings us to where it gets really interesting. Our sources introduced these regression language models, or RLMs, as a general scalable alternative. Essentially, it's a text-to-text regression approach. So instead of trying to force all that messy, complex system data into a rigid numerical box, they treat everything as text. And that right there is the game changer.

5:21It completely avoids the entire headache of manual feature engineering. The RLM can directly process raw text, like YAML configuration file, system logs, you name it. That feeds it straight in. Straight in. The sources show an example string representation for Borg features. It includes things like cell, A, a timestamp like time, 2024, ISX12, T0000000000Z, lists of users, and those deeply nested profiles we talked about. The model just reads it all as text. From a developer's point of view, this is almost like a cheat code for that whole data wrangling nightmare. Ah, I bet. So no more spending countless hours trying to manually group features or define like finite classes for every category.

6:00The RLM just reads it all in as one variable length text sequence, almost like a story describing the system's current state. Precisely. And that flexibility means you don't have to start over from scratch when new hardware or new kinds of jobs pop up. You just fine tune. You just fine tune the existing model with the new data. Exactly. And this directly minimizes what researchers call epistemic uncertainty. All right, what's that again? That's the uncertainty that pops up because a model can't see the full picture, the full state of the world. By absorbing all the available features, especially those tricky nested ones that tabular formats struggle with, the RLM gets a much clearer, much more complete view of the system.

6:39Got it. More data in, less uncertainty about what's missing, and the outcomes. They're also text. How does that work for a number prediction? Yeah, that's clever too. The sources explain that the numeric prediction itself, like the MIPS per GCU value, is also represented as a structured string of tokens. For instance, they might break down the number 72.5 into something like plus 725E1. Ah, okay, so it tokenizes the number itself. Right. Keeps everything consistently in the language models. Language, you could say. Both input and output are text. But what's really intriguing and maybe a bit counterintuitive if you follow the large language models are some of the design choices they made.

7:20Yeah, I was really struck by that in the sources. Some of these design choices seem to, well, go against the grain of what we typically expect from LLMs. For example, they found these RLMs don't necessarily need to be pre-trained on, like, massive amounts of human language like English. What's the thinking there? That's a great question, and it really highlights what's different about this task. The paper points out that regression, well, it primarily requires learning strong correlations between structured tokens. It's about spotting patterns in the data, not understanding the deep semantic meaning or linguistic nuances of human words.

7:52Think of it like this. A chef learning a recipe doesn't need to read poetry right. They just need to learn how different ingredients and steps correlate to a delicious outcome. Makes sense. So a randomly initialized language model can actually learn this tabula rasa from scratch directly from the system's own data. No need for that huge initial dose of human text. Interesting. Okay, another point that jumped out. Yeah. The explicit use of an encoder-decoder architecture. Most of the big headline-grabbing LLMs today are decoder only, which people often say simplifies the design. But the sources here emphatically state that separate encoder layers are, quote, necessary for processing complex inputs.

8:34Why stick with that? That's a really crucial distinction they make. While decoder-only models are fantastic at generating outputs from relatively simple, maybe concise prompts, they tend to struggle when you feed them the incredibly intricate, highly detailed prompts that represent a complex system state like Borg's. Right. The input is much richer here. Exactly. Imagine trying to explain an entire symphony orchestra's arrangement in just one sentence versus giving the conductor the full musical score. The encoder decoder structure is key because those encoder layers are specifically designed to thoroughly process and understand that rich, complex input before the decoder even starts generating the prediction.

9:15So the encoder does the heavy lifting on understanding the input. Precisely. It's built for deep comprehension of complex inputs. Okay, so given these somewhat unconventional design choices, what does this actually mean for performance when they applied it to a real system? Our sources report some truly incredible results for the RLM approach on Google's Borg. You mentioned 100 times lower MSE. That sounds staggering. For someone managing a system like Borg, what's the most immediate, tangible benefit they'd see from that level of accuracy improvement? The results are indeed striking, almost hard to believe initially.

9:47The RLM, even a relatively small one like 60 million parameters, which isn't huge by today's standards, achieves up to a near perfect 0.99 rank correlation on BORG. The average is around 0.9. Wow. You're perfect. And, yeah, that 100 times lower mean squared error compared to traditional tabular methods, the most immediate benefit, unparalleled efficiency and predictability, plain and simple. They can go from those hours-long, compute-heavy simulations to getting predictions almost instantly. Lage difference. Which lets them optimize resource allocation, tweet scheduling algorithms, anticipate problems, all with a precision that was just impossible before.

10:26You can literally see it in the plots in the paper. The RLM's predictions just hug that ideal diagonal line super closely compared to the actual targets, other methods. Their predictions are scattered all over the place. That visual makes it very clear, a massive shift in capability. And it goes beyond just accurate single point predictions, right? The RLM also captures the densities of complex outcome distributions. What does that mean for us, the listener, in practical terms, especially for understanding uncertainty? Right. That's another really powerful aspect. It means the model doesn't just spit out a single number.

11:00It can actually tell you the probability distribution of possible outcomes. Okay, like a range of possibilities. Exactly. And the likelihood of each. This is absolutely critical for understanding and quantifying uncertainty. Think about predicting tomorrow's weather. A single temperature is useful, sure, but knowing there's a 40 % chance of rain, 30 % sun, 30 % clouds, that gives you a much richer, more actionable picture, doesn't it? Definitely. Our sources show the RLM can capture these multimodal distributions, meaning there might be several distinct likely outcomes, not just one average point.

11:34So it might predict two or three common performance levels under certain conditions. Precisely. And you can see this visualized really vividly in what are called kernel density estimate plots. The RLM's predicted density curves closely match the actual spread of target points over time. It even shows multiple humps of likelihood where they actually occurred. That's fascinating. And importantly, it also shows a strong correlation between the model's predicted uncertainty, how confident it says it is, and its actual error, which is incredibly useful for building trust and assessing risk. That makes sense.

12:06You want the model to know when it doesn't know. And the adaptability piece is also really impressive. The sources highlight the RLM adapting to entirely new tasks with just 500 few-shot examples, maintaining strong accuracy even on compute clusters or scenarios that I hadn't seen before. That really speaks to the power of transfer learning. It really does. Pre-training, even if not on human language, but on a diverse set of system simulation tasks, helps the model learn these sort of universal relationships and patterns. Patterns that apply across different systems or scenarios, making it incredibly agile.

12:43Like learning the grammar of system behavior. Kind of, yeah. It's like teaching a kid the alphabet before asking them to read a totally new book. That foundational knowledge is highly transferable. And ablation studies in the paper confirm this. Using more pre-training tasks leads to substantially better results on unseen data, especially when you add just a little bit of fine-tuning data, that few-shot learning. They also found that maximizing feature observability is crucial, basically just letting the model see as much of the complex input data as possible. For instance, they explicitly test it, including the time window feature, adding that significantly boosted performance.

13:21Which makes sense, right? Workloads probably follow cycles. Exactly. It validates known domain knowledge about how workload cycles impact system behavior. The model learned that pattern because it saw the relevant data. So pulling this all together, what does this really mean for us? We've seen how this text-to-text regression using these regression language models is kind of transforming performance prediction, at least in complex industrial systems like Google's board. It really feels like it's about moving beyond the, frankly, severe limitations of traditional tabular data and letting models truly speak the language of the system itself.

13:54Text. I think that's a great way to put it. This work feels like a really significant step towards creating what you might call universal simulators for real-world outcomes. By alleviating that huge burden of manual feature engineering and by robustly handling these diverse, deeply nested inputs, RLMs can give highly accurate predictions almost instantly. Negligible inference time, they call it. Exactly. And that translates to huge savings, not just avoiding those 18-hour simulation runs, but also much faster optimization cycles, deeper insights into how these complex systems really behave. All without the massive computational cost of the old simulators.

14:36Right. And, you know, the simplicity and effectiveness of these relatively small models, around 100 million parameters, you said, that's eye-opening. They're proving that for these kinds of regression tasks, it's not always about building the absolute biggest language model possible. Absolutely. It's more about using the right architecture, that encoder decoder structure in this case, and the right representation, treating data as text. The right tool for the job. Indeed. And this capability, accurately modeling numeric feedback from all sorts of textual inputs, it also sets a really important foundation for future research.

15:07It paves the way for developing more sophisticated reward models. Reward model. Yeah, models that can give realistic, real-world feedback operational experience, you could call it, to other AI systems, which in turn could significantly accelerate advancements in areas like reinforcement learning, maybe even for language models themselves, feeding them better signals. Interesting loop there. It really makes you wonder, doesn't it? If simply treating complex system logs and configurations as pure text unlocks such profound predictive power for system performance, what other intricate, seemingly impenetrable data problems, maybe across totally different domains, could be revolutionized just by letting our AI models read and write their way to understanding.

15:50Something to chew on until our next deep dive.

From the publisher

This paper introduces text-to-text regression as a novel approach to predicting the performance of large-scale industrial systems, like Google's Borg compute cluster. Unlike traditional tabular methods that struggle with complex, non-tabular data such as configuration files and system logs, this method utilizes encoder-decoder Regression Language Models (RLMs). The research demonstrates that these RLMs can achieve high accuracy (up to 0.99 rank correlation), adapt efficiently to new tasks with minimal new data, and accurately capture the densities of complex outcome distributions. The findings highlight the importance of observing comprehensive features, extensive pretraining for transfer learning, and the model's inherent uncertainty quantification, paving the way for more universal system simulators.

More from Best AI papers explained

All 475 episodes
Performance Prediction for Large Systems via Text-to-Text RegressionBest AI papers explained · 16 min
Listen in VO