In short
“PrefillOnly” inference engine for LLM applications where each request outputs exactly one token (e.g., yes/no, option A/B/C, label). It targets long-input discriminative workloads like recommendations and credit verification, improving latency/throughput versus standard text-generation engines.
Guest backgrounds
No guests are named in the transcript; it’s a host-led discussion of a research paper.
Key claims
Pre-fill-only is faster because it avoids token-by-token generation; standard engines waste KV-cache memory and struggle with scheduling because job completion time (JCT) changes with cache availability.
Notable examples
Recommendation tasks with long user histories (11K–17K tokens) and credit verification with 40K–60K token histories. Reported results: up to 4x higher queries/sec at same latency; up to 5x longer max input length without parallelism. Techniques: hybrid prefilling (7.9x MIL boost), suffix KV-cache discarding/offloading, continuous JCT calibration with fairness.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Prefill Only Workloads
0:45 to 3:16
Exploration of the emerging trend of pre-fill only workloads and their implications for LLM applications.
“unexpected roles, the problems they run into, and how this clever new engine boosts their speed and efficiency for these jobs, which, you know, could lead to smarter stuff in your everyday life.”
Advantages of Prefill Only Workloads
3:16 to 6:28
Discussion on the advantages of using LLMs for specific tasks like credit verification and recommendations.
“but even with one token it can still show preference you know like the probability it signs to yes versus no.”
Challenges Faced by Traditional LLMs
6:28 to 10:05
Analysis of the limitations of traditional LLM engines in handling long inputs and single-token outputs.
“And it only gives you a bit more context length, less than double.”
Innovations in Prefill Only Engine
10:05 to 13:50
Overview of the techniques employed in the pre-fill only engine that enhance performance and efficiency.
“Every single time it's about to schedule a new request, it quickly re-estimates the JCT for all waiting requests, considering the current cache state.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Okay, let's jump right in. What if I told you that for a lot of companies, the really valuable thing they're doing with these big, large language models? Yeah. Well, it isn't actually writing long emails or code. All right. It's something much, much shorter, super specific, actually. And it's kind of a hidden powerhouse. It's fascinating, isn't it? We've been looking at some research highlighting this whole new type of LLM job. They're calling it pre-fill only. And it really makes you rethink how these models are best used. And tied right into that is this new inference engine, also called pre-fill only, which sounds like it's purpose-built to make these specific tasks fly.
0:40Exactly. So our mission today really is to unpack why LLMs are shifting into these, well, unexpected roles, the problems they run into, and how this clever new engine boosts their speed and efficiency for these jobs, which, you know, could lead to smarter stuff in your everyday life. Yeah, absolutely. Because when most people hear LLM, they immediately think generative tasks, right? Like chat GPT writing emails or GitHub Copilot generating code, stuff where the output is variable length, often quite long. The chatty stuff we see all the time. But you're saying there's this other kind of growing trend where they're used completely differently.
1:17Totally. LMs are popping up more and more for what we call discriminative tasks. These are things that traditionally used older deep learning models, think recommendations, checking credit worthiness, labeling data. Okay, so less creating new stuff, more making a specific call, like a yes-no. Precisely. Instead of generating paragraphs, they're making a very focused decision or classification. That feels like a big shift for businesses using them. Why make that jump? What's the real advantage there? Is it just easier? Well, ease of development is definitely a big part of it. Traditional deep learning, it often involves tons of data cleaning, feature engineering, model tuning.
1:54It's a whole process, often across teams. Right. That sounds like a headache. It can be. LLMs, though, can often take raw data, and they're general enough that you might not even need fine-tuning. Developers can just tweak the prompts to improve results. It cuts down that cycle time quite a bit. Faster to market. Okay. Faster development, fewer steps. Yeah. Makes sense. Yeah. But what about the quality? Can you really trust an LLM for something critical like, say, credit verification? Is it accurate enough? That's the second piece, and it's crucial, decision quality. It turns out that if you pick the right LLM and you're careful with your prompts, you can get decision quality that's just as good, sometimes even better, than the specialized models companies have used for years.
2:34Wow. Okay. So they're not just easier. They're actually really effective. Yeah. So LLMs are good at these decision tasks, but then there's this kicker, this single token output idea. What exactly is a pre-fill only workload? So the defining thing here is that the LLM for each request only generates one single token as its final answer. That's it. One token. One token. That's enough to give the answer, like yes or no, or picking option A, B, or C, or applying a label. So back to that recommendation example, the prompt might be recommend this doc to this user. answer yes or no and the LLM just outputs yes.
3:10Exactly and you can actually force this you give the engine a list of allowed tokens like just the token IDs for yes and no and it has to pick one but even with one token it can still show preference you know like the probability it signs to yes versus no. And the benefit seems obvious speed latency must drop like a rock. Oh massively processing the input takes time but generating output token by token is slow if you only generate one. Yeah, the paper said like a 2K input, one token output is 1.5 times faster than a 2K input, 256 token output. That's a big deal for anything in real time. Huge deal.
3:46And it also gives you really predictable, controlled output, which you don't always get with the big generative models. They can sometimes, you know, go off on tangents. Okay, so these pre-fell-only jobs sound great. What are their key features that this pre-fell-only engine is built for? You mentioned long inputs. Yeah, very long inputs often. For recommendations, a user profile could have months of browsing history, tens of thousands of tokens easily. All that just to get a single yes or no. Seems like it, yeah. Needs the context. Second thing, and this is kind of counterintuitive, these workloads aren't usually limited by GPU memory like we often think with LLMs.
4:23They're limited by GPU computation. Really? By the processing power itself? The sheer math involved in processing that long input is the bottleneck, not just fitting it in memory. And third, the volume is just enormous. Recommendation systems might need tens of thousands of these decisions per second. Oh. Yeah, which means you need hundreds, maybe thousands of top-end GPUs, like H100s. And you generally can't mix these jobs with generative ones easily. They interfere too much. Okay, incredible scale. But if the standard LLM engines are built for generating lots of tokens, they must kind of stumble on these single-token, long-input tasks.
4:59They really do. They're fundamentally designed for a different job, assuming you might generate arbitrary length outputs. So they make trade-offs that are just wrong for pre-fill only. Where does that mismatch show up most? What breaks down? A big one is the KV cache. You know, those intermediate results from the attention layers? They're stored in GPU memory to speed up generating the next token, and the one after that. Right, the model short-term memory for the conversation or text it's building. Exactly. But if you're only generating one token... You don't need almost any of that cap. Precisely.
5:30Most of those stored values are never reused. But they still take up a ton of GPU memory. Like the paper mentions for Llama 3.1 8B, 100 ,000 input tokens need about 12 gigs just for the KV cache. 12 gigabytes of mostly wasted memory. That must seriously limit how long your input can be, right? Your max input length or MIL? Hugely. It's a major reason engines hit a wall on input length. And another thing is job completion time, JCT. Ah, yeah, predicting how long a job will take. With variable output lengths in traditional LLMs, JCT is really hard to guess. That means you can't use smart scheduling tricks like shortest remaining job first very effectively.
6:09GPUs aren't used as efficiently. So what happens now when an engine hits that input length limit with these jobs? What are the workarounds? Sounds like they'd hurt performance. They do. They're compromises. One is chunked pre-filling, process the input in pieces. But that messes with the attention calculations, slows things down, maybe 14 % throughput hit. And it only gives you a bit more context length, less than double. Not great. Then there's parallelism. Tensor parallelism splits the model across GPUs. Can lower latency if you have super fast connections like NVLink, but it always costs you overall throughput because of communication between GPUs, especially bad without NVLink.
6:50Right, more overhead. And pipeline parallelism is another way to split it, but it can create these bubbles where parts of the pipeline are idle, especially if request lengths vary. Again, less efficient, basically trying to fit a square peg in a round hole. Okay, so existing tool struggle, which brings us to pre-fill only. You said it's built specifically for this. Exactly. First engine designed from scratch ground up to exploit how these pre-fill only workloads actually behave. It's not patching an old system. It's a new approach. All right, let's get into the mechanics. How does it beat that memory limit?
7:22You mentioned a hidden memory hog. Yeah, so technique number one is hybrid prefilling. Turns out it's not just the KV cache eating memory. During the calculation, the linear layers, the MLP modules, they create these huge intermediate tensors. Okay. And these can be enormous, like 14 times bigger than one layer's KV cache for that Lama model. They cause massive temporary spikes in memory use. 14 times. Wow. Wow. Okay, so how does hybrid prefilling fix that? It's clever. It processes the non-attention layers, the ones making those huge intermediate tensors chunk by chunk. So you only have one chunk's intermediate tensors in memory at a time.
8:02Peak memory usage goes way down. But it leaves the attention layers alone, processes them normally. Right, because chunking attention hurts performance, as we saw. This hybrid approach gets the best of both worlds. And the result. It lets prefill only handle much longer inputs. The paper says it increases the maximum input length by 7.9 times. Almost 8x longer sequences without sacrificing throughput. That's huge. That is huge. Okay, what's technique number two? Something about the KV cache. Yeah, suffix KV cache discarding or offloading. Since you only need the one output token, you don't need the KV cache entries for tokens near the end of that super long input.
8:38Those are the suffix tokens. So you just toss them. You can discard them entirely or maybe offload them to cheaper memory if needed. but crucially, you keep the KV cache for the beginning part of the input, the prefix. Why keep the prefix? Ah, because future requests might start with the same prefix. Think of user session data. Maybe the first part is common. If you keep that prefix cache, you can reuse it, saving computation later. Smart. So you save memory by ditching the suffix cache, but you keep the potentially reusable prefix cache. Allows longer requests without needing parallelism. More jobs fit on the GPU.
9:13Exactly. More efficient use of the hardware. And the third technique tackles scheduling. Continuous JCT calibration. Right. Because with pre-fill only, the job completion time should be predictable, right? It's always one token output. It is predictable in printable, which means you could use those efficient JCT-aware schedulers like shortest remaining job first. But you said before that traditional JCT scheduling doesn't work well because of caching issues. Right. The problem is, a waiting request's actual JCT can change on the fly. If the prefix cash it needs suddenly becomes available because another job finished, its remaining time drops.
9:50Or if the cash it needed gets kicked out, evicted, its time increases. Ah, so the shortest job keeps changing. Exactly. Traditional methods don't adapt well, leading to low cash hit rates. Pre-fill-only solution is continuous calibration. Every single time it's about to schedule a new request, it quickly re-estimates the JCT for all waiting requests, considering the current cache state. So it's constantly reprioritizing based on who can benefit most from the cache right now. Precisely. It dynamically steers jobs towards available prefix caches, leads to lower latency, much higher cache hit rates.
10:25Very clever. Like an air traffic controller constantly updating landing sequences based on open gates. And you mentioned it has fairness built in, too. Yeah, it adds a little offset to the JCT based on how long a request has been waiting so jobs don't get stuck waiting forever just because their cache isn't ready. Prevents starvation. Okay, these techniques sound great theoretically, but did they test it? Yeah. What's the actual performance gain? Oh yeah, they tested it thoroughly. Built it on top of VLLM, which is already a very strong baseline engine. Tested on different NVIDIA GPUs, L4, A100, H100s, with and without MVLink.
10:58Used various LLM. And real world type scenarios. Yep, two main ones. A post-recommendation task, lots of cash reuse potential, long user profiles like 11K to 17K tokens. And a credit verification task, really extreme input lengths, 40 ,000 to 60 ,000 tokens of credit history. Okay, the moment of truth. What were the numbers? Pretty stunning, actually. Throughput. Pre-fill-in only handled up to four times more queries per second. Forward I-back QPS and paired to the baseline without increasing latency, average, or P99. Four times the throughput at the same latency. That's huge. It's a massive improvement.
11:35And for maximum input length, it increased that by up to five times without needing to resort to parallelism. Driven largely by that hybrid prefilling technique, I imagine. Yeah, that 7.9x MIL boost from hybrid prefilling is the core enabler there. Basically, at the high query rates typical of these applications, prefill only consistently gives the lowest latency because it avoids all the overheads of parallelism or chunking. It just runs more efficiently. Is there ever a time the old ways are better? Well, they acknowledge that at very low QPS, like if the system is mostly idle, using tensor parallelism across multiple GPUs might give you slightly lower latency for a single request.
12:13Okay. But that comes at a huge cost in overall throughput. It just doesn't scale well. As soon as you have real traffic, pre-fill-only's efficiency wins out. It's designed for high-demand scenarios. So, wrapping this up, why should this matter to, you know, someone listening? This feels technical, but what's the real world impact? It matters because this kind of performance boost unlocks using LLMs for things where they were just too slow or expensive before. Think real time recommendations that instantly adapt or super fast fraud detection or near instant credit checks. Things that need to happen now.
12:46Exactly. It makes LLMs practical for these high volume critical decision points and systems we use every day. It moves them beyond just being chatbots into being these silent, super fast decision engines embedded everywhere. It lets companies build fundamentally smarter, quicker services for you. Okay, so to recap, LLMs are finding this unexpected niche in making super-fast single-token decisions for things like recommendations and verification. Uh-huh. These discriminative, pre-fill-only tasks. But traditional engines struggle with the long inputs and single-token outputs. Right, especially with memory and scheduling.
13:21And this new engine, pre-fill-only, uses techniques like hybrid pre-filling, suffix cache discarding, and continuous JCT calibration. to dramatically boost throughput, handle way longer inputs, and just make the whole process much more efficient for these specific jobs. It really makes you think, doesn't it? If optimizing for this one specific type of LLM workload yields such big gains, what other specialized LLM users are out there, kind of hidden, just waiting for their own tailored engine? Could we see a future with lots of different highly optimized LLM engines for all sorts of tasks we haven't even thought of yet.
From the publisher
The research introduces PrefillOnly, a novel inference engine specifically designed for Large Language Models (LLMs) used in discriminative tasks, where only a single output token is generated. Unlike traditional LLM engines optimized for variable-length outputs, PrefillOnly significantly reduces GPU memory consumption by only storing the Key-Value (KV) cache of the last computed layer and by using hybrid prefilling to manage intermediate tensor sizes. Furthermore, its Job Completion Time (JCT)-aware scheduling continuously calibrates based on prefix cache hits, leading to improved throughput and reduced latency, outperforming existing solutions in these specific workloads. This approach paves the way for more efficient deployment of LLMs in applications like recommendations and credit verification.




