86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic

8 Sep 2026 · 55 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Subquadratic’s approach to long-context “sparse attention” that reduces quadratic attention compute so models can process multi-million-token inputs efficiently, enabling more robust RAG/agent workflows and better long-horizon reasoning.

Guest backgrounds

Alexander Whedon, co-founder and CTO of Subquadratic. ~10 years in AI; worked on language models pre-transformers (LSTMs), then roles including head of Generative AI at Tribe AI (40+ enterprise projects), Instagram (creator monetization), Stitch Fix (first text-based recommendation), research/production at academia.edu (200M users), and research work with World Bank/Blue Cross Blue Shield.

Key claims

  • Benchmarking on SueBench Pro: 86% of coding-agent steps are “read” (context engineering) and only 14% execution.
  • “Not the end of RAG,” but RAG should change: larger chunks, more retrieval results per step, fewer aggressive context-engineering constraints.
  • Their SSA avoids DeepSeek-style separate selection model; single model dynamically selects attention without reintroducing quadratic cost.
  • Reported efficiency: ~40x faster than flash attention and ~64x less compute at 1M tokens on B200/B300; enables more multi-million-token pre-training.

Notable examples

  • Coding: code review planning, long-horizon memory, and multi-step agent workflows.
  • Finance: analyzing hundreds of pages across companies; harder PDF/table parsing.
  • Security: long-context helps find vulnerabilities in codebases (defensive), not necessarily product exploits from outside.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Quadratic Compute Complexity

0:00 to 1:05

Learn about the challenges posed by quadratic compute complexity in transformer models.

“In transformer models, the cost of attention grows with the square of the context length.”

The Role of Sparse Attention in AI Models

1:23 to 2:27

Explore how sparse attention reduces the computational burden of AI models.

“the cost of attention grows with the square of the context length.”

Dynamic Token Relationship Selection

2:27 to 5:08

Understand the innovative approach to dynamically selecting token relationships in attention mechanisms.

“You have this quadratically growing set of math that you're going to perform with an attention.”

Use Cases for Large Context Windows

5:08 to 7:20

Discover practical applications of AI in enterprises that require large context windows.

“so without a separate selection mechanism.”

Challenges of Intelligence in Large Contexts

7:20 to 11:01

Learn about the difficulties in maintaining intelligence within large context windows in AI models.

“Yeah, and that quality offset, or you said earlier a moment ago that it's not only the cost, but it's the intelligence within the context.”

Innovations in Sparse Attention Mechanisms

11:01 to 14:00

Delve into the unique innovations that improve efficiency in large AI model trainings.

“And then finally, I will just add that one of the challenges with large context intelligence and user preference matching is the amount of training that you can do at large context.”

Enhancing RAG with Large Context Windows

14:00 to 19:30

Explore how large context windows can simplify RAG and improve coding efficiency.

“We are around 40 times faster than flash attention before on B200s and B300s.”

Pre-Training and Vulnerability Detection in Code

19:30 to 28:00

Understanding the importance of pre-training in AI models and its application in cybersecurity.

“And you talked about doing multi-million token pre-training rather than just post-training.”

Data Processing in Enterprises

28:00 to 29:43

Learn about the challenges and solutions in enterprise data processing and unstructured knowledge extraction.

“One thing that's great for us is that I think unstructured document processing and knowledge-based products are in pretty much every enterprise out there.”

Future of Coding Agents

29:43 to 32:32

Explore how coding agents are evolving and the implications for efficiency and context management.

“Context engineering step is where the vast majority of human and capital investment is today.”
Show all 20 chapters

Founding Story of Subquadratic

32:32 to 34:45

Discover Alexander Whedon's background and the motivations behind founding Subquadratic.

“will be critical and will also require much longer context reasoning than we have today.”

Memory and Context in AI

34:45 to 37:55

Learn about the importance of memory and context in AI models and the challenges faced.

“This long context, you mentioned that most enterprises don't even fully use 256K context.”

Evaluating AI Performance Across Vertical

37:55 to 40:25

Understand the disparities in AI performance across different domains, particularly financial documents versus coding.

“what user preferences are at scale before this becomes a problem that we can really solve through prompting.”

Theoretical Benefits of Sparse Attention

40:25 to 42:01

Delve into the potential advantages of sparse attention in AI model training and its impact on results.

“So I guess the TLDR here is that long reasoning capability today is very asymmetric across verticals.”

Exploring Sparse Attention in AI Models

42:01 to 44:11

Understanding how sparse attention might improve AI model performance.

“phenomenon that we call the bitter lesson in the ai space uh yeah yeah um and so So that's a big part of it.”

Design Partnerships and Access Plans

44:12 to 45:07

Discussion on design partnerships and future access to AI models.

“We usually go through an evaluation process, scope a design partnership, and then get started.”

Leveraging LLMs in Research and Development

45:08 to 47:26

Insights on the use of large language models in algorithm design.

“One of them is like a higher output to input ratio.”

The Future of AI Architecture

47:27 to 49:46

Looking ahead at advancements and changes in AI architectures.

“to come up with an idea and it comes up with something I hadn't thought of.”

Robotics and AI Integration Challenges

49:47 to 52:07

Challenges in integrating AI with robotics for complex tasks.

“I think that from a go-to-market perspective, we're really excited about working with enterprises on long context reasoning tasks, long horizon agentic tasks as a secondary priority.”

Company Growth and Research Insights

52:08 to 53:16

Overview of company growth and past research efforts leading to today.

“So that's a pretty exciting frontier for us.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00In transformer models, the cost of attention grows with the square of the context length. So doubling context makes the compute roughly four times more expensive. And Alex and his team have solved this issue. When we've done benchmarking of Frontier Models with SueBench Pro, we found that 86 % of the steps were actually read steps, just trying to do that context engineering before execution review, which was only the last 14%. Is this the end of RAG? If you could put everything that's in your RAG knowledge base into the context window, you wouldn't need that more complicated architecture. Is that right?

0:41Yeah, it is. Although I wouldn't say it's the end of RAG, just because it's too extreme. All of these enterprises that are sitting on massive amounts of data, and a lot of that data has yet to be put to use in an AI product, they're getting pushed to do like a$10 million data a transformation project before they could even start to build a product on top. And value prop that we bring to the table is that is no longer necessary. Hi, this is Eye on AI. And today we're talking to Alex Whedon about his company SubQ, which solves a problem that plagues transformer models. In transformer models, the cost of attention grows with the square of the context length.

1:30So doubling context makes a compute roughly four times more expensive. That's what we mean by quadratic compute complexity. And Alex and his team have solved this issue. uh so alex can you start by introducing yourself to listeners and we'll we'll go from there sure i'm alex co-founder and cto here at subquadratic um i've been in the space for about 10 years i started working with uh language models since before transformers existed when lstms were still the hot thing. And our mission here is to always be the first to make the most computationally memory and sample efficient algorithms and build foundation models on top of those.

2:27Yeah. And can you explain, you've tackled this quadratic complexity with sparse attention am i right and can you explain what sparse attention is in simple terms and why it matters compared to standard transformer attention sure so um the quadratic bottleneck of attention is introduced because attention compares two token relationships um so you have 1 ,000 tokens, then there are 1 ,000 squared two-token relationships or a million squared two-token relationships for a million tokens. You have this quadratically growing set of math that you're going to perform with an attention. Sparse attention says we don't actually need all of those token relationships.

3:19And so sparse attention attempts to find which token relationships are important before running attention on only a subset of the token relationships. and one of the main challenges with sparse attention is that there's always been a quality trade-off as you can imagine like um in theory if you're not looking at every token relationship then you can potentially miss some important ones however you're also looking at a lot that don't actually matter um and so the challenge of sparse attention has been figuring out which ones actually matter most sparse attention mechanisms have used fixed patterns so um based off of the distance between the target token and all the other potential tokens.

4:02There's some fixed mathematical formula that determines which tokens should be passed through attention. DeepSeek Sparse Attention upended this by saying we're actually going to use a separate prediction model to determine which tokens are important. And that worked. And they showed that if you dynamically select the token relationships before running attention over a subset then you can do so without a quality trade-off the only problem is that prediction model that selects the token relationships uses full attention and um and so you still have the quadratic problem in fact that uh sparse sorry that uh that separate selection model uses more compute than the sparse attention layers in the full model um at just 52 000 tokens and then continuously grows to an even larger differential.

4:56And so what we've done that's unique is we have a sparse attention mechanism that, like DeepSeek sparse attention, dynamically selects the token relationships, but does so without the massive cost of full attention and does so without a separate selection mechanism. So the single model does both the selection generation at the same time. And can you give us a practical use case where this quadratic compute complexity makes it too expensive to run the model on something? I mean, code bases seems like an obvious, but are there other situations where you want a very large context window and that it would get unreasonably expensive to do that with regular attention?

5:52I mean, regular transformer algorithms. Yeah, absolutely. I mean, I think a very significant portion of enterprise workloads fall into this category. And by the way, I will just clarify now that we're not just after the absolute context window size. We're after higher intelligence within that context window, lower cost of context, lower latency for context. And so these are actually four different axes of context we care about, size, intelligence, cost and latency. and enterprise workloads are dramatically impacted by these. Everybody's making trade-offs. Everybody's trying to reduce the number of tokens that they're using as much as possible.

6:33Everybody is doing context engineering that requires a lot of effort and also narrows the scope of what that system can actually handle. But some of the top use cases that we see are, yes, code is obviously one. Thinking about code review planning and long horizon memory, those are all long context problems. Thinking about the fact that when we've done benchmarking of Frontier Models with Sweebench Pro, we found that 86 % of the steps were actually read steps, just trying to do that context engineering before the execution review, which was only the last 14%.

7:15And then financial document analysis is another really big one. Financial documents can be a couple hundred pages. um and if you really want to analyze trends then it's helpful to have many of those documents over a period of time or several different companies um we've in an enterprise you have hundreds of billions of tokens of textual data and anytime you want to to tackle a problem um it's helpful to have like a broad range of documents um the key point that we're that we're conveying um and working on with enterprises is how do we do coarser grain retrieval more retrieval steps at once um how do we avoid chucking documents as much as possible or make larger more comprehensive chunks etc to offset the quality loss that we often see with a lot of the context engineering that's done today.

8:15Yeah, and that quality offset, or you said earlier a moment ago that it's not only the cost, but it's the intelligence within the context. Is that how you phrased it? Because large context windows, many LLMs don't operate optimally within a very large context window. They don't see things or they get lost in the context. Can you talk about that? How do you solve that? even though you can now do very large contexts efficiently? It's a great question.

9:10So, yeah, the intelligence within the context window is a very important topic. If we just take a step back and think about what Frontier contexts has looked like historically, we didn't get Frontier models with more than a 256 ,000 token context window until March when the Opus 4.6 general access, 1 million token general access came available. And so until then, most enterprises were building with a less than 256K usable context window. You see that, so a lot of people that we talked to, like they're using 100 ,000 tokens or less and they haven't yet adapted their workflows or RAG setups to use more than that.

9:58Sometimes they feel a quality trade-off above that. With particularly hard problems, actually 100 ,000 is a fairly common cutoff, surprisingly. In the coding space, a lot of people set the compaction limit to 400 ,000 tokens because that's where they feel there is a qualitative trade-off within a 1 million token context window. Because of this, in addition to intelligence, I would say there's also an alignment problem within the context window that user preferences with how a model should handle, should leverage 1 million tokens are not super robustly understood today. There are just not enough enterprise applications built to actually leverage a full 1 million token context window to really understand what that should look like.

10:45And so as we push into the multi-millions of tokens, this is very much white space, not even just from a technical perspective, but from our product perspective. what the users want. There are all sorts of benchmarks that kind of measure human preferences at short context, but a long context does not exist. And then finally, I will just add that one of the challenges with large context intelligence and user preference matching is the amount of training that you can do at large context. And so that's one of the main reasons that we set out to create SSA in the first place was that it is cost prohibitive to do large scale training on large context.

11:26As far as we know, we are the only people to do robust multi-million token pre-training, which we think differentiates our models. So like the number of times that a model can see how to handle a certain situation at a reason over some context at large scale, as that number goes up, obviously the model gets a lot better. that's really hard to do with quadratically scaling compute costs so it could go a little deeper on that i mean you've trained uh and evaluated up to 12 million token context windows so explain how the how you avoid uh the problems that large context windows uh create i mean you you go a little deeper on on how how the model approaches that yeah so i mean i feel like i i kind of buried the lead on that with one of my earlier answers but uh i mean the problem with sparse attention historically has been that uh their fixed pattern mechanisms um like sliding window attention look at the last 1000 tokens that closest um and um and uh or before and after but um but these fixed patterns don't robustly model language so like there's no guarantee that the first and seventh token are always going to be relevant to each other um it's highly context or content dependent and so you need a mechanism that accurately models context, accurately determines which relationships are going to be important, no matter what the input looks like.

13:18And so while deep-sea sparse attention uses a separate model to do that, reintroducing the quadratic problem, we do not. We have a mechanism that is much more lightweight and cheaper and leverages the existing model that does the generation to do the selection as well um so it's super cheap super fast um but also dynamic so every input gets a different um attention map and um and then we train our we leverage that the fact that we can do this dynamic selection to achieve high quality results while also having efficiency differentials. We are around 40 times faster than flash attention before on B200s and B300s.

14:13And then we're using 64 times less compute. Both of these figures at a million tokens. So this means that we can do a lot more 1 million token training. with the same compute or two million token training or whatever. And so this enables us to introduce a lot of problems that models have yet to be trained on robustly. The reality is, is we're training on a lot of data that hasn't been including in training to date and introducing capabilities that have yet to be introduced into the models to date because we have this mathematical efficiency differential that uh just makes the math of doing this training so much different just because my audience isn't all um ml engineers um so so the model you use sparse attention um and you how does the model decide what tokens to pay attention to i mean you were saying that deep seek has a separate model that does that prediction but can you go into how your models with these large context windows then can decide which which tokens to focus on Yeah.

15:38I mean, that's the highly proprietary part that we have. But I would say, I guess I could say we creatively use some early information on the inputs in a way that doesn't create all to all comparisons, which is very high level. but i i feel that going deeper than that uh leaves a lot of breadcrumbs well one thing i i wanted to ask you and we've spoken a couple times but uh does this is this the end of rag then of having a vector database that the model searches and i mean could you if you could put everything that's your RAG knowledge base into the context window, you wouldn't need that more complicated architecture.

16:41Is that right? Yeah, it is. Although I wouldn't say it's the end of RAG, just because it's too extreme. So the way it changes RAG. What we're looking to do is simplify the way that people build products with RAG or, you know, agentic workflows.

17:06Examples are today, a lot of documents are chunks that don't necessarily need to be chunked. The chunk sizes are very small. Every 400 tokens is like super common. And they're like, what if you could have a multi-page, multi-section chunk instead? There would be a lot more contiguous context that would prevent misunderstandings. Top K of 10 for the top 10 search results is very common. Well, what if you could pull the top 50? What if instead of doing one type of search, you actually did like 20 parallel searches and dumped all of those search results into the context window? These are ways in which I think we make RAG a lot more, we can enable people to make RAG a lot more robust.

17:56If you have 400 million tokens, there still needs to be some way for the model to process that data. But we're saying don't compress everything into 100 ,000 tokens or less. You can now afford to because of cost and latency. And then also the intelligence and absolute budget of context. You can now pull a lot more context and do a lot more work with every single step in your agent workflow. Like I mentioned before, if we're talking about coding problems, Opus 4.6 required an average of 66 steps to tackle a few bunch of pro problems. And if we assume that we only needed like the last 10%, then we could have gotten away with like six or seven steps if we had been able to just put all the code in the Conics window.

18:48Or even if we took like three or four steps to figure out which parts of the code base are more important, being more aggressive with the context in each step. I mean, that is a much faster, much more efficient process that actually uses fewer tokens in aggregate. So it's more token efficient. And then it's not going to miss the details. We do hear complaints about coding agents missing details and definitely on the longer horizon. So not the end of RAG, but the transformation of RAG as we know it. And you talked about doing multi-million token pre-training rather than just post-training. Can you walk us through how that works and how that changes the model's behavior and practice?

19:47um yeah so pre-training is the stage of training where um you give the model its core knowledge and to some extent some core capabilities um but it's very knowledge focused and so the pre-training simply means you predict the next token over and over again and um and then the model gets penalized based off of how often that next token matches the actual next token in a corpus so it's crude and rudimentary but it's also highly scalable and it's largely seen as a requirement to extract additional capabilities um in the in the post training phase and so um research has largely suggested that your ability to create value in the post-training phase is limited by the pre-training that you've done.

20:44So if you haven't, for example, done a lot of pre-training on code, then you're going to struggle to post-train a model to be really good at code. And so that's why it's important that we've done this multi-million token pre-training, because it means that when we get to post-training, there's already a certain amount of that knowledge there there's a good foundation to build on and in post training we can extrapolate even further and in fact this is something we showed in our technical report for a sub q 1.1 small model um we actually have a model variant where we trained the model we pre-trained the model with 1 million token inputs up to 1 million um but a lot of 1 million token pre-training and we show that that 1 million token pre-training made it so that um when we did post training at largely also only up to 1 million tokens, that the model is able to extrapolate how to do retrieval style problems up to 12 million tokens without ever seeing anything in post training over a million tokens.

21:43So that's the power of pre-training. It enables the model to extrapolate much further than if it's seeing stuff for the first time in post training. Yeah. a lot of people right now particularly with mythos and and fable out now from anthropic and glm 5.2 from yeah jirpu or zai yeah many people worry about large context models making it easier to find vulnerabilities in code bases. Is how do you view that? And and. In in both offensive and defensive cybersecurity use cases. Yeah. I do think that long context windows, long context reasoning capability does help with finding vulnerabilities in code bases, but that's different than finding vulnerabilities in products.

22:48So finding vulnerabilities in code bases enables you to be more defensive. Finding vulnerabilities in products puts your product at risk. So they're two very different things. In our case, if you already have all the code in the code base, being able to reason over a lot more of it at once enables you to find vulnerabilities much faster and also find vulnerabilities you would not have been able to find otherwise. And so we see that as something that's really valuable for people being able to make their products safer and more robust. That is, some of our design partnerships have leaned in that direction.

23:30um i don't think that long context reasoning specifically provides nearly as much an advantage um for being able to find the holes in a product from the outside um because they're more about iterative tests um the context requirement is not necessarily significant um but the way this works is software is built layer upon layer of abstraction, of abstraction. And so, you know, there are probably like 20 layers of different code components underneath the code that you're writing at any given time. And so if I'm trying to hack your product, I'm trying to guess which open... I mean, there are other things too, but this is a big part of it.

24:26I'm trying to guess which open source software you've built on top of. And these open source softwares have known vulnerabilities. And so I'm trying to, I'm testing to see if any of those known vulnerabilities exist in your software. So it's not really a long context reasoning problem. To be honest, I think more than anything, it's actually a knowledge problem. it's it's an existing knowledge of all the the existing vulnerabilities and that's why you've actually seen that people can can get open weight models to to have mythos level capabilities for for hacking because uh the model needs to specifically know to be honest sometimes uh like the number of vulnerabilities that that lead to most of the threat it's lesser than some people realize um because everybody builds on a lot of the same software components so yeah but couldn't this uh i mean to find those well you said they're known vulnerabilities in open source uh code bases but with mythos or or uh you know or i mean with your model that that can handle massive contexts.

25:45Couldn't you then just all the various foundations that are managing open source code bases, they could just dump it all into your model and discover vulnerabilities that maybe they weren't aware of, patch them? I mean, isn't it advantageous for the open source community to have solved the context, compute complexity problem so they can scan large code bases and fix vulnerabilities? Yeah, I think that's a great point. I think partnering with open source communities to find and fix vulnerabilities faster and find it and fix the vulnerabilities have just been there for a long time where speed isn't the issue.

26:39It's the breadth of perspective that's the issue. That'd be a great use case that I'd love to work with folks on. Yeah, you've mentioned design partners a couple of times. what are the most compelling early design partner use cases you're seeing for a large context? Is it finance, code, internal documents? So unstructured document processing and analysis is like the number one right now, and the number two is knowledge-based products. within those two categories finance has emerged as a really important one just because finance companies working in the finance space are building such data intensive applications they're processing so much data that's kind of like the core criteria for whether a design partner early access user is a good fit for us right now is how data intensive is their application So finance is a really big one for us.

27:45We have had a lot of interest on the coding side. We are intentionally moving a little bit more slowly on the coding side because we've kind of realized like how much coding is like a budget game. Like how much data budget do you have? How much label data can you train on? and so uh that's an area that we're really interested in and are working on but with a bit of a longer time horizon and then um and then it's beyond that it's it's like general enterprise document uh processing for knowledge bases and unstructured inside extraction um those are the main use cases that we're seeing. One thing that's great for us is that I think unstructured document processing and knowledge-based products are in pretty much every enterprise out there.

28:42And we're also seeing that because we can work with data that is less structured and we can process more data at once, it becomes a lot easier for an enterprise to get their first value from data. So we have all of these enterprises that are sitting on massive amounts of data, and a lot that data has yet to be put to use in an ai product um or maybe it has but in in a very limited fashion and they're getting pushed to do like a 10 10 million dollar data transformation project before they could even start to build a product on top and a value prop that we bring to the table is that is no longer necessary at least not at that level like we can help you um either work with the data with a lot less structure or create the structure a lot faster because of our ability to process a large amount of data at once with a lot less curation.

29:42That's interesting. And as we move into agents and particularly coding agents increasingly, which is happening, uh most coding agents use very surgical context management i mean they they hop from file to file or uh section to section uh what do you think breaks first when we move to agents that can read everything at once yeah i think that will a couple things will happen um one is we will have systems that can generalize better because every time you add a search engine or you know vector database which is a search engine um or some conditional logic to route between the steps um this human curation really limits the ability of that system to do a lot of different things it makes it so that it works within exactly the scope that you've set it out to to be able to to to achieve and um and uh and the more complicated you make it the less the more narrow that's going to be so i think we'll see agents that can do a lot more um we will also see agents that are a lot faster because you know if you're doing 10 to 60 steps um it's pretty slow so uh i always say agents that are cheaper um we want you know a million tokens to feel like 50 000 tokens in terms of intelligence cost latency etc over time and um and so i think we'll see people build a lot more too, to be honest.

31:36Context engineering step is where the vast majority of human and capital investment is today. And so if we can get rid of that, then we'll see people, like the barrier to entry goes down, people build a lot more things. And when a single step can do 10 steps worth of work what are those 10 to 60 step workflows going to be able to achieve now i mean um it'll be a lot more powerful a lot longer horizon um i guess that's that's another point um we really care about long horizon agents and um the ability to manage context not just through a two-hour session, but across, you know, weeks or months of work is something that I think will be critical and will also require much longer context reasoning than we have today.

32:40Yeah. We didn't go into your background or the founding of the company, and it's interesting. Can you give us that story? Sure. Yeah. So I've been in the AI space for the last 10 years. Most recently, I was head of Generative AI, a mid-market and enterprise AI consulting firm called Tribe AI. I did over 40 projects with a series of mid-market enterprise companies with a medium-sized group of folks underneath me. And that's where I got a really good horizontal view of what generative AI implementation was looking like in 2023 and 2024. And that's kind of what convinced me that it was overdue to move to another model paradigm, that they were bottlenecks in the enterprise implementation layer that were going to be painful for a very long time unless we change the actual mathematics behind the foundation models themselves.

33:46Before that, I was at Instagram working on creator monetization products. Before that, I was at Stitch Fix where I built the first text-based recommendation system. Before that, I was at a research lab where I published some research and built our first production products for our sister companies, Glassdoor and Indeed. Before that, I was at academia.edu where I built some of our first AI products for our 200 million user base. Before that, I did research with the World Bank of Blue Cross Blue Shield. And so I've done a lot of different things. I'm not a domain expert. In fact, that's one of the things I pride myself in is I like to look at things from a very horizontal perspective what's the key value point that's missing across everything everybody's doing and that's really what we're trying to tackle as a company it's like how do we take computational memory and sample efficiency and turn that into downstream business value propositions um context is the first one but there will be many more yeah well and you mentioned memory are you working on memory absolutely so i mean we're working on memory within the context of ssa subquadratics parts attention the thing we talked about publicly we're also working on non-attention algorithms have been for a while um that uh make it totally necessary to have any form of kvcash kvcash is this massive thing right now where at the multi-million token range uh can easily use a lot more memory than the model it's the model weights themselves um becomes very hard to host the context without distributing across multiple cloud data center grade gpus makes it impossible in my view to do long-range modeling for robotics on jets and orans which have about as much memory as your phone and so we see this as a critical problem to solve for long context modeling, for general robotics intelligence, and just the ability to create much better frontier models than exists today.

36:00This long context, you mentioned that most enterprises don't even fully use 256K context. as you introduce them to multi-million token windows what what are you learning about user preferences and behaviors yeah that i mean that they're not well mapped out and they can vary um so we found for example people will tell us um that uh on a news article they'll ask what are the top three insights from here and the model will give three top insights that are technically accurate but they're not the three they cared about um so like this is the level of alignment that becomes necessary uh with long context reasoning gets very complicated so um one of the things we learned is that we are so far from from this area of modeling being mature.

37:09I think people say, well, the frontier models have a million token conics window, but like their ability to deliver what people want at that size is just completely, we are far removed from that today. And, and promptability is still a bit lacking also for these types of problems. And I think to some extent, like this problem can't be solved at the prompt level because it's just so complex and nuanced. We do actually need to solve this through training first, where we see a very wide number of problems that, a very wide number of inputs and outputs or problems through RL that show what user preferences are at scale before this becomes a problem that we can really solve through prompting.

38:01um so that's a super high level insight um beyond that um i think reasoning is pretty unsolved oh here's here's an interesting one actually um we did some we created some evals for code base question answering and financial document question answering um and we ran we did them at 500 000 tokens um a million tokens per opus 4.7 which is like 800 000 tokens for us um and uh and then as close as possible to two million tokens which obviously only we could do and so we found on code-based question answering um um the uh there there were like different levels of difficulties like we could say easy medium and hard and um on the easy and medium problems like the models most frontier models and our model could do pretty well pretty close to 100 percent and uh and there was really not much of a concern there on performance but it is quite easy to create an eval set that on financial document analysis that is well below 50 percent um at 500 ,000 or 800 ,000 tokens and um so it does seem like obviously there's a bias towards code I think that um the model companies have spent like more time in some areas than others and then also I think there are certain types of contexts that harder than others like i've been working with financial documents uh for a long time i've spent a lot of time working with edgar data um and there are a lot of things that make it hard um some of them are intentional but um but the structure of the pdfs can be very challenging for llms to handle um data can be really distributed across the several hundreds of up to several hundreds of pages.

40:12You've got tab viewer and non-tab viewer data. You've got tables that are really hard to parse correctly. All these challenges that I think make these problems harder than code in some ways. So I guess the TLDR here is that long reasoning capability today is very asymmetric across verticals. Yeah. Yeah. On the, you mentioned that you give these models, the standard transformer models, a large context, a large chunk of text in the context window, and ask it to surface the most significant things and it doesn't come up with what you think is most significant. How does this, is the reasoning stronger with your architecture or is it that it sees more of the context?

Read the full transcript

41:21I mean, why is it better with SubQ's model? So the biggest part of it is the training differential that, you know, we can end to end our model training speed is still like somewhere between five to 10x faster at a million tokens. And then as you move to two million tokens and beyond, that's higher. The differential is even bigger than that. and so it's easy for us to do a lot more long context training than other people are doing and we're also very focused on generating a lot more of that data through both synthetic generation and human labeling and so you know focus and scale is a big part of that this is phenomenon that we call the bitter lesson in the ai space uh yeah yeah um and so

42:19So that's a big part of it. I think there's a theoretical benefit that I don't feel like we've fully proven one way or another. But that is that we are reducing the noise for the model. So instead of looking at a million squared relationships, it'll look at, you know, scalar value times a million token relationships. So it has less to consider, less to model on top of. and so the same way that like RAG reduces the noise and can sometimes lead to better results, I feel like sparse attention could also reduce the noise and lead to better results. But again, I don't feel like we've fully proven that one way or another.

43:00It is not going to be easy to prove because A, how do you separate the architecture from the training? And B, what evals are we going to use that are actually helpful today? I think we have a massive dearth of high quality long-conducts reasoning evals. Yeah. And if someone wants to play around with this, how do they do that? So today, I mean, you reach out to us. We'd evaluate you for a design partnership. The number of, we're not giving broad access to the model yet, just because we have limited compute bandwidth ourselves. ourselves and are trying to prioritize access to people that we think could help us make a better version of the product and um and then uh and we are continuously trying to increase the amount of compute bandwidth we have um these additional these early design partnerships are also helpful for that too the market validation enables us to increase our capital budgets and um and thus compute budgets.

44:06So yeah, I mean, that's kind of the TLDR is like, reach out, we'll connect if it seems like an interesting fit. We usually go through an evaluation process, scope a design partnership, and then get started. That'll change over the next couple months. In the next couple months, we do plan to have complete general access. And by then, we will have a very different version of the product too. It will be, our models are continuously undergoing training to respond to feedback that we're getting from partners. So, um, so if anybody thinks the model is good now, it's just going to get a lot better. Yeah.

44:38And when you do have general availabilities, is there going to be a free, uh, layer or a free trial or, or how, how are you going to get people in the door? So, um, yeah, I think we'll have some level of free access for sure. and we'll have tier pricing and tier rate limits and all of that. So it'll be pretty standard. There are some non-standard things that we are considering doing. One of them is like a higher output to input ratio. Today, the output token cost is usually three to five X the input token cost. we'd like the input token cost to be a lot more cheaper oh sorry a lot cheaper still um to make it we really want people to feel like the input tokens are free like just consider the context that you need to for your problem um so i think we'll see higher ratios there um and then secondly we've hesitated saying this a bit um but we've considered um I think our pricing curve as your inputs grow could look different than others because our internal price economics look different for it.

46:00The external pricing for that could also look different. I guess that's all I'll say for now. Yeah, and I have to ask, you know, there's a lot of talk about using AI for scientific research. did you guys use any models to design this? I mean, did you talk to, you know, Claude 4.8 or now I guess Fable 5 and, you know, how do we solve this problem and go back and forth with a model? I'm curious how companies like yours are doing that architecture design, whether they're tapping the knowledge base of the model or whether it's all, you know, iterative trial and error among humans. Totally. So, I mean, we're definitely using LLMs in the process.

47:08and always trying to as much as possible. I don't think you can get there with just an LLM. They're not enough today. I don't think they're creative enough either, but they're definitely super helpful. And I will say sometimes I'll just open-endedly ask an LLM to come up with an idea and it comes up with something I hadn't thought of. So they're super helpful for that. um but um but we've been talking from probably like somewhere between a year and a year and a half about how to get to auto research for algorithms before you know way before auto research was termed way before rsi was termed because of self-improvement um we've been trying to build that product um and so i want to be the first company that creates this high quality flywheel of a model proposing architectural variations creating the experiments um to validate or invalidate those variations um running those experiments looking at the outputs coming up with what the idea is for the next uh set of experiments and architectural variations until it comes up with an architecture that looks vastly different and is vastly better on the axes that we care about.

48:29We will be investing into that very seriously. Yeah, that's interesting. And looking ahead, what's the most exciting frontier for you? I mean, at SubQ, is it bigger contexts? And not only at your company, but across AI research in general? uh new architectures beyond attention or something else I mean yeah a lot of different things um obviously I think the auto research thing is really exciting yeah um I think that uh but again it's it's not about doing the pre or post training better which is where most people are focused it is about the algorithm itself um we want to move from like shifting the algorithmic paradigm every nine years to doing it every 12 months um so we want to replace transformers or attention entirely in the not create not too distant future and then whatever we replace that with i mean we know what we want to replace it with but um when that happens we want to replace that again within 12 months and just shorten this time as much as possible so that's that's interesting.

49:48I think that from a go-to-market perspective, we're really excited about working with enterprises on long context reasoning tasks, long horizon agentic tasks as a secondary priority. And then in the somewhat longer term, though there is active work being done on it, And we do care about robotics and see that it's highly intertwined with our current work, actually. Like the reason why the figure AI demo was on the conveyor belt flipping boxes is because it requires a 15 second memory span. If you think about asking a robot to do the laundry, the robot has to remember the roadmap of the house to be able to navigate.

50:35It needs to remember how to open the laundry machine door. it needs to understand it needs to know what cycles are appropriate for which um for which clothing items it needs to remember what these are preferences are for how to fold each of the items it needs to remember where those the clothing goes and then um and so the amount of context that's required there is significant and then additionally the um the uh we are we are in the same place for robotics intelligence as you were for language models in 2019, where we train models for specific tasks with thousands of task-specific samples that were hard and costly and time-expensive to curate.

51:23And so I think we need a 2020 LMR Few-Shot Learners GPT-3 moment for the robotic space. But if you think about providing a 10 minute video sample so that the model can learn in context how to achieve a task without being trained on that task specifically, that's like a 4 million token input. And people are struggling to do that in the cloud. How do you do that on a phone worth of VRAM, that requires like 100x plus reduction in memory usage. And that's a problem we're thinking about a lot. So that's a pretty exciting frontier for us. There are a lot of exciting things that we're doing here. And some of the stuff is, stuff we're trying to put out in the next six months.

52:25Some of it is things that we're trying to put out in the next two years. So we've got a longer term roadmap. We're not just thinking about incremental improvements in the industry. We want to figure out what problems are going to matter two to five years from now. How do we take a shortcut to get there? And how big is the team? So we have just under 50 people today. uh we've grown a lot in the last couple months um i mean the fundamental research that we did that we've announced was done by just a few people but we've been sitting on a lot of research for a while uh that research was um done in 2025 um and then production our product ties this year um but it was done with a lot fewer resources and people than we have today and so and we've done a lot since then.

53:17And yeah, I mean, it's a little under 50 people today. Okay. Well, this is all fascinating and I'm glad that we've met. I'm certainly going to be following you guys. If listeners want to follow what you're doing, is it subq.ai? Do you have a blog or something there that people can follow yeah uh subq.ai is good um we've released a couple blog posts i wouldn't say that we're prolific bloggers yet uh we do want to try to push out more but it does take some time um yeah i am trying to post at least once if not a couple times a week on twitter and i do want to start going a little deeper on some of our uh product and technical and research vision here in some of those posts.

54:16And I'll call out different things every once in a while. So that could be an interesting place to follow us as well. We have our official Twitter page and you may see some of our other employees being more public as well about some of their work. So social is going to be a pretty good place to stay on top of what we're doing as well. Yeah. Okay, Alex. Well, this is really interesting.

54:42you

From the publisher

Every AI model in production today has the same hidden tax: doubling the context window quadruples the compute. That's what quadratic compute complexity means in practice, and it's the reason enterprises are spending most of their AI engineering budget on context management rather than on the actual problems they're trying to solve. Alexander Whedon, co-founder and CTO of Subquadratic, joins Craig Smith to explain how SubQ's sparse attention mechanism eliminates that tax, achieving 40 times faster inference and 64 times less compute than standard attention at one million tokens, and what becomes possible when that constraint disappears. The conversation covers striking benchmark findings: 86% of what frontier coding agents do is "read steps," just trying to gather and organize context before the actual problem-solving begins, and frontier models drop well below 50% accuracy on financial document analysis at 500,000 tokens, revealing how asymmetric long context capability actually is across industries.

The most commercially important argument in this episode is about enterprise data. Most large organizations are sitting on hundreds of billions of tokens of data they've never been able to put to work in an AI product, told they need a $10 million data transformation project before they can even start building. Alex's core claim is that SubQ's architecture makes that barrier no longer necessary, enabling enterprises to process far more of their data with far less curation, at a fraction of the cost. He closes with what he describes as the most important and underexplored frontier in AI right now: we are still very far from understanding what users actually want from models reasoning over millions of tokens, and the product and alignment work needed to answer that question has barely begun.

Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.

More from Eye On A.I.

All 266 episodes
86% of What Coding Agents Do Is Just Reading — Not SolvingEye On A.I. · 55 min
Listen in VO