In short
Whether “RAG is dead” and why retrieval still matters for high-accuracy AI in tax law; how Sphere’s TRAM system builds, retrieves, and reasons over jurisdictional legislation with exact citations.
Guest
Alex Bowcut, head of engineering at Sphere (revenue-based compliance for sales tax/VAT/GST). Background in engineering and startups: semiconductor/GPU work (CUDA), then web data collection for investment banks/private equity; acquired by Bain & Company; later joined Sphere.
Key claims
For taxability determinations, retrieval can’t be replaced by “grep/agents” without losing citation accuracy. TRAM improves expert throughput by ~2 orders of magnitude with fewer errors. Reasoning models + reinforcement fine-tuning (RFT) and better retrieval (dense+sparse, semantic chunking, reranking, context expansion) drive accuracy.
Notable examples
Maryland/Manitoba changing SaaS taxation (e.g., Manitoba taxing SaaS starting Jan 1, 2026) flagged early and pushed into a deterministic tax engine integrated with Stripe; citations are preserved via passage hierarchy and linked back to source documents.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOOverview of Sphere
1:02 to 1:24
A brief introduction to Sphere and its focus on revenue-based compliance.
“For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters.”
The Tax Compliance Challenge
1:24 to 3:32
Discussion of the complexities of tax compliance due to varying regulations.
“Sphere is a revenue-based compliance company.”
Introduction to TRAM
3:32 to 4:34
Overview of Sphere's TRAM model and its role in tax assessment.
“And so that's been a huge inhibitor to growth for the incumbents and the reason why most of the big incumbents have stayed in the U.S.”
Alex Bowcut's Background and Experience
4:34 to 6:28
Alex discusses his background and journey to Sphere.
“And so in this new era that we're in, we looked at that problem and it was a problem that we thought was screaming to be solved by AI.”
Data Landscape in Tax Law
6:28 to 10:40
Exploration of the data challenges faced in the tax law sector.
“And at the time it was just Nick in a little one-person office here in San Francisco.”
How Users Interact with TRAM
10:40 to 14:00
Explanation of how tax experts utilize the TRAM system for accuracy.
“And sometimes they do have to make adjustments.”
Tax Changes and SaaS Impact
14:00 to 15:00
Explore the evolving tax landscape and its implications for SaaS businesses.
“I remember getting a flurry of emails maybe, I don't know, nine months, a year ago from like every SaaS vendor I use talking about some big change in the way taxes were going to be calculated.”
Legal Review vs. Data Labeling
15:01 to 15:59
Understand the distinction between legal review and data labeling in tax law.
“And so that's something sphere, I think, is kind of at the forefront as well, or these other features outside of just sales tax that are only becoming more common.”
Document Ingestion and Processing Pipeline
16:00 to 18:44
Learn about the process of ingesting legal documents and creating structured data.
“So, yeah, I think it's more akin to a legal review.”
Semantic Chunking in Legal Documents
18:45 to 19:49
Delve into the challenges and solutions of semantic chunking in legal texts.
“I think if you do a naive implementation, you leave a lot of accuracy on the table, essentially.”
Show all 22 chapters
Dense vs. Sparse Representations
19:50 to 23:00
Discover the benefits and trade-offs between dense and sparse data representations.
“The question is, how do you make that happen in a way, you know, doing it for one document is easy.”
Building a Taxonomy for Product Taxability
23:01 to 26:51
Learn how to create a taxonomy for product characteristics affecting taxability.
“that we have a baseline of like these citations, right?”
Retrieval System and RAG Debate
26:52 to 28:08
Examine the role of retrieval systems in the context of the RAG debate.
“of like a document being fed through an ingestion pipeline and ingestion.”
Debating the Future of Retrieval-Augmented Generation
28:08 to 29:55
Explore the ongoing debate about the relevance of retrieval-augmented generation in AI applications.
“Yeah, I think, yeah, I was thinking about this this morning.”
The Role of Citations in Tax AI Systems
29:55 to 32:38
Learn how citations are managed within AI systems for tax law applications.
“Talk a little bit about the citations that you mentioned, how you use those, and how the retrieval system helps you deliver them.”
Reinforcement Fine-Tuning in AI Applications
32:38 to 34:49
Discover the impact of reinforcement fine-tuning on AI model performance.
“which is essentially is fine tuning on their reasoning models.”
Evaluating Model Performance and Changes
34:49 to 37:40
Understand the challenges and techniques in evaluating AI model performance over time.
“And I'm curious your experience with like, I guess what I call undocumented model changes.”
Enhancing Retrieval Processes with LLM
37:40 to 40:28
Learn about innovative approaches to using large language models in retrieval processes.
“changed and improved and maybe degraded actually in some particular ways.”
The Importance of Context Length in AI Models
40:28 to 42:00
Examine how context length influences AI model effectiveness in structured documents.
“And then, you know, where do you in this search for increased accuracy, where do you see like your next jump coming from?”
The Role of Context in AI Tax Solutions
42:00 to 45:40
Explore how context length and model capabilities impact accuracy in tax law AI.
“You need to have a really deep understanding of tax law as deeper than like these models have just like out of the box based on their training data.”
Future Directions for T-RAM and AI in Tax
45:40 to 48:38
Discuss the future developments for T-RAM in tax law and improving accuracy.
“Maybe to kind of wrap things up, where do you see things going for, you know, both T-RAM and kind of AI and fields like techs more broadly?”
AI Tools and Personal Workflows
48:38 to 50:38
Learn about the AI tools that enhance productivity and workflows in tax law.
“Out of curiosity, what are the tools that you use and think of as like your biggest AI unlock from a personal workflow perspective?”
Transcript
Automatic transcript. May contain errors.0:00Sam Charrington:As context windows get larger and larger, one question that keeps coming up is whether retrieval augmented generation or RAG is becoming obsolete. If models can ingest millions of tokens of context and reason over enormous collections of documents, why bother with retrieval at all? The answer, it turns out, depends a lot on the application. I recently sat down with Alex Bowcut, head of engineering at Sphere, which builds AI systems for sales tax, automation, and compliance. exactly the kind of domain where getting the right answer isn't enough. You also need to know where it came from. I asked them this simple question.
0:33Sam Charrington:What's your take on the whole rag is dead argument that some folks make? I think for some use cases, it's certainly true. I think for us, or at least for this particular problem, because we are so sensitive to accuracy and we're so sensitive to the exact right citation, As of today, I don't think agents are just searching over the file system, grepping over it is at a point where we could switch over and not lose accuracy. I'm Sam Charrington, and this is the TwiML AI podcast. For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters.
1:15Sam Charrington:Let's jump in.
1:24A little bit about Sphere briefly, so this makes a little more sense. Sphere is a revenue-based compliance company. So we help companies with all of their revenue-based compliance needs. The main one of those is sales tax in the U.S. and internationally. That's called VATGST. And the way, you know, there's other companies in this space, of course. There's some big companies that have been around for quite a while.
1:51Sam Charrington:Tax is not a new problem. It's not, unfortunately, for companies and for consumers, I suppose. So this isn't a new problem. The incumbents, they face a particular problem, which is in order to support, you know, every jurisdiction in the U.S. because, of course, every U.S. state has different rules. In some U.S. states, even the cities have different rules. And then internationally, of course, every country and potentially province has their own rules as well. And so the incumbents need a way to understand how are products taxed in each of these different jurisdictions. And the way that traditionally they've done these sorts of things is they've hired these massive teams of essentially tax lawyers.
2:39They're tax experts. They'll call them tax content teams. But what these tax lawyers are doing essentially is looking through the legislation in Alabama, for example, and understanding how does Alabama tax SAS? And even more specifically than that, how does Alabama tax SAS that maybe has an API connection and has servers that are hosted within the state itself?
3:04Sam Charrington:So it gets very granular there. And this is a huge, you know, this takes a huge amount of human time to do. And you all, there's also like the moving target of it, which is, of course, like legislation updates that can happen at any moment. And so you have to constantly be updating and looking through the legislation again to see if anything has changed and updating your tax engine, essentially, to make sure that you're applying the correct treatment in all of the jurisdictions. And so that's been a huge inhibitor to growth for the incumbents and the reason why most of the big incumbents have stayed in the U.S.
3:41because, you know, they've kind of tackled this problem in the U.S. and to extend it internationally, you know, it's just like too much of a Herculean task for them to, you know, it's too much manpower, too big of teams to handle it. And Sphere has taken a very different approach. I think the time that Sphere was started as a company was obviously advantageous. We were started during the AI era. And so this is a very classic document-based problem. It just from a high level, you have legislation and court rulings and bulletins from departments of revenue. These are all just documents. And these inform, you know, the answers of how products are taxed are found in these documents.
4:24And it's just a matter of finding the relevant passages and understanding the relevant passages and then assigning a taxability. Right. How is this product taxed? And so in this new era that we're in, we looked at that problem and it was a problem that we thought was screaming to be solved by AI. So what we eventually built is what we call TRAM, which is the tax review and assessment model, which is like a system of a few different things that I'm sure we'll get into. But essentially, its job is to supercharge our tax experts. So what we found is that TRAM allows our internal tax experts to move almost two orders of magnitude faster through this process with less errors than the traditional just fully human focused approach.
5:14Sam Charrington:And did you work in tax prior to joining Sphere? I did not. So I've learned a lot about tax. And I think I was always like, you know, I would read, you know, U.S. Supreme Court rulings kind of just for fun, out of interest of kind of what's going on and how does the legal world work. But no, I didn't. I didn't come from like a traditional, you know, PWCE tax background or anything like that. What was your background? Yeah, just a pure engineering background, mostly in startups. So I started my career before moving into startups. I started in the semiconductor industry and then working on GPUs, writing CUDA kernels, and then eventually left to start a startup with a friend, which was like a web data collection startup.
6:01and we worked with like investment banks and private equity firms and things like this to help them collect really massive amounts of web data at scale to power their internal like analyses and signals. And eventually that company got acquired by Bain & Company. And so I was at Bain for a bit and then started another startup and eventually Nick reached out. He was the founder of Sphere, Nick Rudder. Nick reached out about Sphere.
6:32Sam Charrington:And at the time it was just Nick in a little one-person office here in San Francisco. And, you know, looking back, maybe a little inadvisable, but he convinced me it was a good idea to leave the other place I was at and come and join him at Sphere and kind of chase this dream that was, he had like this nascent idea of what would become T-RAM. And yeah, things have gone quite well. We did our Series A last year from Andreessen Horowitz. and continue to grow at kind of a remarkable level. Talk a little bit about the data landscape that you have to deal with. I'm imagining significant amount of complexity due to the global nature of it.
7:14Sam Charrington:And, you know, whenever I've talked to folks that are doing a collection of legal data, it always surprises me how much of that stuff is in non-friendly formats, like, you know, photo, like image-based PDFs or, you know, I've talked to folks, it's been a little while, but, you know, they had to go send people to go scan stuff. Like, how crazy is it still? Yeah, it's not great. It's not great. And a lot of what we do is working with government systems, obviously, which are, you know, they can be archaic. There's some that are better than others, but yeah, there's a lot that are quite old. so some you know sometimes things are well structured they're you know html pages that we can go and collect the legislation and that's great sometimes they're pdfs and those but they're well structured pdfs and so we can we can parse those easily and so it's always great when the sources are like that and then yeah there's this long tail of you know pdfs but they're just images right so you need to OCR them and use other techniques like that or even spreadsheets or text documents or word documents these things are much more common than than I think I had when I started working on on on the issue and so yeah the beginning of our pipeline like this data collection process accounts for each of these different file formats but yeah not not the funnest problem to solve and and you know trying to especially you know like spreadsheets and things like this quite difficult to try and like pull good information and like retain context when things are stored in a spreadsheet let's take a step back and dig into how your users use t-ram you mentioned that like you, you know, two orders of magnitude near improvement in their process.
9:14Sam Charrington:What is that process and how are they using the system? Yeah. So our end users. So if you look at our website, folks like Lovable or Replit who use Sphere, who are customers of Sphere, they're not direct users of T-RAM. t-ram is an internal tool that allows sphere that has allowed sphere to expand globally even though we're we're a small company we're just a startup and then also have higher accuracy and so um the the way that it t-ram is used is is by our tax experts so there's a web app that the tax experts use where they go in and essentially just review the work that t-ram is done. So there'll be a queue of work that needs review.
9:59So for example, maybe they need to review, you know, T-RAM has done what we call determinations, which is just, you know, this determining of whether a product is taxable or not, and some other features around there in a particular jurisdiction. So the tax expert will go in and say, oh, I see that California digital goods needs to be reviewed. And so there's a list of different types of digital goods and the model's output on whether it's taxable or not, some reasoning the model gives on why it came to that conclusion, and then also importantly, the citations that the model used in order to inform its decision there as well.
10:39And so the tax experts are able to look through really quickly. And sometimes they do have to make adjustments. The model isn't 100 % accurate, but essentially they review the model's outputs. They can lead feedback, things like that. And eventually they click submit. And the submission of that then puts those into our deterministic tax engine. So we have a tax engine that, for example, we integrate with Stripe. We're a first-party integration with Stripe. So our customers, when you're checking out and buying something on the checkout page, we'll be calculating tax. And that part of it is deterministic.
11:15There's no AI there. The AI has been done upstream from there.
11:19Sam Charrington:So when these tax experts are, you know, working on a piece of work, like what's the impetus for that? Is it, you know, they are, you know, solving a problem for a customer and that like, you know, drives work or is the work all driven from, you know, some new, you know, ingested piece of data from a jurisdiction that the system says, oh, this might have a change in, you know, know the deterministic engine so it would be twofold so one thing uh that would be like the impetus for them to go in would be um we're we're expanding uh products that we want to cover so every tax engine has done this uh you you are you can't support every product type from day one right you kind of have to break down and you know we're going to start with with clothing or we're going to start with SaaS.
12:15Sphere specifically, we started with like electronic services. So anything that's not tangible, essentially. So they would go through, we have this backlog, essentially, we're like, hey, we want to add support for all of intangible goods, which we've already done. Once we've done that, now we're like, okay, we want to add support for tangible goods. So we'll move through clothing and servers and things like this. So there's like this backlog And that's kind of driven by customer demand, right? If we want to sell to a clothing provider, of course, we need to support clothing. Then the other thing would be updates in the law.
12:51So when legislation is changed or new bulletins are posted or new case law becomes posted, then we'll scrape that data. It goes through our ingestion process, eventually ends up through the system with, if there's some action that needs to be taken, we'll make that recommendation to our team. tax experts. So an example might be last year, Maryland, or maybe a better example, Manitoba in Canada, they changed, they began to tax SAS at the beginning of 2026. That was flagged very early in our system. And so that was reviewed and essentially pushed into the deterministic tax engine with a start date of January 1st so that we were prepared well in advance.
13:40And that's one of the nice things about having, you know, something like T-RAM, this automated system is a lot of the time, traditionally, these things are done retroactively because, you know, you miss the update and then, you know, you scramble to try and add it, but it's already too late. You know, it's past January 1st or whatever. So, yeah, those would be the two different ways.
14:00Sam Charrington:Got it. I remember getting a flurry of emails maybe, I don't know, nine months, a year ago from like every SaaS vendor I use talking about some big change in the way taxes were going to be calculated. Yeah, I mean, it's especially on the SaaS front, you know, over the last four or five, six years, U.S. states and international jurisdictions are just getting more, you know, they want to they want a piece of the cake, right? They want to share the pie. And so they're changing rules to tax SAS, and they're also changing lots of other rules. There are certain international countries that require real-time reporting to the Department of Revenue or the tax authorities.
14:38So as you're transacting, they want a copy of it. And I think it's kind of like tax rates. Tax rates only go up. They very infrequently will California reduce their sales tax rate. And I think how involved the tax authorities want to be in transactions and how much information they want, that also only increases. It's not going to decrease. And so that's something sphere, I think, is kind of at the forefront as well, or these other features outside of just sales tax that are only becoming more common. And a big part of that is SaaS, kind of like you mentioned.
15:13Sam Charrington:So these human experts, in some ways, what it sounds like they're doing is data labeling. Did you think of it like that? We think of it more as like a legal review. So in some ways, I think what we do is somewhat similar to someone like Harvey. So Harvey AI, that's like a legal AI. And you'll let Harvey draft like a first version of a brief maybe. But then a lawyer or someone at the firm will go in and they'll review that brief and they'll check it for correctness and things like that. They'll do a legal review. And I think that's how we view what our tax experts are doing. It's not necessarily data labeling.
15:52It's a review of correctness from a legal perspective, because these are legal claims essentially that we're making. You know, we are claiming that, you know, SAS is not taxable in Alabama or whatever. So, yeah, I think it's more akin to a legal review. And so talk a little bit about the pipeline in more detail. You ingest this information from lots of different jurisdictions, presumably normalize it in some kind of way, or at least try to extract the information out of those image PDFs.
16:28Sam Charrington:What happens next? Yeah, once we have kind of, you know, the text from the document, and ideally that's well structured and like HTML and some PDFs will give that to you. So we try and preserve as much structure as possible. The next step would be for non-English legislation, we'll do an English translation. So it's been another big unlock is that, you know, we don't need tax experts that speak every language. You know, LLM's great, you know, notoriously great translators. And so they're happy to translate these documents. So we'll create an English translation. And that's kind of the starting point.
17:08From there, you know, we can't, many of these documents are very long. So we can't just like take the whole document and, you know, create an embedding for it and store that in a vector database or even necessarily with like a more with like a TF IDF, like full text search database. You might not want to do that either. So what we do is we break up into sections, smaller sections, and there's a naive way to do that, which is just like every end characters you chop and then you create a new section. And that's obviously not ideal because you lose very relevant context. And again, these are legal documents.
17:47So they're well-structured typically, as long as it's not an image. And so they come in sections and then subsections and bullets. And so what we try and do is our pipeline semantically chunk things into essentially sensible chunks that cut at normal places. And then we still retain the hierarchy of where that chunk came from so that we can reproduce it later. And then we also store like metadata, of course, and things about where this document came from, like the root document. And then eventually we have these text chunks, and we embed those both dense and sparse, and we store them in a vector database.
18:29And that's eventually what we'll then query over when we go to actually make a determination. But I think it is, you know, we don't, I could probably talk for the next 60 minutes about kind of this process of chunking. I think we spent a lot of time there and it's a very important part of this process. I think if you do a naive implementation, you leave a lot of accuracy on the table, essentially.
18:54Sam Charrington:Yeah, I wouldn't mind having you dig into some of the work that you've done to kind of assess the lift on the semantic chunking. And what you've seen there, I think, you know, as you alluded to, like a lot of folks will pull a rag library off the shelf and it'll give you like three or four ways to, you know, chunk number of characters and whatnot. And, you know, folks will do one of that, but there's often ways to take advantage of the inherent structure and the information that you're trying to capture. Yeah. How did you approach that? Was it just obvious that, you know, hey, we're going to do this based on sections because it's a legal doc and they're right there?
19:36Sam Charrington:Or did you like iterate on that for a while? I think it was obvious that like, you know, when you look at one of these documents as a human, it's very obvious how like if you were going to break it up, how you would like to break it up. And so it was clear, I kind of, I guess, what should happen. The question is, how do you make that happen in a way, you know, doing it for one document is easy. doing it for you know the millions of different documents that we've pulled in a way that's general generalizable is much more difficult and so I guess the details there go into you know as as we ingest these documents we have a number of different buckets you could call them of different structures of legal documents that we have parsers and a lot of these are LLM backed parsers, but like bespoke parsers that for that particular type of document.
20:29So we will either from like the metadata of the document or through an LLM call determine, hey, what is this kind of document? Is it a case law? Because a ruling from a judge will look different than like the legislation, the statute law, which will look different from like a bulletin or like a notice that the Department of Revenue releases. And so we have bespoke parsers for each of those. And some of those, yeah, they involve like LLM tool calls. Some of them are fully just algorithmic because the structure is all there and it works fine enough. But I think that's where kind of the devil is in the details on it was, yeah, as a human, you look at it, you know exactly what to do on each of these different examples, but how do you do it in a way, yeah, where it's generalizable across all of, you know, across languages, across jurisdictions.
21:20Sam Charrington:Can you dig into a little bit more detail on the dense versus sparse aspect of what you're doing? We started with just a dense representation, which felt correct at the time. And if we think about the query that we'll eventually run, you know, if we're looking for relevant passages about SaaS, you know every jurisdiction has a different you know ignoring even different languages of course but like every jurisdiction even in English might have a slightly different way that they describe SAS right and especially in legislation legislation reads very old like their description of SAS will be very antiquated it'll it'll might even talk about like CDs and things of this nature and so From the beginning, I think it was a fair assumption.
22:13My opinion was that we should be using a dense embedding, right, that semantically embed these passages. And so that is what we started with. I think what we found and when we brought sparse back into it was there are times, especially with certain when it comes to citations and pulling out certain terms from passages that come from the dense embeddings where you also want to search sparse, where you want to do, you know, a full text search of certain words and certain terms and pull those in as well so that you can. and then compare the two of them. And what we saw was a pretty good increase in accuracy on the citation side.
23:00So we have some evals that we run on the retrieval part that we have a baseline of like these citations, right? These passages are the ones that should be retrieved for these queries. And as we kind of layered sparse back into that, we saw another, we saw an increase in accuracy. and so that's kind of what we stuck with.
23:22Sam Charrington:When you refer to dense and sparse, it sounds like you're talking about embeddings versus full-text search as opposed to like two tiers of embeddings or something like that. Yeah, that's right. So yeah, dense is definitely embeddings, like yeah, semantic embeddings that we use OpenAI's embedding models for and then yeah, when I say sparse, I'm referring to, in our case, we use pinecone to essentially create a sparse representation. So we've loaded a vocabulary and then each passage is fed through and it keeps an index like full text search of the different terms and their different usages across passages.
24:05So it's not quite, you know, elastic search or Apache Lucene, but it's sparse in like a TF-IDF type implementation.
24:15Sam Charrington:Got it. So you've got a predefined vocabulary. And as you pass these documents in the pine cone, it's just flagging which documents talk about which of these terms. Yeah, which passages are talking about which terms. So then you can search. Yeah. So then you can search over them and find, you know, if this one uses this particular term that's not frequently used across the corpus, it'll be a high result. And where are the search terms coming from? Like, do you, you know, is it, I'm kind of getting ahead of your answer, but I'm imagining like document comes in and the first pass is to see if it's at all relevant to the task at hand.
24:59Sam Charrington:And so you're just searching for a bunch of terms to screen the document. Is that the idea? Not quite, but I think, yeah, it's a great question. So the query comes from something slightly upstream, which we also use tRAM for. And that is kind of the first step for us to support a particular product is for us to create what we call a taxonomy of that product. So what that means is we create kind of like a tree structure of for this particular product type, what are the different characteristics across the world that affect its taxability? So an example might be for clothing. Clothing that is made for children versus made for adults can have different taxabilities.
25:52and you can imagine more questions like this where you know maybe pants have different taxability than shirts something like that and we build this big tree and what the query that eventually gets fed into t-ram is essentially a well it's a couple of things but one thing is a description of that particular type of product so in our clothing example maybe it's adult pants. And so we have a description of adult pants. And so that is the main query that we put into the system to then pull out relevant passages. And we'll use filtering, of course. We're doing determinations for Florida. We'll filter to only the passages that come from Florida's
Read the full transcript
26:38Sam Charrington:corpus of tax law. And then we're just looking for relevant portions to this particular product type, which we have an LLM generated few sentence description of. Got it. And so this is, this query is kind of, I guess I'm trying to place this query in the context of like a document being fed through an ingestion pipeline and ingestion. Yeah. An ingestion pipeline. And this is maybe after the pipeline, you've got this, you know, this retrieval system. And now you're trying to use this retrieval system to, you know, update the deterministic model, for example. Is that the right way to think about it?
27:25Yeah. So we built this like big index of law, right? The tax law from every jurisdiction. and then a query comes in which is a description of a product with a little other information around it and then we want to find all the relevant passages in that jurisdiction for that product. So yeah, the index itself is just all of the legislative data and then the query is a particular type of search we want to run to pull relevant pieces of legislation.
27:55Sam Charrington:So you've built this ingestion pipeline and this retrieval system. it immediately calls to mind the you know the r and rag um and you know it may be that's what you're not doing ultimately is generation but you know certainly the idea of like taking a bunch of context and sticking it into an llm and having the llm do the thing uh you know something that you think about um you know what's your take on the whole like rag is dead you know retrieval is dead argument that some folks make. Yeah, I think, yeah, I was thinking about this this morning. I think for some use cases, it's certainly true. And I think we could set up some sort of system where, you know, we just have all this legislation in a file system and then an agent can grep over it and find the relevant pieces that way.
28:51And I think, yeah, for some sorts of problems, that works well. I think for us, or at least for this particular problem, because we are so sensitive to accuracy and we're so sensitive to the exact right citation, essentially we need a more finely tuned scalpel to find us the relevant portion. And we need it to be highly accurate. And so anecdotally, at least, when I use clod code or something and I see it gripping through the code base, there's lots of times it misses. Like I'll go off and I'll find a file that I really wish it would have found. Like this file had the answer I was looking for.
29:34And so, you know, maybe I think maybe we're on a path, you know, five years from now, our RAC system still working the way they are today. I'm sure they won't be. But as of today, I don't think, you know, agents are just searching over the file system, grepping over it is at a point where we could switch over and not lose accuracy.
29:55Sam Charrington:Talk a little bit about the citations that you mentioned, how you use those, and how the retrieval system helps you deliver them. As part of the ingestion process, we carry through a hierarchy of these different passages of text that we end up indexing, and each of them carries a citation. and different passages might share the same citation. But that's very important for us eventually upstream when the tax expert goes to review because those citations also have links which will link the tax expert out to the source document where we collected this. Because a lot of times, you know, they want to review the tax expert that is.
30:40They want to review, you know, a bit more context than maybe the the model gave them in its in its you know breakdown of the citation because the citation you know the model will verbatim give some of the citation back and then a bit of reasoning but sometimes they want to expand on it and so they'll click out and read the citation but essentially yeah the way that we've handled citations is through this hierarchy and tagging of passages which citation they came from which again and i think in theory sounds easy but there There's a process at the beginning with those parsers I mentioned earlier to make sure that we're pulling the actual correct legal citation.
31:22Sam Charrington:You also experimented with using fine-tuning, RFT in particular, for your process. Can you talk a little bit about where it fits in? yeah so we saw a big jump um with 01 open ai's 01 that came out in december of 24 i believe um yeah the first reasoning model pretty much right out of the gate like you know we swapped out the model names like everyone does and we tried out this new model which task in your pipeline in particular yeah this final task of like uh given a certain product type deter and jurisdiction determine its taxability in that region which is what the tax expert themselves review so yeah we have we have evals even then we had evals that would run so we plugged it in it did quite well we ran some through and were impressed so we were already like we're on board with reasoning models it was clear that like our use case was well suited to that extra thinking or those extra tokens that are spent considering the prompt and what the answer might be.
32:32And so we're excited when OpenAI reached out to us to be a part of their alpha program for reinforcement fine tuning, which is essentially is fine tuning on their reasoning models. And what we use, you know, with any fine tuning, you need to provide examples, essentially in like standard SFT. and in rft you need to provide that as well and then you need to provide a grader and what we had that was very useful was feedback from our human tax experts every time the model t-ram had gotten something wrong on a determination so as the tax experts are reviewing when the model is incorrect they leave feedback and they they give that feedback similar to how they would give feedback to like a colleague who had maybe, you know, a more junior colleague that
33:20Sam Charrington:had made this determined text blurb about, you know, what they thought. And explaining, you know, an explanation in a way where you want that person to get better and you want that person to have this, you know, extra context that maybe isn't clear from just the legislation. So some, you know, background information about how Alabama treats a certain vocabulary word, something like that. And what we found was that was a very, so I guess twofold. We had already like a set of questions that we knew the model struggled with today because it had missed them. And then we had a way to give really great signal through the feedback and through the fact that, of course, we had the correct answer.
34:01Like the tax experts fixed the issue, of course, and then they also leave the feedback. So we had the ground truth, we had signal, and we knew that these were hard problems that the model had missed previously. And so that was a really good recipe for RFT. And we saw improvements during the alpha program with OpenAI on RFT. And that's what we use in production today is, while a different model that we've worked with them to RFT, but we've seen performance or accuracy improvements. And that really is the key for us is accuracy. We track it very closely. I'm always checking in on it. We want to know how accurate is the model being.
34:44And accurate means, you know, how often is the tax expert having to make an adjustment to the model's work. And I'm curious your experience with like, I guess what I call undocumented model changes.
34:58Sam Charrington:Like, you know, I think you mentioned either before we started recording or as we've been talking uh your use of clawed code like you know we've seen anthropic document you know some things they do behind the scenes you know tweaking various things that change the model performance like do you do you see a lot of that you know with the with the models that you use like needing uh you know just a unexplicable unexplained change in behavior that you need to run down i think we see that a lot during model generation changes so like we work to not fully rewrite but rewrite significantly a lot of our prompts from model generation change to change i think you know the things that anthropic gets up to on cloud code as far as you know sending your query to a quantized model because you know they they're high traffic i guess they would never admit to something like that but from the outside that that looks like what they're doing i think on the api side because we you know we're using apis with with open ai i think those sorts of changes are are less likely and also um would have even bigger backlash so we don't i haven't seen anything you know in intra model generation but certainly every time the model changes you know things change outside we can't just simply plug into the new version and get the best results immediately is there anything in particular you've learned or specific to your product with regards to the way you approach evals you know beyond kind of collecting a a data set where you know the the models had errors in the pass and um you know running new models through those or that kind of thing i think yeah i think that's been the main thing and i think because yeah and because we have these human experts maybe the one part that's not as standard is you know because the tax experts are reviewing these things we have an ever-growing list of evals uh because it's very easy for the experts there's there's a toggle essentially they can click that says like hey this is a difficult one you should include it in the eval set and they give a description of why so we have this like growing list of evals that we can pull from which i think is important for the model especially because we we do this rft with open ai you know i haven't done this for a while but i think if we went back and ran the evals on like our original evals that were running you know a year and a half ago it would not be nearly as useful as the evals that are running today because the model has changed and improved and maybe degraded actually in some particular ways.
37:47Sam Charrington:Going back to retrieval, you and I previously discussed some interesting things you're doing around reordering and expanding and kind of using an LLM in the retrieval process to enhance your results. And I don't think we dug into that. Can you elaborate on that a little bit? Yeah, so that would be downstream from, you know, we've built this index of all the legislative data like we've talked about. And then when a query comes in, we have a, yeah, we have a multi-step process to essentially build up the relevant context for that query before we eventually send it off to like the final reasoning model to reason through the actual like taxability of the product.
38:37And so what that looks like is an initial search into our database, of course, as far as indents, to pull out relevant passages. We then use LLM as a judge or LLM as a re-ranker to re-rank those into more relevant pieces. We then expand each of the passages because we've retained the hierarchical nature of them. So we can grab, you know, the previous and the following chunks or passages and build out the context of the relevant passages. And then we'll give that back to an LLM again to then reorder and potentially throw away certain things that now seem like they're not relevant as we've added context.
39:19And we repeat this process until we hit either a certain amount of length or a certain confidence that we have the relevant context. And then that goes off to like the final step, the LLM to make the actual determination. But that was a change made a little bit later in the process as well that, you know, in the search for accuracy, increasing accuracy, another wrinkle that I think added quite a bit.
39:44Sam Charrington:You continue until you reach a certain level of confidence. Is that based on an LLM is judged type of scenario, like an LLM's determination of confidence? Yep, that's right. And that's basically by looking back at the previous, we'll give it both the previous passages that were fed in on the last pass before they were expanded, and then the current ones as well. Because at some point, you know, you've expanded too far, and now the legislation is talking about automobiles or something that's no longer relevant. Got it, got it. So you're just asking if there's been a scope change or something like that, essentially.
40:21Yeah. Is the added context actually useful? Like, is it on target for what we're looking for?
40:28Sam Charrington:And then, you know, where do you in this search for increased accuracy, where do you see like your next jump coming from? Yeah. Part of it is model providers. It's great every time. You know, the release cadence has been even faster from OpenAI Anthropics. So that's been great. We see a bump. Once we adjust things with every model that they release, I think it's further as far as like further refinement of the RFT process with OpenAI. I think that's kind of part, that's a big part of the way we'll get to, you know, where we aim to get. And what we want is, like I mentioned, a human expert reviews every one of these determinations today.
41:17They go through every single one. And, you know, we right now that takes them around 10 seconds, nine seconds to review each of them on average. So that's incredibly fast compared to the incumbents who are doing it totally manually. But we like to increase that even further. And one way to do that, the best way to do that is if they could take a random sampling instead. So if we can get our accuracy to a point where we're confident that given a random sample of some number from the determinations the model has done, if those are accurate, we don't need to review every single one of the determinations.
41:54So that's kind of the North Star, at least on this front that we're marching towards. And I think RFT will be a big part of that because chasing this long tail, right, chasing the nines of accuracy, a lot of it starts to become very, to get the correct answer, it's very sales tax focused, right? You need to have a really deep understanding of tax law as deeper than like these models have just like out of the box based on their training data. And so I think that is, you know, we'll make changes to our retrieval process, of course,
42:28Sam Charrington:and then those will be somewhat helpful. But I think to get those last couple accuracy points that we need, it'll be working with, with, you know, the frontier labs to try and do something more bespoke. And I asked previously about kind of this, you know, rag is dead question. But I'm wondering the degree to which context length changes the way you approach the problem. Like, it could be that these documents are so structured, a section is going to be, you know, three to five pages, and it doesn't really matter if you have access to a 2 million token, you know, context window. Or it could be that, you know, there are other ways you can use that context.
43:10Sam Charrington:How do you think about the impact of context window? I think that was actually one of the big reasons why we saw a jump with the release of 01 back in the day, was I think reasoning models are much more capable of reasoning over their full context. Whereas non-reasoning models, yeah, you got real degradation as even if it supported, you know, 128k tokens, when you push that limit, it was not, you know, needle in the haystack wasn't great on those sorts of things. and so i think we saw big improvements there um with reasoning models and so it's still a balance for us uh like we like i kind of mentioned earlier we don't need to fill up and we don't fill up the context window to its max but a big unlock was models where we could give it more where maybe we could be a little less precise on the retrieval portion and expand expand these passages a little more aggressively.
44:07I think before when context was more limited, you know, we were being very selective on which passages we're feeding in because we, you know, we only had so much we could give it before the model just kind of would throw its hands up. And so that was a big unlock. So yeah, we don't push the boundary right on the edge, but I think as reasoning models improve, as the context window gets bigger, again, we won't fill it up all the way, but that's a good sign that the model can handle more tokens than we're giving it today. And that means we can be less precise a bit on the retrieval portion and still get the results that we're looking for.
44:46Sam Charrington:How much time do you spend thinking about trying to reduce token costs either by kind of refactoring from larger models to smaller models or, you know, via other methods? LLMs compared to lawyers, like human tax lawyers, are considerably cheaper, even the most expensive LLMs. So, yeah, this isn't, and this isn't also something, this isn't a process where, you know, we're pushing through billions of tokens. I guess it helps that you're building a deterministic system, and that is the thing that's, you know, kind of the inline, online system as opposed to an LLM inference call. Yeah, exactly. We're not cost sensitive.
45:29And that also means we're not latency sensitive either. So it's very nice. Those are two things that we don't even really have to consider very closely. Quite luxuries, right? Yeah.
45:39Sam Charrington:Nice, nice. Maybe to kind of wrap things up, where do you see things going for, you know, both T-RAM and kind of AI and fields like techs more broadly? Yeah, I think we have a clear path on T-RAM, kind of what I mentioned earlier of increasing accuracy and decreasing human time spent reviewing. So we'll continue to chase those metrics and improve them. And that will allow us to be even more accurate and even more nimble and cover more jurisdictions in the world. So that's certainly somewhere we're going to keep pushing. Then there's other parts of this that, for example, one thing we talked about was these taxonomies that we build that, you know, identify the different characteristics of a product that impact their taxability across the world.
46:32Currently, we do that with our human experts because this is something, it doesn't need to be repeated for every jurisdiction. This is like a one-time thing that we create this taxonomy just for SaaS or just for clothing. So today we're doing that the traditional way with human experts. But if you think about what they're doing and what the question is there, we have all the data sitting in our index to build these taxonomies. For every jurisdiction, we know inherently in that data somewhere holds the answer to what are the different characteristics that affect taxability. And so I think that's another obvious spot that would also allow us to move even more quickly, add more product types.
47:18We'd like to increase the accuracy and the frequency of these ongoing scrapes that we're doing. um could as you can probably imagine there's a huge amount of data sources that we're we're looking at right now and you know not all of them can be scraped immediately or every hour or whatever so we'd like to increase that and increase accuracy of of the outcomes of what those changes do in our system um and then there's some tangential things around like um you know we'd like to make it as easy as possible for customers to move from a different tax solution to Sphere. And one way to, you know, a big reason people don't switch tax solutions or why they become entrenched is because they've spent so much effort in mapping their products to tax codes for a particular system.
48:12And what we're preliminarily doing with TRAM is an automatic mapping from some competitors' tax codes or really any classification system. So if you've classified your products using HS codes, for example, which is what is used for tariffs, we could take in any product classification and map that to a Sphere tax code. And then the switching cost to switch to Sphere is just seriously lowered and you can actually get people to consider making the switch. So I think there's, you know, we haven't talked about e-invoicing and there's lots of other things, but at the end of the day, it all stems from having this index of legislation across the world set up so that we can query over it.
48:57Sam Charrington:Out of curiosity, what are the tools that you use and think of as like your biggest AI unlock from a personal workflow perspective? Yeah. So I've been a subscriber to ChatGPT for a long time. Option space on my Mac, I use it all the time. Cloud code, I have that pulled up all day, every day. That's been a massive unlock for us. well for me personally and I think across the engineering team here at Sphere and we're also beginning work on something akin to like Stripe minions so Stripe put out a paper with with something they called minions which are like AI agents that are running around and looking at the code base and opening up PRs and and working together to kind of improve taking care of things like dependabot PRs that get open and so that's something we're looking at as well to build out?
49:53How can we do that in a sphere-specific way? And kind of related to that, also, what other tools can we add to our internal AI agents? What skills can we add to make them even more valuable for us based on our particular use case? You know, where to pull data, where to look for, you know, these AI agents should be plugged into TRAM's internal index and be able to give answers from the legislation. So I think there's, you know, that stuff is still Mason for us, but yeah I feel like I'm surrounded by LLMs all day every day.
50:29Sam Charrington:Awesome awesome well Alex thanks so much for jumping on and sharing a bit about you know what you're up to at Sphere and how you're using AI. Thank you Sam for having me. Thank you.
50:53Thank you.
From the publisher
As context windows grow into the millions of tokens, many AI practitioners are questioning whether retrieval-augmented generation (RAG) is still necessary. If modern models can ingest entire libraries of documents, why bother with retrieval at all?
In this episode, Alex Bowcut, Head of Engineering at Sphere, explains why the answer depends on the application. Sphere uses AI to automate global tax compliance—an environment where getting the answer right isn’t enough. Every conclusion must be backed by the correct legal citation, and every decision must withstand expert review.
We explore how Sphere built TRAM (Tax Review and Assessment Model), a production AI system that combines retrieval, reasoning models, legal review workflows, reinforcement learning, and deterministic systems to help tax experts move nearly two orders of magnitude faster while maintaining accuracy.
Along the way, we discuss why RAG remains critical in high-stakes domains, how Sphere processes legal and regulatory documents from jurisdictions around the world, retrieval architectures, semantic chunking, dense versus sparse retrieval, expert feedback loops, and the challenges of building AI systems that people can actually trust.
🗒️ Full show notes: https://twimlai.com/go/769.




