Can AI Agents Learn From Expert Corrections?

1 Jul 2026 · 53 min · 23 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How OpenAI’s Codex-based “Tax AI” can learn from expert corrections via a self-improvement loop, turning human review into measurable evals for parsing and preparing complex tax returns.

Guests (backgrounds)

John DeWasseg and Arthur Fernandez, forward-deployed engineers to Thrive Capital (OpenAI partnership with Thrive Holdings). They helped build the joint venture Tax AI with Codex for real accounting workflows.

Key claims

Tax AI classifies messy client documents (PDFs, Excel, handwritten notes/images), extracts and justifies fields with citations to source locations, and routes reviewer attention to complex fields. Expert overrides become evals; Codex investigates traces/code paths and proposes bounded fixes, with engineers reviewing before shipping. The system emphasizes “macro eval” signals from the full user journey, not only raw traces.

Notable examples

extracting and reconciling Schedule E/C/A; handling depreciation/amortization needing prior-year context; using tax-engine guardrails to measure errors; cases where the model flagged a value as wrong but it matched IRS ground truth.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Data Analysis in Tax AI

0:00 to 0:59

Learn how Tax AI analyzes various data sources for tax preparation.

“So basically right now how it works is that the practitioners start with basically grouping all of the data sources that they need to be analyzed.”

Understanding Tax AI's Development

2:09 to 3:39

Discover the collaboration behind Tax AI and its purpose in accounting.

“For listeners who have never seen Tax AI, can you give us a what-it-is kind of walkthrough?”

The Problem Tax AI Aims to Solve

3:39 to 6:35

Explore the challenges in tax preparation that Tax AI addresses.

“So holdings specifically, like the way that they set up the business, as John explained, is that they are working with these verticals and they have a thesis that they are places.”

How Tax AI Operates

6:35 to 9:39

Learn the process Tax AI uses to prepare tax returns and the role of practitioners.

“And it's cool that you try to get it in right before their crunch period.”

Codex's Mechanism for Accuracy

9:39 to 11:55

Understand how Codex maintains accuracy while processing tax data.

“Is it using some sort of like OCR model?”

Challenges in Tax Data Management

11:55 to 14:00

Examine the complexities faced by Codex in managing diverse tax data.

“And for example, if you start looking at depreciation or amortization, like this is also, you know, adds a layer of like understanding the tax workflow.”

The Evolution of AI Models

14:00 to 15:00

Explore how recent AI model updates enhance capabilities and challenges.

“And also because in January, I think we had the 5-2 and now we're like 5-5.”

Challenges in Model Adaptation

15:00 to 17:00

Discuss the limitations of AI models and the necessity for human input.

“is a very different format that we haven't imagined.”

User Interaction and Model Feedback

17:00 to 19:40

Understand how user corrections influence AI learning and performance.

“So it's cool that the model knows like, hey, I need to move past this and I can do this a better way.”

Building a Product for Non-Engineers

19:40 to 22:40

Learn about designing AI products that cater to accountants and end-users.

“But what we realize in our space related to what John said is that this is super messy and it's hard for the model to reconcile.”
Show all 23 chapters

Feature Requests and User Needs

22:40 to 24:20

Explore the balance between user requests and maintaining product focus.

“Because if I'm understanding this correctly, at the point where it's at now, the people who are using it, the accountants at these firms, they are giving it feedback and it's improving kind of autonomously, right?”

Error Recognition in Tax Filings

24:20 to 27:20

Examine how AI distinguishes between unique filings and system errors.

“It might mean that the person, instead of doing 30 minutes there, will just end up spending two hours for something that's helpful.”

The Role of Forward Deployed Engineers

27:20 to 28:00

Discuss the importance of engineers working closely with clients in AI development.

“And then we start seeing like the repetition of like a few tax returns, like are making mistakes on like these specific fields.”

Integrating Engineers with Clients

28:00 to 28:30

Learn about the importance of having engineers closely collaborate with clients to improve workflows.

“And it's able to parse that to some degree.”

The Role of Accuracy in Tax Processing

28:30 to 29:50

Explore how accuracy impacts the value delivered to both engineers and practitioners.

“I think it's very interesting because it's a combination of someone who is a software engineer, but also is potentially a product manager, has a concept of architecture and is able to be ingrained with the customer.”

Strategic Expansion of Product Features

29:50 to 31:40

Discover how a careful approach to product features can enhance user experience without overwhelming them.

“I'd say that that shows up, too, from just the proof points in this project itself.”

Balancing Cognitive Load for Experts

31:40 to 34:10

Understand the importance of managing cognitive load for tax experts in workflow design.

“And we think that in problems that you can really measure the outcome, this is a good way to deploy.”

Navigating the Complexity of Tax Workflows

34:10 to 38:05

Learn how to approach the complexities of tax workflows and the art of strategic feedback.

“Like there was a great article that Dan Shipper released that was about like, you know, evals are all about changing the frame.”

Expanding AI Applications Beyond Taxes

38:05 to 40:06

Explore the potential for AI solutions in various operational workflows beyond tax-related tasks.

“I think that's where it shows, and that's the human part.”

Setting Up Effective Evaluation Frameworks

40:06 to 42:00

Gain insights on how to establish a proper evaluation framework for domain-specific tasks.

“because anyone who is basically working part-time from their computer, doing something can have, you know, on a certain product, can have this group that's there.”

AI's Role in Error Detection

42:00 to 46:01

Exploration of how AI systems can identify human errors in various fields.

“But the structure of the text engine helps us a little bit provide initial structure and comparison compared to if you didn't have this, you just had the end documents.”

The Future of AI and IRS Interaction

46:01 to 47:58

Discussion on potential AI interactions between tax firms and the IRS.

“I think that's a huge power that we overlook, and I love how this works.”

Improvements for OpenAI's Models

47:58 to 51:40

Insights on desired improvements for OpenAI's models in relation to complex tasks.

“about how that might play out, which is interesting.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So basically right now how it works is that the practitioners start with basically grouping all of the data sources that they need to be analyzed. So as Arthur mentioned, this would be, for example, you know, multiples and like huge quantities of like PDF document. It could be Excel, it could be image of climate notes of clients. The model definitely gets better at identifying when it doesn't know something, but basically helping it going, you know, knowing when it is good or not. I think it makes me think about, you know, getting a good way of measuring what is true. I think what is interesting about this is that that part of the product that matters a lot to the practitioners is basically being driven by their use and not by engineer dictating how that works.

0:46The structure of the text engine helps us a little bit get provide initial structure and comparison compared to if you didn't have this, you just had like the end documents.

0:58Corey:Welcome, humans, to the Neuron AI Explained. I'm Corey Knowles, and I'm here as always, joined today by the one and only Grant Harvey. How are you, Grant?

1:07Grant:I'm good. I'm surprised you didn't throw me a curveball today. Today, the curveball is that I didn't throw you a curveball. You've done that curveball before, so it's a double curveball.

1:15Corey:Oh, I've gotten you twice.

1:18Grant:Yeah, well, I'm good, Corey. How are you?

1:20Corey:I'm good. I'm good. I understand we're going to talk about taxes today.

1:23Grant:So today we're joined by OpenAI's John DeWasseg and Arthur Fernandez, forward deployed engineers to Thrive Capital who helped build the joint venture Tax AI with Codex. So people who are familiar with OpenAI know Codex. Tax AI was apparently built for real accounting workflows. So we're talking about messy client documents. We're talking about source evidence, tax software mappings, and human review. And the big idea is using a self-improvement loop where expert corrections become evals. Codex investigates the relevant traces and code paths and engineers review bounded fixes before anything ships.

2:01Grant:It's very, very exciting.

2:03Corey:John, Arthur, welcome to the Neuron. We're so excited to have you guys. We're excited to be here. Yeah, thanks for having us. Well, I guess let's start here. For listeners who have never seen Tax AI, can you give us a what-it-is kind of walkthrough? Yeah, sure. So basically, TaxEye, to give a bit of context, is a platform that Thrive Holdings co-developed with us. And so basically, the context is that you have Thrive Holdings, which is a subsidiary of Thrive Capital. And what they are doing is that they are, through an intermediary company, they are buying roll-ups, basically. And so they are buying companies in different verticals.

2:43And so one of them is TaxEye. The other one is IT services and basically we are working on the tax side. And so the idea is that through those companies, they are helping them by creating and developing a platform. And so what Tax AI does, it basically, it's a platform that allows you, that helps the tax papers to accelerate the tax parsing from the data to basically submitting to the tax engine, but also adding a lot of other features on the side. And it's basically bringing AI and helping the taxpayers do their job. And we co-developed it with them and helped them on the AI part of it.

3:23Grant:Very cool. Very cool. So I understand that it's a product more for, let's say, accounting firms and less for individuals, right? What was the impetus or why did you want to create a product for them specifically and how did you approach it? So holdings specifically, like the way that they set up the business, as John explained, is that they are working with these verticals and they have a thesis that they are places. So they don't hold a short term holdings with these assets. So they think those are industries that benefit from long term AI transformation. So that's why they started with like accounting and IT services.

4:03and a strong belief from Holdings is that in general, like AI, not AI transformation, like technological transformations, they usually hit the industries from the outside in. And they think these specific verticals are ones where this transformation can be driven from the inside out. So they had this strong belief of essentially driving this transformation from the experts. So basically co-building a product with these experts. And OpenAI has a partnership with Holdings where we also hold some stake on these assets. So essentially, we are FDs helping them build solutions around this space. And this problem in particular was the first one we started working on with holdings.

4:47So as kind of like any problem, like you come in there, like what should we build? And the reason why we picked preparation is we identified this is one of the biggest pain points that the experts surface to us. So if you think about accounting with the perspective of the accounting firms, it's very seasonal and they get a huge backlog in the weeks just before the deadline.

5:15Grant:Yeah. And a huge problem for them is basically... That's my fault, by the way. It's all of our thoughts. And the problem is a lot of the work is done, for example, on the more complicated returns on basically going through raw files with all sorts of different formats from more structured PDFs to like spreadsheets, handwritten notes. Sometimes you have to go and pull formal information from the client and you have to get through this with a deadline. Sometimes it takes like eight hours to go through these documents and just input this data into the tax engine. It requires reconciling information from different documents and things like this.

5:56So the problem we defined that had a lot of value. And around the time we started working with them, which was December, we also had a huge opportunity to capture a lot of the value if we deployed a system in production super quickly that we could leverage, like we could get a feedback like January, February, March is something that's super useful to the firm. And we talked about this a bit in the blog post around the overall impact. But in general, we decided to start with preparation because it was a huge pain point for the experts. So they were definitely on our side to like start working with us around like whatever can help us like reduce the gap here and the time.

6:35Grant:Yeah, makes sense. And it's cool that you try to get it in right before their crunch period. Right before tax season. Yeah. That's great. Got it.

6:42Corey:So when you say Tax AI prepares Complex 1040 and 1041 returns, what parts of that work is the system doing versus what the practitioner still owns? Yeah. So basically right now how it works is that the practitioners start with basically grouping all of the data sources that they need to be analyzed. So as Arthur mentioned, this would be, for example, you know, multiples and like huge quantities of like PDF documents. It could be Excel, it could be image of climate notes. Or all of them probably? Yeah, all of them exactly. And depending on who is sending it, you know, for some it might be some very more clean PDFs.

7:24For some, it might be way more, you know, like just screenshots and client notes. And so you get, you know, like you go across the spectrum of all of them. And this makes it, you know, like you can have one process that basically automatically analyzes everything. Okay. And so it starts there. They upload all of this data in the platform. There is some also work that, you know, in some cases you can directly connect them from the email and upload them on the platform. But basically the key part is once it's there, what it does is so there is some, the software is running. and then it extracts.

7:55It's going to first split the files in order to identify and to classify them because you might have everything that it's in one big 200-page document, for example. So you need to basically classify it and know, okay, this corresponds to W2, this corresponds to K1. So this part is also very important. Then they see this split and then they are able to see the data that has been extracted, basically. And so you would see here the split of a W2, of a K1, of a Schedule E. and you're also able to see a justification of why this data is extracted basically so you're able to see okay this value comes from this excel on this page from this cell but also is reconciled with that image that pdfs and that data coming from there and so this is also very important for the taxpayers because it's definitely one point in which you cannot just blindly and you don't want to blindly trust the system and also you want to make sure that you don't don't end up spending more time on actually needing to find why the AI got to that solution.

8:57So what you need to do is make sure that the justification is very correct and very accurate and that you're able to basically see it in details and go back to the files to understand what was there. So what the taxpayers do there is that they're able to review what has been extracted. And so a lot of their time is actually spent on focusing more on the review process and making sure that the more complex fields are correct compared to basically the more simple ones, which they would still usually need to spend that time. This is correct all of the time by the AI, basically. And so they are able to focus on those harder ones.

9:31And then they review this, and so they spend much more time on the review process. And then they are able to make the submission on the tax engine.

9:38Grant:How does Codex keep it straight? Like all of the facts? Is it using some sort of like OCR model? Is it like doing like loops on the back end? Just like to be like, okay, I've already, does it have like a running list of everything it's already checked against? Like, I'm just so fascinated how that works. Yeah, so basically what John described, and I think Codex is very good actually at following instructions in general. So what John described is sort of a workflow, right? In terms of like how you process a file. And so in general, I think to summarize, like you essentially have like a set of steps and you have a set of durable artifacts that you're following through these steps because also some of these steps might also like fail and you might need like retries and things like this so you need a durable execution but codex is good at following instructions and it can basically drive like like new facts like new files coming in and getting like stuff into into that like end step around basically the the reviewer ready so it's basically a system that works through like a like churning through these durable artifacts and the harness around it is essentially the the set of instructions that were that were catered and when we're talking about self-improving in the article is essentially that harness the set of instructions skills that like everybody's building their own skills uh as well like these days so it's it's the set of things that you're basically giving the model that helps it for example deal with certain edge cases so let's say uh in the example that john uh gave that the like a few tax preparers are having to override a specific field so basically the system is starting to identify basically a pattern here that like we're always extracting information incorrectly for a specific field in the tax engine or or maybe it's related to a specific type of document that we classify that we perform more, like, more poorly.

11:38So those are kind of, like, the things that we try to get the self-improvement loop to contribute changes back to is essentially, like, the mechanism around handling these edge cases and also making sure that, like, because we can measure things well, like, making sure that, like, we don't regress, like, around what we supported before. What are the, what's kind of the hardest real-world

12:00Corey:messiness you've found to handle and and is there anything that was easier than you expected so one thing that's definitely one of the hardest i would say so i would say on the personal side you definitely have for example things like schedule e schedule c and schedule a's those are i would say are usually very complex because you need to reconcile many different files together but then also you need to sometimes in those files you don't have the full context itself And so usually you would need to look at prior year context. And for example, if you start looking at depreciation or amortization, like this is also, you know, adds a layer of like understanding the tax workflow.

12:41And so, for example, in those cases, it's also very useful to give context on taxes. So, for example, you know, if you give IRS documentation to the loop that extracting it, then it's also very useful. And so this is on the personal tax side. And then you have also the entity tax side, which is, you know, kind of separate and works differently. And there is also like a huge amount of messiness because you need to start looking into trial balance. And this means that, you know, usually you don't have two same companies which are doing things in the same way. And it can be even more messy because the quantity of data or the number of rows can be even added.

13:15Grant:Oh, wow. Yeah. Well, were you surprised with how Codex was able to handle a lot of this? Like, were you like, oh, it actually can do this better than we even thought it could? Or did it take a lot of that tooling and self-improvement loop to finally get it to a place where you're like, yes, it can do this reliably? I think what's interesting here as well is the pace of some of the models that we are deploying and also like Codex app itself. And a lot of the stuff, like some of the concepts, if you think about like the app was released in January, if I'm not mistaken. And then we have like multi-agent and we have like more things that enable you to like Codex now can manage like different threads.

13:59So you have all these different capabilities that enable you to handle a lot of the orchestration, have more like independent reviews and also like other processes that you can incorporate, which in the past you would need to like build by hand. And also because in January, I think we had the 5-2 and now we're like 5-5. And the pace of changes and intelligence that is added to the models is frankly quite impressive. So I think we could only get to the point we got when we talk about in the blog post is the models. Is that a state that given a well-bound task that is measurable, it can do that reasonably well?

14:39So I think it's around like the main problem is around like now, I guess, defining the right objective. There will still be instances where the model potentially is not going to be able to automatically climb that loop because it's something that might be potentially ambiguous still and requires like intervention and going through the original document and seeing like, okay, like this is a very different format that we haven't imagined. or this is completely misclassified. The document lacks some capability around the visual language model side of things. So there are still some capabilities around that space where it was still going to be a source of errors or a source of problems.

15:21The model is not going to be able to automatically heal climb. But with every generation, we are seeing massive gains in intelligence. And with every new generation as well, like some of the things where you had to provide a lot of instructions on it reduces your need to do some of these things so we also think a lot about basically like what is like some of the things that you can shed away like like that you don't need anymore because the model is going to have that in distribution or or is going to be smart enough about not making like not requiring that specific like uh prompt it might even have like a more creative way around resolving that task then what you're right oh yeah i think a story on what arthur just described that one thing that we noticed for example around i think it was february is that at some point we had like basically in the extraction loop we were giving a skill to the model to really look for uh you know previous like frequency data of older tax forms and what we observed is that at some point the model wasn't even using the skill anymore it was basically to fetch that zeta by itself and combining it with other data in a way that the skill wasn't allowing it to do or wasn't designed to and so you know it's proposed basically okay i'm not you know you can it ends up basically creating a pull request to remove it or to you know like in this case it like added some things into it which is interesting as well and so you want your system to be able to you know propose changes that you know keeps it interest like keeps it focused and allows to to do more of it

16:50Grant:Yeah, that's a good point because even my own skill creation that I do, sometimes I'll notice like, you know, maybe there's a model change or something or I start to shift what I'm looking for from the skill and the skill can sometimes be holding me back. Right. So it's cool that the model knows like, hey, I need to move past this and I can do this a better way.

17:09Corey:Well, you know, in building on that, something that I think is a good call out here, and I think this is not possible if the models haven't improved to a point they can do this, is that they do – and I would say this is probably since 5.2 forward even – is the ability to recognize, oh, I can't do this and come tell me I can't do that instead of just winging it and doing it anyway. And there's a lot less of that than there used to be. did that play a big role in making this work? Because, I mean, I feel like with the loops we're talking about with the practitioner in place, its ability to recognize, hey, this is a problem, you need to check this out, has to have played a big role.

17:53I think that's totally, like, an important point. And one thing that, you know, from the experience and from what we have seen in the past months, one thing that helps on this, and, you know, like the model definitely gets better at identifying identifying when it doesn't know something, but basically helping it going, you know, knowing when it is good or not, I think makes me think about, you know, getting a good way of measuring what is true, basically. Giving the model datability is super important. And so one thing that, you know, we spend quite a lot of time making sure that was in the product, and this I think is more on the architectural side of how you define your product, which, you know, the AI doesn't always do great.

18:36this is yet it's basically you want to make sure that when the user is going to spend time on the product and fix something so the thing we described was for example correcting a value or saying it was written in the bad field or you know like splitting it saying that the file had been incorrectly split at a point and should be split another place those things and letting the user do it in a way that the output of what they are doing can be used afterwards to guide the model and say, hey, look, if you thought that, you know, the split was there or the correct value was this, this is untrue and use all that data from the user.

19:10And so, you know, it's not exactly about recording everything that the user does. It's more about knowing exactly where you want the human input to be and the expert input to be, such that afterwards you can help the model know itself, you know, find out by itself that it was wrong. So I think a lot of people are also thinking about what's the information you track around, like agent traces, for example. So I think in one end of the spectrum, you have a building in eval with just the ground truth and the prediction. But what we realize in our space related to what John said is that this is super messy and it's hard for the model to reconcile.

19:50Like you can measure where the errors appeared, but what exactly is the source of this thing? And then on the other end of the spectrum, you can capture, let's say the agent traces, like the raw application traces, but this is a lot of information for the model to derive context on. So I think one of the things when we designed the product that John mentioned is trying to capture more of the, instead of the micro, but more of like the macro eval. So getting the whole user journey on what matters. So if you build your product around capturing that, so like the interactions where the user has to steer, the same way that you steer your codec session sometimes to get it on the right track.

20:30So getting the right steering is basically how you can basically build the user journey for the relevant signals. So this is way more useful data for the model to do who climb on the errors, essentially, and triage.

20:44Corey:And I guess ideally you need this to be a tool that any accountant is capable of using, not just a computer engineer, is part of a tricky element, I assume. No, yeah, for sure. So I think one thing they're very proud about this project in particular is that John and I, even though we're FDs here at OpenEd, we come from a product engineering background. So we built products before, like the folks at Thrive Holdings they're building. We did that in the past. And in the traditional software engineering lifecycle, it's usually the engineer's job or the product manager to structure things in a way that you collect feedback, you triage, and then the engineers are going through maybe if it's like more driven by logs and things like this, basically fixing issues.

21:29But I think what is interesting about this is that that part of the product that matters a lot to the practitioners is basically being driven by their use and not by engineer dictating how that works. Of course, the engineer is going to set up like the architecture, like how maybe some of the UX things are going to work in the front end. But basically the automation that immediately impacts the time they're going to be spending on the review is being driven by the more they use the product essentially, which I think is like an interesting antidote on like deeply integrating with the experts. Like I think as FDEs, we have a very humbling experience in general because we come to these deployments, like John and I were not even based out of the US.

22:14So we don't even file taxes in the US. So instead of becoming tax experts, you need to basically give the people that know what they're doing. You should not trust us to do your taxes. Please don't do that. But give these people the power to shape the product and the way that fits their use case.

Read the full transcript

22:39Grant:It's sort of like the next generation or next evolution of that feedback loop with the software development cycle, where you're almost letting the product itself get the feedback from the user. Because if I'm understanding this correctly, at the point where it's at now, the people who are using it, the accountants at these firms, they are giving it feedback and it's improving kind of autonomously, right? Is that more or less what's happening at this point? so i would say like so they are giving feedback and the feedback is triage so like when there is a new feature request it's definitely triage across the others because on the other spectrum you don't want to directly authorize that you know when someone asks for a request it automatically become becomes a new feature you know that could be interesting for that very specific person but that actually could be a pain for the rest of the users using it so this is also something so there is kind of a human review at some point in that early feedback phase afterwards you know in that iteration loop this one you know can be self-improving but on the feedback side from the user this is definitely important because one thing that's very interesting here is that on in one way you want to give you know to make the product as smooth and as easy to use as you know Arthur was describing before but on the other hand you don't want to allow any kind of feature to be added because you You know, in some way, if you create this, you know, you have a tax expert that instead of, you know, spending eight hours by manually reviewing the files and, you know, for example, switching to like reducing it to 30 minutes on the platform.

24:12If you start by adding some tools that, you know, for example, a game to play while the data is being extracted or a chatbot to click through without any intention, for example, It might mean that the person, instead of doing 30 minutes there, will just end up spending two hours for something that's helpful. So I think this idea of making sure that you stay focused on what is actually truly useful for the user and for the end task is quite important.

24:41Grant:You basically don't want them to spend eight hours just messing with the tool instead of actually being able to serve more clients. That makes sense. Yeah. How does it separate?

24:51Corey:And I guess this is where the human comes in as well. But I'm curious as to separating between, you know, a true system error of sorts and normal workflow noise, like just the difference from tax return to, I say return, from tax filing to tax filing. I assume big ones aren't getting returns usually. But I'm curious because there's a certain amount of uniqueness to every filing, even though the forms are the same. How does it recognize this return is different from this return, but not in a way that means something is broken?

25:28Grant:So the uniqueness. Yeah.

25:30Corey:I hope that makes sense. No, it definitely does. So I think in our problem space in particular, there are a few things that give a bit of structure that helps sift through these different things. I mean, there's still a certain amount of ambiguity, which might lead to not the rights. As I mentioned before, around like maybe it's a completely new type of document that was misclassified and things like this. But there are a few things that basically make some of the focus for the test. So classification is one of them. Also, the mechanism at which we deploy this product, we don't support the entire surface area of filing on all of the states and all of the aspects of federal tax and entity tax.

26:13So the way we deploy the product was very targeted. So we started even with the simplest of forms. And of course, when you do this initially, the probability of this being more correct is higher. So that kind of gets us the right shape for understanding exactly how to capture things. And we don't have the loop at that point. But also the tax engine gives us some guardrails around the fields that are filled in there. So the tax engine that the firms use, it kind of already encodes a lot of opinions on how you put in a value. So it's different, for example, if you ask Codex around your return in general, like Codex is going to be taking care of parsing the file and also doing the calculations and things like this.

27:01But the tax engine, for good and for bad, like it provides like a lot of guardrails around like exactly like what to expect. Why is that field filled in a specific way? And that gives us basically a way to more strongly measure errors. So, for example, the error is going to be anchored on a few fields, for example. And then we start seeing like the repetition of like a few tax returns, like are making mistakes on like these specific fields. It does create some problems around grading the task correctly because some of the fields might have more of a human preference. And like some of the fields like an accounting firm might do that slightly different than the other.

27:41So the way that we encode some of these things is by going through these initial interactions with the practitioners. So we get some of these things encoded in our grader. So we only get to the more self-improving loop once we get some of these things right.

27:58Grant:I guess if I understand this correctly, it would be like the tool that the companies use, it already kind of has details of like, this is how we report this type of revenue or this type of situation here. And it's able to parse that to some degree. So it's like already there's like, this is how our company does it. Tell us a little bit about what that is like and if this process of what you're doing now would be possible without forward deployed engineers or how integral that is to being able to work this closely with the clients. I think it's very interesting because it's a combination of someone who is a software engineer, but also is potentially a product manager, has a concept of architecture and is able to be ingrained with the customer.

28:48And so instead of, you know, having the people who actually go with the customer, separated, fully separated from the people who are designing the software, the idea is that by being very close to the customer, you're able to understand the intricacies of, you know, what you need to crack to basically allow a new workflow to be solvable. And so in the case of OpenAI, what is really interesting is that you are very close to research you're very close to engineering in order to shape how the models are going to improve, you know, where they are currently lacking in terms of results and where they are not able to do some tasks.

29:24You're also close to engineering because you work directly on the product that engineering is working on. You help them go in certain direction. You add features when there is a need to. But then also, you know, you're making sure that what the customer is trying to achieve, you're able to, you know, get them through. And so you work on very specific programs which are, you know, and so and which basically you try and go to see from zero to one if it's actually.

29:51Corey:I'd say that that shows up, too, from just the proof points in this project itself. I see you guys did like 7000 returns processed, like a third of prep time saved, roughly 97 percent draft accuracy. That's a huge number. And I wonder where is that the most meaningful to a practitioner and where is it most to an engineer? Do you mean like in terms of the value we give the engineers that we work with and the practitioners? Yeah, I think the way I meant that and I realized it was worded really poorly is like, is that high accuracy the most important point to the end user, I assume, when it comes down to it?

30:35Corey:97 % strong, but I assume there's always a desire to go farther. Yeah, so I think that is an interesting number to anchor on because if you think about from the perspective of a practitioner, and also we talked about the iterative, like basically expanding the perimeter of the product more strategically, so starting very small. And I think this is very important because if we supported a lot of coverage, So we covered 100%, but our accuracy was very poor. This is a complete failure, right? Because they will be spending a lot of their time. They will have no trust on the product. But if we deploy something where it will cover some of their work that they will have to do, like things that we classify that we can't support, like you still have to go through these manually, and that's completely fine.

31:23But if you support those with strong accuracy, I think that's a huge win. So we really focus on trying to expand the perimeter strategically and not really surfacing a lot of errors for people to fix. Because we imagine this would just generate a lot of friction to practitioners. And we think that in problems that you can really measure the outcome, this is a good way to deploy. Like in a lot of fields, sometimes like you build, for example, like a Go dataset. set but when you're when you're working through this with with an expert you're basically making them commit extra time to basically build this with you but if you can deploy something that is immediately valuable and can basically give them like a like a basically it helps them with something that is like lower complexity or mid complexity there's already a huge win so we i don't think we should we should ever track like getting to 100 because this is a world that's always going to have a lot of nuance and some ambiguity so the point that john made like some time ago on surfacing the the citations back to like what the numbers mean we think that's always going to be an important capability of the product because in it would you to understand precisely i would say it's kind of a model receipt it gives you a receipt on exactly like here is your bill it costs you$100.

32:49But here is the breakdown. So it's, I think it works in a similar way. So you can understand exactly where.

32:56Grant:As you approach this process, you know, from an engineering standpoint, from product standpoint, how did you think about the cognitive load of the expert? Like, how did you balance like how much they could actually process at a given time? Because I imagine with some of these accounts that are more complex, and you're offloading a good amount of the thinking to Codex in this case or TaxAri, like how do you still keep them in the loop so that they can still keep track of all of it in their head? Yeah. So basically what we, in order to kind of, you know, make sure that you, yeah, you need to calibrate the product to make such that what, you know, the preparer does and see, you know, it doesn't overload them.

33:37But then like my feeling is that this is a lot about also calibration and making sure that when you iterate and when you develop the tool, you stay very close to them. So here, the Thrive team and us, we also, you know, spend quite some time with the tax preparer and close to them. So I think there is a first point of making sure that you understand very well how they are doing things currently, what they are to process on a day-to-day basis. And there is this thing where you don't want to, you know, disrupt totally the new flow by saying, you know hey look just drop everything in there and then you know you trust it and then you submit it you know like there yeah it could be way higher they could you know process with many packages but it doesn't work right and you want you know you want to make sure that you um keep the expertise uh where it is in you know like account on that uh all expertise and so staying close to them by you know making sure that when the you know like having this closed loop iteration and you know delivering and shipping quite quickly to make sure that every time you do something you're not disrupting their flow and it's still in that capacity and it makes the their expertise and their knowledge to go into the the most interesting tasks basically is what we is definitely important um also maybe taking a step back on you know the kind of you know achieving so basically when we started this workflow there was not a goal of you know trying to reach 98 percent of accuracy for those packages you know i think what this is you know it's like how far can we you know push the taxes workflow and how best can we make it across individual taxes entity taxes and there is no real limits you know the same way you would say it's done for models in evals where basically you know you might say that you know two years ago you had some way more simplistic eval compared to now where it was okay you need to solve all those math math problem you need to solve those you know computer science problem now the evals have been saturated and so you focus on much more long-running tasks i would say there is you know some meta point that's a bit similar here where it's you know you have at some point solved or almost solved the w2 form then you start focusing on the more complex one but then you know once you have solved most of those then you know how can you help even out of the way you know and how can you take a step back and you know make sure that you're starting a non-running task even more.

35:59Grant:It's all about changing the frame. Like there was a great article that Dan Shipper released that was about like, you know, evals are all about changing the frame. Once you change the frame, you like start from zero again. And it's like, now you have a whole new thing you can work on. I have a question for both of you. So after working on this project, in your personal opinion, do you think taxes in general are a verifiable, solvable, like quantifiable task? or is it more of an art? I guess I'm asking, is there more art to taxes now that you've been working on this or is it more like something that at some point we'll be able to fully solve?

36:36I think for our problem in particular, we were solving a bounded part of it, which is essentially just the data entry part of things. There is something that I imagine there's more nuance around ambiguity around tax rules, which I think that's probably a space which requires more of the expertise. And also some of the expertise in our case is already encoded in the tax engine in a way, which provides some of the structure and some of the things related to this. But I think the art part, which is also what Holdings is trying to give the companies by giving them more time to spend on strategic work is how can you provide more tailored feedback throughout the year, for example, so the clients can basically better optimize based on prior information, how the next year is going to go.

37:37And I think that's an area that has a lot of ambiguity, a lot of context-dependent things where that's where I think the art is and that's probably the more strategic work that the moment you unlock experts from doing some of the things that they definitely don't want to do around just churning through documents, they can focus more on their time around that strategic work. And that's where the expertise shows the client management, the previous client relationship. I think that's where it shows, and that's the human part. That's the art part, I would say. You know, in this, I remember OpenAI and Thrive in the blog post,

38:17Corey:I believe it was, mentioning that this is a pattern that could apply to bookkeeping or audits or IT help desks or even other operational workflows. What workflows? What has to be workflows, sorry, or other operational workflows? What has to be true about a domain for this kind of agent loop to work? So the interesting thing is that so right now what we did with Stripe currently is focused on taxes. But so Thrive Holdings has other verticals on which you're focusing on. And another one is IT services. So basically where you want to solve IT tickets or help engineers solve the IT ticketing that they have.

39:01More generally, the loop that we have is usable on, you know, I would say many kinds of problems. And, you know, currently I wouldn't see a limit. I would say whenever you have some experts, you know, that know the field or that are doing something, when you have a product that is being developed where you are actually able to track, basically track what the true value is or what the correct value is, is a perfect setup for this group. And it's great. So in our setup, we made it specific to the taxes workflow, but it's totally expandable because things are centered around codecs, the way you structure basically your software and all the components and things that are saved throughout time.

39:44this can be replaced with the ground truth, the evals. The evals is something which can also be generic. And so what you end up seeing, like what you want at the end is you want to make sure that there is some users using a product where the ground truth that they have can be used beneficially in the long term. And this kind of setup, I would say, applies to an infinite possibilities of sectors because anyone who is basically working part-time from their computer, doing something can have, you know, on a certain product, can have this group that's there.

40:15Grant:So how would you, if you were advising a company to try to build its first, like, domain agent, kind of like following the same roadmap, what would be the first thing that you would tell them to do or set them set up so that they can get that, like, kind of same process going? I think for us as FDs and OpenAI and maybe for other people doing these applied applications, it might be uh it might be pretty straightforward what i'm gonna say but uh you need to be able to measure exactly what you're working against so basically building the the evals and they are very specific to uh to the domain you're working on so i think that's the the the first place to start so as when we go to a new new project we have the the scoping and all of this but we always start from evaluating first because it's basically the first building block that you need to leverage.

41:12It also helps you think about exactly like what is the right hill to climb. So I think, and that's very domain specific. So that's the first thing you have to answer, I think.

41:22Grant:How do you approach the evals? Is it like a goal? Do you have like a golden set where you're like, this is what a perfect version of this task looks like? Or is it a little bit more broad? like what's your advice on just like how to set up a good email so i i think a gold set is definitely very helpful in the especially in the initial discovery but in our domain we have the finalized returns done by the experts so in a way we already have the we already have the the the gold set like there is some complexity as i mentioned around making sure that we remove some of the like things that might be filed in different places and things like this.

42:03But the structure of the text engine helps us a little bit provide initial structure and comparison compared to if you didn't have this, you just had the end documents. But I would say depending on the domain, you might need to go the direction of doing more of the manual curation with an expert. You might have no other way to get through it. Or in other cases, basically you have a lot of data being generated on the field already that you can leverage.

42:28Corey:Much like in mathematics recently where you all discovered – I say you all, not you specifically – discovered an error in a known problem. I suspect we'll see that approach at some point too where it's finding problems in human-produced returns. And we'll probably see that in every field I would assume at some point. Do you have any thoughts on AI? And this is maybe a little off topic. Do you have any thoughts on AI in that role, in catching mistakes we've made versus just us trying to catch mistakes from it, which has been the approach so far? Yeah, I think it's a very good – like the math problem you mentioned is a very interesting example because it – I mean it relates to something which we in our case for the tax case, we all have also observed.

43:19is that for some of the cases, as Arthur mentioned, when you're building your golden set, you basically are using all the data that has been submitted from last year to the IRS. And so the quick part is you know this is true and it should be true. It's supposed to be true, right?

43:36Grant:Yeah. You might be able to find some more flaws. The IRS accepted this. Yeah. Does the IRS have a whistleblower fee? You might want to hit them up. And so, yeah, this is definitely a great luxury because, you know, like not, as Arthur mentioned, like not every use case does have this. And so in our cases, what we sometimes saw is that the model, like this loop, basically the software was extracting some fields. And so we would see some prediction when we were rerunning the evals. What you often need to do is deep dive into specific things that has been extracted to understand and try to see, okay, is it good or not to understand a bit the errors.

44:14And so we noticed that it had extracted some data and that it was actually true. but it was marking it as false because in the ground truth from the IRS, it had this value. And so diving into this, we realized that the actual submission that had been done was off by almost not a lot. So it was not very important in the end, but it was a very interesting case of, hey, look, the model found something that it's actually the real ground truth and the prediction, like the thing that we thought was a ground truth was not. One thing that's interesting there is that our take on this is basically you have you know when you have an agentic loop that's running it's going to be up and you know in the software it's going to be up with the same energy all the time and so there is with humans you can have you know energy and concentration and ability to do things that you know is you know can be more modular or you know some you can have some lows and so in the end there is this idea of you want to make sure that you concentrate what people are best at in some specific case so in In the case of like a huge tax return, you want to, you know, if they are spending most of their time on the more complex stuff that's more interesting and maybe even more funny to do in some way.

45:22And, you know, removing that part of, you know, just copy pasting a W2 from a box to another one. Then, you know, you potentially can decrease those errors. And that's what we hope as well to see in this case. It makes sense because in the in the AI discussion, a thing that I know, especially people who who aren't as as micro analyzing it as as we are, there's this temptation to believe that if a human's done

45:48Corey:it, that means it's perfect. And and the truth is, we are not perfect. You know, we are very fallible. And I love the idea of a backstop or a safety net behind me to to catch my own errors. I think that's a huge power that we overlook, and I love how this works. I think this is a really, really cool approach to a problem that everyone here dislikes. Like even for accountants, this is a tedious thing. Like it makes a lot of sense because with such a complex tax code, with so many variables that come into play in taxes themselves, themselves that this feels like a thing that if it can knock this down flawlessly, there are a lot of other things it's going to do really well with just because.

46:35Also, I think this is one of the type of tasks, and I think there is a bunch of tasks that are similar to this, which is like producing the extraction into the text engine requires a certain amount of effort. Reviewing requires a different amount of effort. So if you get sufficient accuracy on the things one-shotting and you have to fix a few things, reviewing is way less effort than you would do that job yourself. So the overall accuracy of the system increases overall because a lot of the work, as John mentioned, they'll be consuming a lot of your attention, which is maybe the causes of some of the errors and also like fat fingering or or doing like some of some of these mistakes

47:22Grant:yeah just the amount of stuff you have to review and then you get lost in the details and you miss like the most important thing because you were too busy like i had to double check all of the you know we're so distractible yeah well i had two two more things before we wrap up here i'll make them quick the first one is this almost makes me think like i wonder if this will lead to sort of like a adversarial game where the IRS will have their own AI, tax AI, and then the firms will have their tax AI. And then they'll be, you know, trying to one-up each other, you know, kind of like how today's accountants sort of are trying to do that with the government.

47:57Grant:I wonder if you have any thoughts about how that might play out, which is interesting. Yeah, that's very interesting. I actually never thought about that case, but that, yeah, you know, in some, what you could have with this is maybe you would end up having a conversation between someone's tax software and the IRS and each of the software AI could post a challenge, like I said on the channel saying, okay, this is true, no, this is true because of this. And then the other one saying, I agree, what I thought was wrong. And then the other one I think you need to have some framework if you want this to be working but you know in the future you could have something like this and i think it would definitely help solve some use cases where sometimes you know you might notice a mistake that was easy to fix and if you would have you know such a direct interaction you might fix it way earlier than you know having to spend way much more time months afterwards i would think so yeah this would be yeah because a lot of people in in the u.s are often say well why why do we have to you know figure this out on our own right at the end why can't they just tell us how much we owe

49:07Corey:they know the tax code you know why are we trying to do their own returns in the u.s is a thing too yeah we're not we're not tax law experts yeah exactly so it would be kind of nice if you could

49:19Grant:just like call them up and be like hey i think this is correct but is this correct and your

49:23Corey:agent could just kind of chat it through codex talks to their agent and uh and i just find out when it's finished yeah yeah exactly i think some of the problems as well are the different reporting mechanisms so your receipt you'll be leveraging like your accounting firm to do that like i think for some basically maybe you have the same fact being reported like you're guaranteed to have the same fact like your employer but i think for some facts it's slightly harder like like charitable contributions and and other things like this so i think that's where like i think things gets like more complicated around like what are the facts the IRS holds and what are the facts that you hold and make you sure like that those reconcile correctly.

50:08Corey:Yeah.

50:08Grant:Yeah. And then my last question to both of you is, you know, what do you want to see, you know, OpenAI? How do you want to see OpenAI improve its models from here? Like what do you really want to see from the next generation of models? What could really improve this use case and really help you with what you're trying to do? I mean, I think for like on this specific use case, I think what we see generally is that the ability to do a long running task definitely helps on a lot of other things. Because usually, you know, like if you're able to do more long running tasks, it means that the model is able to, you know, look up in some of the data.

50:45In our case, it would be, you know, the evals, the traces, also being able to do some tool coding. you know imagine at some point being able to do some part of the tax calculations itself in a in a way better way this would be interesting as well i mean i think generally speaking like the more it becomes able to do things that are related to you know the economy in general is better because i think in taxes you have a lot of concepts which are human invented and you know which are very potentially hard to grasp for a model. And so, you know, the more it understands the economy, and so I think we have, for example, evals that are well-known on this, such as, you know, GDP evals, which are more broad.

51:27I think expanding into those domains means that you have taxes that improve directly, but then also has consequences of other things improving. And so, yeah, this is quite exciting and will be useful for the next model version.

51:41Corey:It absolutely will. John, Arthur, thank you so much for taking the time to join us today. It's been really interesting. Thanks a lot. We're really excited to be here. Excellent. That's awesome, yeah. Great conversation. We'll share the link to the blog post about Tax.ai as well and make sure that it's there for anyone who wants to check that out and learn more and see what's going on. Please take a moment, if you haven't yet, to subscribe to the channel so you can check out all of our interviews with the people building and impacting AI every day. Also, please pop by the Neuron.ai to subscribe to our daily newsletter.

52:15Corey:And that's all we have for today, ladies and gentlemen, but we will be back soon with more. Farewell for now, humans.

From the publisher

OpenAI and Thrive Holdings built Tax AI, a Codex-powered agent that helps prepare complex tax returns while preserving evidence for accountant review.


In this episode, Corey and Grant talk with OpenAI’s John de Wasseige and Arthur Fernandes Araujo about how expert corrections become structured signals, how Codex turns repeated failures into evals and scoped engineering tasks, and why the best AI deployments still need humans close to the work.


They also dig into what this pattern could mean for bookkeeping, audits, IT help desks, and other expert workflows where the system can measure what “right” looks like.


Relevant links:

OpenAI Tax AI case study: https://openai.com/index/building-self-improving-tax-agents-with-codex/

OpenAI Codex: https://openai.com/codex/

Harness engineering: https://openai.com/index/harness-engineering/

Thrive Holdings: https://www.thriveholdings.com/

Crete: https://www.cretepa.com/


Subscribe to The Neuron newsletter: https://theneuron.ai

More from The Neuron: AI Explained

All 106 episodes
Can AI Agents Learn From Expert Corrections?The Neuron: AI Explained · 53 min
Listen in VO