Inside Reducto: YC to Series B in 18 Months, From Pivot to Fortune 10 Customers, Lessons in Founder-Led Sales | Adit Abraham, Co-founder and CEO of Reducto

23 Oct 2025 · 1 h 19 min · 34 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode topic: Reducto’s journey building “agentic OCR”/document-to-data infrastructure—how it grew from YC to Series B in 18 months, landed Fortune 10 customers as a two-person team, and scaled founder-led sales to multi-million ARR. It also covers why PDF/document ingestion is still hard for LLM workflows (layout, tables, charts, spreadsheet structure) and what Reducto does to achieve human-level accuracy.

Guest background

Adit Abraham, co-founder and CEO of Reducto. Reducto started as an API for reading documents for language model use cases; today it provides a toolkit/API layer for parsing, structuring, editing, and extracting insights from unstructured human data (PDFs, images, spreadsheets, faxes, medical/financial documents). Company status mentioned: ~20 people; incorporated Sept 26, 2023; processed over 1B pages; ~9 “Mount Everest” stacked pages; ~75M Series B led by a16z; very low burn (~$1M).

Key claims

Ingestion errors compound downstream; off-the-shelf PDF pipelines fail on real-world layout/table/chart cases; Reducto gates/chooses customers to stay “best in class”; chart extraction requires iterative verification/editing; spreadsheet clustering is unsolved and needs custom techniques.

Notable examples

Fortune 10 hedge fund parsing petabytes; soil analysis lab reports; JFK files transcription; Bengali family history digitization; chart extraction from line charts; education use cases for messy handwritten math.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Reducto's Problem-Solving

0:45 to 1:30

Discussion on how Reducto extracts data from documents and its market position.

“Andy told me they've only burned$1 million of capital so far to get here.”

Explaining Reducto's Functionality

3:26 to 4:49

Adit explains how Reducto works and its capabilities in document processing.

“Really quick for you who don't know, what is Reducto?”

Challenges in Document Processing

4:49 to 7:16

Discussion on the difficulties faced in processing complex documents and accuracy issues.

“Like this should have been solved a while back.”

Real-World Applications of Reducto

7:16 to 10:50

Overview of Reducto’s impact in various industries and the significance of document accuracy.

“when you just think about what computers and software do.”

Reducto's Business Model and Growth Metrics

10:50 to 12:22

Adit discusses how Reducto generates revenue and its growth in processed documents.

“So yeah, how many Mount Everest are you at now?”

The Broader Impact of AI on Document Processing

12:22 to 14:00

Discussion on the implications of AI in automating document processing across industries.

“Is it like, do you pay some kind of a process, like page process type of fee?”

Exploring Use Cases for Reducto

14:00 to 15:03

Learn how Reducto is utilized for unique document processing tasks.

“You can imagine all the documents that people are dealing with in supply chain.”

The Impact of AI on Education and History

15:03 to 17:48

Discover the potential of AI in education and the digitization of historical documents.

“Like that's a significant amount of their time that they could be using, doing more productive things, teaching the kids better.”

Challenges in Document Interpretation

17:48 to 19:23

Understand the difficulties AI faces in interpreting documents and charts.

“enterprise data ingestion that is what we sell, but it is like more generally, how do you, how does human data get reasoned on with synthetic intelligence?”

The Complexity of Spreadsheet Data

19:23 to 24:29

Examine the challenges involved in processing complex spreadsheet data.

“We don't really spend too much time thinking about things that are already really easy to do.”
Show all 34 chapters

Developing Chart Extraction Technology

24:29 to 27:11

Learn about the advancements in chart extraction and data visualization.

“that require some amounts of semantic interpretation of the contents.”

Origins and Evolution of Reducto

27:11 to 28:05

Explore the foundational insights that led to the creation of Reducto.

“Like did you guys recently, are you announcing something?”

Origins of Reducto

28:05 to 29:05

Learn about the initial insights and motivations behind founding Reducto.

“Yeah, so when we applied to YC, we had built long-term memory for language models.”

Challenges in Document Processing

29:05 to 30:55

Discover the technical hurdles Reducto faced in document management and processing.

“We had one of those classic, like you post on Twitter and your calendar's booked out for weeks.”

Landing a Fortune 10 Customer

30:55 to 33:45

Understand the steps taken by Reducto to secure a major enterprise client early on.

“about how, Hey, like this is better than what I'm seeing from companies like Textracts and others, even folks that were focused on the space.”

Building Trust with Clients

33:45 to 35:15

Explore how Reducto gained credibility despite being a small startup.

“And obviously it was still a multi-month process because this is a very technical company and they had a whole team of people working on intelligent document processing.”

Focus on Product Quality

35:15 to 36:49

Learn why Reducto emphasized product excellence over rapid growth.

“Like this isn't even like a big viral thing.”

Iterative Improvement in Offerings

36:49 to 38:41

See how Reducto gradually expanded its capabilities based on customer needs.

“Like we just would say no to people if we didn't think that we were at a point where we could solve their individual use case.”

Founders' Influence on Company DNA

38:41 to 40:06

Understand how the backgrounds of founders shape a startup's direction.

“And I talked to them of like, you know, what didn't work?”

Ronak's Academic Influence

40:06 to 42:00

Explore Ronak's early academic experiences and their impact on his career.

“And like for some of these competing companies, if you look at the founder's background, it's like a go-to-market and marketing and sales background.”

Research Missteps and Consequences

42:00 to 43:16

Learn about a significant academic paper's flaws and its repercussions.

“The short summary is that that professor had published a paper called GPD for I Can Solve MIT.”

Strategic Product Pivots

43:16 to 47:24

Discover how Reducto approached product pivots to ensure market fit.

“Obviously you got to get fair usage rights of the stuff they're using.”

Understanding Market Dynamics

47:24 to 48:56

Gain insight into the existing market for document processing and AI.

“Yeah, it sounds like it's almost like figure out like a core problem or trend or customer base to serve.”

Fundraising Journey Insights

48:56 to 51:44

Explore Reducto's fundraising experiences and strategic decisions.

“I think maybe we could talk now about the fundraising journey you guys have been on.”

Navigating the Series B Process

51:44 to 55:44

Understand the unique approach Reducto took in their Series B fundraising.

“You're right that when we raised a Series A, we hadn't even spent our YC funding at that point.”

Decision-Making on Series B Funding

56:00 to 57:30

Learn about the strategic decision-making process behind raising a Series B round.

“So we very much did not want to do that.”

Finding the Right AI Talent

57:30 to 59:41

Discover the journey of hiring a key AI researcher and its implications for Reducto.

“went with the two firms that we were most excited about.”

Building a Relationship with a Key Hire

59:41 to 1:02:56

Understand how personal connections and shared vision led to hiring a top candidate.

“You mentioned the first AI researcher that you hired.”

Transitioning from Founder-Led Sales

1:02:56 to 1:05:15

Explore the shift from founder-led sales to a more structured approach as Reducto grows.

“It was partially, it wasn't even like necessarily us interviewing him.”

Navigating Enterprise Procurement Challenges

1:05:15 to 1:10:03

Learn about the complexities and challenges of enterprise procurement processes.

“I think you told me before you got to about 5 million in ARR, before you hired them, before you brought anyone on the sales side and you are, you describe yourself as like not a salesperson.”

Understanding Value in Software Sales

1:10:03 to 1:12:08

Learn how to determine the value of software and the importance of empathy in sales.

“It's just, are you delivering enough value to make it worth them paying that?”

Building Relationships with Clients

1:12:09 to 1:14:16

Discover effective strategies for maintaining communication and building rapport with clients.

“So like the email, maybe it's a more semi formal follow up and the text is like, I don't know, like a funny meme related to like their product.”

Adapting to Rapid Changes in AI

1:14:17 to 1:16:38

Explore how companies like Reducto must stay agile in the constantly evolving AI landscape.

“hey, this is better at, let's say, checkbox detection, whatever that subtest might be.”

The Future of Document Processing

1:16:39 to 1:17:43

Understand the potential future of document processing and AI's role in automating workflows.

“You had some corpus of information and you wanted to be able to ask questions from it.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:02Turner Novak:Welcome to The Peel, where we explore the world's greatest startup stories. I'm your host, Turner Novak, founder of Banana Capital. Before we jump in, a quick thank you to our sponsors who make this show possible, Numeral and Hanover Park. Today's guest is Adit Abraham, co-founder and CEO of Reducto. Reducto's products takes PDFs and physical documents and extracts all the data, just like a human would if they were reading it. It's a slick initial product that solves a huge problem. Essentially, using computer vision to extract data, from blatant hidden sources all around us. At the time of recording, they processed over 1 billion pages, grew 6x for the past five months, and are fresh off a$75 million Series B led by A16Z, which we'll talk about later.

0:44Turner Novak:The craziest part, Andy told me they've only burned$1 million of capital so far to get here. Anyone building an AI agent for a real-world industry probably sees Reducto as essential infrastructure, and our conversation gets into how they built the best product in the space, landing a fortune 10 customer as a two-person startup, getting to 1 million ARR within their first few months, lessons doing founder-led sales to get to 5 million in ARR, and what the future of PDFs and human computer data looks like. Thank you to Liz Wessel at First Round, Chatham Pudigunta at Benchmark, and Adele Wu at Reducto for helping brainstorm topics for Adit.

1:17Turner Novak:A quick reminder, I publish two episodes of The Peel every week. Check out the back catalog of over 100 episodes exploring the world's greatest startup stories, just like this one. Now, a quick word from Numeral and Hanover Park. This episode is brought to you by Numeral. Numeral is the fastest, easiest way to stay compliant with U.S. sales tax and global VAT. It's easy to set up, and they automatically handle all registrations, ongoing filings, and their API provides sales tax rates wherever you need them with all the integrations you need. Numeral supports over 2 ,000 customers in both the U.S.

1:50Turner Novak:and globally, and they pride themselves on white-glove, high-touch customer service. Plus, they guarantee their work and they'll cover the difference if they mess anything up. They're fresh off a fundraise, closing a$35 million Series B from Mayfield, which they're going to reinvest back into the product to make it even better. If you want to put your sales tax on autopilot, check out Numeral at numeralhq.com. That's N-U-M-E-R-A-L-H-Q.com for the end-to-end platform for sales tax and VAT compliance. This episode is also brought to you by Hanover Park. Hanover Park vertically integrates fund admin, portfolio management, and the LP experience for finance and investment teams.

2:32Turner Novak:Most of you have probably interfaced with a fund admin provider in some way. They're a necessary evil for every type of asset manager across not just venture, but also private equity and private credit. They provide bookkeeping and accounting so investment firms can report to their investors on a quarterly basis. What's crazy is they charge hundreds of thousands or millions of dollars per year to basically not screw up your accounting. They sit together third-party software like QuickBooks, Bill.com, Salesforce, and Excel, and then throw a bunch of bodies at you. And that's where Hanover Park comes in.

3:01Turner Novak:They built their own accounting system from scratch, which ingests all your firm's data and documents, and their AI-native solution automates all the manual work that drives private market investors crazy. Head to hanoverpark.com slash Turner and try the AI native ERP for private market funds. That's H-A-N-O-V-E-R-P-A-R-K.com slash Turner and 10x your fund admin. Adit, welcome to the show. Thanks for having me. Yeah, I think this should be fun. Really quick for you who don't know, what is Reducto? So Reducto, it started as an API to be able to read any sort of documents for language model use cases.

3:40Today, more broadly, you can think of Reducta as a layer that helps you connect all sorts of unstructured human data, your PDFs, spreadsheets, whatever you might have, and interact with them for whatever you might need to do. That includes the parsing that we started with, but it also includes a lot of newer things like editing documents, getting structured insights, splitting and classifying. Everything you would need as a toolkit.

4:02Turner Novak:So what does that actually kind of look like? Practically, like what kind of things could someone do with it? every single person watching this has probably at some point uploaded a PDF into something like ChatGPT. If you're building those sorts of products, that context is really important, and you need great accuracy, not just for the simple text PDFs, but things that maybe have tables and charts, the more complex cases, because a lot of the more important points of human data is things like faxes, you know, patient self chart. So we will go through, we will parse all of that data out really accurately.

4:37We'll structure it effectively for language model inference. We'll do a lot of the post-processing that people would need so that when people pass it into language model, they're getting the best answers that they can.

4:46Turner Novak:And why is that so hard? Because I would kind of think, you know, we have LLMs, ChatGPT came out a couple of years ago, like this should be just fixed, right? Like this should have been solved a while back. Like why, why was it, was it, and is it still such a big problem for people? I talk about this a lot where it's not even like a LLM specific problem. Kindly, people have been working on PDF processing longer than I've been alive. I know people that built the original drivers for printers to print PDFs back in the 1990s. And there've been many generations of companies since. But the context of the problem is very different, which is historically, you would probably have a document processing vendor that's templated out one type of document, like just your bank statement for one type of bank or just your W-2s.

5:36And these were brittle pipelines that you'd spend a lot of time maintaining and even deploying in the first place. But today's context is completely on the other side, which is that people want to be able to reason on any form of unstructured data. Right. Like we're not far from a point where when you apply for a mortgage, there's going to be some loop that goes through all of your transactions and makes the judgments. When you go say hi to your doctor, there's going to be some summary of your medical records. And when you need to read all of those, the hard thing is at the end of the day, these documents are almost like a whiteboard that you could canvas any which way.

6:17Like we often when we think of documents, we think of just like that canonical paragraphs next to each other. but you could have things like patents where you have two columns and you have line numbers and you start to see things like the line numbers mess up the data in the actual paragraphs of text. You'll see things where like you'll have financial tables and if there aren't clear grid lines, sometimes the language model will choose data from the wrong cell in the table. All of these downstream problems, that's obviously since this is the first step of that pipeline, every mistake that you make at ingestion is going to compound throughout your pipeline.

6:52And so one, this was an issue when people were building chatbots that you needed to ask questions to. But two, now that people are trying to do end-to-end workflows, like you're trying to automate your entire invoice parsing. You're trying to go from documents to like a completed work product. You just need human level accuracy for that data. And that's a much harder problem than sort of the canonical traditional OCR.

7:15Turner Novak:Yeah, it kind of reminds me of when you just think about what computers and software do. like back in, I don't know, the 40s or the 50s with like accounting. Like you just have to write all this stuff by hand or whatever. You know, you do the books, credits and debits, et cetera. Maybe you have to like erase things or whatever. Tons of different pieces of paper. And we basically just kind of took it and like did the same thing on the computer. It's faster, but there's still a lot of like, you know, when you're an accountant, you'll like get a piece of paper and you'll like type it into the computer.

7:50Turner Novak:And it's still a lot of this kind of manual human labor. And it's interesting. I mean, if you just think about what's AI going to do, it's just going to automate a lot of this manual human labor that doesn't really necessarily require a lot of skills or decision making. It's just kind of like a yes or no or like a route tree type of thing. But to your point, really hard. So it's kind of interesting when you just think about your kind of like infrastructure of like doing this harder, you know, manual labor that we've done for, I don't know, all of human history, basically. One, we see it across literally everything.

8:30As somebody that has been in tech for my entire career, I think I almost didn't appreciate how deep the problem goes. But, you know, when a large insurance company receives a claim submission, they receive just a folder with a bunch of different files. And it's all the supporting evidence. And they don't even know what's in that folder. Like it's not well tagged and listed as like, here's the data that you'll find in this folder or something else.

8:55Turner Novak:It's literally like a claim holder or like a policy holder that's like, okay, here's like a bunch of pictures, a video, my title for my mortgage, or like the deed for the car, whatever the claim's for. Just all over the place. Yeah. And so that was something that historically you would need to get humans to look at. And I've heard crazy stories here where I was meeting with the CIO for a really large insurance company, like multiple billions in revenue. And he was talking about how for every 10 large claim submissions that they received, they work with really large claims. They only have enough humans to look at three out of those 10.

9:33And of those three, like only one or two will actually be applicable to what they want to cover. And so there's this other 70 % where it might be relevant, it might not, they just don't know, because it's been a human process up until now. And that's the kind of thing where even from the beginning, you could see the promise of language models in that space. Like from the early days, there was a scope of what it can be. And we're trying to make it faster for them to get that sort of thing to a production ready state.

10:01Turner Novak:So do you know that maybe this is kind of outside the scope of what you talk about with them, but do they just not look at those other 70 % of claims? Like, do they just deny them? Because it's just like, oh, it doesn't seem like it. Really, wow. Well, it's not that they're denying the claims, that they're choosing to underwrite the claims that they think are relevant for them. So it's almost like they're bidding against other firms, but you can think of it more as like revenue leakage for them. There are opportunities that they're not considering that they should be. Interesting. Okay. So I guess one general question I have, like what's kind of the current state of Reducto?

10:39Turner Novak:Like I know you have some people like say ARR, some people like say employees. I heard you mention once that if you stacked every piece of paper that you processed, it was like three and a half Mount Everest. This is like five or six months ago. So yeah, how many Mount Everest are you at now? We're probably at like nine Mount Everest today. That number is changing really quickly. Even if you look at the start of the summer, like June, July-ish, we are at roughly three times as many, or sorry, five times as many pages processed per month as what we used to be at the start of the summer. And that's growing quickly too.

11:17We have this tracker in the office that's just the weekly pages that we process. And it updates every 30 minutes and the team will constantly look over and it's like a new record for how many pages were processed in that 30 minutes. The company is still fairly young. So we were incorporated September 26, 2023. So right around the two-year mark. But we have grown incredibly quickly. Teams about 20 people today. And despite that, I think we're lucky to say we work with what I consider to be some of the best teams in the space. That includes both the newer AI native companies. Like if you look at LegalTech, RV, Legora, and a bunch of other folks use us there.

12:03And Vrogo, great companies that are building great vertical AI applications. Across horizontal platforms too, like Mercor is the customer and ScaleAI and a bunch of other folks there. But also really, really large enterprises. So I mean like FANG companies, Fortune 10 enterprises, some of the biggest hedge funds in the world and so on across different industries.

12:25Turner Novak:So then how do you make money? Is it like, do you pay some kind of a process, like page process type of fee? It's based on the number of pages you process. And we have different API endpoints now. So roughly 40 % of our customers actually use Reductor for two or more endpoint products. So like parse their documents with Reductor and then a subset of those documents they'll want to do something like grab structured data. Sometimes they'll use that structured data to then fill out a new document. So edit documents with Reduct as well. It's just like this series of API requests. And it's kind of interesting when you just think of, you mentioned things like legal, financial services, insurance, healthcare.

13:08Turner Novak:That's like 30, 40 % of like GDP or whatever. Like I think financial services, something like 10%, Healthcare, something like, I don't know, around 20 % in the US. Legal is like, I don't know, surprisingly higher than anyone would realize. So it's like, it's basically like most industries rely on a lot of this stuff. Yeah. And I mean, all of those four have been a really big part of the early traction for Reducto. You can almost think of the importance of Reducto scaling with the cost of a mistake. Like when you make a mistake when reading a healthcare document, it's not like, oops, I messed up this period versus a comma.

13:49It's like, no, this is somebody's health that you're making decisions on. And then same thing in finance, insurance, etc. But it does, this is a universal problem, not just in those four. You can imagine all the documents that people are dealing with in supply chain. Like you have bill of ladings, all of the things that you would just have people looking at manually, historically. And also a bunch of long tail cases that I just wouldn't have considered before. Like we have a customer that uses us to process soil analysis lab reports. And that's the kind of thing that we would never have thoughts to even test ourselves on.

14:29That's just esoteric. But out of the box, they found that Reducto was just the best thing they could use for that. And I think that's the value of building this as not a like, we didn't want to build a financial statement parser or like a health record parser. We wanted it to be, how can you read anything that comes in with the accuracy of a human reading it?

14:50Turner Novak:Yeah, it makes me think of maybe like teachers too. Like my, I was in the orientation for one of my daughter's classes and, you know, they write everything down and then the teacher has to look at it and grade it. And you just think of like a high school teacher grading hundreds of math papers every week. Like that's a significant amount of their time that they could be using, doing more productive things, teaching the kids better. We actually do have people that are building AI for education products where like students will upload photos of their homework. And that's hard because students have messy handwriting.

15:24You know, if you're dealing with math homework, you have equations in there. That's the kind of thing that you just wouldn't have been able to do with traditional OCR. that now with the combination of both the LMs and also Frontier CV techniques, you can get there. Like you can read what the person had and downstream of that, you can imagine all the great things people are going to build with AI tutors and so on.

15:46Turner Novak:What's been sort of the, I don't know, either craziest or coolest thing, use case or like product that you've seen someone build on top of Reducto? It depends on how you define cool because I think there are, there are use cases that are just like shareholder value. Like you see things and, you know, people will talk about how there are repositories of documents that's, you know, date back a decade and now they've been able to digitize that. And that's crazy. Like there's a lot of value in that. One of the largest hedge funds in the world came to us with, they wanted to parse petabytes of data because it was all the things that their analysts have looked at historically.

16:27but the really interesting use cases to me are actually just devoid of the the like practical ones like we see really interesting things come up where when the jfk files were released those were really hard to read like i as a human would have struggled to read those and we were finding cases where reducto was transcribing the documents and it was only once I saw the reductive parse outputs that I was like, oh, that's what that word was trying to say. And it made sense when I looked at it side by side. My co-founder has this really crazy story where his great-grandmother had written their family's history in Bengali.

17:11And he doesn't read Bengali himself. And he tried uploading that into chat products historically, and it was just gibberish. like it just didn't make sense. But when we first released Agentic OCR, which is I think one of the biggest steps forward for us as a company, he ended up finding that that was like the first time that he could digitize that like family history transcription. And that's the kind of thing where like obviously that's not the focus of the company, but it's a good reminder of how deep the problem goes. And it's not just like enterprise data ingestion that is what we sell, but it is like more generally, how do you, how does human data get reasoned on with synthetic intelligence?

18:00Turner Novak:Yeah, that's pretty cool. My wife was just looking at her grandma had like some handwritten recipes from, you know, over the past 60 years. And, and I'm, it reminds me of like, when I, whenever I get cards from my grandparents to be written in like the most intense, impossible to read cursive. And I would give it to my mom. I was just like, can you read the card? I can't read this. It's just, even if we're a human, that's hard to read sometimes. Yeah. I don't know if you saw this, but there was this big historical challenge of people were looking for humans to go through and rewrite all of the historical archives in cursive because a lot of younger people cannot read cursive anymore.

18:42Our struggle too. And that's the kind of stuff that we want to make sure that we can just do out of the box, no matter what the content is.

18:51Turner Novak:Interesting. Yeah, total, total side tangent. But I remember back in third grade, like learning cursive. I'm like, this is why I don't learn cursive. Like who's going to use cursive? And everyone's like, oh, you got to learn it. Everyone's going to write in cursive when you when you get a job and then now no one writes anything anymore. We're all like AI slot videos just like straight into our brain. Like words don't even exist anymore. Are there any things that like the product still kind of struggles with or just AI and LLMs in general seem to not quite be able to do yet? Well, a lot of our work is around...

19:28We don't really spend too much time thinking about things that are already really easy to do. Like digital text, for example. If you had clean digital text in English, you would have been able to OCR that well before VLUPS. And that's fine. Like, obviously, that's table stakes for our products. Everything that we do is, like, find the things that are challenging and find ways to help address them. And so you'll see all sorts of cases where, for example, there's no product today or there was no product, that could do chart extraction effectively. Because when you have like graphs, like line charts, bar graphs, all of that, bar graphs are maybe a bit easier.

20:15Line charts are incredibly difficult. If you just imagine the things that you would see in like a financial research memo, a line chart could have all sorts of inflection points. That's really hard to read. And if you wanted to turn that into, table representation of the underlying data, historically, you just wouldn't have been able to do that. And I'm sure you've seen sort of the... VLAMs are not as bad as the tweets would sometimes make it seem, but I'm sure you've seen those cases where you ask a model, what does the speedometer read? And obviously, the pointer is at like 20 miles per hour, and it says something like 80.

20:55It's just soft. Imagine applying that to hundreds of data points to a chart where it's choosing where the data points is, but with some certainty, you know, that set of those data points was just like off by a non-tributable amount. Suddenly that data isn't useful anymore. Like you don't want to do anything with it. And so when we think about the types of things that we are doing, we tried everything that we could in terms of just training models for using other frontier models. and there was nothing that was good in a single shot image of a chart to markdown table. So instead we started thinking of it as how do we create an environment where we can iteratively catch our own mistakes?

21:37So how do we not just make the initial representation but give it tools to be able to re-render the chart, look at the mistakes that it made and then edit individual data points again and again and again until a verifier model is happy with the end chart extraction.

21:51Turner Novak:And then it gives the output to the end user. To the end user, exactly, yeah. And it's slower. Like there's a clear latency cost there. But for the people that wants to be able to use that data, this is the only way that's possible today. And there are plenty of other things like that where spreadsheet clustering is this like weirdly deep problem. I didn't know this before we started the company, but there are people who have done their entire PhD thesis on how to cluster Excel spreadsheets. And when I say cluster, I mean like you have a spreadsheet and you could have set that up any which way, right?

22:28Like you could have one massive table, you could have multiple sets of data. It's just free form for an arbitrary number of rows and columns. You need to know what is like the disjoint sets. You need to separate that piece of information apart. And that's a hard problem because you can't rely on formatting. The person might not have given you like an outline around what they need. So you're looking at all sorts of things like what is the homogeneity of the data? Like, does it seem like there's a dense cluster here and a dense cluster here? And like, how do we separate that out? Those sorts of things are unsolved problems.

Read the full transcript

23:06We've done really deep literature reviews and ultimately end up having to come up with our own techniques to be able to do it at the bar that people need. because you need to be able to do this if you want to have a banker upload a spreadsheet into your AI for finance products.

23:23Turner Novak:So this is essentially like someone doesn't design their spreadsheet in like a best practices type of way and they just have weird, like their data is organized weird and like weird column and rows setting. And so the computer, when it reads it, it just doesn't understand what it's looking at. That's essentially what the problem is. Yeah, there's a lot of meaning that I think ultimately these documents were made for humans like you and I to read. And there's a lot of meaning that is really easy for us to interpret that when you try to codify is a lot harder. Like every time you have a gap between two paragraphs, that's me telling you, hey, this is a new semantic unit of information, right?

24:01Like that's just encoded and we don't think about it. It's just there. But if you have something like a spreadsheet, maybe they have two tables and they added one column of space. But it could also be the case that maybe that column was just supposed to be empty. Like it was actually one table, but they didn't have a row of values there. So you can't hard code a value of like, if there's one column of space, separate this out as data. Like there's all of those sorts of things that require some amounts of semantic interpretation of the contents. It requires some amounts of just like looking at that individual sheet and getting a sense of how it's formatted.

24:38But there's no like, or at least I don't know if a guidebook that tells bankers how to format a spreadsheet or anybody else that's working with Excel. And you need to be able to do that on the flag.

24:49Turner Novak:There's actually insane, there's training consulting companies that will train the first-year analysts. Literally, you spend the first three months on the job where it's learning to code, but learning how to format a spreadsheet. Inputs are blue. Formulas that are linked to the same page are black. Links to a different tab. The font needs to be green. and like there's so many different rules of like you know sometimes you'll do you know that little the tilde squiggly thing I forget what it's called but like you'll if you want to make it so you can like control arrow jump between like across empty columns because you know how when you're in Excel and you you jump to the end and it'll take you to like the the furthest non-blank cell if you want to be able to like navigate an entire book they'll like add the little tildes in blank columns that You can traverse across them faster with your keyboard without using your mouse.

25:43Turner Novak:Anyways, yeah. There was one semester in college, I helped this guy. He came to our school and trained people on Excel. It was like a crash course over a week where you paid them like$1 ,000. And they trained you how to format these models. Anyways, but yeah. The banks are even more intense. It's literally your first month or two on the job. All you'll do is you'll go in the basement and learn how to build models. So anyways, but yeah, it's like a very manual thing, right? Like it's tons of rules and people still mess it up. And I mean, ultimately, Excel is not just used by the banks, right? Like it might be the case that bankers are more rigorous, but we'll have cases where the spreadsheet is actually just, we had an enterprise customer run their historical social media engagement data.

26:36They just had an export CSV that they were like, oh, I need a language model to reason on this. And that was just like a dump. It was messy and it wasn't clearly codified with the blue formatting for inputs, all that. And you'll have things like, you know, logistics vendors where they'll just list off the items that were purchased and still kind of do it in a freeform way because they're just making it on the fly.

26:59Turner Novak:Yeah, they think of it like a Word doc. They're just like, they're just like kind of throwing them into the cells and like they're just whatever it is. Exactly. Yeah. You just hinted earlier that no one had done this until recently. Like did you guys recently, are you announcing something? Yeah. So we've been working on chart extraction for a while. We are about to do like a really thorough write up on exactly how we achieved the results that we have. Will that be up by the time this episode came out or? It should be. Yeah. It's actually available. Like customers are using it today. Oh, cool. it sounds yeah it sounds like it was a big undertaking based on what we just talked about a lot of work went into it and art that have been honestly it's one of those things where they went so deep into the problem that even once we hit the good enough points and we were like okay we need to do other things one of the engineers working on it was like can i at least work on this over the weekends um i was like yeah of course but he has been grinding through like it's a side project now to just have the best chart extraction possible and the work he's done is just incredible interesting well well thinking about that or talking about that then going beyond just good enough i know back when you guys kind of started this there was kind of good enough products that were out there right like they they kind of worked right yeah for sure so then what was sort of the the insight to start doing this and going back to the beginning like why did you guys start working on this?

28:29Yeah, so when we applied to YC, we had built long-term memory for language models. It was, I think, the first API to do that. But it was really early. Like, this is before GT4 Turbo had come out. The very first wave, you know, chat applications was coming off the ground, ignoring chat GPT, of course. And so this is one of those things that naturally goes viral on Twitter because it's cool to talk about memory and how we'll remember context that the user gave, but just wasn't needed. Like you'd hop on a call. We had one of those classic, like you post on Twitter and your calendar's booked out for weeks.

29:11You'd hop on a call and you'd be like, oh, like what are your users struggling with as a result of not having memory? And it's like, oh, people are barely even sending messages. I think this would be interesting in a few months kind of thing. But one of the things that did really resonate is people started saying, hey, if you're managing the user's chat history, can you also manage the files that they're uploading? Like a classic just managed RAG pipeline. And we thought that that would be this low lift, like we would use off-the-shelf tools and just add it on as a feature and it would be this fully managed platform.

29:42But surprisingly, two things. One, that ended up being the most interesting part of the platform. Like people got really excited by that. But two, it was the hardest part of the platform because the things that we were using, we tried pretty much everything on the market, would just keep falling apart in random ways. Like reading order would mess up when you have two columns of data, read left, right, instead of column by column. You'd find all sorts of cases where the table wasn't parsed correctly and so on. And so it just became this thing that we were spending so much time building custom logic around.

30:16We eventually had to start training our own models around to segment the layout, all of that. And where Reductive is today wasn't really us deciding like, oh, this is a$10 billion opportunity. Let's pivot into it and do everything with it. I actually started as a technical marketing blog where we had this really ugly Streamlit app. You just drop a document and we would draw boxes on the document. I cannot exaggerate enough how simple it was, but that just immediately got a ton of interest. Like we started seeing a lot of founders for, but at the time were early stage companies, but are now kind of household names talking about how, Hey, like this is better than what I'm seeing from companies like Textracts and others, even folks that were focused on the space.

31:06And we would hop on calls with them and realize that they were going through the same process that we had. Like they were spending a ton of their time post-processing outlets, building, you know, custom fallbacks. It was just this thing that was holding them back from doing the things that they really once. Because if you're a legal tech company that parses contracts, it's not the parsing that you're interested in. It's the intelligence on top of contracts that your customers actually care about. But if you parse it incorrectly, that's the bottleneck that holds you back from being able to give good insights to the lawyer.

31:38So we thought of it as like, what would it mean for us to almost be like the ingestion team for our customers? Like, what if we could actually solve those last mile problems such that ideally they shouldn't even be thinking about PDF processing. And that's how we kind of went down that road and very quickly started seeing really, really intense adoption. Like I think I, we were talking on email and I mentioned, we ended up getting a fortune 10 customer as like a full annual enterprise contract as a team of two.

32:07Turner Novak:Yeah. So tell me about that. How did, how did, how did you land that? Like, that's just, you know, I think you guys were still pretty early on. Like the product was pretty, pretty new at the time. Yeah, I can, I can send you a screenshot of how ugly our website was. Like it wasn't this nice enterprise ready, you know, professional page. It wasn't meant to convert. It really was like a super simple text tagline. And the important thing was we had, again, a really ugly front end application that would at least show you what the product does. So from day one, we had a playground where you could test reductive for yourself because we knew that it was a really noisy space.

32:48Like it's been around for a while. People have seen decades of PDF frosting companies. And even today, like every week you hear about something, not just in our space, like across the AI landscape, you hear like the new best or like the first AI agents for blank. And it's actually the 20th. And so we wanted people to just be able to kind of see like the proof is in the pudding when we say that we're the most accurate. And I don't know exactly where they found this from. I think they might have seen a LinkedIn post or something like that. But they decided to upload one of the, what I call gotcha documents.

33:26Everybody has some set of documents that they expect to fail. It's like a thing that they've been trying to solve and it just hasn't worked. And that out of the box ended up working in the playground.

33:39Turner Novak:That'll get their attention pretty quick. Yeah. I was like, okay, interesting. let's have a conversation at least. So they came inbound. We had a conversation with them. And obviously it was still a multi-month process because this is a very technical company and they had a whole team of people working on intelligent document processing. So trusting a two-person startup just almost doesn't make sense. Borderline feels irrational, I'm sure. Yeah, they're like, who are these kids? Like, what are they lying to us? Is this like fake? Like, yeah, honestly, there was a lot of skepticism when we first got started.

34:18And I understand it's because, you know, the company had just been founded. We hadn't even announced any sort of like a seed round at that point. And so, yeah, like there's a lot of just intense back and forth, just stress testing everything that we've done. At some point, they had like 15 of their engineers meet with us for seven or eight hours across the day. We were just in a conference room, stress testing everything about Reducto, what it could do, what it couldn't do.

34:50And I really don't think it was like, there are sales lessons that I have coming out of this, but it wasn't like, oh, we're incredible at sales and that's why we landed this customer. I genuinely think it was just we had an incredible products from day one. We spent a while working on it before we even released it as a product. And the product kind of spoke for itself.

35:11Turner Novak:you mentioned you just kind of made a post. I think it was on the YC internal forum. So it's like private. Like this isn't even like a big viral thing. Like it's kind of the antithesis today is like everyone needs to launch with a launch video, right? It's kind of the meme. You guys did like a private non-viral opportunity, like just like post with like some pictures and screenshots and stuff. And it was basically just, you showed that you solved this problem that people had and people resonated with it. And it was really that simple. Yeah, at least for the early points. And obviously, it didn't stop at that internal YC post.

35:51Like we started posting a lot more publicly over time. But for a while, we didn't even... For the launch video category of companies, I think there are cases where you just want to go big really fast. Like you want everybody on your platform. maybe you're a BGC founder and you need that. For us, we actually gated demand a bunch. We didn't have self-serve onboarding, which for an API product is kind of weird. But we kept that true, I think, until a few months ago. So for probably a year, you could not sign up for Reducto without us onboarding you. And part of that was we, from day one, have been adamant like the only reason the company has the right to exist is if we are the best product in the space.

36:43Like that needs to be true. We didn't want to be the 200th PDF processing company that fails to deliver on the promise. And so we can talk a lot about this. Like we just would say no to people if we didn't think that we were at a point where we could solve their individual use case. Like it's such a deep problem that you can't solve everything from day one. And so if you came to us with like a chart extraction problem two years ago, even though we can do it now, like historically we would have said, Hey, we don't think we're the right vendor for you. Like we'll do it someday, but we're not doing it yet.

37:22And all sorts of things like that. To start, all we said we would do is the very first thing that we decided to train our own model for was we wanted best in class layout detection and understanding. And so we spent a lot of time building that and just being incredible at layout understanding. And then from there, we decided we wanted to be incredible at table parsing and table detection. And so we spent a ton of time on that. And these are broad horizontal problems. Like if you have great table detection, it ends up applying to a bunch of different industries and same with layout and the other things we did.

37:56But it was this iterative process of growing the scope. But for every individual stage, We made sure that we felt like we were incredible at what we were doing. And I know you mentioned this earlier of like, there were pretty good solutions available when we started. I think one of the closest comparable companies to Reducto when we first got started, like along this idea of unstructured data for LLMs, had this thesis around, you can upload any file to this platform. They had like 35 different file types.

38:30Turner Novak:I think it's 64 now, isn't it? Or something like that? Something like that. Like the number grows week after week. Yep. And a lot of our early customers were actually churned customers of theirs. And I talked to them of like, you know, what didn't work? Like whether they needed us to support that many file types because there was no way we were going to do that on day one. And they were like, no, actually we only really have two to three types of files that really matter. like this is the 80-20 or 95-5 in our case. And we just need those done incredibly well. And we can't spend all this time with like okay-ish pipeline for things that we don't even care about when the things that we do care about don't work.

39:15And that was really important. It's like, it was okay that we weren't competing on number of file types. We were just competing on for the people that needed PDFs and images. Great, like productive solution for you. and then gradually expanded from there.

39:28Turner Novak:So it's almost like 80-20 rule, but it sounds like it's more of like 95-5 rule or whatever, where like, yeah, it's like people maybe only wanted a couple things. It was like the vast majority of like actual customer demand versus like fancy marketing or, you know, like, oh, we can do everything. But, you know, the whole going super deep versus being super wide. Yeah. Yeah, and this is actually something that's taking a step back from Reducto. I just generally feel is true in startups. I think you see a lot of the founder's DNA in the ethos of the company for any company. And like for some of these competing companies, if you look at the founder's background, it's like a go-to-market and marketing and sales background.

40:15And that has its merits. Like I'm sure it helped a bunch for the early days. Whereas Ronak and I really started off more on the product side. Ronak's been doing research since quite literally he was 12 years old. Like he published in high school when he was 14 or 15. Like as a first author on a computer vision paper. And so when we came into this, like it wasn't the case that we were incredible at sales. It wasn't a case that we knew a ton about marketing. But what we did know is we can build something incredible that once people use it, still understand. And everything else has been a learning journey along the way.

40:53Turner Novak:Yeah, I think you mentioned, just like talking about Ronik for a second, I think you were in some kind of class at MIT and he was a freshman, like teaching the class or something like that. Am I remembering that story right? He was like a learning assistant for, it was my first graduate course. and I remember I had a ton of imposter syndrome because I was a junior in my undergrad and most of the people in the room were phds there's a course on meta learning so like teaching models to learn and yeah Ronik did the like he did the walkthrough of how we would approach the first psets like his approach to the problem was like the guide for solving it there's a whole different rabbit hole that we can talk about between like the professor at that class like ronick ended up getting that professor fired the really viral twitter post but yeah that's a deep rabbit hole oh no hopefully oh man i don't i don't know if we're gonna go down that rabbit hole right now if we had an extra an extra maybe 15 minutes.

42:03But oh, man, that's crazy. The short summary is that that professor had published a paper called GPD for I Can Solve MIT. This was one of the really famous papers around when there was all this language model hike. And then Ronak did a breakdown of the paper, talking through all the reasons why the results were invalid and all of this. And then when he's a freshman or sophomore, And there were just like crazy aspects to it where like there would be a multiple choice question. And the way they were evaluating was they would have GMC4 being prompted in a for loop again and again and again until I answered the correct answer in the multiple choice.

42:46And then I was like, oh, it got a 100 % off the score. So it just was nonsensical. And in the process of, you know, Veronica doing that write-up of all the reasons why it wasn't good research, it ended up coming out that the professor also hadn't gotten permission to scrape a lot of this data. And so it just became this like bigger thing that blew up and eventually the professor got fired from MIT. Wow.

43:16Turner Novak:And that's a pretty big deal. Obviously you got to get fair usage rights of the stuff they're using. One slightly different topic, but I don't want to miss it. you mentioned that you guys like kind of this whole concept sort of pivoting the product. You kind of made like, it sounds like a pretty bold pivot at the time. How did you decide to do it? Because I know you've said, I've heard you mentioned in the past, you were like pretty cognizant of you don't want to kind of be aimlessly wandering around, trying all these different things. How did you kind of approach the process of like, this is our like specific sniper shot, you know, organized pivot that we're going to do.

44:00Yeah. I think one of the worst things about PivotHell as a founder is that there are so many surface level interesting opportunities that have a like one layer deeper for whatever reason are hard or not viable, whatever it might be. Like if the way that you're approaching pivoting, And we've done this at some point before we land on Reductive. If the way that you're pivoting is you pull up a Google Doc and you're like brainstorming all these cool startup ideas and you're like going back and forth of like, oh, shoot, that would be so sick. You end up in this horrible position where you land on an idea that you think is, you know, incredibly insightful.

44:44It's like your eureka moment. And then as soon as you start Google searching, you find that somebody's tried it before. Because markets are smart. like if you thought of it in a brainstorming session somebody in that space has probably thought of it too and then you're in a horrible loop where you just you never make progress in any given thing like you're just bouncing from thing to thing so we knew we didn't want to do that because all of the interesting work is you have some sort of insights you find the reason why the insight is not as obvious as you thought or like they're blockers to it not just being free money or like a market inefficiency.

45:23But then you go one layer deeper of like, okay, other companies exist. What are they not doing well? Like what is the one step deeper reason why something could be better? And so I think the space that we're working on hit a few different checkboxes for us if it made sense to do. One is we felt we had an unfair advantage in computer vision. Like I mentioned, Ron has been doing computer vision in particular for basically his post-adolescent life. And so it just seems like a space where we could offer a lot of expertise. Two is we had spent a lot of time around this area, so we understood the problem well.

46:08The way that we actually started working together is we competed in the Anthropic hackathon where they first released Quad 2.0. We got early access to Quad 2 before it was publicly released. And we won that hackathon together. And one of the things that's, you know, one of the components of what we built there was the ability to upload PDFs. And it's like we saw the details of where foundation models struggled here.

46:33Turner Novak:And three, I think the, the like two steps removed version of the company was really really exciting on face value working on pdf processing felt like you know not sexy problem but looking at it as this as this broader what does it mean for language models to interact with human data like what is the scope of that clearly even back then felt like something that had to be huge like if it wasn't going to be us somebody would figure out this problem. And that was like the trifecta that was necessary of at least trying to go deeper in it. And we committed at the start of YC that we wouldn't pivot during YC, at least not like a meaningful pivot.

47:19We wanted to see YC through with like really pushing on one thing with a lot of focus and dedication for that thing. And that paid off.

47:27Turner Novak:Yeah, it sounds like it's almost like figure out like a core problem or trend or customer base to serve. It sounds like what you described doesn't quite fit in any of those buckets, but it sort of does. But it's kind of like find the one thing and then you can kind of like bounce around in it and then until you like dial in the specific part. And then it sounds like you kind of expanded from there once you actually figured out the core initial thing that people wanted. Yeah, and I think our space was especially interesting because there's... people ignoring the AI market, like imagine it didn't exist.

48:10There's already billions of dollars spent per year on document processing. Like there's an existing markets where if we were trying to build a marginally better solution, like there's probably some money to be made. It's not an exciting company if that was the market. Like it wasn't us trying to create something that new that was exclusive to language models. To us, it was this two-sided thing. There are legacy use cases that we can just be an order of magnitude better at. But also as this new generation of companies are being built, we can be core infrastructure for them to build better products.

48:47Turner Novak:And I think you, I've heard you say before, I think as a public, you got to about a million in revenue within a few months. So it sounds like you had some startups who were using you, obviously, you signed a couple of big customers. I think maybe we could talk now about the fundraising journey you guys have been on. I think you mentioned two years old, raised a total of it. I think you said 108 million, maybe I'm not remembering right. And you've only burned a million throughout the whole course of the company. Something like that. Yeah, maybe between now and when we publish, you'll have hired a couple more people or something.

49:21Turner Novak:And momentarily, maybe you'll sign a new deal though. Maybe it'll go back up. So just in terms of fundraising, what's that been like over the past two years? Yeah, we've been really lucky in that I think at each round, we've gotten to partner with whoever we were most excited with or about. and because we weren't raising as a like shoot we're about to run out of runway sort of position it also meant that we weren't approaching funding rounds as this needs to happen and more so as like a opportunistic there might be room for us to really accelerate the company as you know alongside this person.

50:05So we announced our seed round, I believe, in August last year. Raised a Series A a few months later. And most recently, a few days before this podcast was released, we'll be announcing our Series B from Andreessen Horowitz. And at each one of those stages, it really has been a question of what is the next phase of the company? I remember when we raised our Series A, we were four people. I don't know how many benchmark series A's are done at that team headcount. But they saw across their... Some of the benchmark portfolio companies were already customers of Reducto before Benchmark invested. And they saw how much those companies loved the product.

50:56And that, I think, speaks really, really strongly. The number one thing that has always mattered to us is not VC interest. It's how strongly do our end customers feel about what we're doing? Does it feel like something that is just an apathetic part of their stack that they could just replace at any point? Or does it feel like something that they are truly excited to build alongside? And I generally believe that funding is almost like after effects of the first thing being true. whereas I think a lot of founders try to look at it the other way.

51:34Turner Novak:Yeah, so it sounds like I think you probably only raised a couple million in the seed. You've only burned a million. So in theory, you didn't actually need to raise any money. Why did you think you needed to raise a Series A? So there are a few things. You're right that when we raised a Series A, we hadn't even spent our YC funding at that point. we really liked chatham at benchmark like i don't take this the wrong way but it's not often the case that you meet with a vc and it's like the way that they're thinking about your business is genuinely insightful and i i told chatham for for months that like hey like we're not going to raise like this doesn't make sense we'd received a lot of preempt interest outside of benchmark before that as well.

52:24And we just kept saying like, no, this doesn't make sense. Like we're not gonna raise now and so on. But we would just continue to have dinners and every dinner, it just ended up being a really productive conversation where over time we started to really think about it, not just as like cash in the bank thing, but also as we want this person on our board. Like we think that this would help in the next phase of the company. And obviously, I think I've heard the warning stories of raising money when you don't need it. To me, a lot of that comes from being undisciplined with the money that you raise.

53:05Like there's a great way to spend VC dollars. They're a tool like any other. And we've had the luxury of any time we want to work on a frontier model, we don't think in terms of compute costs. We've never told our ML researchers, hey, you have this cap on the number of GP hours. So we're purely thinking in terms of the end outcomes that we want. And that's the luxury that we get as a result of VC dollars. But the thing to me is we didn't want to see raising money as a mandate to immediately spend the money. We're obviously ramping up our spend, especially as we're scaling go-to-markets. But we were on the same page as the investors that preempted us about what the next year of the company looked like.

53:53And at that point, it just made sense to raise because terms were also sensible.

53:58Turner Novak:He's another benchmark portfolio company, Bobby DeSimone at Palmerium. He was telling me one of the board meetings they had, they were like accidentally profitable. And it was kind of like, it's kind of a joke that they're like, you got to be spending money. Like, why are you not burning cash? We gave you a bunch. We don't want you to be profitable right now. But that's good. I mean, it's good that you have a real business that makes money versus and like solves problems for customers being a, you know, a vehicle that transfers LP dollars to the landlords in San Francisco. Even now coming out of the series B, I remember I sent out our first investor update following the series B, I sent a note saying like, hey, like, we're, you know, really ramping up.

54:47And I do expect burn to ramp up. And so we started being more aggressive with spending, but also the following month, even before us announcing the round, ended up being the most six-figure contracts that we've ever closed. And so our spend went up, but our net burn didn't. So it kind of looks like we are just this flat net burn curve, even though that's not intentional. That accidentally profitable line resonates.

55:15Turner Novak:That's a good skill though, because when you're a public company, They want you to give these exact targets and being able to continue to make it look like... Even though every company is kind of a shit show, everything's always crazy. They want it to look from the outside. Like, oh, these guys are so good. The CFO is so good at guiding the street and hitting the exact target and look like it's super smooth. So you guys got it down two years in. That's a good sign. With another finance background, I guess. Yeah. Well, you got all those finance customers. Maybe you're learning something, talking to everybody.

55:49Turner Novak:So you mentioned that it took you about 48 hours to do it. What was the process of doing a Series B? Like, I think, don't people usually, you know, big, long process, tons of meetings, etc. What was it like for you guys? So we very much did not want to do that. One of the problems with having 20 people working at the scale that the company is at today is that I really think that every hour of time really matters. And so we didn't want to distract ourselves from the things that really matter, sales, products, and customer support, to just go on like a goose chase. So well before we decided to raise a round, I opportunistically would take meetings.

56:38There is an investor, or actually as of this podcast, this will be public. I'd known Jennifer since our seed round and really enjoyed our conversations. We'd kept in touch across a year plus. And so we had a pretty strong relationship, which I think maybe is not always true when people start a process. But when we came to a decision on whether or not it might make sense to raise a Series B, this is something I went back and forth on with a lot of our existing investors. We just came into it with like a clear mindset of we're going to run a very tight process. Like this is not going to be put the company aside for the sake of chasing around.

57:29And so we went with the two firms that we were most excited about. We had like a short list of five firms that we would want to race from on sunday i kind of just told the two that we end up actually pitching to like hey like something's happening if you're still interested like now's the time and then the following 48 hours were really really intense like we obviously weren't officially planning to do this so we hadn't done like a really thorough data room but needed to put one together because they needed to see everything about the company. So it was like the scramble of exporting all of our Stripe data and all of that kind of stuff.

58:15That was my full Sunday. I probably called the two investors a few dozen times in that 48-hour window where it was like every 45 minutes we'd be on the phone because everybody was trying to condense everything. the following day we did the two partner meetings we got two term sheets and both of these firms were firms that we were really excited about so it wasn't like if we get a term sheet from one we're just immediately saying yes but at the end of the day like i think there are there are a few things that matter in a funding ground and i don't think that valuation is the number one thing there.

58:55So we weren't too concerned about just artificially bidding up just the raw numbers in the rounds. The things that we really did care about is like making sure that a given firm was right for exactly what we were trying to accomplish in the next 12 months. And especially in the spirit of what I mentioned, Reductose had incredible traction without a sales team, primarily almost entirely from word of mouth if our customer is talking about the product. Now's the time when we were ready to really work with the broader range of customers that could benefit from Reducto. And Andreessen felt like the right firm for that.

59:35Turner Novak:And it sounds like you're going to start to hire up a little bit more on the team. I think you mentioned there's about 20 people. You mentioned the first AI researcher that you hired. I know it's an interesting story. What happened there? So Ife is incredible. We knew about Yifei from very early on in our company history. So just for context, he was his PhD student at Purdue. He was doing his entire PhD on vision models for document processing. Like one of the, probably one of the best experts in the world at exactly what we do. And it's rare that you find that sort of fit. Yeah, it's probably like the most boring PhD to like think of like, I want to do a PhD on this, but like the perfect fit for you.

1:00:19Turner Novak:for reductile. It's rare that you find somebody like that, but he knew all of the nuances of why the problem that we were working on mattered because he was spending all day thinking about it. And so as a side project from his PhD, this isn't even the frontier research that he was doing. He had open source models for the space and they became number one on Hugging Face. I remember we even had a competitor who has since shut down whose whole product was built on Mifei's side project open source work. They were wrapping that model. Great work, honestly, especially considering it wasn't his core focus.

1:01:01And so we spent a while just trying to get on the phone with him. Ronik had DM'd him. And I think generally speaking with somebody like him, there's always a ton of companies that want to convince them to come on board. I DMed them again like two months later and eventually just got him to take one 30-minute call.

1:01:25Turner Novak:How'd you convince him? It was just like us texting on Twitter. So it was Twitter turned into Google Meet.

1:01:35And I think even on that first call, you could kind of see that a lot of the ways that we were thinking about the problem, what it could be, were just very much aligned. but he still had some reservations because we hadn't even announced the seed rounds. We were a two-person company still. Just made our first founding engineer hire. And so we decided to tell him that, hey, actually we have some meetings in Indiana. And I know you're based in Indiana. Can we just grab lunch on the way? I only found out recently that he actually believed thought we were in Indiana for no reason.

1:02:18Turner Novak:Amazing. The most random place we had to be. Yeah. He's a very sincere person. So yeah, we told him, hey, we're going to be there. Can we get lunch? And as soon as he said yes, we took a red eye over, flew to Indiana just for this, like classic brushing in the airport bathroom, and then just going directly. And it wasn't just lunch. Like we ended up spending seven or eight hours together. We were just walking around the Purdue campus, talking about everything that we were doing. It was partially, it wasn't even like necessarily us interviewing him. It was also him interviewing us, really digging into why we made the decisions that we'd made and all of that.

1:03:08And I think the thing that really sealed it for him, going back to what I mentioned, And, you know, our customers choose us because the product is best in class. Yifei had offers from some of our competitors as well. And he tried everyone's products and ended up seeing the best results from Reducto. And again, like there were competitors that were larger teams. And the same way it applies to our customers. I think he just saw this as a place where he'd be excited to build that sort of state-of-the-art. And so, yeah, we're lucky to have him on the team. He's been on the team for quite a while now.

1:03:44That's for quite a while in the context of a two-year-old company.

1:03:49Turner Novak:Yeah, percentage of the time the company's been alive, yeah. What else are you guys hiring for? I guess just like, I mean, we can plug it obviously right here really quick. But for people listening to this right now, things that sounds interesting, like who are you looking to join the team? Yeah, we are hiring across the board. where like on engine research, that is a role where we are always looking for exceptional people for all the reasons that we've talked about today. Like these are really hard problems and we need exceptional people to be able to solve them. We also are hiring for design as we build out the platform for Reducto, but especially the thing that I'm really excited about is we are newly hiring for go-to-market.

1:04:31For a long time, we would get salespeople reaching inbound and we just weren't ready to transition away from founder-led sales. But even in our early hiring, like our first AE in his first four weeks, I'm including ramp up time here, closed a six-figure deal from zero, like initial discovery call to contract signed. And like he was talking about how in his entire career so far, he has never had a deal cycle move that quickly for a greater than 100K deal. And that to me is like the signal for, okay, we clearly have way more demands than we have capacity to sell. So we're growing our AE team, hiring a head of marketing and so on.

1:05:15Turner Novak:I think you told me before you got to about 5 million in ARR, before you hired them, before you brought anyone on the sales side and you are, you describe yourself as like not a salesperson. Like I'm listening to you talk. Doesn't sound like you're selling me anything. like the classic snake oil salesman. How did you get good at sales? Like what's like a general, we can unpack this a little bit more, but like at a high level, like how did you learn how to do this? I know that one of the things that sales leaders always say is like, you're not supposed to demo on the first call. Like don't do it.

1:05:48You should just spend your time doing discovery. That's been historical advice that I've been told many times because I guess the nature of the advice for a long time was like, You need to spend a lot of time understanding the customer's context before you come to them with a nice polished demo. And part of this also is the traditional structure. If you have a non-technical AE that then needs to bring in an SE for the demo, that sort of work. I think that the benefits that founders have when they're selling is founders know every little nuance of the product on a really deep level. And to me, it would be a shame if that wasn't captured, even in the first call.

1:06:34Because the way that I've kind of mapped out sales, I know there's traditional sales, like milestones. But to me, there's really only two phases that matter. The first is the inspiration phase, like getting the prospect to a point where they want to buy. And as soon as that is done, the job of the salesperson becomes facilitation. Like it's you going through the steps of, you know, walking through procurements, talking to the rest of the team and so on. I'm being redundant here, but you get what I mean. And so the thing that we have really optimized for is I try to get people to that inspiration phase as early as possible.

1:07:14And oftentimes that's in the first call. Like we've had cases where people will bring in C-suite on the second call for really large companies just because the first call was something that they were just buzzing about. that kind of thing is something that I think if I came from a traditional sales background, I wouldn't be approaching it this way as like a benefit of being an amateur in the space. With that being said, like there's plenty of things that I didn't know about. Like I'd never gone through a procurement loop and that's something where no amounts of product expertise even matters. There's just like traditional knowledge that you need.

1:07:51Turner Novak:Oh yeah. So I guess for somebody who's never gone through that, can you just describe why that's a big deal? Like the shock that you get? So for people that haven't done sales before, there's like two components or every large enterprise will have the sales process that you're going through to convince the end buyer, like the person that's using your product. And that, if you're aligned on product and what it does is hopefully the easier one of the two. If the value of the product is there, like people should be aligned. The second part is that enterprises have so many things in place to prevent bad purchases.

1:08:33That includes their security review team. There'll be dedicated teams that are just there to make your contract as difficult as possible. Like they'll negotiate everything. It's not just pricing. It's like the renewal terms. It's the time that they take to actually send you the payment for the invoice, like all of that kind of stuff. and that can be to me that seems like something or at least when I came into this I would have imagined that would be like you know a one week process in practice enterprise procurement can be a month plus process if not longer in many cases and so that sort of stuff I ended up just needing to lean a lot on our investors like first round was incredibly helpful here Emery at first round actually guided me through the process of our first procurement loop And it just became something that we got better at over time with more and more reps.

1:09:26But the challenging thing, I think, as a non-seller is similar to how, as a founder, you're at a disadvantage against VCs because VCs are seeing thousands of deals. And you don't even know the scope of what you can really negotiate unless you have a lot of founders you can lean on. procurements is seeing tons of deals done and the moments they sense weakness some procurement people will just like jump in and they'll be like you know you're a two-person company we need to access deal unless you can meet this price things like that where that just aren't true that's as a first-time founder like can sometimes fall for someone described it to me as basically selling software, there's no cogs.

1:10:14The price can be$1 or$1 million.

1:10:18Turner Novak:It can be anything in between. It's just, are you delivering enough value to make it worth them paying that? There's a 10x ROI or whatever. I don't know what the number you pick, but you got to make sure it's actually valuable and they'll justify paying whatever this pretty big price tag is at the end of the day. Yeah, and I think the thing that I've realized over time is when I first got into it, procurement was painted as this like antagonist. And I think that on some level, the incentive structure is there. But for a lot of our enterprise customers, I'm actually on really good texting terms with their procurement leaders now.

1:10:59Because the same applies to them, right? Like I'm sure they deal with vendors that are really difficult, vendors that don't live up to the promise of what they put in a contract. and so everything that a seller might say about how difficult procurement can be procurement can talk about of how painful vendors have been for them i think if you approach it with that sort of empathy of like they are trying to do what's best for their company you find that like they're not actually trying to screw you over generally they're just trying to make something work that works on both ends and people are willing to be reasonable if you're reasonable.

1:11:36Turner Novak:Yeah, so one of my portfolio companies, his name is Chris Hlagic. He has a company called Hanover Park. I don't know if you've come across them. I've seen them. Yeah, so they do. It's like fund admin for investment firms. I don't know if they use Reducto, but maybe they should because they do a lot of document ingestion. Maybe I'll text them after this. But one of the things he told me that works really well is like, just you need to immediately get people on a second channel, right? Like you need to be able to have maybe you get introduced on email, but like you got to get their phone number and just have like a just like a different way to talk to them.

1:12:09So like the email, maybe it's a more semi formal follow up and the text is like, I don't know,

1:12:16Turner Novak:like a funny meme related to like their product. I mean, just like another way to stay top of mind versus like, hey, just following up. Did you see the thing or like, do you have any questions? Like you kind of need a reason to stay on top of people that's not bugging them, that actually kind of seems natural. And like, oh, you know, just thinking of you, love you guys so much. Like, you know, so that's something I've always kind of kept in mind. Like immediately get their number. I think that's actually something Jathan does too at Benchmark. When I've learned it's like, get their number, get in their calendar, plan the dinner, plan the next dinner after that.

1:12:52Turner Novak:You know, you kind of like, you just get in their flow almost. He's really good at that. It's like what we were talking about earlier of like me telling Chaitlin, hey, we're not going to raise like I'll go to dinner, but no expectations for dinner. One last question I had for you. So with the models kind of always changing, just like I mean, everything's changing. I think like the classic in AI is like, you know, you can't make plans because in a week everything changes. Like I'm sure by the time we publish this, everything we talked about will be irrelevant two weeks later. So what does that look like behind the scenes?

1:13:26Turner Novak:on your end? Like, just how do you stay on top of it? I think one of the mandates for Reducto is not even just train great models, it's help our customers not have to think about PDF processing. And so generally what that translates to is anytime a new model comes out, like in GT5 or someday like a Gemini 3, our customers shouldn't have to think about, oh, like, do I need to swap this in for some subset of my documents, all of that? We should be doing that for them. And so we've put a ton of time into just being really rigorous that's evaluating on all sorts of different subtests. And we will swap in models, you know, post-train models really quickly whenever we find that, hey, this is better at, let's say, checkbox detection, whatever that subtest might be.

1:14:21But the other side of that of just making sure that your product is always at the forefront, is the most interesting thing as a founder in this space is you end up coming across new opportunities that just weren't possible before more often than I think you would have pre-language models. Like, I remember we wanted to do document editing around this time last year. Like in 2024, we were curious of how can we fill out documents that are scanned, which is a really hard problem. It's almost like the reverse of everything we've talked about. You need to figure out what the empty space actually translates to in terms of putting something in.

1:15:03Or if there's a square, you need to figure out, oh, this is a fillable checkbox, like all that kind of stuff. And that just wasn't possible to do well with models that existed then. And we tried a bunch of different things, but it just wasn't a product that we wanted to ship because it wasn't good. It didn't live out to the promise. and it wasn't until recent advancements that that did become possible. We released what I think is the first, you know, truly horizontal document editing API and it requires everything. Like we're using reasoning models, we're using Frontier VLMs, we're using traditional CB and all this put together is like the first time that you can do something like this.

1:15:45And I think that'll continue to happen maybe on the order of, you know, when this podcast is released, maybe on the river for a few months, but I think you just constantly have to be looking at not just what is the problem today, but what will people need in the next few weeks? Like what is now possible that just wasn't possible a few weeks ago?

1:16:03Turner Novak:Interesting. I mean, and it begs the question, maybe like last point is just like, what is the future of just kind of all this sort of look like? Like when you just fast forward to the end, I don't know if you thought that far, but like, what does it, what's it all look like? The end state of PDS and documents is, I think, Ultimately, these documents are proxies for human work. They're generated by humans as an artifact of the work that they did. They're the source through which people get information for their work. And the first wave of AI was sort of like synthesis and summarization tools. You had some corpus of information and you wanted to be able to ask questions from it.

1:16:44You wanted to be able to pull insights, whatever that is. but the place where people are clearly headed as you see more of these AI employees and you know agentic workflow products and all that is people are doing that end-to-end like it's not just about answering a question that a person answered or asked it's how do you fill out the end-to-end process like we have customers that do security questionnaire automation and they'll just take the context about the customer like they'll parse all that with producto they'll extract the questions from the questionnaire and they'll actually fill out the questionnaire for you so that you just upload that completed questionnaire.

1:17:19That sort of thing, I think, is really interesting because that's the point where it's not purely a tool to help you do small subtests. It's a tool that takes things off your plate.

1:17:32Turner Novak:Yeah, it's kind of like when you think about this whole agent-to-agent promise that we have, it's just like instant transfer of tons of information. Yeah, well, this has been a lot of fun. Thanks for coming on the podcast. Yeah, this was great. Thanks for having me. And thank you for listening. Thanks again to Numeral and Hanover Park for supporting this episode. Head to numeralhq.com for the fastest, easiest way to stay compliant with U.S. sales tax and global VAT and to hanoverpark.com slash turner to upgrade your fund admin to the 21st century. If you missed it, make sure to check out last week's episode with Eric Bernhardtson and Modal on building AI native infrastructure for developers.

1:18:11Turner Novak:If you like this conversation, please like, comment, subscribe, and name your next PDF after me. If you want to miss a future episode, subscribe to my newsletter, The Split, linked in the description to get each episode plus a transcript emailed directly to your inbox every week. Thanks again for listening. See you in the next episode.

From the publisher

Sign-up here to get weekly episodes + transcripts in your inbox: https://www.thespl.it/


Adit Abraham is the Co-founder and CEO of Reducto. Reducto’s product takes PDFs and physical documents, and extracts all the data, just like a human would if they were reading it.


At the time of recording, they’ve processed over 1 billion pages, grew 6x over the past five months, and are fresh off a $75 million Series B led by a16z. And Adit told me they’ve only burned $1 million of capital so far to get here.


And the craziest part, Adit told me they’ve only burned $1 million of capital so far to get here.


Anyone building an AI product probably sees Reducto as essential infrastructure. Our conversation gets into how they built the best product in the space, landing a Fortune 10 customer as a two-person startup, getting to $1 million in ARR within a few months, lessons doing founder led sales to over $5 million in ARR, and what the future of PDF’s, and human / computer data looks like.


Thank you to Liz Wessel at First Round, Chetan Puttagunta at Benchmark, and Adel Wu at Reducto for helping brainstorm topics for Adit.


Thank you to Numeral and Hanover Park for sponsoring this episode.


Numeral: The end-to-end platform for sales tax and compliance. Try it here: https://bit.ly/NumeralThePeel


Hanover Park: Modern, AI-native fund admin at https://www.hanoverpark.com/Turner


Timestamps:

(3:35) Reading unstructured human data

(10:44) Growing 5x in four moths

(12:38) Insurance, healthcare, legal, logistics

(19:13) Where LLM’s still struggle

(28:23) Starting Reducto from a blog post during YC

(32:01) Landing a Fortune 10 customer with two people

(35:48) Limiting the product and growth early on

(40:57) Getting an MIT professor fired

(43:50) How to avoid pivot hell

(49:00) $108M from First Round, Benchmark, a16z

(51:48) Chetan convincing them to raise a Series A

(55:50) Raising a Series B in 48 hours

(59:36) Redeye flight to hire the 1st AI researcher

(1:05:42) Lessons hitting $5m ARR with founder-led sales

(1:13:09) Staying on top of changes in AI models


Referenced:

https://reducto.ai

https://reducto.ai/careers


Follow Adit

Twitter: https://x.com/aditabrm

LinkedIn: https://www.linkedin.com/in/aditabraham


Follow Turner

Twitter: https://twitter.com/TurnerNovak

LinkedIn: https://www.linkedin.com/in/turnernovak


Sign-up here to get weekly episodes + transcripts in your inbox: https://www.thespl.it/

More from The Peel with Turner Novak

All 64 episodes
Inside Reducto: YC to Series B in 18 Months, From Pivot to Fortune 10 Customers, Lessons in Founder-Led SalesThe Peel with Turner Novak · 1 h 19 min
Listen in VO