How Ramp built an AI agent that can think outside of tokens | Alex Shevchenko

7 May 2026 · 44 min · 21 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Ramp’s “Ramp Sheets” agentic spreadsheet editor and the surrounding agent architecture, eval/testing, and self-improving monitoring loop; plus two R&D experiments (latent briefing with KV-cache communication and steering vectors inspired by Golden Gate Quad).

Guest

Alex Shevchenko, Head of Applied AI Research at Ramp (RAMP). He leads applied AI research; discusses internal finance automation origins, agent tooling, and experiments.

Key claims

  • Spreadsheet-native agent actions beat codegen for finance users because it avoids a “black box” and matches how accountants read formulas.
  • Ramp Sheets uses an agent SDK plus a SpreadJS sandbox; ~95% of work uses Excel tools (~10 tools), Python only as an escape hatch (~5%).
  • Inspect enables a self-monitoring loop: automated monitors run in shadow mode, then get “promoted” if low-noise; it can also open PRs.

Notable examples

  • Month-end close/reconciliation workflows from Loom videos; agent reads specific cell ranges and writes new sheets with cell references.
  • Latent briefing: orchestrator is closed-source (Anthropic-like), workers open-sourced; KV-cache communication reduces token usage.
  • Steering vectors: concept obsession (Ramp-themed) using synthetic contrastive pairs; applied to ~five Gemma layers; sometimes causes “self-aware” off-topic loops.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring RAMP's Sheets and Internal Agent

0:16 to 0:28

Discussion on RAMP's AI spreadsheet editor and its internal coding agent.

“One of the approaches that's easiest and probably more accurate is still doing CodeGen and then just plopping it back into the Excel-like view.”

Experiments with Recursive Language Models

0:28 to 1:18

Insights into the experiments done with recursive language models.

“and how RAMP's internal coding agent, Inspect, fits into it.”

The Origin of RAMP Sheets

1:18 to 1:47

Alex shares the evolutionary process leading to the creation of RAMP Sheets.

“one of the things that I want to dive into is Sheets Agent, or that's what I call it.”

Markov Diagrams and Automation

1:47 to 2:16

Explanation of using Markov diagrams to document finance processes.

“And we would do like this video to like Markov diagram process to try to just map it out so that engineers could then pick it up as a piece of documentation and try to automate it on the finance person's behalf.”

Using Loom Videos for Process Documentation

2:16 to 3:38

Discussion on how Loom videos serve as effective communication tools for documenting processes.

“Can you talk a little bit more about that process?”

Challenges of Creating Artifacts from Loom Videos

3:38 to 4:28

Exploring the difficulties in producing artifacts from Loom videos for automation.

Feedback on Black Box Automation

4:28 to 6:18

Alex discusses feedback received regarding black box automation approaches.

“that are like well documented off of these Loom videos.”

Building the Agentic Spreadsheet Modifications

6:18 to 8:00

Discussion on the decision to create agentic modifications in spreadsheets based on feedback.

“And like 99 % of the time, they're in a spreadsheet.”

Realization to Package and Share RAMP Sheets

8:00 to 8:10

Explaining the decision to prepare RAMP Sheets for external release.

“And we looked at it and we realized like, maybe we ship this out into the world.”

User Experience for Non-Financial Companies

8:10 to 8:58

Discussing the user experience for early-stage startups in relation to RAMP Sheets.

“That's not necessarily true for a lot of companies, especially like early stage startups.”
Show all 21 chapters

Evaluating the Process Mining Outputs

8:58 to 9:32

Alex evaluates the outcomes of the process mining efforts related to RAMP Sheets.

“And so that's what was launched in November, which we called Ramp Sheets, which is like this agentic spreadsheet editor.”

Graph Representation and Future of Agents

9:32 to 12:00

Discussion on the outputs of the process mining and the future of creating agents.

“I mean, it depends on, on the task at hand.”

Architecture of the RAMP Sheets Agent

12:00 to 14:00

Exploring the architecture and tools used in the RAMP Sheets agent.

“Maybe talking about the RAMP Sheets agent a little bit more.”

Building an Excel-Integrated AI Agent

14:00 to 17:26

Learn how Ramp created an AI agent that leverages Excel tools for finance.

“It's like still better with Python and they're like less reluctant around it being like a black box thing.”

User Interaction and Performance of the Agent

17:26 to 21:46

Discover how users interact with the AI agent and its performance metrics.

“Like it was very, very overfit to the task that we ended up like landing on.”

Self-Monitoring and Automation in RAMP

21:46 to 28:00

Explore the self-monitoring system developed for RAMP's AI tools.

“Could you compare them exactly and like these need to be exactly the same and if they're at all wrong, then this is wrong or is there some gray area?”

Experiments with Sheets and Memory Management

28:00 to 31:15

Learn about various experiments using Sheets and memory management techniques in AI.

“What other wacky ideas have you run on Sheets or experimented on Sheets, whether they've seen the light of day or not?”

Exciting Recent Projects at RAMP Labs

31:15 to 37:19

Discover two recent experiments at RAMP Labs focusing on context management and user interaction.

“Yeah, there's two, uh, that we published like last week and the week before that were really exciting.”

Reviving Golden Gate Quad Experiment

37:19 to 39:35

Explore the revival of the Golden Gate Quad experiment using user-defined concepts.

“So it's not going to be like an exact perfect like replication of that like Jeep steering vector, but it's going to be roughly broadly aligned to it.”

Steering Vectors and Model Interactions

39:35 to 42:01

Understand how steering vectors are applied in models and their impact on interactions.

“And we would just pass it over to like Opus as part of that like evaluation process to try to figure out, okay, at these layers it does end up working pretty well and at these layers it creates degenerate output.”

Understanding AI Engineer Roles and Skills

42:01 to 43:50

Learn about the evolving role of AI engineers and the skills needed in the industry.

“So like interpretability or like we have one guy that was doing an RL startup previously and so he's interested in doing a bunch of RL experiments within Ramp Labs.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Alex Shevchenko:Maybe we can just finish the loop so that they just record the video and create the automation themselves. Today I'm talking to Alexander Shevchenko, head of applied AI research at RAMP. Sheets, their AI spreadsheet editor, has been getting a lot of traction, and we dive into the architecture decisions underneath it.

0:16Max Agency Host:One of the approaches that's easiest and probably more accurate is still doing CodeGen and then just plopping it back into the Excel-like view. We decided not to do that. We go deep into the self-monitoring loop behind Sheets

0:28Alex Shevchenko:and how RAMP's internal coding agent, Inspect, fits into it.

0:31Max Agency Host:The agent deems them to be, like, good enough, and they get, like, kind of promoted to start bugging the engineers.

0:37Alex Shevchenko:Alex then goes into some of the experiments they did around recursive language models.

0:41Max Agency Host:By having, instead of them, communicate in token space, the actual orchestrator in the ROM is closed source, so it was like an anthropic odd family orchestrator, but then the worker agents were open sourced.

0:53Alex Shevchenko:We also get into some of the experiments that Ramp is running around steering vectors.

0:57Max Agency Host:It would almost become self-aware around like its interactions of like, why am I talking about spaghetti when the question is about like the meaning of life?

1:04Alex Shevchenko:Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you. I'm really excited to chat with you today because I think you guys are famous at this point for doing a lot of really cool stuff in the AI engineering space. one of the things that I want to dive into is Sheets Agent, or that's what I call it. I don't know if it has a more official name.

1:27Max Agency Host:Ramp Sheets.

1:27Alex Shevchenko:Ramp Sheets, yeah. So I think you guys launched that in November. What was the origin story of that and why build Ramp Sheets?

1:37Max Agency Host:Yeah, so we actually arrived at it through a very long evolutionary process of trying to address the needs of our own internal finance team. And so initially it started out as this process mining project where we were just trying to understand what were the things that our own finance team was doing, what our accountants were doing as part of the like month end close tasks. And we would do like this video to like Markov diagram process to try to just map it out so that engineers could then pick it up as a piece of documentation and try to automate it on the finance person's behalf. and that turned out to work super well.

2:18Max Agency Host:Can you talk a little bit more about that process?

2:20Alex Shevchenko:Like you would video to Markov diagram.

2:23Max Agency Host:Yeah.

2:24Alex Shevchenko:What's the video of and then how did that process work?

2:28Max Agency Host:Actually, one of the like richest ways of communicating in any company still to this day, but especially like a year ago is like recording a loom video of you going through something and just sharing it because it's like the like kind of easiest way of creating an artifact. It takes a lot less work to just show around and click around the spreadsheet and show how you do something compared to like trying to write out a nice clean notion doc of the process or trying to like formalize it as an SOP. And so a lot of communication between like accountants, but also like engineers and kind of like everyone is like you record a loom, you send it off and then the person watches it.

3:09Max Agency Host:And would that include like a voiceover or the person talking as well? Yeah, exactly. Video of them walking through like an Excel and audio of them narrating it. It is very easy to produce, but it's kind of harder to consume than an ocean doc. And it's like the more effort you put into producing that artifact, the easier it becomes to digest for the person, but the more work it is on the person producing it. and our finance team is like very very busy and so they don't have necessarily that much time to like produce these artifacts for our software engineering team for us to automate it and so we try to to basically build something out so that to facilitate our work of consuming those looms that are very very like highly technical but in the finance like world and so getting that financial context into our own heads by getting this pipeline working that would take that loom video of them speaking and explaining a process and then generating basically a text document and this like process diagram of just like node starting like I opened this bank account I pulled the invoice from it I opened this other system of record maybe I opened like the GL or something and I tried to reconcile between the three and they have a bunch of these like one-off tasks that they run every month.

4:26Max Agency Host:And so we started building out kind of like this library of tasks that are like well documented off of these Loom videos. And the idea then was, well, we'll take a software engineer and sit them down and they would just like automate them one at a time by writing like a Python script or something. At that point, we thought, well, maybe we can just like finish the loop and take what we have built and do like cogen or something to automate it away for for the financial professional so that they just like record the video and create the automation themselves. And we tried a few different approaches of it, like just Python code gen.

5:04Max Agency Host:We tried N8N, um, we tried like retool automations and kind of the feedback that we got was all of these are kind of two black box for the finance professional. Like, obviously this is like very high stakes, kind of like a asymmetrical thing where like saving time is great but making a mistake is like disproportionately bad and so this black box approach really like was not good enough for them it would be black box because they would

5:30Alex Shevchenko:pass in an excel spreadsheet and get back something and have black box in the from the

5:37Max Agency Host:like finance person's perspective like it's not a black box for us because it's like python code that we could probably verify but if we're trying to automate it as much as possible we want to get it end-to-end automated and Python is like a great medium that's very expressive but to them it's a black box because they're not going to understand that they're not going to understand like the pandas import and the data frames and how they're they're being like uh transformed and so we took a step back from that and like the feedback what like how can we work around that feedback and we started just watching the looms ourselves um of them doing the work and And it's interesting because you would drop in your cursor at kind of like any point into one of their videos.

6:22Max Agency Host:And like 99 % of the time, they're in a spreadsheet. Like there's some amount of time at the beginning where they like loaded up from like a bank account or like the GL or whatever. But then the majority of the time, it's just them going through an Excel spreadsheet and like making modifications or creating like a worksheet off of it to make the reconciliations or something. And so we decided, well, we should probably meet them where they already are, which is the spreadsheet. And so we took that process mining, like video, like process generation pipeline, and built out like the V2 version of the automation.

7:00Max Agency Host:But instead of doing code gen, we decided to replicate the exact same way as they perform the work, which is opening an Excel spreadsheet and making the modifications directly in it. and taking those learnings of like not making it a black box like one of the approaches that's like easiest and probably like more accurate is still doing code gen and then just like plopping it back into the excel like view we decided not to do that we decided to make it like agentic spreadsheet modifications in the sense of like read this column read this row write this row and with excel formulas in the spreadsheet to really like try to capture the exact same workflow as they do.

7:41Max Agency Host:And that was much, much better received because, well, they right away get a lot more visibility into what's happening because they are professionals at parsing Excel spreadsheets. They're better at reading all those formulas and how they fit together than we are at reading code. And so they became a lot more receptive of that medium. And that was used somewhat internally. And we looked at it and we realized like, maybe we ship this out into the world. And as we prepared like with Ramp Labs to like package this up in a nice way and send it out into the world, we thought to ourselves, well, Ramp has like these amazing processes for like month and close and like all our like ducks are in a row basically from the finance side.

8:28Max Agency Host:That's not necessarily true for a lot of companies, especially like early stage startups. They don't necessarily have a process to automate yet they are in like they need to create that workbook from scratch and so we started looking at like maybe we just take that automation piece and we package that up without the process mining piece because well there is probably not that much to process mine for a lot of companies as is and it's like a more complex setup and just like the first time user experience is like nicer if you can just pop in and just be like build me a dcf or something And so that's what was launched in November, which we called Ramp Sheets, which is like this agentic spreadsheet editor.

9:08Alex Shevchenko:I want to talk more about Ramp Sheets, but first I want to talk a little bit more about the process mining stuff. Like how far did that get you for Ramp Sheets? Like did you have this process down where you would get these looms, drop them into the system, some system you created, and out would come an agent that could do all of these things? Like what did that process look like? And if it was an agent at the end, like how good was that agent and how much like more tuning did you have to do on top of it?

9:34Max Agency Host:Yeah. I mean, it depends on, on the task at hand. Um, so for like very complex, like if it's like FPNA and if it's something that's like out of the distribution of things that the model was RL'd on, it's not going to perform super well. Even if you, if, even if you give it a really good context, like LMs kind of like tend to go towards like middle of distribution outputs. Like we didn't get super far in those strategic finance workflows. But accounting and like reconciliations and bookkeeping, those are types of things that are very repeatable. The context, like you can probably get most of the context in and you don't need to like any cleverness or originality about it.

10:15Max Agency Host:It's just like a very strict process. So those types of things are like, if you map out the process and you get the right pieces of context to the agent, it's able to replicate it pretty well.

10:27Alex Shevchenko:And so what was the output of this process mining thing?

10:32Max Agency Host:Yeah, so it was a couple of things. It was like, we kind of from the same video, I'll put it in multiple modalities. One was just like a text description of the process itself. that's like as simple as like it's like the equivalent of like today if you go into like i don't know like gemini uh like studio and you drop in a video and you ask it to like summarize it that's kind of like the textual representation of like as far as it went and there was like a more complex thing where we would create like this kind of like directed graph of nodes of like well first you open like your bank account to pull invoices and then that invoice will get used for like the sheet like sheet one in this workbook and then you need to get your ledger and this is going to be like sheet two and so kind of like this directed graph of dependencies and like work to be done and did you have some DSL for this for this dag or how is that represented this is like almost naive but it worked out well we we basically generated like graph whiz language for it so like the library for for making like dot graphs and we would just have it generate that but with a lot of constraints around like the graph that we wanted it to output and that's what that's what would like generate that process line

11:49Alex Shevchenko:diagram do you think this is how all agents will be created in the future you make some loom video you drop it in it creates a skill or a graph whiz thing or something like that and boom there you go

12:00Max Agency Host:yeah nowadays i would replace like a few of those pieces with something like kind of like more modern more 2026 of uh building out skills or building out something that is like more cloud code native i think the process mining piece which is still interesting um we actually haven't been like doing as much like experimentation to develop it further but it's so good for like you take a video and you get you produce like an artifact of documentation that you can have some like store somewhere and then have it consumed by like cloud code or like cloud co-work maybe not end-to-end automations, but just like as you are working and as you're like pulling in context, that ends up being like a pretty good workflow as well.

12:42Alex Shevchenko:Maybe talking about the RAMP Sheets agent a little bit more. So what did it look like under the hood? You already mentioned some of the tools it had, but like what was the architecture? Was it similar in architecture to like Cloud Code, this LLM running in a loop with tools? Did it have access to a file system? How is the spreadsheet represented?

13:01Max Agency Host:We have like an agent SDK that runs the loop. And then we spin up kind of like a modal sandbox in which we have SpreadJS, which is like this Excel manipulation library that loads it up and then is able to like perform modifications on it. And the agent will have like specific tools like read range, set range, format range. And those are given like as actual tools,

13:29Alex Shevchenko:not as like functions in a REPL or scripts in a coding environment or something.

13:34Max Agency Host:Yeah, directly as tools. But they interact with like the spreadsheet inside of the sandbox. The reason for the sandbox is that there is like still at the end of the day, like an escape hatch for doing some code gen stuff. So sometimes people want to do things that aren't necessarily like Excel shaped or Excel native, but we still want them to be able to do it in RAM sheets. So like doing like a large cleanup or something that still lends itself better. And when it's not like a month in close task or something that someone has like a list of vendors and they want to just like filter out easily on a bunch of clauses.

14:08Max Agency Host:It's like still better with Python and they're like less reluctant around it being like a black box thing. So it's still like running in a sandbox. We can do code gen. We can execute it against it. But most of like the actual like finance or accounting shaped stuff is done from pure like function call of like read range. And then the like agent will try to systematically as a human would like kind of like read in diagonals with like these kind of like blocks of it's not going to read the entire thing so that we don't pollute the entire context. It reads in one spot, then it reads further down just as a human would.

14:44Max Agency Host:And then it would produce like another sheet, for example, by writing to it with references to cells from the sheets that it discovered or like columns.

14:53Alex Shevchenko:Do you know, do you have any clue how often it used the Python tool versus the normal Excel tools that you gave it?

Read the full transcript

15:00Max Agency Host:We tried to like forcibly like make it use it the least possible, like just as a last resort escape hatch for things that are not possible just through pure Excel formulas. like we really tried to bias it towards Excel formulas. There's stuff that's like outside of the pot, like things that are possible within Excel. So as a percentage, it would be like 95, 5 % maybe.

15:24Alex Shevchenko:For the Excel tools that you gave it, was this like order of magnitude, like five tools, 10 tools, 20 tools? How many different kind of like,

15:33Max Agency Host:yeah, specific tools for specific capabilities

15:35Alex Shevchenko:did you give it in Excel?

15:37Max Agency Host:Excel-wise, I think it was around like 10-ish tools maybe. so not super super large tool set but different than maybe like bash which is one universal tool we found that like the especially the entropic family of models is like really good at agentic spreadsheet manipulations especially back when we launched it actually the latest models from like open ai have been like catching up around it but you could like kind of clearly tell that the entropic models were like really are out on on being able to like decompose excel tasks into specific actions to take on the spreadsheet one by one.

16:15Alex Shevchenko:I think there's some idea in coding that you want to line up. If you're building your own harness, you want to line up the tools that you give it similar to what the models have been RL'd on. And so that's kind of like similar tools to whatever's in cloud code or something like that. Do you know if the 10 or so Excel tools that you gave it were like, did you talk with the Entropic team and were like, hey, what did you RL it on? And can we write the same type of tools? or was it like close enough probably and so it probably just worked?

16:41Max Agency Host:We launched it in November. That was like before we got access to like Cloud for Excel or anything. So we kind of had to like rediscover it from first principles ourselves. And the initial version was kind of like more open AI leaning. And then we discovered like the Cloud family just tends to perform much better on the way that we had developed the tools and I think it ended up replicating it quite well in their internal environments.

17:03Alex Shevchenko:Did you confirm that with them or is that just a guess because it worked so well?

17:06Max Agency Host:Yeah, it's a guess because it works so well, but it does work very well.

17:10Alex Shevchenko:You mentioned an agent SDK. What agent SDK did you use?

17:13Max Agency Host:So it was like OpenAI agent SDK that we started out on. But it was kind of like this ship of DCS thing almost as like we went further and further where we had to like rip out pieces one piece at a time and just like customize it. Like it was very, very overfit to the task that we ended up like landing on. so it's kind of like this this frankenstein harness of like we started out thinking that it was going to be like relatively simple but then as as the progress like as the project grew

17:42Alex Shevchenko:had to like replace a bunch of stuff question about the ux uh how do people interact with

17:46Max Agency Host:ramp sheets is it chat based yeah so you have a spreadsheet editor on the left left hand um you can basically do anything you would in excel write formulas etc in there the one thing that's agentic in there is like you can select for example a range of cells and like add to context that specific range and then there's a chat interface on the right that general like experience is just like you prompt it and you ask questions for for things to be done there's like some templates that are customized for specific tasks um that someone will do like reconciliation there's uh some some like nice cities in the chat where we'll like ask you follow-up questions or try to plan where it's going to be like a slightly generative interface, which is like a form to fill out to answer the questions.

18:37Max Agency Host:But for the most part, it's just like a chat interface.

18:40Alex Shevchenko:How long does the agent take to run? That depends on the task at hand.

18:44Max Agency Host:We have like a fast and expert mode. And the user toggles that? Yeah, the user decides. And yeah, it depends on the complexity of the task. If you give it like some very complex task in FP &A of like get data from the web around like SSE filings or something aggregated together and produce a model for me, it can run for like 30, 40, 50 minutes. For simpler things, if you have like the process already defined, if it doesn't need to like go and get out and get like information from anywhere, it can be very quick. Like it can be 20, 30 seconds.

19:20Alex Shevchenko:Do you have any sense of what the distribution is? Like what types of questions are people asking are they asking the complicated ones or the simple ones

19:26Max Agency Host:or somewhere in the middle it's a it's a mix of everything i think the average session is probably

19:31Alex Shevchenko:around like 10 ish minutes and is a session like does that include multiple back and fours like i could kick off the agent multiple times within a session or is a session like one run of the agent

19:42Max Agency Host:from start to finish what one conversation tends to be around like 10 minutes um but it's dominated kind of by like the question answering piece of it, where out of the 10 minutes on average, I think it would be like seven minutes of like just the agent grinding on producing an output for you.

19:59Alex Shevchenko:There's this architecture kind of like question or debate of like agent in a sandbox or agent like with a sandbox as a tool. It sounds like, actually, yeah, which one did you guys go for and why?

20:11Max Agency Host:Ours is the agent is outside of the sandbox. I don't know. architecturally, I don't know if like our answer is even the correct one. I think that's why everyone's still debating around it. I think for us, it was like, often we will like just spin up the sandbox again. So to constrain, like if there is a cogen step, we don't want it to be polluted in any way from the previous one. And having the agent outside of the sandbox thing, I mean, you can, it's kind of like this gradient of like, you can probably do it in any way you want but the one that we landed on is like the agent just keeps track of the state outside of the sandbox and then if we need to like the cogen does not get polluted by previous runs because it gets spun up and the sandbox like applies the coding steps uh the cogent code onto the like worksheet and then just goes back and doesn't get polluted every time and you've got a sandbox

21:09Alex Shevchenko:per thread, basic per session?

21:12Max Agency Host:So like per kind of like agentic action. Oh, okay. So even more granular. I don't know if that's like the correct term. Like you have a conversation, which I guess I would call a session. A user tells it to do something and that will have its sandbox based on like the input received from the user. And then it will grind out on producing some output. And then that output goes out of the sandbox. And then if the user tells it to do something again,

21:38Alex Shevchenko:that will like get another sandbox a different sandbox yeah okay but for multiple tool calls within the same grinding it out those all use the same sandbox yeah how do you think about testing just like gaining confidence that this thing actually like works yeah i mean we have

21:57Max Agency Host:some like some label data sets from like experts we like a lot of it you could call it kind of call it like vibe evaling to some degree but it's like very educated vibe evaling because we had like input documents from the product like someone doing a month on close task for example from a previous month and we know that what it looks like so it was like kind of like the the vibe eval was in the sense of like a human would or like an engineer would go in drop in that file see what it outputs and compare it to like the large spreadsheet um without like using an lm as a judge or anything along those lines.

22:34Alex Shevchenko:Could you compare them exactly and like these need to be exactly the same and if they're at all wrong, then this is wrong or is there some gray area?

22:41Max Agency Host:That's where the complexity of it is and that's why like it was mostly done by an engineer that kind of like tries to understand like this is the way that we do like an AR reconciliation. You are able to evaluate it with an LM as a judge, but the amount of effort to produce like the serious eval of like this is the kind of like the data set this is the task and this is the rubric becomes like pretty extensive and pretty expensive to produce and so you can just kind of like use the human's judgment instead uh like if they have like a good understanding the end of the day like an AR reconciliation is just like you go through line item by line item and you try to find a discrepancy and then try to explain the discrepancy but like evaluating that like automated with an LLM as a judge with like a rubric that becomes like much more work to build out.

23:32Alex Shevchenko:One of the very cool related things to Sheets was this self-improvement, self-monitoring kind of loop that you guys wrote about. Would love to hear more about that.

23:41Max Agency Host:Yeah. So this actually leverages like this internal coding agent that another team at RAM has been developing called Inspect. And so Inspect is like this large coding agent that is plugged into all of our systems, all of kind of like the MCPs pre-configured for you, snapshots, kind of like all the environments like RAM sheets so you can spin it up in like a couple of seconds and have the back end and front end running inside of it. And because it's like so well customized to how we do coding at RAMP, there's some niceties around making automations on Inspect. And so one of the experiments that was very successful was one of our engineers, Alex Levinson, has created this like self-monitoring loop where he set up inspect to run an action whenever there's a new PR that gets created.

24:35Max Agency Host:Also, like as a cron job at night to just instrument ramp sheets with extra data dog monitors so that we get like more signal of like things that break. And it's like a pretty large workflow and viewers at home can go and read the blog posts around all the intricacies of it. But yeah, basically either when a PR gets created or on the cron job, it will go and it will like look at the code base and try to see if like there's a monitor that is missing or there's like a metric that we should be adding. And then it goes into kind of like this shadow mode where the monitor alerts won't like bug anyone yet.

25:17Max Agency Host:But then there's like another agent that will review them. And if they're like too noisy, then they get pruned because there's not enough signal. But if they run for a while and they're not noisy and the agent deems them to be like good enough, then they get like kind of promoted to start bugging the engineers and also like kind of like slacks us the results.

25:38Alex Shevchenko:How do you know if the signals, I guess, initially are like noisy or not? Like you could set up new monitors and there could be issues and that might not be noise. Like how do you determine whether that's?

25:49Max Agency Host:Yeah, I mean, the handmade monitors that an engineer would go and create, those are obviously like they have like the human judgment in them. So those will just like alert right away. but the ones that are automated they like in order for us not to get too much noise because like the v0 of this was extremely annoying and we just like spam our our channel that where we like do a lot of work or like no this is too noisy like create a new channel and like set up some filter and the blog post does go into detail around how like that filter was created to try to well minimize the amount of noise and there's just like a couple of like heuristics around like the types of things of like if the like it's like a latency thing and like you get alerted for the latency but it's like something that you really don't care about or it's like something that is like the p90 or the p95 is like too large but in the grand scheme of things it's like expected sla for end-to-end interactions you're like well that's probably noise and opus is actually like pretty good at just like deciding what is noise and what isn't.

26:54Alex Shevchenko:And what is the end result of this? Is it kind of like a ping in Slack or does it open up a PR as well? Does it try to like fix the issue?

27:01Max Agency Host:It's both. So it will ping us on Slack. There's like this channel where there's a bunch of these alerts and then, yeah, it will do like a first pass at trying to create a GitHub PR to try to fix it. And they do get merged in quite a bit, especially when they're like on the simpler end. Obviously, it's not going to do like, oh, your architecture has this misstep and it's going to create something large. Those tend to not be mergeable, but it's very good for alerting and for fixing very small bugs that people run into.

27:32Alex Shevchenko:Have you seen this applied to any other teams or projects at Ramp, or is this unique for Sheets?

27:37Max Agency Host:So Sheets is our kind of playground for testing out these wacky ideas, but now Alex Levinson has been spearheading this effort. He's productionizing and platformizing this so that it can be reused by other projects. And he's like running another pilot pilot on an internal product that is like zero to one that another team is building right now.

28:02Alex Shevchenko:What other wacky ideas have you run on Sheets or experimented on Sheets, whether they've seen the light of day or not?

28:09Max Agency Host:Yeah, I mean, just like from experimenting, I feel like like Sheets is like this ship of Theseus thing that we keep like rebuilding. A lot of like context management of like how much data should you be giving it? Like we looked at like foveating what it sees. So like instead of like it having like a small five by five like cells that we show it, it's like it decides on like the range that it should be reading as an argument, kind of like very specific things like that. I don't even know. There's like just so many small experiments that go into this and a lot of them get discarded over time. there's like another one in in kind of like memory management where we were like doing embedding based memory to try to help it out there's memories that just like a naive tool call for memory creation kind of like chat GPT does it like remember I want to always have like the first row be blank or the left like left column be blank for comments and it's like a tool call that would just go create a memory that gets passed in I was like very similar to how chat GPT does it hasn't been like super super useful for some reason um so like looking through like user interactions hasn't been like bringing any like super helpful interactions and under the hood

29:27Alex Shevchenko:for that memory stuff was is it just a list of strings that is memory or is is it some other

29:33Max Agency Host:data structure for that specific like function call it was yeah it's like a string so it has like a function call basically of like if a user like explicitly says like this is how i do uh I like to keep my left column blank for comments. It will just go and decide to make that tool call. That tool call just populates a memory bank and that memory bank gets injected with context for every future agent run.

29:59Alex Shevchenko:How does that memory bank get updated? Like if I say, you know, keep the first row blank and then the next day I'm like, actually I want it to be two rows blank. What happens then?

30:07Max Agency Host:Well, that one is like relatively simple. So like if it does get triggered in that specific interaction, it should like realize that it should go and update and that function call will go and just modify the kind of like scratch pad that it has for it but to your point it does run into like kind of like a lot of issues especially for more complex flows of like because it's like such a rudimentary memory system if you want it to be like oh when i am preparing my dcf this is like the specific ways that i want to be linking cells and then they come back and they like add more like to that it because of how rudimentary it is it's not necessarily gonna know that it should go and like modify um that specific dcf memory in an appropriate way i'm sure that there's like ways to improve it but also like from user interactions it doesn't seem to be like just like a flow that makes that much sense where people like want to have it remember that much stuff.

31:08Alex Shevchenko:You guys have been doing a bunch of other experiments as well at, at RAMP labs. Uh, what are some of the recent ones that you're most excited about?

31:16Max Agency Host:Yeah, there's two, uh, that we published like last week and the week before that were really exciting. Uh, the one from last week was called latent briefing. Um, and this is kind of like a way for doing, um, we're very excited around like RLMs and like, kind of like more intricate ways of doing context management. And so this was an experiment of like taking RLMs and giving like a way to do like KVCache communication between the subagents. And so there's a very nice write-up by Ben Geist that we published last week that got some very good Twitter attention, basically trying to reduce the amount of tokens that get used by the RLM by having instead of them communicate in token space of like, this is the output, I'm going to take this piece of text and send it over.

32:05Max Agency Host:They just take the KV cache and there's an interesting kernel trick that is used to reduce it and pass it over directly. Kind of like KV cache communication. KV cache to KV cache.

32:16Alex Shevchenko:Can you do that with closed source models or does this have to be done with open-weight models?

32:21Max Agency Host:This needs to be open-weight models that you have access to. So the actual orchestrator in the RLM that was used is closed source. So it was like an anthropic cloud family orchestrator. but then the worker agents were open-sourced.

32:39Alex Shevchenko:Interesting. I'd love to understand. Okay, so my basic understanding of RLMs, and this is great because you can correct me. So you've got the top-level agent, and basically it's got a REPL-like environment and a few functions in there and a function for kicking off a sub-agent, basically. And it's prompted to, and I think some of the, and it can inspect some of the variables, but like it's mostly prompted to use the sub-agents to break things down and calls them programmatically, which is one of the big benefits, I think. So rather than using the subagent tool like Cloud Code has to call 500 subagents in kind of like the tool calling way, it will just write a script that will run over 500 segments or something like that.

33:21Alex Shevchenko:And so you're saying the different subagents would communicate in the latent space. So how exactly does that work? Because in my mental model, I say, I don't know, foo equals subagent and then pass in some input text and then i guess foo is a return value and then like what happens if the model tries to inspect foo like is is that in the latent space because it can call like print foo or something like that like can it do that or and then i guess like you're saying like if if in the next time i'm like okay bar equals subagent foo and foo is like the input like well that is that where it takes over kind of like some of the latent space and Like, what if it, like, appends a string to foo or something like that?

34:03Max Agency Host:I mean, for those specific, like, examples in the REPL, it's kind of pre-populating the next, like, subagent with relevant context from the previous subagent by getting its, like, KV cache in latent space.

34:20Alex Shevchenko:do you always want that like if i if i have like a book of like a hundred pages and it's like find the the tenth word on every page or something like that my understanding of why rlms are great is because they'll split it up and they'll run each page and like you actually want that isolated context and you want to start from scratch basically how would that work here would it

34:41Max Agency Host:would it always pass it in or do you yeah well for splitting stuff that is just like you want it to be very parallelizable. I guess this is not something where it makes sense. It's more about like you split out the work and you have like a piece of work that needs to have had some progress on it. And then you need to pass over the learnings from that piece of work over to another sub-agent. That's where it kind of like makes more work. From like the write-up, you can actually see that like on the kind of like benchmark that it does like for the same level of accuracy, it produces token count you can kind of like use it in the other direction as well for like if you keep the token count at the same level there's like an accuracy gain that you can benefit from it okay so that's the first thing that you guys launched last week or one of the things you guys launched last week what's the other one yeah the one from the week before was like this funner experiment around like we wanted to revive golden gate quad uh which like two years ago anthropic published golden gate quad which was like this using steering vectors they took like a sonnet model and they made it obsessed with the golden gate bridge to the point where like a lot of interactions were like you ask it like i have 10 bucks how should i spend them and it would answer like oh you should take the golden gate bridge twice and pay the toll so like it couldn't get the concept out of its head um and we kind of like just wanted to revive that and we're like oh maybe we just like do it around the concept of ramp um and then we're like oh what if we just like let the users decide what concept it should be obsessed over.

36:11Max Agency Host:And so that's kind of like what we built out. So using steering vectors and like synthetically constructed contrastive pairs, the user would pass in like just as one text box, what concept they want the LM to be obsessed over. We would take that concept, generate like 80 contrastive pairs synthetically, and then get the steering vector out of them, and then apply it to a Gemma model on specific layers to try to get it to be obsessed over that concept without degrading the performance too much. And the interactions were pretty fun, like very similar to the Golden Gate Claude interactions where it's like you give it the concept of a Jeep car or something and you're like, my girlfriend left me and it would go and be like, this is a very touching and like heavy topic but before that let me tell you how capable a vehicle the jeep grand trochee is or whatever um so just kind of like this goofy experiment but also like it lets us to experiment with interpretability a bit um so interpretability is not necessarily super super like actionable for capabilities yet but it's always nice to to get like this black box lm that we always treat as a black box and kind of like peer behind the curtain at least somewhat and get some sense of like okay at these layers in the gemma model this is like responsible for like the core decisions and this is the responsible for like the stylistic output and here where broadly speaking where the concepts are stored what exactly are contrastive pairs and how do they generate the steering vector the idea is like you take a specific concept and you take two strings of text that one of them will have that concept and then the other one doesn't have that concept and you pass them through the LLM and look at what the layer activations look like and then just subtract the layer activations and so then you have like one data point and then if you like average over like 80 contrastive pairs it gives you a pretty good like rough steering vector around that concept.

38:22Max Agency Host:So it's not going to be like an exact perfect like replication of that like Jeep steering vector, but it's going to be roughly broadly aligned to it.

38:32Alex Shevchenko:And what's the shape of that steering vector? Is it the shape of the weights of the LLM itself, or do you narrow in on a subset of them?

38:40Max Agency Host:Yeah, so we had to do quite a bit of work to figure out which layers we want to like put it on. So the first version of this was on a Quen model, and the Quen model, one of the smaller ones, was like working pretty well, but normalizing the magnitude of the steering vector was like pretty hard and it would like degrade into speaking Mandarin because it had a lot of Mandarin in its retraining corpus. And so just from switching from Quen to Gemma, we actually like had that entire problem of switching to Mandarin like completely evaporate. But then Gemma was like a lot harder because it was a lot more kind of like sensitive to being steered from the architecture of the LM itself.

39:25Max Agency Host:And we had to basically do a bunch of sweeps of applying those steering vectors at different like magnitudes and at different layers to try to see what like text would degenerate or wouldn't. And we would just pass it over to like Opus as part of that like evaluation process to try to figure out, okay, at these layers it does end up working pretty well and at these layers it creates degenerate output.

39:48Alex Shevchenko:What was the final result? Did it end up being like 10 % of layers that you applied it to? 90 %?

39:53Max Agency Host:I think it's just like five layers, middle-ish, that were applied to. So like go check out the blog post and write up for it for the specific places where it was applied. But yeah, the Gemma model ended up being like a lot more sensitive than the Quinn model that you can just like steer kind of like anywhere with like less regard to like the magnitude.

40:16Alex Shevchenko:How exactly does it get applied? Does it replace the weights? Is it like in addition to some of the weights?

40:23Max Agency Host:For kind of like each pass through the LLM, you have like the layers that get activated and we just like slap it on top at inference time. And so for each next token that gets predicted, you just have that steering vector that nudges it. You slap it on top like addition, basically. Yeah, at inference time. So it's not like added in. was actually like one of the problems that we needed to solve when we were serving this. You had to retrieve the steering vector and apply it to the model at inference time. But yeah, that's what's kind of like, because it's at every token that gets produced, it's also what like creates some of these interesting interactions.

41:00Max Agency Host:Sometimes like it would almost become self-aware around like its interactions or like it outputs something that is like irrelevant to the question and then be like, wait, why am I talking about like spaghetti when the question is about like the meaning of life, let's get back to the meaning of life. The meaning of life is tomato sauce. And it's gonna, like sometimes it would get into these loops of like self-awareness. It would produce the next token and the path, like the kind of like the circuit that it would take in the LLM was not responsible for like spaghetti. It was responsible for answering the meaning of life, but then it would like degrade into the wrong direction and then kind of become self-aware about it.

41:38Alex Shevchenko:You guys do a lot of really cool experiments at Ramp Labs. What does the team makeup look like? Like who is a good fit for Ramp Labs? What do you look for when hiring?

41:47Max Agency Host:There's no like perfect persona. We try to look at just like people that have like very large spikes around things that don't necessarily fit for the rest of Applied AI and people that are interested in like these specific things. So like interpretability or like we have one guy that was doing an RL startup previously and so he's interested in doing a bunch of RL experiments within Ramp Labs. And so it's kind of like this.

42:10Alex Shevchenko:What is the rest of applied AI look like in terms of backgrounds and focuses?

42:15Max Agency Host:Very, very strong AI engineers, very, very strong ML engineers. What is an AI engineer that didn't exist until three years ago? I mean, I feel like it's like a mix of everything. I feel like we're like as an industry now, like starting to narrow down and starting to like understand what it looks like more and more. But like people like honestly, the thing that we look for the most in general is just like high level of agency and like high rate of like learning of like having a very very good slope of like I might not have experience around this specific thing right now but I will pick it up very quickly um that's like kind of like the the most abstract like um persona that I can give there but yeah like tend to to be someone that has done some like LLM related projects that were in prod and used by a bunch of people people that have started thinking more seriously around evals and environment building because that's still like a very complex uh topic that like i mean as an industry we still don't like fully grasp around like what's the best way is it like you always build out a very complex world with very complex like very exhaustive list of tasks with which for each task you have like a very exhaustive list of rubrics or do you just go and like kind of like vibe eval your way out and it's like both like for some things it works out well for some things it doesn't work out well.

43:37Max Agency Host:And people who have started to develop like this intuition around like which types of project or which types of product needs this exhaustive approach and which ones don't. It's kind of the intuition that you want to capture as well.

43:50Alex Shevchenko:Thanks for listening to Max Agency. If you liked this episode, leave a review and subscribe. Send feedback or questions to maxagency at langchain.dev. We want to hear from you.

44:05Thank you.

From the publisher

Alexander Shevchenko is the head of applied research at Ramp, where he leads Ramp Labs – the team behind Ramp Sheets and a steady stream of public AI engineering experiments. Ramp Sheets started as an internal process mining tool that turned Loom videos of accountants into Markov diagrams, before evolving into the agentic spreadsheet editor that shipped in November. In this conversation, Alex walks through the architecture under the hood, why Ramp biases the agent toward Excel formulas over Python code gen, and two recent Labs experiments: Latent Briefing and a user-steerable revival of Golden Gate Claude.


We also discuss:

  • Under the hood of Ramp Sheets
  • Inspect, Ramp's internal coding agent, and the self-improving monitor loop it powers
  • Why finance professionals rejected code gen as too "black box"
  • Why Anthropic models tend to excel at agentic spreadsheet manipulation
  • The case for putting the agent outside the sandbox, not inside it
  • The Loom-to-Markov-diagram process mining pipeline
  • RLMs and how subagents can share memory in latent space
  • Latent Briefing and KV-cache communication between subagents
  • Reviving Golden Gate Claude with steering vectors on Gemma


Referenced:


Where to find Alex:


Where to find Harrison:


Where to find LangChain:


Send feedback or questions to maxagency@langchain.dev


Timestamps:

(00:00) Introduction

(01:13) The origin of Ramp Sheets

(02:27) The Loom-to-Markov-diagram process mining pipeline

(04:28) Why code gen approaches felt too "black box" to finance

(06:13) Meeting finance where they already are: inside the spreadsheet

(09:08) How far process mining got them

(10:31 )Text descriptions and Graphviz DAGs as output

(12:41) Under the hood of Ramp Sheets

(14:52) Why the agent uses Python only as an escape hatch

(15:47) Why Anthropic models excel at agentic spreadsheet manipulation

(17:12) Frankensteining the OpenAI Agents SDK

(17:43) The Ramp Sheets UX and fast vs. expert mode

(19:58) Agent in a sandbox vs. agent with a sandbox

(21:55) Vibe evals with expert humans

(23:40) Inspect, the internal coding agent

(24:13) The self-monitoring loop and auto-PRs

(28:01) Other wacky experiments on Sheets

(28:43) Memory experiments that didn't pan out

(31:16) Latent Briefing and KV-cache subagent communication

(35:13) Reviving Golden Gate Claude

(37:47) Contrastive pairs and steering vectors

(39:47) Picking the right layers in Gemma

(41:37) What Ramp Labs looks for when hiring

More from Max Agency

All 11 episodes
How Ramp built an AI agent that can think outside of tokensMax Agency · 44 min
Listen in VO