In short
Benchling AI’s agent design for life-sciences R&D, emphasizing tool/context design, SQL-first data access, production-trace evals, and multi-model workflows to improve answer quality and speed.
Guest
Nick Larus-Stone, head of AI at Benchling (Benchling/BenchLink platform). Benchling has recorded life-science experiments for 14 years; Benchling AI launched ~6 months ago.
Key claims
Agents can speed drug discovery “from initial discovery to bringing a drug to patients by 2x” by reducing lab-to-analysis delays and improving experiment planning (e.g., DOE). Quality rises when models are grounded in the right Benchling data. Benchling uses SQL heavily (embedding table names/descriptions) rather than relying on generic coding-agent harnesses. Evals are hard for novel science, so Benchling leans on production traces plus user feedback.
Notable examples
IND report writing (thousands of pages; AI drafts in 15–20 minutes; humans verify before FDA submission). Data-entry agents that cross-compare multiple model providers; disagreement flags errors. Background/event agents triggered by instrument uploads to run analysis automatically.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Benchling AI
1:14 to 2:06
Discover what Benchling AI is and how it enhances data management for scientists.
“Benchlane has been a company for 14 years.”
User Interaction with Benchling AI
2:06 to 3:09
Explore how scientists interact with Benchling AI through a chat interface.
“So it helps them find data, design experiments, analyze data, kind of accelerate each part of the scientific journey.”
Task Complexity in AI
3:09 to 4:35
Understand the range of tasks Benchling AI can handle and their complexities.
“You get something back in five seconds or is it, you know, kick something off 15 minutes later?”
Data Advantage and Model Quality
4:35 to 5:48
Learn how having a decade of data impacts the quality of AI-generated answers.
“I would say that's one of our core advantages, actually.”
Agent Architecture and Improvements
5:48 to 7:37
Find out how Benchling's platform is adapted to improve its AI agents.
“We think kind of having that data up front can both speed up agents to get you the answer faster, but then also just give you much higher quality answers.”
Building a Custom Agent Harness
7:37 to 10:01
Delve into the unique design of Benchling's agent harness for scientific tasks.
“I imagine search is pretty important for a lot of these research style agents.”
Multi-Model Approach in AI
10:01 to 12:28
Explore how multiple AI models are utilized to enhance data accuracy in science.
“You could probably, you know, force a coding style harness into that.”
Challenges in Verifying AI Outputs
12:28 to 14:00
Understand the difficulties in verifying AI outputs in scientific research.
“Because what we saw was if two models kind of disagree, there's usually an error.”
Model Selection and Verification Challenges
14:00 to 17:03
Discussion on the selection of AI models for specific tasks and the challenges of verifying their outputs.
“but like we're going to do exactly what you said earlier.”
Evaluating AI Performance in Science
17:03 to 19:36
Exploration of how Benchling evaluates AI performance and the importance of production traffic analysis.
“At the end of the day, we talk to users and we look at kind of production traffic because it's pretty hard to build evals for kind of novel scientific discovery.”
Show all 24 chapters
Tool Design and Context Engineering
19:36 to 22:46
Insights into tool design and context engineering for effective AI agent performance.
“think that's necessarily the right kind of like solution for the space.”
User Experience and Skills Integration
22:46 to 25:59
How Benchling integrates user skills into the AI experience and the importance of user education.
“And are these skills right now at a user level, at an org level, at some, the user can choose whether to share it?”
Limitations of Coding Agents in Science
25:59 to 28:00
Discussion on the limitations of coding agents and the differences in application for scientific discovery.
“And so we spend a lot of time working with our users, showing them what they can do, listening to their problems, making sure we're building the right things there.”
Challenges in Adapting Coding Agents to Biology
28:00 to 29:00
Learn about the limitations and adaptations needed when applying coding agent principles to life sciences.
“It seems like a lot of the architecture is similar in a lot of ways to coding agents and how that's kind of like brought into this new domain.”
Importance of Life Sciences in AI Development
29:00 to 30:30
Explore why life sciences are becoming a focal point for AI applications and research.
“I think there's certain tasks that maybe are more kind of, as you're saying, verifiable in silicon.”
Understanding LLMs Through Biological Metaphors
30:30 to 31:36
Discover how working with LLMs resembles scientific experimentation and understanding.
“So I think it's a combination of both of those things.”
The Role of Agents in Drug Discovery
31:36 to 34:06
Examine how agents are revolutionizing the drug discovery process and their potential impact.
“They're kind of like poking the system that they don't fully understand and watching how it responds.”
Barriers to AI-Driven Drug Discovery
34:06 to 38:10
Analyze the challenges AI faces in drug discovery, including model limitations and experimental systems.
“Maybe historically, like you would actually take a flash drive, plug it into the computer, right?”
Future UX for Human-Agent Collaboration
38:10 to 41:20
Learn about the ideal user experience for integrating agents into scientific workflows.
“I think you're going to have humans and agents working kind of in concert to push that forward for quite some time.”
Monetization Strategies for AI Usage
41:20 to 42:00
Explore how Benchling AI approaches pricing and usage models for scientific applications.
“But the other thing I would say is actually, I think most scientists really like the part of their job where they're thinking about hard scientific problems.”
Usage-Based Model for Benchling AI
42:00 to 43:35
Learn about the usage-based pricing model for Benchling AI and how it accommodates scientific organizations.
“You can kind of think of the agents like that today.”
Agents and Underlying Models
43:35 to 45:28
Explore the structure of Benchling AI's agents and how they handle various tasks through unified interfaces.
“and we'll build evals around that and we'll monitor production traffic to make sure that that's true.”
Future of AI in Scientific Research
45:28 to 47:28
Discuss the future of AI harnesses in scientific research and the potential for novel idea generation.
“A lot of this we think, you know, smarter models will actually get better at science.”
Intelligence Constraints in Biology
47:28 to 49:26
Understand the intelligence constraints faced by AI in the biology domain and the challenges of model training.
“I'd say like, we talk a lot about building an AI scientist as kind of the overall goal here.”
Transcript
Automatic transcript. May contain errors.0:00We think actually agents will be able to like speed up the time from kind of initial discovery to bringing a drug to patients by 2x.
0:08Nick Larus-Stone:Today I'm talking to Nick Larus-Stone, who is head of AI at BenchLink, the platform that life sciences teams have recorded experiments on for 14 years. Now they're building agents on top of it. What we've seen is when you kind of put these models on top of the right data, the quality answers go way up. He breaks down how Benchling's agents for scientists are unlike typical coding agents. We actually still rely very heavily on SQL, right? We do some tricks around embedding table names and descriptions so that we can have kind of like faster paths to get them to query the right thing. And he explains why and how, when it comes to evals, Benchling leans on production traces.
0:52The types of questions that people ask are so specific to their science, it is very hard for us to build evals. We spend a lot of time actually looking at traces in production.
1:02Nick Larus-Stone:Throughout, Nick shares some unique perspectives on agents that may challenge your assumptions. Some people say like LMs can't do anything novel, right? Their next token predictors. I think people just like haven't experimented enough with this. Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you. Benchlane has been a company for 14 years. You guys launched Benchlane AI six months ago. What is Benchling AI and how does it fit into the overall platform that you guys have been building? Yeah, so Benchling is a platform for life science R &D organizations to kind of store and manage their data.
1:40So this is everything from the experiments that they're running to the actual kind of like physical things that they're building in the lab to helping automate their instruments to helping do kind of data analysis on top of that. And so the core of Benchling is this platform to help manage data. And Benchling AI is the layer on top of that to bring intelligence to scientists who want to interact with their data. So it helps them find data, design experiments, analyze data, kind of accelerate each part of the scientific journey. What does it look like? How do I use it? Yeah, so it's a chat interface in Benchling.
2:19So scientists log in, and then we have an agent kind of backing this chat interface that can interact with objects within Benchling. So, for example, let's imagine you're a scientist who is kind of working in a normal drug discovery company and you've run some experiments. The main question on your mind is, what should I do next? You log into Benchling, you can type in your question. It might be, hey, has anyone done a similar experiment to me in the past? It could be, hey, help me design this kind of next experiment. It could be, hey, help me analyze the data from my last experiment. And kind of our agent will go kind of do whatever the task that you've asked for and then return the results in an interpretable way and kind of seamlessly integrated.
3:03So you can kind of save it back into the platform and kind of continue on with your job.
3:08Nick Larus-Stone:How long horizon are these tasks supposed to be? Is it like question answer? You get something back in five seconds or is it, you know, kick something off 15 minutes later? it ends. Yeah, it's a range. It depends on what you're looking for. Sometimes it's simple, just data retrieval, right? Go find something and kind of give me that answer really quickly. Sometimes they're much harder tasks that take, yeah, tens of minutes. A good example of this is around report writing. A lot of what scientists do is kind of generate data in order to, you know, try and get to the truth of the underlying biology.
3:43But eventually, you need to kind of prove that to someone else. That could be kind of a partner organization that you're working with. It could be internal to your manager. It could be a regulatory agency like the FDA. In order to do that, you need to kind of consolidate kind of large amounts of data that you've maybe gathered over years and produce it in a often standardized format. So for example, there's something called an IND, an Investigational New Drug Application that you submit to the FDA. It's very standardized. it's usually thousands of pages you know compiling all of that historically takes months AI on top of the right data crucially the right data can you know give you parts of that within as you're saying 15-20 minutes so it really depends on the task but we see a wide range of difficulty in the
4:35Nick Larus-Stone:tasks that people are asking about you mentioned data like how big of an advantage is it that you have had this data platform for a decade plus and are now starting to build agents on top of it, as opposed to every other agent startup, which I feel like has come out in the past two years and is starting from the agent and then probably trying to get data in after the fact. I would say that's one of our core advantages, actually. So I spend a lot of time talking to scientists and, you know, asking them about their opinions on AI, trying to get them to use AI. And a lot of scientists tried ChatGPT when it came out.
5:06They asked about their specific science, right? and they got pretty bad answers. Maybe it was hallucinating. Maybe it was just like not relevant or maybe it was kind of like ignorant answers. What we've seen is when you kind of put these models on top of the right data, the quality answers go way up, right? And because scientists are like focused often on a very narrow area, but somewhere in the pre-training data. But if you can contextualize that and kind of put the right data in front of the agent, it's going to do much better than a generic agent, even with tools to query outside systems, right?
5:41So like MCP, obviously, you know, it's very powerful, but there's some limitations there. We think kind of having that data up front can both speed up agents to get you the answer faster, but then also just give you much higher quality answers.
5:56Nick Larus-Stone:Interesting. This maybe gets to some of the agent architecture in the hood, but like, how does it speed that up? What is the difference between just giving it via MCP to some other agent? I mean, part of it is just, you know, we're building our tools with a knowledge of our architecture such that, you know, the agent can make best use of that. But part of it is MCP is just kind of a limited view into other systems, right? So if you're taking like a very naive view, you only get what those kind of third party systems allow you to get. That's often not all of the tools or all of the data that you would need in order to do well.
6:28So kind of by building it from the inside out, we have the advantage of being able to access any kind of relevant piece of data and also being able to change the underlying system such that we can make it better and we can build new tools on top of the system we actually change. So I think like MCP as a protocol makes a lot of sense and it's great. And we are both an MCP server and an MCP client. But there's some limitations versus what you can do kind of like inside the agent.
6:55Nick Larus-Stone:Have you found that you've adapted the platform as you build your agent to change the platform in some way to make it better for the agent? Yeah, I mean, maybe a simple example of this is, you know, pre-AI, we have a standard search platform. But obviously with the, you know, AI coming out, we've done a bunch of things to kind of change our underlying search platform to make it better for agents. We've kind of changed our chunking strategy. We've changed some of our storage on the back end. obviously like embedding search is very important. So there's kind of like underlying search platform improvements, which also go to the humans, right?
7:30They actually get better search a lot of times. But we've kind of made those underlying platform changes because we know that agents will be able to take advantage of it.
7:38Nick Larus-Stone:I imagine search is pretty important for a lot of these research style agents. How much have you focused on improvements to the search versus improvements to the agent? And has that changed over the past six months? Or I feel like there's a big trend recently in like agentic search and like really changing the agent's architecture for search as opposed to making the search go from like 97 to 98 or something like that. But it sounds like you have made changes. So I'm curious how you've approached that. One kind of benefit of us being 14 years old is we have lots of engineering teams working on lots of things, not just agents, right?
8:08So our team very much focuses on the agents, but we work very closely with our search team. So we have a dedicated search team, right? Who's managing our search. And so that lets us both focus on improvements to the agents as well as working with other teams who are actually improving the search directly.
8:24Nick Larus-Stone:Okay, let's talk about the agents. What does it look like under the hood? Honestly, pretty standard agent loop, right? LLM calling tools. I think a lot of the interesting aspects of it are how we build the tools, the skills that we put in, and some of the context management. But at the end of the day, right, it's kind of just LLM in a loop with a bunch of tools and a kind of coding sandbox attached. Cool. Okay, you mentioned a few things. You mentioned skills. Is this skills in the same sense as like coding skills? Like, do you use that same paradigm? Yeah, so we conform to the agent skill spec, and we both support users bringing their own skills, as well as us kind of building what we call kind of like built in or native skills to the agent.
9:07Nick Larus-Stone:Do you use a coding agent harness under the hood then? No, we've built our own harness in part because I think coding agent harnesses are not the right solution for science. We haven't done as much kind of experimentation here as I'd like. We're starting to do a bit more here. But I think there's a lot of problems in science. And a lot of the questions that our users are asking are not good kind of like one-to-one comparisons to what coding agents are good at. For example, Google last year came out with their co-scientist paper where they had like a radically different harness where, you know, in order to come up with more novel hypotheses, they would have this like tournament style play of LLMs kind of proposing ideas and kind of having a greater look at two different proposals and then choose the best one and doing that iteratively, which is very different from like a coding style harness.
10:01You could probably, you know, force a coding style harness into that. But I think that's kind of an example in my mind of like, oh, actually, this is like a specific scientific task that looks very different from, hey, I have a spec and I want to like write a bunch of code against it. I'm of the belief that there is still value in having kind of your own harness and not being kind of wedded to the decisions that make something really good at writing code. I would say writing code is an important part of this, but a lot of the kind of queries that we get involve like writing no code, actually.
10:36Nick Larus-Stone:And so you mentioned code sandboxes, and so it sounds like your agent does write code. You're just saying it does things beyond writing code, and that is what requires the other ones. Yeah, and writing code is very helpful, right? It's very helpful to manipulate files. It's helpful to do data analysis, which obviously is kind of an important task in science. But that is only some fraction of what people are asking for. And so I see it more as, you know, coding is an important ability for this agent, but it's not the driving force of this agent. You mentioned like tournament style things. Do you guys do that?
11:10Nick Larus-Stone:Do you have like multi-agent architectures in play already? We do a few different things. So one kind of example that we are quite interested in is how to take unstructured data and turn it into structured data. So a lot of scientific data, unfortunately, comes in the form of PDFs, Excel spreadsheets, files from instruments that are not well structured. And, you know, a big kind of advantage of Benchling as a platform is we help companies define their data model and we give them this place to put structured data. Unfortunately, the disconnect is it historically has taken a lot of work to go from, you know, maybe a 10 or 100 page PDF and kind of like enter that data into a kind of structured table.
11:56So we obviously look to AI to help with this kind of bringing in the right data is really important. Like if you get that wrong, everything downstream is kind of like irrelevant. And so the companies we work with are very sensitive to, you know, having high quality data import. And so actually one of the first things we ever built as a company on the AI front was this kind of data entry agent that uses multiple models under the hood, where we give the same problem to multiple different model families, and then we cross compare the results. Because what we saw was if two models kind of disagree, there's usually an error.
12:35Maybe it's sometimes one of them was right and one of them made an error, but oftentimes they both get it wrong or there's something that a human should look at. And if they both agree, that usually is good enough data quality. And so we've kind of explored that specifically in this space, but we started exploring it for kind of non-data entry tasks as well. Maybe harder scientific questions where, you know, sending the same question to multiple different model families and cross comparing the answers can give you a higher quality result than just asking kind of the either one model or one model multiple times.
13:09Nick Larus-Stone:You mentioned like model families. Are you basically saying that like Haiku and Sonnet and Opus, they all tend to make the same types of mistakes? Yeah, I would say when I say model families, I mean different model providers, not just different size models within the providers. I think of this as kind of like spiky intelligence, right? Like Claude has different advantages to GPT, has different advantages to Gemini, right? And so each of them will make slightly different errors, whether it's Opus versus Sonnet Haiku. Obviously, there's different quality within that. But actually being able to ask different model providers, we found gives us kind of much better performance.
13:45especially on these data entry tasks, but we're starting to see it in other tests as well.
13:50Nick Larus-Stone:Do you guys know like which types of things certain models are good at and route there accordingly? Or is it more like we're going to route to all three? We don't really have a great taxonomy on which ones are better, but like we're going to do exactly what you said earlier. Yeah, right now we route to all of them and we kind of cross compare. We think there's advantage there. There are certain tasks where we'll choose the right model family. So Gemini tends to be very good with images is what we found. And so, you know, for specific tasks involving images, we might route to that. There's obviously like speed cost considerations where we'll route to specific models.
14:25But specifically for this aspect where I'm talking about like improving performance on a task, I think you actually need to go kind of multi-model and spend more tokens.
14:35Nick Larus-Stone:How many of your tasks are, like data entry seems like one that's pretty verifiable in terms of like, did it extract the right thing or did it not extract the right thing? Writing a research paper sounds maybe way less verifiable. What percentage of them are really verifiable versus not really verifiable? Yeah, I'd say most tests in science are less verifiable. And even the data entry one, if you get sufficient data, it's pretty hard for a human to verify. I'd say it's actually easier for them to verify than it is for them to enter and then verify. So we definitely think there's an advantage there, but it can be somewhat challenging to verify.
15:14Even the report writing that I was talking about earlier, you will have a human verify that right before you submit it to the FDA. It's just like so high stakes that that will happen. And so what we've seen is AI is very good at producing that first draft. You're still going to have a human verify it. You can actually have another AI verify it. So we've seen this pattern a lot where like maybe kind of you use Benchling AI to write the report. And then you open a kind of fresh session, put in that report and have another agent kind of double check it. But there's still a lot of kind of manual verification.
15:47And then there's the kind of set of tasks which are basically unverifiable because it's you're on that bleeding edge of science. And like the way to verify it is you go run an experiment that takes weeks or months and hundreds of thousands of dollars. And people are doing that, right? We see people asking questions to the agent. You know, maybe it's suggesting some experiments that they should run. they have some intuition, right, about like what might work, what might not work. But there's no ground truth, right? The ground truth means you actually have to go run that in the lab. And that's actually quite expensive to verify.
16:18How do you guys do evals then? We do do evals. I'd say we spend a lot of time kind of thinking about evals in terms of tasks that we know we can verify, right? So can you find the right data, right? So if you're finding the wrong data, kind of everything downstream of that's not going to be good. Entering data, right? There's some verification there. But then we spend a lot of time actually looking at traces in production because the types of questions that people ask are so specific to their science. It's very hard for us to build evals. We are working actually on some stuff that we're hoping to release soon about kind of building better evals that are more realistic for the types of problems that scientists see in industry as opposed to maybe some more academic benchmarks.
17:03At the end of the day, we talk to users and we look at kind of production traffic because it's pretty hard to build evals for kind of novel scientific discovery.
17:12Nick Larus-Stone:I feel like looking at production traffic is so key to like tuning the agent once you release it. What is your guys' process for that? Who looks at the traces? How do you decide which traces to look at? Like, what is that process? So everyone on the team looks at the traces. I would say we have a few different patterns for this. So one is we have like a weekly fire chief rotation who's kind of addressing issues. As part of that, they are asked to look at traces and bring them to our kind of like weekly tech operations meeting and share that with the team. So there's kind of like a kind of like recurring, hey, everyone look at traces.
17:49We also get feedback from users. So they say thumbs up, thumbs down. And that's kind of like a good external signal to look at traces. And then I would say people who are working on specific features, right, are going to go look at the traces and, you know, our product managers, our engineers who are building something will actually go and see how people are using that feature after releasing it or kind of putting it into beta or something like that.
18:13Nick Larus-Stone:Early on when you were talking about the agent, you mentioned kind of like tool design and context engineering. What have you guys learned over the past few months, year in that vein? I would say this is in some ways like the key part of making your agent good. It's challenging. Longer context windows help, but they're certainly not a solution. We also have seen the industry move more towards the harnesses that are very popular, right? These coding agent harnesses, you know, having virtual file systems because it mimics like how a coding agent works on your local file system. So we've started to adopt some of those techniques as well.
18:51We also, as you mentioned, are 14 years old, right? And we've kind of built, you know, a whole infrastructure around our SQL database and taking advantage of that. And so I think sometimes it makes sense to follow the crowd and do maybe more file based things. But we actually still rely very heavily on SQL, right? Because, you know, the models are quite good at writing SQL. And so we do some tricks around embedding table names and descriptions so that we can have kind of like faster paths to get them to query the right thing. But I would say it's kind of a mix of trying to keep up with the latest tools where we think the models will be good at while still kind of taking advantage of some of the architecture that we've built over the last 14 years.
19:35Nick Larus-Stone:You mentioned before that you're not using a coding agent harness, and you actually don't think that's necessarily the right kind of like solution for the space. I feel like the labs have been talking a lot about how they're RLing models for their harnesses, which are coding agent harnesses. Have you observed any negative or adverse kind of like effects because they're being overly tuned for those harnesses yet? I don't think I've seen it yet. It's possible it's too early, right? We might see that in the next generation of models, but I don't think I've actually observed any of that. And I think, you know, when we eval kind of our overall system, we look at different models, right?
20:17And so if you start to see future generations of models getting worse with our harness, I think there's like an important discussion for us to have internally of like, do we think we moved to a different model that's going to do better with our harness? or do we change the harness to better fit the model? And I actually don't know what we would end up doing here.
Read the full transcript
20:36Nick Larus-Stone:You mentioned that you guys use a bunch of different models before. When someone kicks off a run in BenchLink AI, how many different models might it hit under that run? Depends what it's doing. I'd say like upper limit is probably seven or eight. I'd say most of the runs are maybe only one to two. We typically have like one main driver model for our kind of core harness, but it'll certainly call out to usually at least one other model. And then depending on the complexity of the task, maybe several other different models. I feel like a big part of Bio is running experiments and learning from that.
21:17Nick Larus-Stone:Do you also think of these agents as learning from what they're seeing their users do? Or are you thinking of memory for your agents in any way? We're, I'd say, still early here. One of the things we've recently released was this ability to have agent skills within the platform. And we think this is actually really powerful, primitive for scientists. Of course, we've added the ability for agents to be able to create and update those skills. So I would say that's like maybe the first, maybe simplest layer of memory is just, hey, you did a task. The next time I need to do this task, I want you to be able to do it in a similar way and hopefully do it much faster and with higher fidelity and better.
21:58Nick Larus-Stone:And are those skills, like will the user say that exactly what you said? Like, hey, remember this? Or is the agent kind of like, you know, seeing that the user did it three times in the background and doing thinking to itself that it should create its own skill? Yeah, so the agent actually has the ability to create its own skill. It's been tuned and we've prompted it such that it's pretty biased towards a user instantiating, though it could do that. We are also in the process of adding the ability for agents to run in the background. There, I think you could start to do a little bit more of what you're talking about of, hey, I've looked through your last 10 chats.
22:32I think, you know, I should create these three skills. We've been a little conservative there because we don't want to create a bunch of skills that are not helpful. So it has the capabilities, but it is kind of prompted to be pretty responsive to what a user wants.
22:46Nick Larus-Stone:And are these skills right now at a user level, at an org level, at some, the user can choose whether to share it? How do you guys think about that? We have a concept of personal skills so users can create them for themselves. And then, you know, since we have been around 14 years, we've built a lot of permissioning into our platform. So we have the ability for users to then share those skills with other users or teams organizationally. And then we also actually have, as Benchling, have been building skills and sharing them with kind of all of our users. so that one, we think they're kind of general purpose and useful.
23:23And we've done a bunch of work to, you know, learn from how people are using our agents. And two, they provide kind of an example for people who are maybe less familiar with skills of like, what should a skill look like? What could go into a skill? So yeah, we think the organizational piece doesn't seem to have been solved broadly as like a ecosystem, I would say, as part of having a platform where people are already sharing things and working with teammates. we think it's important to build that directly in.
23:51Nick Larus-Stone:You mentioned background agents and potentially using those for memory. Are you also thinking of using them for other things? We think of two types of background agents. One is kind of like schedule-based. You know, every week, give me a report because I have a discovery meeting. I think that's super useful. Scientists spend, again, a lot of time, you know, documenting things, building PowerPoints, building kind of like documents to share with others. If you can get a first draft out every week, that's going to be very powerful. The second one is event-based. So based on something that happens within the platform, can you trigger an agent to run?
24:26And I think that, again, is one of the advantages we have for being around for a lot of years. People can do a lot of things in bench link. And so we have a very rich kind of interface where lots of different things are happening. So you could imagine, for instance, if you're a scientist, you're running something in the lab, You have an instrument in your lab that actually takes a measurement. We've built over several years the ability to like take data from that machine, bring it to the cloud, right, and kind of automatically register it within BenchLink, register those results. You can imagine that automatically triggering an agent to do analysis.
25:04So by the time you get back from your lab to your laptop, you don't just have the kind of clean data. You actually have the analysis of like, hey, here's what happened in your experiment. And here's what I think you should do next. I think I'm kind of maybe most excited about some of those use cases where, you know, you as a scientist are already doing kind of your normal work and the agents really are in the background seamless, but just like getting you to the next stage that much faster.
25:30Nick Larus-Stone:You've mentioned a lot of really cool and advanced features that are like reminiscent of a lot of things in coding agents. Your users are scientists. Like, how are you adapting? Like, are they are they do they understand all these things? How are you thinking about the UX? Like, do they understand what skills are? Do they know how to think about skills? Do they know how to think about events, background agents? How are you thinking about the UX for these for these agents? Yeah, we spend a lot of time on user education, I would say. I think, you know, sometimes being in Silicon Valley, right, you get caught up in, you know, everything that's going on and you don't realize that, like, not everyone has maybe the time or the inclination to experiment with all the latest and greatest that's coming out.
26:11And so we spend a lot of time working with our users, showing them what they can do, listening to their problems, making sure we're building the right things there. So take skills, for example. When I go and talk to scientists, I'd say it's maybe like 50-50 of whether or not they've heard of skills at all. Right. So there's some people who like, yeah, I've heard of skills or maybe kind of my, you know, company has some enterprise coding agent or chat client and there's some concept of skills there. But rarely have they spent a lot of time building skills. And then there's, you know, half of people who just like have never kind of touched skills at all.
26:48And so it's a matter of us saying like, hey, this is this thing, right, this way of working. Luckily, I think there's actually like a good parallel in science, which is an SOP, a standard operating procedure. I'll sometimes describe skills as like SOPs for agents, right? You aren't guaranteed that a human following an SOP will do the right thing. But, you know, you write the SOPs with the goal of them kind of doing exactly those steps. And, you know, similar idea, I think, for skills. So we spend a lot of our time doing that translation, I would say. So do you guys have a forward deployed engineering team?
27:21Is that basically what this is? We might not call it that, but Benchling actually from the beginning has had a large service organization because science is really complex and it's really varied. So no two companies are really doing exactly the same science. And so even when we deploy kind of core Benchling, we spend a lot of time with the customer customizing it to their needs. And so we're, I'd say, like pretty familiar with doing that. And I think in the AI era, we've quickly seen that it's not just give people AI and they'll know what to do and they'll immediately be more productive. Sometimes that happens, but oftentimes it's a matter of, you know, deeply understanding their workflows, their bottlenecks and kind of showing how that can map into the tools that we're building.
28:06Nick Larus-Stone:It seems like a lot of the architecture is similar in a lot of ways to coding agents and how that's kind of like brought into this new domain. Have there been cases where you've seen something in a coding agent or coding harness and been like, that is wrong. That won't work for this domain. Like other things might, like skills, sure, great. Let's take those. But this thing, that doesn't transfer. Some of the rules or spec-based development kind of testing-driven doesn't work because you don't have a way to verify in silicon, right? So sometimes it feels like coding agents are really good when you can give them a defined end condition.
28:42Maybe like this goal setting idea, the Ralph loop that has been more popular recently. Your goal is like make a drug that is safe and efficacious in humans.
28:53Nick Larus-Stone:Unfortunately, you can't verify that in silicon, right? And so like that doesn't map super well to what we're doing here. I think there's certain tasks that maybe are more kind of, as you're saying, verifiable in silicon. But in general, I'd say like those types of patterns just don't map well because the speed of your loop is, again, like you physically have to go do the experiment in real life that's going to take days, weeks, months. I saw a tweet the other day. You might have retweeted it or something about how coding is the first domain that a lot of the model labs are kind of like building in.
29:27Nick Larus-Stone:And it seems like bio might be the second. If you know that tweet, like why do you think that might be the case? Two answers. One is those of us who work in the life sciences are kind of blessed with working on problems that are really important. Like human health is kind of like universally important to everyone. Food security, right? Like there's kind of like many positive externalities. And I think the model labs, you know, if they're building this kind of general intelligence, this super intelligence, you're going to have to like apply it to something that is really worthwhile. And I think like at the end of the day, to some extent, it all comes back to life sciences and kind of like, you know, improving human health.
30:14So I think there's that aspect of it. I think there's also an aspect of a lot of researchers come out of the life sciences. So like Dario did his postdoc in neuroscience. So like there is, I think, some kind of just like inherent interest in familiarity with that domain. So I think it's a combination of both of those things.
30:33Nick Larus-Stone:One of the things that you said before this to me that I thought was interesting was that working in life sciences is kind of similar to like working with agents in some way, like the mental aspect of it. Could you expand on that? Yeah, I think LLMs are funny, right? So, you know, in the history of software, we've been building kind of deterministic software, right? And I think there's a certain mindset that people who are good at writing code have. And we call ourselves software engineers, right? And so there's this engineering mindset that is very productive when you're writing code. People have been trying to engineer biology for a long time.
31:07That's actually how I first got interested in it. Biology just doesn't work that way. We are not good enough at understanding biology to make it do exactly what we want. And we spend a lot of time building assays to understand what's happening in the underlying biology. and we probe it and we perturb it. And when I see some of the research that's coming out of model labs around interpretability, it feels like they're actually doing biology research in some ways, right? They're kind of like poking the system that they don't fully understand and watching how it responds. And this is how biology works, right?
31:43Like scientists have been doing this in labs for over 100 years. And for me, it's kind of funny to think about actually sometimes the people who are maybe best at, you know, understanding LLMs or using LLMs are much more similar to scientists and potentially biologists than they would be to software engineers.
32:05Nick Larus-Stone:It's like this black box that you're kind of like probing and see what happens and then using that as experiments. Exactly, and many advantages in working in LLMs, right? Like in theory, we understand it all the way down, even though kind of in practice, there's still some lack of understanding. We understand biology far less, right? Like it's a much harder system. It takes much longer to do experiments on. But if you look at, you know, some of these large training runs, right, they take weeks or months, right? And then you only see what happens at the end, right? Like it's it feels more like biology than traditional software engineering ever has.
32:38Nick Larus-Stone:What's the profile of builder that you found most successful in building agents on your teams at BenchLink? Yeah, I think it's really product minded engineers. At the end of the day, right, like we're trying to solve problems for our users who are scientists and it just requires someone who's, you know, one, willing to talk to and listen to scientists about their problems and kind of deeply understand that. And then two, someone who's willing to kind of like work with LLMs, really deeply understand where they're strong, where they're weak, and then kind of merge those two things together. So the people who are best at that don't necessarily have a science background.
33:16I actually don't think you need a science background to be effective here. But I would say it's more of this like product engineer
33:21Nick Larus-Stone:archetype. When will agents discover a novel cure for a disease? That's a loaded question. In the biotech industry, there's lots of discussions about what is an AI discovered drug. I would say right now, I think we're in an era of agents making it cheaper, faster to discover drugs. We think actually agents will be able to speed up the time from initial discovery to bringing a drug to patients by 2x. How? Why? There's a lot of white space when you're doing drug discovery. So take the example I talked about earlier where maybe you're running an experiment in the lab and your results are on that machine.
34:09Maybe historically, like you would actually take a flash drive, plug it into the computer, right? You'd copy the file over. You'd bring that back to your laptop, you'd email it to a statistician or a bioinformatician for them to do the analysis. They might take a day or two, email you back. And maybe that's actually like, sounds like it's only a day or two, but maybe like, you know, the next time you can get into your sequencing core, that window has already passed, right? And so like now you're actually a week behind. And so like a little thing like having the data ready for you to analyze and knowing what experiment to do next when you get back to your laptop doesn't just save a day, it actually saves a week.
34:46And we think there's like lots of those little things that just add up over time. So that's one aspect. The second aspect is I think actually like agents can make you better at doing science and like do fewer experiments over time. So one example that we've actually seen some of our customers do is using Benchling AI to help them with DOE, design of experiments. So taking a more structured approach to designing their kind of next set of experiments rather than maybe just tweaking kind of like one experimental parameter at a time. You do like a factorial approach of, OK, I'm going to change all of these parameters, but I'm going to do it in a structured way that I get statistical power.
35:26And so I can run 12 experiments instead of running six experiments, then changing something else, running another six, running another six. And I ran 24 in the end. Right. So I think between those two things, we'll get some benefits to your original question. When will like agents discover a new drug? I think we're still years out from that just because, you know, I think the underlying technology of the models are like not good enough at understanding biology right now. and like kind of our kind of overall knowledge base of biology is weak enough that they'll be helpful, but you're not going to have like a solely agent discovered drug in the sense of find me a cure to cancer for quite some time.
36:10Nick Larus-Stone:And so is the blocker that the models aren't good enough or we're still building kind of like the systems around the models to let them run experiments? Because I imagine even if I like if I ask a human, find me a cure for cancer, like they're going to run a bunch of experiments and that's the system around. And so like, how would you, yeah, how would you break it down between the model versus the system? Yeah, that's a great question. Both. We see a lot of cases where the models lack some understanding of both the physical world and kind of the underlying science. So maybe one example of this is, I'll go back to automation here.
36:46If you're a scientist in a lab and you've designed this experiment, typically the way you do this is you then go run the experiment yourself, right? You'll kind of pipette, you'll move liquids around kind of by hand. There's a whole class of technology called lab automation where you have robots do this for you. But the translation from an experiment that you do by hand to lab automation is really hard. There's a whole career called lab automation engineers who their job is to basically do that. And there's a lot of complexity in that. So in theory, that is actually just a code writing task, right?
37:21Like you could just program these robots. In practice, understanding that this specific robot has kind of this tolerance to liquids and actually, you know, the liquid class that it needs is going to be slightly different than what you did by hand. It's very hard. And I don't think that the models are very good at that because there's not a lot of public training data on this. That's kind of like a class where we don't think LLMs are that good. Overall, kind of like experimental design, I think they're still getting better. The models need to get better. The second piece is like, yeah, like even if the models get better, I find it very unlikely that they'll just be able to think really hard and give us kind of an answer.
37:59They're still going to have to kind of interface with the real world and generate data because we haven't generated enough data to understand all of the science yet. And so I think that actually, unfortunately, will move slower. It's very hard to build kind of robust, multi-purpose lab automation. I think you're going to have humans and agents working kind of in concert to push that forward for quite some time.
38:21Nick Larus-Stone:As you think about that future where models and agents work with humans to do longer and longer horizon things, what is the right UX for that end up looking like? Yeah, I think this is, again, kind of shaped by our opinion as a 14-year-old company where scientists have been doing science for the last 14 years. Sometimes what is good for a human is good for an agent. Sometimes it's not. But one example of that is kind of a core part of our product is the lab notebook. So scientists record what they're going to do and then kind of fill in what they actually do in the lab. And that actually turns out to be a reasonable abstraction for a scientific agent as well, right?
39:06You want it to have a surface where it can say, like, here's what I'm going to do, have other people look at it. And then, hey, here's what I actually did. And we built out a whole review process on that. And so I think sometimes actually like those concepts do translate, which is nice. Sometimes they don't. And this is where I think, you know, having agents that can generate code, we've seen, you know, agents be able to build kind of dashboards, HTML dashboards kind of like in the app and have custom interfaces, be useful for certain things. But at the end of the day, you don't want that for everything.
39:39You don't want your agent kind of vibe coding your structure prediction viewer or your plasmid editor. We think it's going to be some combination of, hey, there are things that are useful for your humans to record. Some of those will work for agents. Sometimes you'll want something more vibe coded that an agent can interface with. But in my experience, most humans don't want something that changes on them day to day, week to week. And so having at least some continuity is quite important.
40:08Nick Larus-Stone:One of the things that we've added into LinkSmithFleet, which is our no-code thing, is the concept of like an agent inbox, where you've got agents running the background, doing things, but then you still want the human to approve certain actions or certain steps. Have you guys thought about like a similar like notification center or things like that? Yeah, we're actively doing that right now because we think, again, both these background agents will be really important, but still having that like human in the loop and ability to approve things and monitor what the agents are doing is crucial, especially in our kind of like both like regulated and non-regulated domains.
40:43Nick Larus-Stone:One of the things I feel like we've seen in software is that everyone's becoming like a manager or reviewer of these agents. And on our software engineering team, we've seen that some engineers are getting like review fatigue. They don't want to be reviewing code all the time. They want to actually be building an engineering. Have you seen the same thing happen with scientists or do you think the same thing will happen where like, hey, I'm now a manager. I have to review all these experiments. I don't get to think about what experiments to run. And that takes some of the fun out of it. I don't think we've seen that yet, in part because I think the agents that are really good for scientists are still behind the coding agents that are really good for software engineers.
41:16We sometimes think of it as like roughly a year behind. So if you can imagine we're a year ago in the software industry, I'd say that's actually kind of a good analogy of where these scientific agents are. But the other thing I would say is actually, I think most scientists really like the part of their job where they're thinking about hard scientific problems. You know, they're thinking about what experiments to run in that design and like whether or not this is the right experiment rather than the kind of like ticky tack, like I need to fill in the data from this PDF or I need to like write some Python code, right?
41:53Like there's like a different part of the job that people get enjoyment from. And so I think that actually lends itself much better to being more of a reviewer here, kind of a senior scientist, right, where you're giving feedback to your RAs. You can kind of think of the agents like that today. So I don't think we've actually seen that type of fatigue yet.
42:14Nick Larus-Stone:BenchLam AI sounds like it's doing a lot of things. How do you guys think about charging for the work that it's done? Yeah, so we have a usage-based model. so we charge based on the amount of kind of usage of Benchling AI and then we give a monthly allotment of free credits to each customer based on the number of seats that they have with us. We think this is actually like a good model for a lot of these scientific organizations because you might need to like spike up on like your credit usage but like maybe most of the time you have a baseline level that you can kind of, you know, buy credits for, use free credits for.
42:55But then let's say you're writing an IND, you're going to use a lot of credits to write all the reports, right, and get that done. So you can kind of like spike up rather than having kind of any sort of like flat fee or seat-based model.
43:06Nick Larus-Stone:Have you seen that people care about the underlying models at all? Like, will they say, hey, I need to use Opus for this task because I'm spiking on a really hard task and this is like data entry on something that I think is easier. so let's use a cheaper model for that. We sometimes get that question. We don't actually expose the models. I would say like we see that as part of our job is to make sure you don't have to think about that where we will do the work on like, hey, these types of tasks, we think this model is sufficiently intelligent enough and we'll build evals around that and we'll monitor production traffic to make sure that that's true.
43:41But I think like unlike software engineers where they're a little more willing to experiment and they, you know, follow every model release and kind of have this intuition of what kind of different models are good at. We don't think our users should have to care about any of that. Like that is actually our job to make sure that it works.
43:57Nick Larus-Stone:And how much of Benchlane AI is like a single agent versus like under the hood, it's actually like one agent here, one agent here, a different, maybe just like an LLM call here, but it's all packaged up as Benchlane AI. We initially launched kind of multiple different agents and have been combining them into a single interface because there was at least some confusion with users about like which agent should I use for which task. So we think there's a lot of value in having like one interface that a user can go to. Under the hood, we are still mostly one agent. We do have sub-agents, right, for kind of certain tasks, not even sub-agents, but like sub-workflows, which might be kind of a series of LLM calls.
44:38We've been doing a lot of work to kind of unify it as much as possible. though there are some tasks again in some of these like harder scientific tasks where we've been doing more experimentation of okay maybe actually a different harness is the right solution here and so like there may need to be kind of like multiple agents under the hood.
44:58Nick Larus-Stone:Do you have any hot takes on like what the future of harnesses looks like here or like what's the most like are you experimenting with things that you haven't seen anywhere else because it is so different from coding or things like that? I don't know if there's anything that hasn't ever been done. I'd say we're often inspired by seeing what other people do and then just trying to translate it to our domain. I think actually the kind of like overall driving force behind this is we think our tests are often intelligence limited, which is not true in other domains. A lot of this we think, you know, smarter models will actually get better at science.
45:35And so to some extent, it's a question of like, how can you spend more tokens fast enough? but also like reasonably enough in order to get to better answers. So I would say like a lot of the current thinking and experimentation is around like, how can you spend a lot more tokens because we think it's worth it for some of these harder scientific tests?
46:00Nick Larus-Stone:And does that mainly look like kicking multiple model families off in parallel as we discussed earlier? Yeah, that's kind of where we are today. But I think there's more creative things that we could do here. you know some combination of background agents in that or you know potentially like entirely different model harnesses that we haven't even thought of yet so for example you could imagine building a harness to to really focus on like novelty here so some people say like lms can't do anything novel right their next token predictors i think people just like haven't experimented enough with this they haven't cared enough about kind of novelty right like a lot of software engineering tasks are well scoped, right?
46:42But in our case, actually being able to like make connections between different parts of kind of your research data as an organization and come up with new ideas is really valuable. And so you could imagine kind of building something, some combination of this kind of novel harness where you're, you know, spawning a bunch of different ideas and kind of like cross comparing them, kind of combine that with some of the background agent work that we're doing, where you're just kind of like looking at kind of new science that is getting done as an organization to come up with novel ideas. And then actually, you know, as a scientist, your day changes from, hey, here's the, you know, list of tasks that I have to do to here are some ideas that are kind of proactively suggested to me by my system, because it understands the totality of my organization's science.
47:31I'd say like, we talk a lot about building an AI scientist as kind of the overall goal here. We think a lot of the interesting work there is actually like getting it to work in the lab, right? Some of these like harder physical challenges. But I think part of it also is being able to kind of understand the totality of your organization science and kind of proactively suggest ideas
47:53Nick Larus-Stone:to you as well. And is that the main part that's like intelligence kind of like limited, as you said earlier? Yeah, I don't think like data entry is necessarily intelligence limited, but some of the understanding of the physical world and the kind of what should you do next, right, is kind of the core question here. I think a lot of that is actually intelligence limited. Bio is a very specific domain. And if you guys are intelligence limited, have you have you guys considered kind of like training your own models or post training as a way of putting more of this like very specific intelligence into the model weights themselves?
48:24So our customers are very security and privacy conscious. So we would not kind of like cross train like across customers because leaking any data between customers is big no no. So anything that could lead to that is something we wouldn't do. In terms of like kind of fine tuning open source models or anything like that, I don't think we've seen a lot of success with that. There have been a number of people who have, you know, tried fine tuning models on this domain. In general, they don't appear to be much better than kind of state-of-the-art frontier models. And so then, you know, I think you lose out on a lot of the other model capabilities, and that's often not worth it.
49:06It's not to say we won't ever do this, but right now I think that there is benefit in just having the smartest model that there possibly is. The other aspect of this is like science is made up of like a thousand little tasks, right? And so if you get something that's really good at kind of a specific task, that is only one of the other 999 things you need to do. So I could imagine like having kind of a core frontier model as the main driver and then kind of farming out to task specific models. But it's not something we've seen a need for yet. It doesn't mean it won't happen.
49:38Nick Larus-Stone:I feel like we've started to see more of it recently, but I feel like the main difference is people are doing it for cost reasons because they're not intelligence constrained. So they've got something working and now the volume is going up and up and up and they're just eating all these costs. And so they're trying to drive costs down. But you guys are actually interestingly in a completely different boat, where at least for some tasks, you're intelligence constrained. Yeah. And I think honestly, like it's not a solved problem on like how to make these models smarter at biology because it's not a very verifiable domain.
50:05Right. And so like, you know, there's a lot of work on doing RL for specific bio problems. And I think we will see the models get smarter there. It's not quite as straightforward as some other tasks that you could imagine kind of RLing a model for.
50:18Nick Larus-Stone:Thanks for listening to Max Agency. If you liked this episode, leave a review and subscribe. Send feedback or questions to maxagency at langchain.dev. We want to hear from you.
From the publisher
Nick Larus-Stone is the Head of AI at Benchling, the R&D data platform that life science companies use to store and manage their experiments, samples, instruments, and analysis. Benchling has been around for since 2012. In October 2025, it launched Benchling AI, an intelligence layer with a chat interface, backed by an agent, that helps scientists find data, design experiments, and write reports. Nick came to Benchling through its acquisition of Sphinx Bio, the analysis startup he founded. In this conversation, Nick walks through what it takes to build agents for scientific work, and where the playbook from coding agents holds up and where it breaks down.
–
We also discuss:
- Why Benchling invests so heavily in getting clean data upfront
- How they cross-check answers between models to get more out of each one
- Why and how Benchling leans on production traces
- Where AI actually helps science today, and where it still gets stuck
- Why understanding LLMs is closer to biology than software engineering
–
Timestamps:
(00:00) Intro
(01:22) What Benchling AI is, and the 14-year data platform underneath it
(04:36) Why a decade of structured data is a core advantage
(05:57) The architecture under the hood
(08:28) Similarities and differences compared to a coding harness
(11:14) Benchling’s multi-agent architectures
(14:36) Dealing with verifiable vs non-verifiable tasks
(16:19) Doing evals when clean benchmarks aren’t possible
(18:13) Context engineering: SQL vs. file-based harnesses
(22:11) Memory: agents that create and update their own skills
(25:30) What user education for scientists looks like
(30:33) Why understanding LLMs is closer to biology than software
(33:28) When will agents discover a novel cure for disease?
(44:58) The future of harnesses in science
(48:13) Why fine-tuning on biology hasn't beaten frontier models
–
References:
- Agent Skills (Claude Docs)
- Benchling’s Deep Research Agent
- Claude (Anthropic)
- Design of experiments (DOE)
- FDA Investigational New Drug (IND) application
- Gemini (Google)
- Google AI co-scientist
- LangSmith
- Model Context Protocol (MCP)
- The Ralph (Wiggum) Loop (Geoffrey Huntley)
- Sphinx Bio
–
Where to find Nick:
–
Where to find Harrison:
–
Where to find LangChain:
–
Send feedback or questions to maxagency@langchain.dev




