Position: The Pre/Post-Training Boundary Should Govern IP in Industry–Academia ML Collaborations

25 May 2026 · 13 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How to resolve IP deadlock in industry–academia ML collaborations using a contract template (PBOS) built around a “pre/post-training boundary.”

Guests

No guest names are provided in the transcript; it’s a host-led discussion of a Yale University + eBay working paper.

Guest/author backgrounds

Not specified; the episode references “researchers at Yale University and eBay,” plus “lawyers” vs “scientists” as roles.

Key claims

Current collaboration fails due to incentive misalignment: academics must publish; companies must protect proprietary trained models. PBOS classifies pre-training artifacts (architecture, training code, untrained weights) as open science, and post-training artifacts (weights trained on company data) as business IP. Enforceability comes from three pillars: define the science with scientists, protect business by restricting rights to trained weights, and open source the science by committing to release pre-training code.

Notable examples

A spring 2024 case where a university and a peer-to-peer marketplace built an “AI negotiator” but stalled for months over lawyer drafts about the publish/proprietary line; cited internal tech norms like DeepMind’s AlphaGo (publish methods, keep trained weights proprietary).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Bottleneck in AI Collaborations

0:45 to 2:38

Exploring the challenges of collaboration between academia and industry.

“Yeah, because we generally like to think of these things as engineering problems, because engineering problems have engineering solutions.”

Incentive Misalignment Explained

2:38 to 3:56

Discussion of how misaligned incentives affect AI collaborations.

“Oh, this case study is a classic example.”

Systemic Flaws in Negotiation

3:56 to 5:52

Analyzing the flaws in how legal departments handle AI contracts.

“I mean, having lawyers negotiate a machine learning contract without scientists at the table, it feels like having accountants write the rules for a new sport.”

Introducing the PBOS Framework

5:52 to 7:46

Presentation of the PBOS framework for resolving IP issues in AI.

“Yes, and that boundary is so powerful because it is legally clean and entirely auditable.”

Defining IP Boundaries in AI

7:46 to 9:45

How to establish clear boundaries for business IP in AI projects.

“The university doesn't get to own it or sell it.”

The Impact of PBOS on Collaboration

9:45 to 11:35

Understanding how PBOS can resolve tensions in academic and industry partnerships.

“The biggest tech companies in the world are already doing exactly this.”

Future Implications of AI Data Practices

11:35 to 12:57

Speculating on the future of business IP as AI evolves.

“Bring the scientists in, draw a technical line, and you satisfy both sides.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, let's unpack this. Welcome to the Deep Dive, everyone. Our mission today is to, well, we're exploring a really fascinating working paper from researchers at Yale University and eBay, and it tackles this massive bottleneck in the world of artificial intelligence. It's a huge issue. It really is. Like, imagine the current state of industry and academia trying to collaborate on AI is like two people trying to bake a revolutionary cake. Okay, a cake. I like that. Right. So one person has the secret recipe and the other person has the only working oven in a town. But instead of, you know, actually baking the cake, they've just been six months paying lawyers to argue over who gets to keep the pan.

0:43That is honestly a perfectly accurate way to visualize it. It's an incredibly frustrating paradox. I mean, it's wild, right? Yeah, because we generally like to think of these things as engineering problems, because engineering problems have engineering solutions. You know, you can build a bigger cooling system or just optimize the math. Right, just throw more compute power at it. Exactly. But the bottleneck we're dealing with today is fundamentally about human incentives. It's a structural incentive misalignment. So break that down for us. Why is the structure failing so badly? Well, look at the universities.

1:16They produce these brilliant algorithmic ideas and the top graduate talent. But academics absolutely must publish their methodology and results for their career advancement. Right. They need tenure. They need grants. Precisely. A collaboration that ends in total corporate secrecy is just a catastrophic failure for them. It's years of lost time. Because they can't exactly put on their resume. Did some super groundbreaking but highly classified stuff for a tech giant. Just trust me. Right. Nobody's granting tenure for just trust me. But then, you know, you look at the company side of the table. They have these massive, incredibly rich proprietary data sets.

1:55We're talking like transaction logs, user behavior, stuff like that. Yeah. Highly specific, real world data. And they absolutely must protect any models trained on that data to keep their competitive edge. If a model that perfectly mimics their users is just leaked to the public, their entire business model is at risk. So it's like an immovable object meeting an unstoppable force. Neither side is actually acting irrationally. Not at all. They both have valid needs. But without a shared framework to reconcile them, that friction just kills the collaboration before it even begins. Which brings us to this specific real-world breakdown from the spring of 2024 that really just perfectly illustrates how paralyzing this gets.

2:38Oh, this case study is a classic example. Yeah. So a university research team and this major peer-to-peer marketplace decided they wanted to build an AI negotiator together. Which, on paper, is the ultimate strategic alliance. Totally. The company had millions of transcripts of real people haggling and negotiating on their platform. And the researchers, well, they had the advanced methods to actually study and model all that complex behavior. It should have been a massive win-win. Exactly. But instead of building the AI, the project just sat in purgatory for months. And they weren't fighting over money.

3:09Right. Not at all. They weren't even fighting about the scientific scope. They were just trading dozens of document drafts between their lawyers, trying to define this tiny line between what could be published and what had to stay proprietary. And what's fascinating here is that this isn't just an isolated, unlucky negotiation. It's a systemic flaw. How so? Well, historically, these contracts were negotiated solely by legal departments. The technical reality of how machine learning actually operates is completely invisible to the people writing the terms. I just don't get the tech. Right. Lawyers look for precedent.

3:42They try to mitigate risk using contract law. They simply cannot correctly diagnose these highly technical machine learning boundaries on their own. Wait, but shouldn't legal departments be the experts at defining intellectual property? I mean, having lawyers negotiate a machine learning contract without scientists at the table, it feels like having accountants write the rules for a new sport. That's exactly what it's like. They just see the liabilities. They don't understand how the game is actually played or what pieces of equipment even matter. So they just try to bubble wrap the players and ban running altogether.

4:17They try to protect everything because they literally don't have the technical expertise to know which specific digital artifacts actually contain the company's secret data. Okay, so this realization that defining IP in AI is fundamentally a technical claim, not a legal one, that leads us directly to the solution proposed in the text. Which is such an elegant framework. Yeah, it's called PBOS, protect the business, open source the science. It's this community adoptable contract template. And the entire framework rests on this one single technically grounded rule called the pre and post training boundary.

4:56Right. The pre and post training boundary. Let's really dig into this because this is the golden rule. Under QBOS, pre-training artifacts, so that's like the architecture, the training code, the untrained weights, all of that is classified as open science. Yeah. But the post training artifacts, the weights that have actually been trained on the company's proprietary data, those are classified as business IP. Think of a neural network like a massive stadium size mixing board with a billion little volume sliders. OK. A giant mixing board. Got it. Before it sees any company data, those sliders are just set randomly.

5:28That blank layout, the wiring, that's the architecture, that is open science. But once you feed the secret data into it and it adjusts all those billion sliders to learn the patterns. Here's where it gets really interesting, though. A completely blank neural network is open science. But that exact same network, once it reads the company data, suddenly becomes business property. It's tracking the actual information content, not just the file type. Yes, and that boundary is so powerful because it is legally clean and entirely auditable. Auditable because of where the files live, right? Precisely.

6:03The trained weights, those adjusted sliders, they are massive data files. By design, they require specialized infrastructure and they stay on the company's secure servers. So it's not like the researcher can just accidentally email them to a buddy. No, absolutely not. We're talking gigabytes or terabytes of data. And meanwhile, the untrained artifacts, the blank mixing board, contain literally zero proprietary patterns. The company can just look at their own IT logs to prove the trained weights never left the building. Okay, but what if they use, like, fine-tuning? Say, the researcher brings in a public model they downloaded, puts it on the company's server, and fine-tunes it with the private data.

6:42Does that break the boundary? The framework is robust enough to handle that beautifully. The base model, since it's public, remains open science. But the newly fine-tuned weights the specific adjustments made after exposing it to the company's data, those belong strictly to the company. Ah, okay. So the rule holds up. Drawing that technical boundary is a great concept, but how do you actually enforce it in a legally binding way? I mean, without triggering another six months of lawyer emails. That is where the three pillars of the PBOS contract come in. Right. So pillar one is define the science.

7:15The contract has to explicitly define what the science even is for that specific project. Meaning the mathematical equations, algorithmic descriptions, any data-free visual exhibits, or models trained only on public or synthetic data. And the crucial part here is that scientists actually have to be in the room to write this part. Lawyers can't just look at code and decide if it's data-free. Exactly. You need the technical experts to define the technical boundary. Makes full sense. So then pillar two is protect the business. Basically, anything trained on the company data is exclusively commercial property.

7:50The university doesn't get to own it or sell it. Right. They only get a really narrow research license. So they can use the trained model on the company's servers just to evaluate it, figure out how it performed, and then publish those academic results. But they can't deploy it or distribute the model itself. It gives them the empirical proof they need for their paper, but the asset stays locked down. Okay, but pillar three is open source the science, meaning the university actively commits, in the contract, to releasing their pre-training code under a permissive license. Yes, they do. Why would a university actively commit to just giving their work away?

8:27Doesn't that feel like a massive concession to the company? Like they're just surrendering their IP. It sounds like a concession, I know. But if we connect this to the bigger picture, it isn't a concession at all. It's exactly what academics need to do to get published anyway. Oh, because modern journals want to see the code. Right. Top tier peer review basically demands open source methodology now. But more importantly, putting it in the contract legally prevents the company from getting cold feet later. Wait, really? How does it prevent that? Well, say a PR executive suddenly panics right before publication and tries to suppress the paper, citing vague confidentiality concerns.

9:03Oh, I'm sure that happens all the time. It does. But with PBOS, the contract structurally guarantees that the pre-training then is open science and will be released. It shields the academic. Wow, that's smart. And I guess on the flip side, it structurally guarantees to the company that the academic researcher won't just take the code and go launch a rival startup. Exactly. Because it's open source, everyone has it. There's no secret sauce for a stealth startup. It aligns everyone's incentives perfectly. That is genuinely brilliant. But, you know, this whole contract structure might sound like this radical new way for universities and companies to interact.

9:40But the text reveals this super surprising secret. The hidden industry norm. Yeah. The biggest tech companies in the world are already doing exactly this. It's true. This raises an important question. If this boundary works so well, why hasn't it worked between disparate institutions? Because internally, it's the standard. Like when DeepMind published AlphaGo in 2016, or, you know, Google publishing Paul M, or OpenAI publishing GPT-3. Right. When they published those massive papers, they allowed their internal researchers to publish the full architectures, the methodology, the science, basically.

10:15But they kept the trained weights strictly proprietary. They kept the competitive asset locked up. Exactly. PBOS is simply taking this proven internal norm, which works flawlessly behind closed doors, and making it portable. Portable across completely different institutions. Yes. So a university and a corporation can just adopt it without reinventing a wheel. So for everyone listening right now, you might be wondering, why should you care about contract templates? It sounds super dry. It does sound dry, but the implications are massive. Yeah, because right now there is a literal market failure in scientific production.

10:51There are entire fields of AI research, like agents learning from real human negotiations or modeling large-scale human economic dynamics that are just sitting empty. Because academic labs just don't have the data. Exactly. And the transaction costs, the legal friction to get that data from the companies is just too high. So what does this all mean? It means that if technology transfer offices and the ML community adopt PBOS as a default. The jam finally breaks. Yes. We are going to see a massive flood of new AI breakthroughs in areas that intersect with real human economic and social behavior.

11:27The bottleneck isn't the tech. It's the paperwork. And PBOS clears the jam. It really just comes down to having the right people at the table. Bring the scientists in, draw a technical line, and you satisfy both sides. So to recap, we started this deep dive looking at that paralyzing structural friction between academic publication and corporate secrecy. And we discovered that the solution isn't more legal drafting. It's actually bringing scientists to the table to draw a technical pre - and post-training boundary. It's an elegant solution to a very human problem. It really is. But before we wrap up, I want to leave you, the listener, with a final really provocative thought to chew on, something that builds on everything we've talked about.

12:06Oh, I'm curious. So if this entire boundary between open science and corporate IP hinges solely on whether a model has been exposed to proprietary human data, what happens in the near future as AI models increasingly train on synthetic data? Ah, data generated by other AI models. Exactly. If the proprietary human data is removed from the loop entirely, because AI is just learning from simulated environments, does the whole concept of business IP in AI eventually vanish? That is a fascinating question. If there's no proprietary data, is it all just open science? Right. Does the boundary just dissolve completely?

12:44It's something to think about. Is this tech move so fast? Thank you for joining us on this deep dive. We love your curiosity. And we invite you to keep questioning the invisible structures, the contracts and the norms that shape the technology around us. Until next time.

From the publisher

This paper proposes a new contractual framework called PBOS to resolve persistent intellectual property conflicts in industry-academia machine learning collaborations. By involving scientists in legal negotiations, the authors suggest a clear division based on the pre/post-training boundary of a model. Under this model, pre-training artifacts such as code and architectures are treated as open science, while post-training weights derived from proprietary data remain protected corporate assets. This approach ensures researchers can fulfill academic publication requirements without compromising a company's competitive advantage. Ultimately, the framework aims to reduce the high transaction costs and legal delays that currently prevent many valuable large-scale research partnerships.

More from Best AI papers explained

All 475 episodes
Position: The Pre/Post-Training Boundary Should Govern IP in Industry–Academia ML CollaborationsBest AI papers explained · 13 min
Listen in VO