#273 A Big New Market for Dubbing and Accessibility Solutions with 3Play Media co-CEOs

11 Dec 2025 · 51 min · 21 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

3Play Media’s evolution from expensive English captioning to a scalable “AI + expert-in-the-loop” platform for video accessibility and localization (captions, subtitling, dubbing, and multilingual audio description), driven by unit economics, quality control, and the European Accessibility Act (EAA).

Guests (backgrounds)

Josh Miller and Christopher Antunes, co-CEOs and co-founders of 3Play Media. They met at MIT (over 15 years ago) and started 3Play to solve MIT OpenCourseWare’s accessibility problem (captioning thousands of engineering videos). Both had consulting backgrounds.

Key claims

3Play has processed 30M+ videos; 100k+ videos in a month returned within 8 hours with experts “in the loop.” They avoid “retrofit” workflows by designing quality/scale from day one. Audio description is a “gateway” for synthetic voice because it’s simpler than dubbing (no lip-sync), and script quality/judgment matter more than voice alone. EAA creates a budgeting and scale opportunity for multilingual accessibility.

Notable examples

MIT OpenCourseWare (Hewlett-funded, no early web-video accessibility regulation); Perkins School for the Blind research for audio description; Amazon’s early synthetic audio description feedback; Netflix captioning backlog; ASR research findings: big leap with Whisper/generative models, then plateau; speech recognition is scenario-dependent (studio single-speaker vs noisy multi-speaker).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to 3Play Media

0:00 to 0:13

Learn about the impressive capabilities and recent performance of 3Play Media.

“We've processed more than 30 million videos through this thing, 100 ,000 videos plus in the last month alone delivered back to customers within eight hours with the expert in the loop.”

Founders' Background and Company Origin

0:39 to 1:00

Discover the backgrounds of the founders and the inception of 3Play Media.

“So, first, I want to talk a bit about the background.”

Early Problem-Solving and Business Model

1:00 to 4:09

Understand the early challenges faced and how they shaped the business model.

“So I'm Chris, like Florian said, one of our co-founders.”

Growth and Funding Journey

4:09 to 6:08

Explore how 3Play Media grew from its inception to achieving profitability.

“Yeah, I think some of that's even the function of being at MIT in the business school there.”

Unique Human in the Loop Model

6:08 to 7:12

Learn about the significance of the human element in their technology solutions.

“pretty much from the beginning, since like you pointed out, we had customers first and after MIT, many customers in the education segment and the higher ed segment right afterwards.”

Platform Development and Customer Segments

7:12 to 8:13

Examine how 3Play Media's platform meets diverse customer needs across industries.

“You need to start there because unlike pure technology businesses, you need to be thinking about cost.”

Scaling Operations and Quality Control

8:13 to 11:42

Delve into the challenges of scaling operations while maintaining quality standards.

“So I think you're hitting on something really important here because I think there are two constituencies that we have to think about when it comes to scaling in the platform.”

Trends in Multilingual Audio Description

11:42 to 14:00

Discover recent trends impacting the growth of multilingual audio description.

“But we need a process that scales and scales consistently across all those every day.”

The Future of Live Events and Localization

14:00 to 18:08

Explores the challenges and opportunities in localization for live events.

“So you see big streaming platforms start to own sports rights, event rights, and those events, they want to get out to the world really quickly because they have an expiring clock in lots of languages.”

The Evolution of Audio Description

18:08 to 21:51

Discusses the growth and development of audio description services at 3Play Media.

“You still need to fit the descriptions into the empty spaces in the audio, but it's a simpler task.”
Show all 21 chapters

Script Writing for Audio Description

21:51 to 24:57

Delves into the intricacies of script writing for audio description and the role of AI.

“So I think they take a stance, and I can't speak for them.”

Dubbing Solutions and Market Dynamics

24:57 to 28:00

Examines the current state and trends in the dubbing market, focusing on client needs.

“I think I checked in Q2 2024, so roughly 18 months ago, you launched your human in the loop dubbing solution.”

Exploring Complexities in Dubbing and Captioning

28:00 to 29:58

Learn about the intricate challenges in dubbing compared to captioning and the importance of scaling.

“We're figuring out even the price points because there's a wide range.”

The Challenges of Demos in Dubbing

29:58 to 31:20

Understand the pitfalls of demo processes in dubbing and the unsolved problems that arise at scale.

“I mean, we saw some of that issue just in captioning, But I agree, it's very much magnified in dubbing when there are so many different steps.”

Navigating the European Accessibility Act

31:20 to 36:18

Gain insights into the European Accessibility Act and its implications for companies in North America.

“where are we at the various stages of its rollout?”

Insights from ASR Research Findings

36:18 to 38:26

Learn about the recent advancements in Automatic Speech Recognition and its performance metrics.

“which would be the next phase of our growth in terms of the EAA and kind of where that goes.”

AI's Role in Dubbing and Growth Strategies

38:26 to 42:00

Explore how AI impacts dubbing processes and insights into the growth strategies of companies.

“with the generative models entering the mix with the whisper model.”

Growth Strategies and Investment Plans

42:00 to 43:59

Explore the company's growth strategies, including M&A considerations and organic growth opportunities.

“So we have like sort of built in rapid learning environment.”

The Importance of Dubbing for YouTube Creators

44:00 to 45:25

Discuss the significance of dubbing solutions for YouTube creators and the international market.

“And as we look out into next year, we're making a lot of investments on the data science side.”

AI-Driven Solutions and Customer Analytics

45:26 to 49:10

Learn about the integration of AI in their services and the importance of customer analytics.

“What YouTube solves for everyone immediately is distribution.”

Future of Localization and Company Visibility

49:11 to 51:29

Insights into future localization projects and the need for better company visibility in the market.

“And even when you upgrade, giving optionality there.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Christopher Antunes:We've processed more than 30 million videos through this thing, 100 ,000 videos plus in the last month alone delivered back to customers within eight hours with the expert in the loop.

0:13Christopher Antunes:Hey everyone, and welcome to SlaterPod. Today on the podcast, we welcome Josh Miller and Christopher Antunes. So Josh and Chris are co-CEOs and co-founders at 3Play Media. Now, 3Play Media is a leading language solutions integrator for video accessibility and localization. They combine proprietary tooling, advanced AI, and expert-in-the-loop solutions. So, hi, Josh and Kristen. Thanks so much for joining today. Thanks for having us. Hi, Florian. Thanks for having us. Very, very cool to have you on today. So, first, I want to talk a bit about the background. We always do that on the podcast, get a bit of a feel for the company, for the founders, you know, what's kind of been the trajectory so far.

0:50Christopher Antunes:So tell us a bit more about maybe your own professional background, origin of the company, and then kind of take us to the elevator pitch of what you're doing now to kind of set the scene. So I'm Chris, like Florian said, one of our co-founders. I met Josh, man, more than 15 years ago at MIT. We both had, and all four of our founders, so we had four originally. So I don't know if we'll get into this, but the co-CEO model is much easier than having four founders. That is for sure. but we all met at MIT and only had a few years of professional experience before that. Mine was in consulting, Josh's was as well.

1:27Christopher Antunes:But at MIT, got started on 3Play, focused on a very specific problem for a group at MIT called OpenCourseWare. And this was a group that was putting engineering courses at MIT online for free. This is a common model now, e-learning. There's companies like Coursera that I know was in the news on the dubbing side recently. But MIT was putting courses online and they were funded by the Hewlett Foundation to do that. And it was their funding, not any regulation at the time that required their video be accessible. So at that time, that meant closed captioned. So the thousands of videos that they wanted to put in, all technical, had to go up online and be accessible.

2:10Christopher Antunes:And at the time, back in 2007, 2008, closed captioning, and we'll get into the translation and dubbing and audio description, but captioning was very expensive, right? There weren't solutions that were leveraging AI meaningfully that could drive the price down and make this economical and feasible for MIT. So their group, like legitimately, 75 % of their budget would have been absorbed by captioning this video and it was just untenable. So we were presented with this problem of, can you help? And that was really it. So business problem first, customer problem first. And Josh can jump in in a second, but exuberance of youth, four of us at MIT, we thought, we'll solve this problem by the end of the afternoon with AI.

2:53Christopher Antunes:It wasn't called AI back then, data science or speech recognition in our world. And so we worked with a professor, Jim Glass, in the computer science and artificial intelligence lab to try to solve that problem. And by the end of that afternoon, we realized AI wasn't sufficient. Speech recognition wasn't sufficient, but it was a really good start. And literally, we have speech recognition in Excel, sitting there, cell by cell, looking at the cells and thinking about what we might need to do to clean up. And there's a lot of steps to fast forward through here. But to solve that problem for MIT, day one, first customer, first problem we were introduced, we needed AI and we needed tools because no tools existed out in market that could do what we wanted to do, which was take the speech recognition and produce captions efficiently.

3:37Christopher Antunes:So we had to build it. And then we needed people. And first it was us. And then it was our friends that we could wrangle into it. But very soon it became evident we needed lots of people to address this issue at scale. So it was AI, it was tooling we developed, and it was a marketplace of people at massive scale to solve this problem. And it was from the beginning. And we'll come back to this, I'm sure again and again, but none of this is retrofit. Our idea of how do we do this at scale and maintain quality and maintain consistency across all these services was built into it day one.

4:09Josh Miller:Yeah, I think some of that's even the function of being at MIT in the business school there. And everything's about technology and scale and avoiding manual workflows and things like that. So it was kind of drilled into us from the beginning, but it was even kind of why we went to MIT in the first place is to do something interesting with technology at scale. So from the very beginning, as Chris said, we were looking at this a little bit differently and kind of how can we take what, you know, is a solution that's out there, right? There were captioning providers out there and there were language providers out there.

4:43Josh Miller:So what were they doing that we either wanted to emulate or wanted to avoid? And we thought very, very carefully about that. But I do think it's worth pointing out even the context and doubling down on what Chris said of the timing. This is end of 2007, beginning of 2008. There's absolutely no regulation around web video for accessibility. The only regulation is around broadcast video. Netflix has not started streaming yet. They're still shipping DVDs. YouTube is just getting going. So, I mean, it's early days of online video. And one of the most important things we did was actually avoid the media market in the beginning.

5:19Josh Miller:because there were providers and we didn't have the right to win yet. We didn't have the differentiation that we were aiming to build. So we very, very purposely focused on this emerging market of web video and kind of saw the tea, I guess you kind of saw the tea leaves that there was going to be regulation or certainly hoped there was going to be regulation, but also recognize that that's where there was going to be the potential for explosive growth.

5:44Christopher Antunes:And then was it bootstrapped since you had a customer kind of from day one and you leveraged that into growth? Those were the days. There were four of us in an apartment just after we graduated from MIT. We did soon afterwards raise angel investment, but we were able to build a profitable growing company pretty much from the beginning, since like you pointed out, we had customers first and after MIT, many customers in the education segment and the higher ed segment right afterwards. this trend continued. And then e-learning emerged. And then, like Josh mentioned, online video and streaming came afterwards.

6:25Christopher Antunes:And then, you know, we helped Netflix caption a backlog of content. Like it's sort of stacked one after another. So we raised some angel and then much, much later in the journey, raised more institutional capital. Got it. Yeah. I think I listened to another podcast, accessibility podcast you did a couple of months ago, where you said it's been profitable very early on. And then you kind of like, you You know, never, never am back to the series A, B, C, D type of. It's in the DNA. And I do think an important part of these human in the loop or call them expert in the loop models is the, when we'll get into the AI, I'm sure.

7:01Christopher Antunes:And we'll get into the tooling and we'll get into the orchestration and the infrastructure, which is all magic in all. Certainly a force multiplier in your ability to scale and in drive efficiency. But at the end of the day, 95 % plus of this is still people and it's still experts and It's still artisans across all of these tasks and really understanding the unit economics and understanding of how to do that in a way that's profitable for us, but also makes sense for your customer. You need to start there because unlike pure technology businesses, you need to be thinking about cost. You need to be thinking about the economic model from the beginning.

7:38Josh Miller:I think what's still important there is that even though it's still heavily, heavily people to get it right, it means we're able to process 10, 20, 100 times the amount of content in the same timeframe. That's the big difference is kind of what kind of scale we can unlock.

7:55Christopher Antunes:So you started with kind of English closed captioning, and then you really pushed towards a unified platform where you now offer accessibility, localization, workflow. So it's not just a solutions and services business with thousands of people, but you have a ton of tech now. Can you just tell a bit more about the platform and then also what kind of customers, maybe top two, three segments that you have?

8:18Josh Miller:So I think you're hitting on something really important here because I think there are two constituencies that we have to think about when it comes to scaling in the platform. One side is the customer experience and making sure that scales. And the other side is the operational experience and making sure that scales. And so it's almost like two platforms in one or like a two-sided platform in a way. And so one of the things early on that we had to address was the reality that web video publishers, in many cases in the beginning, educational institutions, enterprises, where you've got people who've really never dealt with video before, all of a sudden being tasked with do this new thing.

8:58Josh Miller:And there are these video platforms coming around beyond YouTube, things like Brightcove and Kaltura and many more, it was very clear that we had to simplify the workflow. And that was a combination of having a very easy to use platform to self-serve almost in the sense of feeling like a web app. So there might be customer success and support like you would with any web application, but ultimately you as the user could do everything yourself. No Excel files flying around about when files were coming or anything like that. And then similarly, integrate directly into the video platform and make that a fully automated workflow.

9:39Josh Miller:So those were things that we were thinking about very, very early on that would actually be part of the scale story. And also just a really great differentiated customer experience because we knew that the providers that did exist weren't doing anything like that. And then we had to extend the same concepts to the operational side and figure out how do we actually make that scale too.

10:02Christopher Antunes:Well, and Josh, I mean, in terms of end markets, our three primary end markets are media and entertainment, traditional media entertainment, higher education, e-learning, and then corporate. And corporate is a bit of a catch-all. But at the end of the day, if you have video, you're a potential customer of ours. I mean, full stop, right? because our goal, our mission is to make video accessible. And that can mean to a community like the deaf or blind community, or it could mean in any language to anyone in the world. So if you have video, so we had to build a process. We'll get, and this I think is a good segue into the platform that could handle the scale we anticipated, but also the quality expectations and variability in quality expectations.

10:45Christopher Antunes:Because different segments expect different quality. Absolutely. That is a very tricky one. The variability, some where you need perfections, other where it's like, well, I'm actually okay with 98%.

10:58Josh Miller:And everyone defines it differently. That's the reality. So what one person says is perfection, another segment entirely might say that's not usable.

11:06Christopher Antunes:And this maps directly to the platform or the assembly line that we need to build. And I use that term endearingly, not derogatorily in any way. And what I mean by that is we have to process, I mean, we've processed more than 30 million videos through this thing. 100 ,000 videos plus in the last month alone delivered back to customers within eight hours with an expert in the loop. So not just AI. And these are customers with really high quality expectations and discerning requirements. But every video is different. Every set of expectations is different. Every industry is different. Every use case is different.

11:42Christopher Antunes:But we need a process that scales and scales consistently across all those every day. So tens of thousands of videos flowing through the system. And what the platform allows us to do by owning the sort of integrating the AI models, owning the tooling, owning all of the instrumentation and data collection inside that allows us to have observability and allows us to constantly improve it, enhance it, and customize it scalably for all those different use cases. I want to talk about AI, but I don't want to talk about AI super broadly right from the start. So, let me try to kind of pick one area, which is multilingual audio description, one angle.

12:21Christopher Antunes:So, A, I think multilingual audio description has become a pretty good area for growth for you. And it seems like it's an area where kind of major media companies feel more comfortable adopting AI-generated voices. So A, would you agree? And B, is this like a gateway for AI voice adoption versus maybe like AI dubbing where it's a little bit trickier? We spent the first 10 plus years or longer of 3Play as an accessibility company. So audio description and captioning centrally focused on English and Spanish, basically North American customers. And three big trends over the last, call it 18 months, have gotten us really excited about global multi-language audio description and localization, subtitling, and dubbing.

13:10Christopher Antunes:Subs and dubs aren't new, as you know, better than anyone. They've been around for a long time. We were not centrally focused on that. So I'll just run through these three trends quickly, and then one of them will connect directly to audio description. I'll let Josh take that. So the first is what you said, the European Accessibility Act, which is a regulatory change in Europe, which is driving a higher standard of accessibility. video is a big part of that. And it maps directly into dubs. And again, I'll save that for Josh on the audio description side. We saw that coming and we are really good at audio description at scale.

13:39Christopher Antunes:We do a lot with media companies already. Global language was on the way. Two, voice technology trends, which we'll get into. We were not interested in building traditional brick and mortar studios and scaling dubbing in that way. But with voice technology, we are good at building tools that people can pilot from home and use our marketplace to scale it. That technology trend made us really interested in what we can do in dubbing. And then the third is live events. So you see big streaming platforms start to own sports rights, event rights, and those events, they want to get out to the world really quickly because they have an expiring clock in lots of languages.

14:19Christopher Antunes:The idea of high quality fast is very exciting to us. And it's not there yet. State of the art is not high quality Netflix theatrical or high end theatrical subtitles or dubs with human in the loop in hours. It's still days, right? But we have a dream of saying, could we make that hours? And those three things, one, a regulatory trend, one, a market trend and one, a technology trend made us super excited about bringing the three play platform into localization. And then Josh EAA and dubbing and sorry, audio description to specifically,

14:51Josh Miller:I think the audio description storyline at 3Play is an interesting one here because if we go back 10 years before we launched any audio description at all, we had a couple customers come to us and say, hey, can you do audio description? We said, no, we're not doing that. And a lot of it was because it was really complicated. Audio description is not straightforward. And so we were avoiding it for a long time until all of a sudden some things were happening with some regulations online and there were some things happening with some of our university customers. And we started getting the requests more and more.

15:23Josh Miller:And so we had a conversation at one of our annual retreats, and it was still a small team. And we looked at each other and said, do we need to do this? And we said, okay, well, if we do it, we got to do it differently. And so that was one of the things that I think prepared us for what we're doing now. Because from day one, we were using synthesized voice, even for English audio description. And so we built that out with a really thoughtful approach to script writing. So the human is still heavily involved in the script writing. And we can talk about AI script writing as well. But we put together a really thoughtful process.

16:00Josh Miller:We were interviewing Perkins School for the Blind. We were setting up studies with blind students and testing it really carefully to bring it to market. Recognizing, just like Chris was saying, we're not going to replace a lot of the audio description being done today for theatrical content. That's not the point. There is just so much content out there that is not accessible. And the schools are struggling to figure out how to make content, classroom content accessible for a blind student. And there's just no way they're going to pay$30 a minute for audio description. So it was very clear there was something that we could disrupt.

16:38Josh Miller:So we built that out. We started to really bring that to market. And it was actually a conversation not quite a year ago. we were really excited about what we had started to build with dubbing and subtitling. And we were in LA having some conversations with some of our customers and even some prospects. And we said, hey, we want to get your feedback on this new localization stuff we're doing. And they said, both conversations, I will not forget this. Both of them said, I'm sure it's great. We have a different problem we need you to help with. And that was about this EAA thing. and they said, we need to understand what you can do with multilingual audio description.

17:17Josh Miller:And we're open to synthesized voice. And a lot of this is around, again, a lot of content, limited budget, limited timeframes, and the perception, and I think an accurate one, that the synthetic voices have gotten a lot, lot better. And so there is, again, not all content's equal, and we need to figure out how do we make content accessible? So I think, Florian, to your original question, is this the gateway? I'd say yes.

17:43Christopher Antunes:And Florian, it's the technical answer to that question of why is it a gateway? It's a couple of things. Audio description is simpler than dubbing. It just is. It's typically not filled with emotion in the same way or as much acting. It can still be well acted, but not as much. You don't need to worry about lip syncing, right? Because it's more of a voiceover style description there. You still need to fit the descriptions into the empty spaces in the audio, but it's a simpler task. Also, audio description is principally consumed by blind users. It is not consumed as widely, let's say, as a German dub in Germany might be.

18:25Christopher Antunes:So it's safer because the audience is smaller. and also anyone, the reality is, what are you used to? Anyone, the reality of how the internet works today is anyone who's blind and using a screen reader is using synthetic voice all day. There is not a voice actor riding alongside you as you operate the internet. It is synthetic voice all of the time. So you're used to that synthetic voice already and there's not as maybe big of a bar to climb or a hurdle to climb. Whereas in media, traditional dubbing, obviously, aside from even the column political issues to navigate, there's just an expectation of a certain bar and a certain quality.

19:08Christopher Antunes:And that's real. And that's been there for a while. It's different in the audio description setting.

19:11Josh Miller:I think that point is so critical for the first gateway, which was synthetic voice with informational content. So whether it be educational or corporate, like a corporate product video, you know, that's where having synthetic voice five years ago wasn't as big of a deal. And there was, you know, there was definitely conversations around, you know, is it ready for media content? And a lot of people said no. So we were working with media companies back then, and we're certainly doing so more now. But the improvement of the synthetic voice actually does make it sound human now. And again, without the emotional issues, it sounds really good.

19:55Josh Miller:And so it's harder to even make the argument that it's robotic or anything like that anymore. The other thing that I think is so critical here is the reality that we're not cloning. We're not matching. We're not mimicking somebody's likeness when it comes to audio description. There's no source where we are creating that audio track from scratch. Whereas obviously with dubbing, you're creating a copy of sorts from some original source, which just opens up a different set of issues.

20:27Christopher Antunes:Is there any research on how much longer people can now listen to an audio description without tiring because the quality has gotten so much better? I would assume like, let's say the 10-year-ago robotic voice, I would have just zoned out after maybe 10, 15 minutes. But now I'm sure, I guess I could listen to the new ones for a little longer. Although if, let's say, you have these notebook LM podcasts from Google that you can turn your notes into a podcast. After five, ten minutes, I'm like, oh, my God. It is good, but it still has a little bit of this kind of repetitiveness to it. It's a good question.

21:07Christopher Antunes:For audio description, I think it's still early days. The improvement in voice is happening so rapidly. I don't think there have been any good studies yet about, as far as at least the studies I know of, that map back to the improvement in the quality and the emotion in the voice. But it's an interesting question.

21:24Josh Miller:I mean, Amazon was very early in being public as a major streamer doing synthetic voice audio description. And they got some negative feedback early on. And they kind of went back to the drawing board, figured out why was it a problem? And they actually, yes, there were some technical things that they could tweak and things they could do to improve the voice. A lot of it came down to good script writing. And ultimately, is it a good description or not? That's what's going to keep people involved and engaged. So I think they take a stance, and I can't speak for them. But the script writing is so critical for good description, more so than the voice even.

Read the full transcript

22:03Christopher Antunes:Tell me more about the script writing. I'm really ignorant about it. So you're saying AI is maybe helpful, but it's not fully kind of used yet. Or like, what's a good script writing look like? And then if you have that original script, is the translation also slightly different? Or is it just kind of just a normal translation done by a pro? Audio description, and we'll talk about translated audio description for global language. Audio description has three components. You have a video that you receive, and then there's three steps. You write a script to describe what's happening, only the relevant things happening in the video for someone who can't see.

22:41Christopher Antunes:So imagine you're watching that movie and you need the scenes described. So script writing, step one. Step two, voice generation. So take that script and translate it to voice. And then three, mixing. So ultimately mix that back with the original video. All three of those steps can be done by people. They can be done by money rounds of people, or they could be done entirely by AI or a mix. I would say most commonly today, we will have an expert write the script, maybe assisted by AI. So we might use an AI draft first in some cases, but ultimately someone is responsible for editing or creating from scratch or approving that script.

23:23Christopher Antunes:And then we'll use voice technology for the voice, and then we'll automate the mix. So two out of three steps, automation or AI. So where's the state of the art in AI script writing? I would say, like in all tasks, the AI is better than people at some things and much, much worse at other things. And one of the things that's worse at is judgment. And we certainly want someone with really good judgment to be the final arbiter of what stays in that description. I mean, we know about hallucinations and things where the AI can sort of go off the rails a bit there. But I think one area where it is very good is in factual identification of things in videos.

24:12Christopher Antunes:The font might be really small and hard to read, or there might be a logo. Imagine someone doesn't know without a lot of research. An AI model can usually detect that thing and describe it quite well. and the judgment part and then we can move on but like the the judgment part would be it would describe things that a human just would know that okay don't describe the tree behind the building

24:33Josh Miller:because nobody cares exactly behind the building it's the pertinence it's the it's the the context and kind of how important is that item or that object or that movement in the scene for someone to understand what's going on just because a car is driving by it's not the fact that it's a car it's the fact that it is the red car that was in the scene before right and just understanding that that's the important part of it is really important.

24:56Christopher Antunes:All right. Let's go to the dubbing you mentioned. I think I checked in Q2 2024, so roughly 18 months ago, you launched your human in the loop dubbing solution. So how has it been past one and a half years? Traction, clients, experiences?

25:13Josh Miller:We've got all of the above. I think we've learned more than we had in the last year and a half. I think we've learned more than we had in the last five years about, you know, this industry, certainly. And just even in product development, I mean, we've, we've, we built a product that I think was really good for some of our customers, but not all of our customers. And I think we, initially, and in the last few months, we've completely changed that to, you know, to really go after what has become a really emerging media demand, which I think is fascinating. So to make that more clear, we have a ton of customers in the e-learning space and the corporate space who are looking at this as an opportunity to distribute content when they had never been able to really localize their content in this type of engaging way.

26:06Josh Miller:What their standard is for good dubbing is completely different from what a media company would consider good dubbing. And a lot of that comes down to their anchoring. What are they used to? Nothing versus something really good. And so that was a very real, both a good thing and a bad thing for how we were building the product.

26:27Christopher Antunes:If you zoom out with this whole dubbing market and kind of what's happening here, on the high end, you have obviously traditional dubbing, theatrical level dubbing, where obviously, Florian, you know, you could be paying$400, $500,$600 per minute to dub something into German or Korean or a language that's difficult. and on the other extreme with full AI and AI dubbing is as a category means a lot of things but AI only with no human intervention at any step you might be paying cents. Our bet is there's a big market in the middle and there's sort of two categories in that market I'd say there and Josh alluded to them.

27:10Christopher Antunes:There's first time dubbers. There's companies, There's whole use cases. There's whole markets that never could afford the economics. We started with talking about unit economics and building a profitable business could never afford 300, 400,$100 a minute, just out of scope. But their brand matters to them and they're not interested in cents. So we're discovering, I'd say together, and I think the market is largely in a discovery phase now on both directions, builders and buyers. We're all discovering together. We're discovering what the right price is. We're discovering what the right quality target is.

27:46Christopher Antunes:And we're discovering how we build that scalably for them. And scalable and repeatable quality is a big topic. You can do it once for a demo. Can you do it a thousand times? Right. And so I think that market is still in the discovery phase. We're figuring out expectations. We're figuring out even the price points because there's a wide range. And what's exciting about 3Play, or for me at 3Play, is the dynamism in this market. it. There are so many use cases, so much configurability, and it's so much more complex. And I love this problem and I can geek out over it for hours. The product and recovery engineer in me will love to do that because it's so complex.

28:25Christopher Antunes:It's not an AI or person choice. There's literally 10 different steps strung together with multiple AI and tooling and expert stops along the way. And I think ultimately that's a big advantage for 3Play because we love complex problems that require orchestration like this. And dubbing is, in order of magnitude, more challenging, I'd say than captioning more description for that reason. And what you said is really true, that a prototype, a demo, a pilot, it's just one thing, but then you want to do it at scale, it all breaks down, right? Even some of the platforms, I log in, I spin up my little video, and then it takes, even that little video takes 30 minutes to render right now.

29:03Christopher Antunes:That's one five minute clip. I guess if you want to do this at scale and 30 language becomes a problem that you seem to like. I struggle with this. I'd love to like, you know, grab a beer someday and dive into this. I struggle with this as across any business. I think demos in a lot of cases, whenever, if there's certainly if there's human in the loop, or if there's expert in the loop in like a model like ours, the whole demo process is broken. What do you learn dubbing one video that's three minutes long into one language. You learn that someone can perform all sorts of magic that's unnatural to produce that three minute video, but you don't learn anything about what the experience will look like in production at any sort of real scale.

29:51Christopher Antunes:So how do you figure that out together in some sort of testing or trial or POC process? That's an unsolved problem. And it's one I think we all need to think about together. Yeah.

30:00Josh Miller:I mean, we saw some of that issue just in captioning, But I agree, it's very much magnified in dubbing when there are so many different steps. And there are really experienced providers out there who know how to throw a lot of people out of problem. And again, going back to our origin, we're not here to throw a lot of people out of problem. We'll throw people out of problem, but in really, really thoughtful ways. And the whole premise here is how do we look at those 10 or 15 tasks that are required to create a good dub and understand one for the customer in the use case, which tasks apply first, and then two, which tasks require human intervention.

30:42Josh Miller:And really, really think that think through that carefully. I mean, we're seeing it now in some of the demos with media companies, and some of them are starting to figure out, and I applaud them for this, how to try to weed out, you know, kind of some of the fake it till you make it, which is, we're not going to tell you what kind of content we're going to send you. And we're going to give it to you on this date and we're going to put a deadline on it, right? And so there's only so much fake, throw a lot of people at it that you can do when you can't really prepare for it.

31:13Christopher Antunes:All right, can we talk about the European Accessibility Act, the EAA, so I don't know, the TLDR, where are we at the various stages of its rollout? I'm based in Europe, but I'm in Switzerland, so I'm slightly detached from the European Union. So tell me a bit more, how big of a deal is it? How does it impact the North American market, if at all? And how much of an opportunity is it for companies like yours?

31:40Josh Miller:It's a great question. I mean, the quick answer is, I think we're in the early days, but we're learning a lot. And so it's certain, we're seeing different approaches from different companies. We started to hear early on from certainly our enterprise clients who have business in Europe, they need to be prepared. And I think for them, it's a little more straightforward, to be honest, it's it's a little easier to say, Okay, we do business here, we have to be accessible, our websites have to be accessible, the videos that support our product launches have to be accessible. Pretty straightforward. That's pretty easy.

32:11Josh Miller:Now, what we don't know yet is the enforcement around that, that's still very early, you know, it just went into effect, we've seen a couple of sightings, but there's not a whole lot of activity yet, because I think there's, you know, there's a reasonable approach here that like, let's give people a chance to get it right and not go too crazy. But I think we'll see soon, more examples being made of companies that ignore it. The media industry is a little bit more diverse, I'd say in their approach. We have a number of companies who we're working with today who've said, we're not screwing this up.

32:49Josh Miller:We're not going to make a mistake. Let's go. We've got to get ready. And we've got to start implementing. We had a couple companies that we were working with literally on day zero. They were ready for the launch of the EAA in June this year. We have other companies that we work with who say, look, we need to know what's possible. We need to be ready, but we're going to wait and see. And a lot of that is that there's some different interpretation of the law itself, that this is where I'm going to say we'd like to play Switzerland and just be not take a side necessarily, but try to be an educator and share what we understand.

33:28Josh Miller:And there are some legitimate ways to interpret the law the way it's written to the point that the lawmakers have actually gone back and said, we messed up, we weren't clear enough. And so what that means is there's some discrepancy around, is it the supporting tools around the media that must be accessible or is it the media itself that needs to be accessible? It is vague enough to interpret either one. Now, a number of media companies have said it's not really up to the EAA to make that decision. It's actually it should fall under the European Media Freedom Act, the EMFA, to make that call. And there are conversations right now about that legislation being updated to kind of take over where the EAA left off.

34:22Josh Miller:But I'd say we're early days. But I think there is a common thread that pretty much every media company is looking at this and trying to figure out how can we be ready.

34:32Christopher Antunes:And Florian, I mean, obviously, like the regulatory environment is more complex because every member nation has their own policies and enforcement practice. And there's a lot to figure out there still. But just like to make it concrete in the localization space, what we're talking about around video that I think still needs to be, you know, ultimately determined. And I think there'll be more guidance coming from the body itself. And I think there'll be enforcement that helps to set precedent. But really, it's if I'm a streaming platform and I've been distributing content in pick any European country and I've been dubbing into 40 languages.

35:08Christopher Antunes:Well, all 40 of those dubs now, in theory, create a requirement to be closed captioned and audio described. That is a significant problem at scale for a lot of companies that have significant distribution in Europe, even if they're North American based.

35:24Josh Miller:That's right. And so what's important there is that this is an unplanned budgeting exercise. And so we go back to what we were talking about before, this idea of synthetic voice being used for multilingual audio description. This is where it's not a square peg round hole anymore. It's square peg square hole because it is so important that they can continue to distribute this content because it's a moneymaker for them and they need to figure out how to do it. And it can only continue to be positive ROI if they have a reasonable approach budget-wise to make it all accessible. To your initial question, Florian, the thing for us that's so interesting is we're still based in North America predominantly.

36:07Josh Miller:We are really working with U.S.-based and Canadian companies who are extending their reach into Europe. we have barely scratched the surface of working directly with European-based companies, which would be the next phase of our growth in terms of the EAA and kind of where that goes. So, you know, we're still early ourselves.

36:27Christopher Antunes:Come to SlaterCon in London. It's London. It's not the EU anymore, but there's a lot of Europeans going to be there. So you guys are a bit of a competitor to ours because you also publish research. You published the state of ASR in 2025 report. report and I looked it up and it says you looked at more than a thousand video files, over 200 hours of content and 1.7 million words. So tell us the key findings of this. So this is something we do, like, you know, and we're definitely not a competitor on the research side. So this is something we did internally. This is a theme with 3Boy. This is something we did internally for a long time.

37:08Christopher Antunes:And we just thought it might be of interest to the market. And we're really biased, but in a really good way. And I'll explain what that means. We consume ASR all day. Like I mentioned, 30 million videos processed, tens of thousands of videos every day. And on the captioning side, our first step is ASR. And then we put that ASR in front of people, and they clean it up. And we are more invested than pretty much anyone in the world. and getting the absolute best ASR so we can speed up that task for people. So we are religious about testing every model. And this is true for all of our human in the loop or expert in the loop services, because the real cost isn't in the AI.

37:54Christopher Antunes:The real cost to us and ultimately to our customers is in the time it takes to get the AI to an acceptable level of quality. So we measure that everywhere and we're religious about driving it down. So we're testing all the engines anyways, so we know which engine to use and in which use case. And so really all we do is take the results of that test and share them with the world because we think it would be of interest. I would say this year, the main finding is that there was a material leap forward, not like an earth shattering one, but there was a material leap forward about a year ago with the generative models entering the mix with the whisper model.

38:32Christopher Antunes:So the LLM or generative AI models entering the mix assembly and whisper most notably. This year, the results kind of plateaued. So we had a big leap forward, but largely the results were the same as they were a year ago. And for English recorded audio, speech recognition is pretty good now. Pretty good. pretty good it's pretty good and i think we all experience that and it it still is highly dependent on scenario so the example i always give which you know i resonates with sports fans at least is if you have you know jim nance at the masters this golf tournament professionally recorded in a studio single speaker that speech recognition is going to be pretty close to perfect um if you're transcribing a focus group for let's say a major cpg company and they've hidden the microphone in a plant somewhere and there's 10 people speaking over each other, the speech recognition is useless.

39:33Christopher Antunes:It's still situation dependent. But for the clean audio single speaker cases, the speech recognition has gotten really good. It just didn't get much better last year.

39:41Josh Miller:Yeah, I think what we're also seeing is, you know, we have the large language model phenomenon, which is great. There's been a lot of talk, and I think appropriately so, about like small language models, right? And so kind of getting more focused on the content itself. And we do this automatically already, the idea of can you build customer-specific kind of vocabulary sets and models and domain-specific models, things like that, that can actually target the type of content you're getting. And I think that's where we're seeing improvements. It's not the broad ASR off the shelf that you can get now.

40:17Josh Miller:So that's where we go back to and believe deeply that ASR is a great tool. AI is a great tool, depending on how it's implemented. And it takes someone pretty sophisticated, I think, to actually get a ton of value out of it, as opposed to some value out of it.

40:33Christopher Antunes:For the AI that you're using generally across the business, for ASR, for dubbing, for all of that, I think you're agnostic, right? Engine, model agnostic, the infrastructure is built, so it doesn't really matter which model leads you can plug in. Yeah. And I think that's, I mean, we thought hard about this. We have a data science team. We have a head of data science who has deep experience in NLP, built speech recognition models in a number of stops before joining. And we thought hard about this, but we really think of it as a feature and not a bug. And I think that's proven true, especially in the last 18 months with the like just massive and rapid improvement and dissemination or propagation of new models.

41:13Christopher Antunes:So the first reason we do it is we, like I said before, all of the cost really here for us is in the people, time and the time on task. And so we need to have the best engine for every task all the time to minimize that time on task to produce the quality we need. And that so we invest in that. Second, because our whole business is starting with AI and sort of cleaning it up, we have tons of data naturally. So we don't have to invest in very expensive evaluation harnesses. We can literally take a new model, plug it in, an A-B test, model A against model B, and we'll get 10 ,000 data points tomorrow.

42:00Christopher Antunes:Because we literally process that much content. So we have like sort of built in rapid learning environment. And we have an ecosystem, a marketplace of thousands and thousands of experts who can essentially do that evaluation for us in a real world context. So it makes a lot of sense in our environment. So let's close on, I don't know, growth strategy. So initially, you said you took on some angel and then, you know, a lot of profitable growth. And then in 2019, you took on some outside investment, I think PE firm, growth PE. What about M &A? like anything in the past and any plans for the future?

42:34Christopher Antunes:What would make sense for you to acquire?

42:36Josh Miller:It's interesting. Literally on day two of working with a private equity firm for us, which was six years ago, they're starting to show you opportunities. So it is built into the process of how do we grow fast, right? So organic or inorganic, their game for anything. So that was a big change for us is to think that way. And so over the years, I mean, we've had plenty of conversations. You know, for us, it needs to be a really good fit. It has to make sense for our initiatives. So there's no perfect answer to say absolutely. Like what we can say is we don't have a mandate to go acquire companies to grow.

43:14Josh Miller:We have a lot of great organic growth in front of us right now with everything we're doing, which we're really excited about. More so than we've seen in a long time with everything that's happening, which is really great. but we're always open-minded. We always look at what we're doing and try to look at the landscape and even lots of companies put themselves up for sale. We are always looking to be thoughtful and if it can add a ton of value, we're going to think about it. If it can't, then that's fine too but it's something that we're always evaluating but it must be one thing or another.

43:49Christopher Antunes:And I would say our sort of growth strategy is not an organic growth strategy where we want to buy a lot of things and that's the way we'll sort of generate revenue and generate profit. But if we see interesting opportunities that align with our roadmap, we're going to look at all them. And as we look out into next year, we're making a lot of investments on the data science side. We're making a lot of investments on localization expertise to surround real localization subject matter experts around this thing we're building so that we can make it as exceptional as we possibly can. That's an area of interest.

44:23Christopher Antunes:there's specific use cases you know we're pretty wide in terms of the use cases we serve for all these products but there's some like that um like you know the the youtube content creator use case is really interesting on the dubbing side now interesting so there's areas like that where there's some companies that are small doing really interesting things that might be interesting targets for us as well i've been waiting for this to like become like almost a mainstream i don't know vertical or horizontal or whatever you want to call it the youtube creator like you know some of them are very big they will eventually want to go international the youtube ai dubs aren't great and uh but the logic is just so obvious like yes we could talk for another two hours just about that this is a this is definitely a focus area for us i mean just just thinking like you can go to youtube and you click and you have 10 audio tracks uh like like on netflix like how obvious can it

45:15Josh Miller:be that and from a content publisher perspective i mean when you think about dubbing and localization at the core of it to monetize the localized content, you have to have a way to distribute that content in another country. What YouTube solves for everyone immediately is distribution. It is a phenomenal concept that people maybe don't even appreciate enough, but literally you have to do nothing to start monetizing other than localizing your content. Whereas we talk with some of our existing customers who say, yeah, we're thinking about expanding into Europe we have to set up a sales team there and sell the content there and find partners there i mean that's gonna be two years before they're ready to actually dub their content whereas with the youtube creator or even a media company publishing on youtube it's it's ready to go it's it's you need

46:06Christopher Antunes:2 000 subscribers i think the thing to solve there and i think you hit the nail on the head like um another another nice feature here is there's a single decision maker it is the content creator are almost always making a call, not a committee of buyers. And they are moving fast, certainly relative to traditional media companies, and they'll take on risk, right? And they'll iterate and they'll learn quickly. And a lot of them operate like almost like at a bigger scale, maybe not the 2000 subscriber ones, but the ones with 20 million operate like little tech, like technology companies. They're rapidly learning.

46:39Christopher Antunes:They're making really bold and quick decisions. I mean, like you said, like if you can prove that this makes sense at a certain quality level and opens up a market so they can monetize it in one language, why not go into 50? There's an economic model to figure out who pays. Do the advertisers pay? Do the content producers pay? Does a vendor like 3Play pay? And there's a licensing arrangement. There's a lot to work out. But I agree with you, Florian. It's definitely a no-brainer and it's coming. So you mentioned a couple of things you're planning for 2026. Anything to add? Any technical cool things you're planning you can you can share yeah so the thing i'm most excited about you know into 26 and as we go out further and like this is more like commentary on the ai economy kind of writ large here that's still that's still emerging but what we're seeing uh with with a lot of our customers i think it's really resonating is um and we think we're going to figure out how to what to name this thing but i'll describe what it is which is our system today like we said it's AI first, and then it's tools and people, and then we sort of deliver those results back.

47:47Christopher Antunes:But what we don't often talk about between AI and tools is analytics and scoring. And so what I mean by that, let's take it in the captioning setting because it's the simplest. Dubbing's a little more complex. But if you have an AI model, and I just mentioned AI can be really good for English captioning if you have a single speaker, we score it. And we score it for internal use. we need to know how much to pay that captioner. If they only have to do a very little to get that to the quality level we need, you don't have to pay as much. If they have to clean up lots and lots and it's going to take them hours, we need to pay them more.

48:24Christopher Antunes:So we've built all these predictive models that tell you how good the AI is and also tell you where the AI fails and where it doesn't. And we use those internally. Well, the realization we've had over the last year or two, as our customers have had macroeconomic climates, made them more budget sensitive, things like that. They would love to see those analytics themselves so they could use them. So in a captioning setting, what this looks like is we run an AI model for you and then we give you a score. And then we give the operational decision and keys over to you. So you decide, is the AI good enough?

48:58Christopher Antunes:Is it not good enough? Do you want us to upgrade? So this concept of AI plus scores that we give you In analytics, we give you to have you make a decision, help you make a good decision about how to allocate your budget and sometimes upgrade with our solution. And even when you upgrade, giving optionality there. So we started this conversation saying there's a big difference between back catalog e-learning and Netflix original theatrical. We call that the dial. So even when you upgrade, you should be able to set the dial exactly where you want to optimize your budget. So shifting this model to AI and then scoring and letting you decide, do you need experts at all?

49:41Christopher Antunes:And if yes, where do you want to set the dial? And doing that across all the services. I would say with captioning today in English, there is a lot of experimentation with just AI. I'll just say. And there is lots of upgrades still. On the dubbing side, it's all expert in the loop. It's all human in the loop for any serious buyers. But we still want to give you the analytics. We still want to give you the scoring so you can see and you can learn. And as it evolves and gets better, we want to be transparent about that with the customers so they can see the evolution of the technology in real time.

50:18Josh Miller:I think that, which I'm very excited about, it applies to even other markets where education, for example, with ADA Title II in the US, they now have to caption and describe way more thousands of hours more content than they're used to this is the only way to do it is is to have more of an ai pass with with intelligence to understand what really matters the other thing i'd say is we are building something that i will i'll speak for myself i'm very excited about in terms of what we're doing with localization dubbing is a great example we've we've not done a great job of telling that story yet and i think the world is going to learn what we've built um this is a great start to that.

50:59Josh Miller:It was a good start. I was going to say it. But we have not really shown everyone what we are capable of in localization, and people don't know who we are as a localization provider still. And I think that's going to change over the next year, and I think people might be surprised what we've been able to build.

51:15Christopher Antunes:All right. Thank you so much for taking the time today. This was a good first step in telling the story. That's right. Thank you, Florian. Thank you for catching up in person. Thank you.

51:27you

From the publisher

Josh Miller and Christopher Antunes, Co-Founders and co-CEOs of 3Play Media, join SlatorPod to talk about the company’s trajectory as a leading language solutions integrator (LSI) in multilingual video accessibility.

The duo explains how the two met at MIT, where an early challenge from OpenCourseWare revealed that captioning thousands of technical videos was financially impossible, leading to the company’s founding, where they developed proprietary tooling, leveraged AI, and incorporated expert-in-the-loop solutions.

Josh describes how their platform evolved into a dual system supporting both customers and large-scale operations. Chris notes that the LSI now serves media and entertainment, higher education, e-learning, and corporate clients.

Chris explains that three major trends — the European Accessibility Act (EAA), advancements in voice technology, and the rise of live events — drove their expansion into global localization. 

The co-CEOs detail their dubbing journey, noting rapid learning over the last 18 months and the emergence of a big mid-tier market between high-end theatrical dubbing and low-cost AI-only output.

Josh explains how the EAA is pushing companies to prepare for large columns of multilingual captioning and audio description. He notes that interpretations of the law still vary, but major media firms are already investing to avoid disruption. 

The duo shares findings from their 2025 State of ASR Report, where they found that accuracy initially improved sharply with generative models but has now plateaued.

Looking to the future, the co-CEOs are working on shifting their model to incorporate AI-generated scores and analytics, allowing customers to decide on the level of expert intervention.

More from SlatorPod

All 39 episodes
#273 A Big New Market for Dubbing and Accessibility Solutions with 3Play Media co-CEOsSlatorPod · 51 min
Listen in VO