#296 Yeop Lee: How Coxwave is Redefining AI Evaluation

26 Oct 2025 · 43 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Episode Summary

Episode Title

#296 Yeop Lee: How Coxwave is Redefining AI Evaluation

Podcast Overview

  • Host: Craig S. Smith (New York Times correspondent)
  • Focus: Interviews with individuals making significant contributions in the field of artificial intelligence, analyzing incremental advances within a global context.

Episode Description Yeop Lee, Head of Product at Coxwave, discusses the innovative evaluation platform, Align, which aims to move beyond traditional accuracy metrics in AI evaluation. The conversation delves into user satisfaction, trust, task completion, and how Coxwave's solutions can guide product teams in enhancing user experience with AI.

---

Key Discussion Points

Introduction to Coxwave

  • Coxwave’s Product: Align, an analytics solution for AI-powered conversational products.
  • Often described as "Google Analytics" for conversational AI.
  • Focuses on conversations between users and AI to assess their effectiveness.

Transition in AI Evaluation Strategies

  • Shift from Accuracy-Only Metrics: Moving towards outcome-focused evaluation to provide deeper insights into user satisfaction and trust.
  • Traditional evaluation methods focus on objective metrics that assess whether AI outputs meet predetermined criteria.
  • Coxwave introduces subjective evaluation—assessing whether AI responses are contextually appropriate for varying user needs.

Key Features and Metrics of Align

  • Evaluation Focus:
  • Measures satisfaction, trust, and task completion across multiple communication channels (chat, email, voice).
  • Utilizes both LLMs (Large Language Models) and human reviewers to assess AI performance.
  • Important Metrics:
  • Completion Rate
  • Customer Satisfaction (CSAT)
  • User Retention
  • Cost per Resolution

Analytical Capabilities

  • Search Engine for Conversational Data:
  • Allows querying to identify user sentiments, such as distrust or dissatisfaction, even if not explicitly stated.
  • Real-Time Monitoring:
  • Near real-time updates for SaaS customers, while on-premise customers receive batch updates due to hardware constraints.

Deployment Options

  • SaaS vs. On-Premise:
  • Availability of both a SaaS solution for general customers and an on-premise option for enterprises with stringent security and compliance needs.

Client Interaction and Continuous Improvement

  • Feedback Loop:
  • Continuous improvement based on client feedback to enhance the accuracy and efficiency of the evaluation process.
  • Consulting Services:
  • Offering guidance for clients on fixing identified issues, adapting AI models, and improving user interactions.

Market Insights

  • Growing Demand:
  • The need for AI evaluation tools is increasing as more companies adopt generative AI for internal and customer-facing applications.
  • Differentiation:
  • Coxwave distinguishes itself through its subjective evaluation approach, focusing on personalized user interactions.

Future Directions

  • Multimodal Analysis:
  • Future development plans include analyzing interactions across different modalities (text, voice, images) to provide comprehensive analytics.
  • Personalization Challenges:
  • Addressing how to implement user feedback effectively in AI systems, ensuring that the user experience remains seamless and intuitive.

Conclusion

  • Global Expansion:
  • Coxwave is planning a more active push into the North American market, building on its existing presence in Korea and India.
  • Strategic Partnerships:
  • Collaborations with foundation model companies like Anthropic to enhance the capabilities of its evaluation tools.

---

Key Takeaways

  • Coxwave’s Align is a paradigm shift in AI evaluation, focusing on user experience rather than just accuracy.
  • The integration of subjective evaluation metrics is crucial for understanding AI's effectiveness in real-world applications.
  • Continuous interaction with clients helps in iterating and improving the product, fostering a collaborative growth environment.
  • The AI evaluation market is expanding rapidly, with a growing emphasis on personalized user experience as the industry evolves.

---

Stay connected

  • Craig Smith on X: [@craigss](https://x.com/craigss)
  • Eye on A.I. on X: [@EyeOn_AI](https://x.com/EyeOn_AI)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Coxwave, our product Align, is an analytics solution for LLM powered conversational AI products. So the way we think about it is we're the Google Analytics for products like Chachi, Petir, Claude. We've always like visited the U.S. to understand, OK, what are the sort of developments in Silicon Valley? How are builders approaching building these new types of products and so forth? But I think this year and then especially next year is when we'll more actively push into the North American markets. Build the future of multi-agent software with agency. That's A-G-N-T-C-Y. Now an open source Linux Foundation project, Agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework.

0:50All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and more than 75 other supporting companies to build next generation AI infrastructure together. Agency is dropping code, specs, and services, no strings attached.

1:49Visit agency.org to contribute. you. That's A-G-N-T-C-Y dot O-R-G. Hey, I'm Yop. I'm the head of product at Coxwave. I joined Coxwave, I think, a little over two years back, but I knew the team since they were founded in 2021. The way that I got to know Coxwave was actually, so I met the founder of Coxwave, his name is Geejung, at a government competition. So this was when Coxwave was just starting off. And at that time, I had just left Reed, and I had started my own company. And so we were both just getting started off. And that was sort of during the peak of COVID as well. So this government program was like a three-day program, but none of the participants were allowed to interact with each other.

2:42Everyone was sort of locked in, everyone's sort of separate room and we had to watch sort of each other's presentations over screen. But, you know, Gijang had a wonderful presentation sort of outlining the vision of Coxwave. And I had given my presentation. And during the award ceremony, we connected just, you know, swapping business cards. And during that competition, Gijang Coxwave got first place. My company at that time got second place. And so that's sort of how the relationship started. And then we sort of went our separate ways building up each, you know, I built up my own company, he was continuously building up Coxwave.

3:22And then I had the opportunity to sell my company in 2023. And that was also when Gijang got the offer from a public Korean company to sell two of the Gen AI products that he was building. So he had built Gen AI products since 2021. And this Korean public company, because they saw sort of the rise of Chad Chippetee and the potential of Gen AI, wanted to acquire these two products to sort of leverage internally for themselves. And that's sort of when I joined Coxwave, because Gijang and Coxwave was transitioning from B2C to the B2B space. So that's sort of the journey. Yes. Yeah. And the products that he sold, you mean Coxwave no longer owns them or do you mean they were their customers of Coxwave?

4:15So we no longer own those two products. So I can give a brief introduction of those two products as well. The first one was a product called Hama, which is an image editing tool. So, you know, like sort of background removing, et cetera. And so it was used by a lot of companies in Korea that were sourcing, like, for example, foreign products. And they wanted to edit out different sort of labels and, you know, edit it in Korean. So that was sort of a big use case. And the second product was a product called Enerpix. Enerpix was a search engine for generative AI, you know, created images. So it was an image search engine.

4:56And so the company that acquired these products, they're Korea's largest font company. So they're sort of like the default font company in Korea. And they wanted to leverage this technology for their own sort of B2C use cases. Oh, and so the premise of CoxWave today is a platform that LLM companies can use to evaluate their models or explain exactly what CoxWave is doing. So Coxwave, our product Align, is an analytics solution for LLM-powered conversational AI products. So the way we think about it is we're the Google Analytics for products like ChatGPT or Claude. So what we focus on specifically is the conversations between a user and an LLM.

5:54So we actually focus on the interaction that's happening to help businesses and product managers understand, is the chatbot working properly? To our users having a good or bad experience? If not, why are they having a bad experience? How can we fix this and so forth? And so I'm sure you've seen a lot of these evaluation platforms that have been coming up. And I can sort of explain a little bit about the way we differentiate with these sort of other evaluation platforms. So these evaluation platforms that exist today, they focus on an area called objective evaluation. So this is sort of, you know, a product manager intends the LLM to say a certain thing and you evaluate, does the LLM actually say that in a bunch of different scenarios, right?

6:39Because you don't want the LLM to just, you know, go off rails. where we focus on is okay let's say that the LLM is already saying what the product manager intends the LLM to say but that response may be a good response for a certain user but also may be a really bad response for someone else right I'll just give you a quick example like let's say there are two users they ask the same question what is the best food in Korea and the LLM responds, oh, the best food is samgyepsal, which is a type of pork, right? But one user says, oh, thank you so much for letting me know. But the other says, oh, but like what restaurant sells this?

7:22You need to tell me the restaurant, right? So in that case, for the second user, the response wasn't sufficient enough. So it wasn't the best sort of most ideal response, right? So that's sort of the areas where we pinpoint and analyze to help businesses understand, okay, is this certain LLM response really a good response for this user? What is an optimal response? And so forth. So that's sort of an area we call subjective evaluation and feedback analytics. Yeah, it's an interesting area because initially, you know, there's reinforcement learning with human feedback to try and guide LLMs to answer appropriately and not hallucinate.

8:13And that's been successful to a degree. And then there are, you know, now there are systems, multiple models that talk to each other for providing an answer and the reasoning models and all that. And for a time, evaluation relied on benchmarks. And there were all these benchmarks that got increasingly complex. But users figured out or the market figured out pretty quickly that you can train to a benchmark. So, you know, Grok3 comes out and they show their chart and they do, you know, so much better on this or that benchmark than any of the others. It doesn't mean that the user experience is better.

9:13So this is, as you said, did you call it a subjective evaluation? Correct, yeah. Yeah. So, you know, from the point of view of a user, do you train, I would guess, that this is done by evaluation LLMs that talk to the primary LLM, and then there's some scoring or something. Is that right? Or how much human intervention is there in the judging? Yeah. So we also leverage LLM as a judge as well. But then we do have human interactions that are built into our evaluation system. So for example, when our customer is onboarding, our system will do initial judging. And it'll show those examples. Like, is this an example of user dissatisfaction or not?

10:17And the reason why we do this is even if, for example, we just use sort of like a standard model, depending on the conversational product, right, it's very different. And so what we realized is this area of subjective evaluation ultimately is personalization, right, to the user. But that also means our platform and our product has to be personalized to our customers as well. And so that sort of type of interaction is built in. And one of the areas where we focused on a lot is our search engine, which is a search engine for conversational data. So we have a natural language interface where if you type, for example, I want to find messages where the user is expressing distrust against the chat box, then our search engine will go across the conversational data and identify that.

11:10And this is for even conversations where the user isn't explicitly saying, I don't trust you. But any sort of implicit feedback that you can extract from the conversation will identify as well. And so that's sort of the system that we've built in place for our analytics system. Yeah. And when you say there's a conversational interface, is that for the customer, for Coxway's customer to query your platform about the LLM it's evaluating? So that's another layer? Is that right? That's correct. That's correct. Yeah. And you would have to evaluate that LLM as well. Correct. Yeah. So what we do is. It's kind of a house of mirrors.

12:05That's true. Yes. So we work very closely with our customers to help them understand the performance of our search engine as well as our Copod so that it ultimately is up to par to their expectations. and initially like honestly like right out of the box like for example the accuracy is let's say 80 percent but we all but the system gets better and better as our customers leave more feedback as well because our search engine right for example same example of user is expressing distrust right there may be new cases of user distrust but our system needs to learn as well to understand okay for this company, for this use case, what is this trust?

12:50What does it look like, right? So we always tell our customers, if you use our product more, if you give more feedback, it'll get better and better, and so forth. Yeah, and how does someone use this? Is this a SaaS product that they access through the web, or is it an API that they build into their system? describe how that works. Yeah. So we have a SaaS product. As you mentioned, they can just go on the web. They can log in and just leverage our platform. We have an SDK that helps our customers ingest the conversational data into a line so that it's available for analytics. The other is we have a on-premise self-hosted solution as well.

13:39So these are for our enterprise customers, you know, who value privacy, security. They have more rigid data governments requirements, etc. So we'll provide that option as well. And we are looking to expand into an API option too, but it's in our product roadmap. Yeah. And how many conversations, does this run live as well? If you run a customer service AI system, Will it monitor the reactions or the exchanges live and then maybe signal to somebody that it's not, the customer service bot is not answering properly or is creating problems? And then the more important question is, after the customer knows this, what do they do?

14:39How do you correct it? Yeah. So for our SaaS product, it runs near real time. For enterprise ones, we usually run it in batch, especially because there are some limitations on the hardware, etc., right, of our enterprise customers as it's on premise. for the fixing part. So our job is one to sort of identify, help them monitor, analyze the impact to help them understand and prioritize what are the issues that they should really be focusing on. Right. And then in terms of the fixing part, because like our customers, they have many different ways that they're building these systems. Some are just fine tuning all the models.

15:23Some are just, you know, using prompt engineering and so forth. So we can provide guidance. We have a separate consulting service that helps provide this guidance. But other times, our customers just go on and fix it themselves. And they'll understand when that sort of fix has been made. And they can add that data into the metadata that they're logging into line. So they'll understand, okay, after this change has been made, did the sort of problems disappear? Or are these problems still happening and so forth. Yeah. How many conversational interactions do you typically screen with a customer? And then it's an ongoing thing, presumably from your point of view, that's what you guys want it to be.

16:15Or do companies get to a point where they feel pretty confident in their model and stop evaluating.

16:28It's such a huge range from thousands of conversations to millions of conversations per day, especially because companies that are running contact centers, for example, with AI, there's a lot of volume that's happening. There are other companies that are running character chats that have a lot of daily usage as well and so forth. Up till now, the customers that we've engaged in, they haven't left us. And the reason is this sort of evaluation and analytics is it's never over, especially because new LLMs come up, things that are working before stop working, right? And as users, more and more users come into the platform and there are new sort of user engagements, this type of new user behavior leads to new problems that need to be analyzed, need to be fixed, and so forth.

17:22And again, these sort of new insights can propel new feature ideas, new sort of interaction ideas, and so forth. And these require constant monitoring, analytics, and evaluation. Yeah. Yeah, this issue, you're working, I presume, primarily with application companies, companies that are building applications on top of either open source model that they've fine-tuned or that they've hooked up to a vector database with the RAG system? Or do you also work with the foundation model companies themselves, companies that are coming out with their own models, listing them on hugging face and hoping to get a piece of that pie that's dominated obviously by OpenAI and Anthropic and a handful of others?

18:33Yeah. Yeah. So as you mentioned, we mainly work with the application layer companies. And these are also enterprises that are building these applications for their internal use cases, too. But we've begun to expand our work with foundation model companies as well. You mentioned Anthropik. So we're one of Anthropik's closest partners in Korea. So a couple of months back, we actually hosted Asia's first builder summit. So we had the CPO of Anthropik, Mike Krieger, in Korea as well. And we had this sort of huge event that we had hosted. So that sort of work is something that is ongoing as well. As Anthropik is also looking to expand its presence in Asia and Korea, as well as a company called 2AI.

19:20They are a foundation model company with a presence in Korea, India, the US. And they have a consumer product called Chat Sutra. Also, they offered their foundation model as via API to enterprise customers as well. So we work closely with them too. Yeah, yeah. I've had two AI on the podcast. Yeah. Yeah. And are most of the applications that you're working with, are the models conversing in Korean? Or is it predominantly English, even though your market presumably is East Asia? So most of the language is actually in English. And very interestingly, when we first launched Align, we made the decision initially to not support Korean.

20:14And so our product was just in English. And our main target market was actually we started off with India. And so most of our customers, enterprise customers as well, are based out of India. And our sort of support for the Korean markets, including like localizing our product, supporting the Korean language and so forth, started early this year. Like before that, we have Korean customers, but they were kind of forced to just use the English version of our product. But yeah, now we're seeing more and more, you know, Korean customers come on board. And the reason for this was actually quite, quite clear.

20:57It was because when we first launched our product, there just weren't as many Korean Gen AI products in the market. And so we had to, as a startup, right, pick and choose and really focus. And so that sort of led us to the Indian markets. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux foundation project, Agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting.

21:56Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. collaborate with developers from cisco dell technologies google cloud oracle red hat and more than 75 other supporting companies to build next generation ai infrastructure together agency is dropping code specs and services no strings attached visit agency.org to contribute that's A-G-N-T-C-Y dot O-R-G. And you mentioned Align Align AI is the flagship product from Coxwave. Is that right? And what LLM do you guys use on the conversational side and then on the evaluation side in in looking at the conversational data that comes in from a customer?

23:22Yeah. So we use a bunch of different models, but then our main is actually Anthropics cloud models. Also explains kind of why we have a good relationship, partnership with Anthropic. We also have our own custom embedding models that we've trained to be very specifically good at retrieving conversational data. So that's sort of our own internal technology. And also, there are like a bunch of different like smaller analytics tasks. And depending on the task, we've sort of found which model works the best. And we've sort of put that LLM inside. And so some do run on open AI models as well. Yeah. And the foundation field is changing quickly.

24:16I mean, it seems to have slowed down a little bit, But do you keep sort of rotating in the latest foundation model so that you're keeping up with the market and not getting stale? I've always wondered how companies handle that. Yeah. So that's always like a big challenge as well. But what we do is we don't always switch out for the newest model, but we do like very rigorous internal testing to understand, OK, with these new models, how do they perform against the current models that we're leveraging on these specific analytics tasks? But because what we found is, you know, like even these like newer models, like sometimes they're not the best at the analytics tasks that we need completed.

25:08Also for our enterprise customers, switching out for the newest models every single time, that change always introduces new risk as well. So we'll always go through that internal evaluation process. And sometimes we'll make that switch over, and sometimes we'll stick to the model that we're using. Yeah. You've been around, what did you say, since 2021?

25:38Yeah. So with the company, I started off, I think, early 2023. Yeah. And then this company has, Coxwave has been around since 2020. That's what I meant. Yeah. Yeah. The, initially Coxwave, the two products that they sold, was that Hama and Enerpix? Yes, correct. How am I in Interfix? That's what you said, yeah. And you've recently raised money and are putting all of your focus on Align AI. Is that right? That's correct. Yeah. And is this a crowded market? I mean, everyone's talking about evaluation, but I haven't spoken to many companies in the evaluation space. Is it a growing market? I mean, how do you guys see the market globally?

26:47So it definitely is a growing market, especially as like enterprises and more and more companies start leveraging generative AI not only for their internal use cases, but for customer facing use cases. So it is growing. And the need is very clear, especially because as a company and builder, if you can't trust the output of your LLM, then you can't really introduce it to your customer, right? Especially for a lot of your brands. The way we differentiate, though, is that area of subjective analytics. And honestly, we haven't seen that many players in that space yet. But as the sort of world moves more and more into hyper-personization, where these LLM systems are hyper-tuned to that specific individual, right, and very flexible, adaptable, this area will become more and more important.

27:39Yeah. Yeah. Is it, how affordable is the platform for companies? Is it something that's within reach of a relatively small startup or is it really focused on, you know, Fortune 500 or Fortune 1000? Yeah. So we do have a SaaS plan that's$499 per month. So it is more affordable And then we also do have like a design partner plan for earlier companies, especially because at that time, these companies are just starting to understand how their users are interacting with their products. And we want to be good sort of growth partners for them as well. Yeah. Why did you settle on Anthropic as your base model?

28:36is it through an evaluation or was it based on business relationships and that sort of thing? Yeah. So it was on both. So one, we internally evaluated the cloud models and they performed really well on the analytics and evaluations tasks. And then the second was the sort of shared philosophy that both our companies have on, you know, building safe, responsible, reliable AI, Especially Anthropic, you know, they have the concept of constitutional AI as well. Right. And so those are very important values, especially when we were going to our enterprise customers. Like we wanted to have a foundation model partner that shared the same values as us that we could trust and so forth.

29:19So that also played a big role. Yeah. Do you work with like Hugging Face to provide evaluations there? I mean, it seems their leaderboard is really crowdsourced, but there's all kinds of problems with crowdsourcing. I mean, it can be gamed pretty easily. Do you work with Hugging Face or do you publish rankings of LLMs that are in the market? So we don't work with Hugging Face yet or publish that, especially because what we're analyzing is not a specific model, but we're evaluating the product itself. So, yeah, not yet. But one thing that we do do, though, is we work very closely with consulting companies.

30:14And so one of our closest partners is PwC India. They were our first enterprise customer as well as enterprise partner. especially because, for example, I'll just give you the example of PwC India. They have 30 ,000 employees in India alone, and they have so many verticals that they're working on. And it was very clear that generative AI can be very transformative, not only for their internal use cases, but they're also seeing a lot of demand from their customers as well on how they can leverage generative AI. And so we've been a very close partner to them, providing not only analytics and evaluation, but also, as I mentioned earlier about, okay, if these problems are there, how can we fix it?

31:00So we provide that sort of consulting layer service as well. I see. And so where is this going? I mean, this field is, as you said, it's growing. It's becoming increasingly integrated into the global economy. The number of LLM-based applications has exploded. I mean, is it just from your guys' point of view right now, just working to build market share? Or is there something on the horizon that you see as a new product that Coxwave is looking to launch? Yeah. So increasing market share, of course, is very important. But another piece is a multimodal analysis. So nowadays, these right now we focus mainly on text.

31:59So most of these are chatbots. But then now, you know, even with like ChatGPT, you go across different modalities. And so providing that layer of analytics for these multimodal products, helping businesses understand, for example, okay, for this user, like when do they prefer one modality over the other? Are there some modality features that aren't performing as well and so forth to help these companies make those sort of feature design choices as well? So that's a very important thing on the roadmap. Another thing is this thing about personalization, right? The reason why I think it's such an interesting problem but also very complex is let's say you get all this sort of user feedback that's coming in and you understand this user.

32:47but with that information you can't always like make that change immediately into the product right so for example I'll just like for example when I use chat GPT sometimes there's this problem of with chat GPT memory where it's just remembering everything there are things that I don't want it to remember and actually make that change and some of my colleagues they've made the decision to just shut off the memory feature because that gets really annoying, right? So that type of decision and how we can sort of help product managers make that decision in a more informed way, right? What are things that should be remembered and changed?

33:28What are things that users probably don't want changed? And actually helping them make those decisions, that's a very important thing on the horizon for us. Yeah, actually, that's interesting you mentioned that. That problem seems to be relatively new and getting worse. even if you start a new instance or new conversation or a new session, sometimes it'll mix the answers from your previous conversation into the current one, which you're right it can be uh can be annoying and you have to repeat to the model that i've stopped talking about that i'm talking about something new yeah uh yeah forget about that i didn't realize you could turn off the memory feature but uh yeah uh have you guys looked at excuse me like uh grok 3 i mean do you have customers that are building applications on uh on the grok models um so far off the top of my head i don't think we have customers building on the grok models uh mainly open ai anthropic and then for um our enterprise customers they are usually building on top of like llama models or models and then there are like large enterprises in korea that have their own models so lg has their own model samsung has their own gauss model um skt has their own model as well so they've been trying to figure out okay how can they leverage these internal models that they've built and just bring up the sort of performance on that Yeah.

Read the full transcript

35:27Does Korea have a national model? Some countries are building sovereign models, whatever you want to call them. Yeah. So right now we don't, but our recently elected president, Lee Jae-myung, he's pushing for this as well. So I think there's a new, almost like$100 billion AI fund, like$80 to$100 billion AI fund that'll be used to build sovereign AI as well as just supporting startups, ecosystems, and so forth. Yeah. Yeah. And I mean, I've been impressed by the Korean startups that I've spoken to because there's so much government support for startups, as you were saying. That's how you met the founder.

36:21Is that increasing? And do you pay much attention to how things are organized in the United States where it's largely a private exercise or private initiatives that fund AI startups? Yeah. Yeah. So we pay attention to the U.S. as well. And then the government support has been increasing, too, especially with the sort of rise of AI. The government is looking for new partners that they can work with to help, you know, not only build sovereign AI, but just provide this level of support to the startup community. So recently, there was a new government position created. I think it's like a policy advisor position for AI.

37:16But the person who's currently in that advisory position, he used to work for Naver. And I believe he worked on the foundation model at Naver. And so he has an understanding of the role startups should play, how this ecosystem can play out and so forth. So just by having that new position created in the government, that also provides like a very strong channel and voice that startups can, you know, present to the government as well. Another thing that we're doing as well is we and a couple of other companies, including Ritten, which is Korea's largest B2C Gen AI company, Liner, etc. We created the Generative AI Startup Association.

37:59And so this is an association of Gen AI startups and companies coming together. So we organize a bunch of different events to share knowledge, learnings across one another, as well as discussing with government policymakers to understand, okay, what are their priorities? What are things that us as startups really want to see changed in the market? How can we work together to create sort of the good future that we all envision? Yeah. Are most of the startups focused on the Korean Japanese markets or are most of them focused on the English language speaking market? Right. Yeah. So I think traditionally it has more been focused on Korean, Japanese, Asian markets.

38:52But now you're seeing more and more sort of English speaking, U.S. facing products as well. And I think the reason for that is actually because of these like foundation models and how good they are with languages. In the past, like the language of Korean has been a big barrier for large companies from abroad coming in. But now with these LLMs, like Claude speaking Korean is just so natural that language is no longer a barrier. So for these Korean companies, now if you don't go to the U.S. markets or these larger markets, there is just no way of surviving. So it ultimately comes down to survival.

39:31Yeah. And on that, on the multilingual LLMs, I use, because I spent a lot of my life in China, I work between English and Chinese with various LLMs. is the Korean to English as, I mean, can Claude, is it 2.7? 3.7. 3.7, I always get that wrong. Does Claude 3.7 speak Korean as fluently as, for example, it does Chinese? Yeah, it's so fluent that last time like an anthropic researcher came over for the Builder Summit, we asked him, does Anthropic just have a lot of Korean teammates? How is Claude so good at Korean? And even they're surprised that it's so good at Korean. But where in China did you live?

40:37I lived for like 13 years in China as well. Oh, yeah. I saw that you had spent time in Shanghai. I was, yeah, I was the New York Times bureau chief in Shanghai for quite a while. And then I was in Beijing running the business side of the New York Times. Since my audience is not only, I mean, there are 150 countries or something that listen. But most of the listeners are in the U.S. and Canada and the U.K., I mean the English-speaking markets. How big has your push been into these markets? So up until now, we've been focusing primarily on the Indian markets as well as Korea. But now we're starting our push for the U.S.

41:38markets. So up till now, like we've always like visited the U.S. to understand, OK, what are the sort of developments in Silicon Valley? Like how are builders approaching building these new types of products and so forth? But I think this year and then especially next year is when we'll more actively push into the North American markets as well.

42:04Thank you.

42:34Thank you.

From the publisher

This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents. 

Visit https://agntcy.org/ and add your support.


How is Coxwave Redefining AI Evaluation?

In this episode of Eye on AI, host Craig Smith is joined by Yeop Lee, Head of Product at Coxwave. Together they explore how teams move beyond accuracy-only metrics to outcome focused evaluation with Coxwave’s Align. We look at how Align measures satisfaction, trust, and task completion across chat, email, and voice, how LLM as judge pairs with human review, and how product teams search conversations to find hidden failure patterns that block adoption.

Learn how leading companies design an evaluation stack that guides prompts, agents, and UX, which pitfalls to avoid when shipping updates, and which metrics matter most for success, including completion rate, CSAT, retention, and cost per resolution. You will also hear how to run experiment tracking with model and prompt change logs, set up governance that prevents regressions, and choose between SaaS and on premise deployments that meet security and compliance needs.

Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI

More from Eye On A.I.

All 266 episodes
#296 Yeop Lee: How Coxwave is Redefining AI EvaluationEye On A.I. · 43 min
Listen in VO