964: In Case You Missed It in January 2026

6 Feb 2026 · 25 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

January 2026 “In Case You Missed It” roundup (episodes 962–957, 961, 959). It covers (1) “secret cyborgs” using AI at work, (2) disappointment with enterprise “agents,” (3) distributed artificial superintelligence (DASI) via language, (4) AI evaluation frameworks beyond accuracy, and (5) incentives to retain top developers/data scientists.

Guests/backgrounds

Ethan Mollick (best-selling author of Co-Intelligence; Wharton/Worden professor). Sadie St. Lawrence (Founder/CEO, Human Machine Collaboration Institute). Dr. Vijoy Pandey (Cisco; distributed ASI). Sinan Ozdemer (O’Reilly educator; AI entrepreneur; author). Ashwin Rajiva (co-founder/CTO, Excel Data; raised $100M+ VC).

Key claims/examples

Over 50% of Americans report using AI at work; “secret cyborgs” save 20–70% time and may boost quality, but organizations lack processes and incentives for disclosure. Agents hype exceeds enterprise readiness (examples: Salesforce “Agentforce,” SAP). DASI thesis: intelligence scales horizontally through language enabling shared intent/knowledge/innovation. Evaluation: use task-specific metrics (generation vs retrieval/classification) and precision vs recall based on false-positive/false-negative costs (LinkedIn engagement example). Retention: engineers need mission/innovation freedom; leadership should ensure teams can build.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

AI Usage in Workplaces

0:45 to 2:37

Discussion on Ethan Mollick's insights about AI use in workplaces.

“Do you think that this 20 % to 70 % has continued to accelerate in the past two years since you originally started writing about secret cyber?”

Disappointments of 2025

2:37 to 3:41

Discussion on the highs and lows of 2025, with Sadie St. Lawrence.

“My next clip is from episode number 955.”

The State of AI Agents

3:41 to 5:48

Conversation about the effectiveness and implementation of AI agents in enterprises.

“No, so I wrote a post this year as my best performing stuff called agents of disappointment.”

Apple's AI Performance

5:48 to 7:03

Discussion on Apple's AI advancements and disappointments in the past year.

“I once made an emoji and then didn't do it again.”

Distributed Artificial Superintelligence

7:03 to 11:41

Dr. Vijoy Pandey discusses the concept of distributed artificial superintelligence.

“And now we're jumping straight into some of the most exciting frontier tech for 2026.”

AI Evaluation Frameworks

11:41 to 14:01

Sinan Ozdemer explains the importance of evaluation frameworks in AI.

“A multi-agent human society sounds great, so let's hear from the people who are helping us to reach that goal.”

Evaluating ML Models: Metrics Matter

14:01 to 15:12

Learn about the importance of task-specific metrics in evaluating machine learning models.

“because how I evaluate a child's essay on a catcher in the rye is going to be different than how I evaluate the embeddings that this embedding model is producing.”

Precision and Recall: Key Metrics Explained

15:13 to 17:37

Understand the differences between precision and recall and their implications for model evaluation.

“I'm now going to walk through 20 case studies that are all very different from each other, all with different metrics.”

Discussion on Incentives and Risks

17:38 to 18:11

Explore the relationship between risks of failing tasks and the incentives for developers.

“If you say this factory part off the line is good, but it's not good, is a plane going down or is someone's light going to break?”

Finding and Keeping Top Talent in Tech

18:12 to 22:30

Discover strategies for retaining talented programmers and fostering innovation in tech.

“full circle to incentives with my final clip, which I'm taking from episode number 957.”
Show all 11 chapters

Hiring for Passion: What to Look For

22:31 to 24:48

Learn what qualities to seek in candidates to identify those passionate about technical work.

“I think it's easy to tell in some sense.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:This is episode number 964, our In Case You Missed It in January episode. Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. This is an In Case You Missed It episode that highlights the best parts of conversations we had on the show over the past month. In my first clip from episode number 962, I talked to Ethan Mollick about the actual number of people using AI to help them at work. You might know the highly sought-after Worden professor as the best-selling author of Co-Intelligence. And in this extract, you'll hear Ethan Mollick address how both individuals and companies can benefit from their workers being upfront about their use of AI.

0:42Jon Krohn:A couple of years ago, you identified secret cyborgs, individuals in organizations who leverage AI for time savings of 20 % to 70 % on many tasks while maintaining or increasing the quality of their work product. Do you think that this 20 % to 70 % has continued to accelerate in the past two years since you originally started writing about secret cyber?

1:03Ethan Mollick:And we have some evidence on this. Over 50 % of Americans said they used AI at work, right? And probably more actually have. They self-report that on a fifth of tasks that they use AI for, that they are seeing a three times performance improvement. That's the self-report. Whether that's true or not is hard to know. What's slowing that down from organizations is we don't have the process needed to make that, you know, to like, what do you do? You're running agile development, someone gets all their code done right away. Like what's the point of their stand up? What's the point of, how do you work sprint planning around that?

1:32Ethan Mollick:How do you think, like, what do we do about that stuff? Or they're just not telling you they're using it because they're just misincentivized.

1:38Jon Krohn:Yeah, yeah, yeah. So if these people, if these secret cyborgs are in your organization, what can we be doing to surface them and to be taking advantage of what they're doing? maybe have what they're doing be taught, be less secret and be taught to other people who have a whole army of cyborgs.

1:55Ethan Mollick:Well, this is where the leadership and lab come in. So your secret cyborgs come out of your crowd, the people in your organization doing things. Your leadership needs to incentivize people to actually tell you this. If people think that they're going to be fired or punished or other people will be fired because they're showing productivity gains, they're just not going to show you. If they're working 90 % less, they're not going to want to give that up for free. So leadership needs to think about the incentive plan that puts this into place. And then you need the lab because you need somewhere for these people to go.

2:19Ethan Mollick:And then, you know, either it's a reward to actually work in the lab or else to say, hey, I've got this prompt that kind of works and saves me five hours a day. Could you make it good and get it out to everybody? So it's not just a one component piece. You need the other pieces.

2:31Jon Krohn:Incentivizing workers rather than punishing them sounds like the best way forward for an AI positive future. My next clip is from episode number 955. This episode rounded up 2025 and marked the beginning of 2026. Together with the brilliant Sadie St. Lawrence, we shared thoughts on the highs and lows of the year that was. And Sadie, who serves as founder and CEO of the Human Machine Collaboration Institute, made one disappointment of hers very, very clear. All right. Disappointment of the year. It's interesting because you earlier in this episode, you used a company name and the word disappointment in the same sentence.

3:13so is that this one may be controversial i think you're gonna disagree with me on this john but i'll explain myself and so i'm just gonna say for me actually agents were disappointing and i wrote a sub stack this year called agents of disappointment

3:29Jon Krohn:take her off the air take her where's where's the abort button if all of a sudden i get muted you'll know why right all right and that brings us to the end of the episode. Yes, exactly. No, so I wrote a post this year as my best performing stuff called agents of disappointment. But really what I was talking about was the divide between like the hype of agents. And really where I saw this come into disappointment was from the enterprise companies. You know, we have the obviously Salesforce has been talking about its agents force for some time. You have SAP, you have all of these enterprise companies who have been talking about how to implement agents into your existing enterprise tools and it it just doesn't work right like or at least a lot of companies aren't set up for it properly or in my mind don't know how to think about how to structure them properly and what to have agents do and so from that standpoint from like an enterprise agent standpoint i think the hype and the practicality of the implementation the divide between those two was too great.

4:38And so that was my disappointment of 2025.

4:41Jon Krohn:I totally get it. And I don't disagree with you. It makes a lot of sense to me. There's too much talk about agentic AI relative to the impact that it's making. No question. I do think that a lot of it is related to people not having their data silos set up in a way or their security set up in a way where they're comfortable with it. But yeah, a lot of tinkering with agents, not nearly as many enterprise deployments. But I do think it will come. That is not my disappointment of the year. My disappointment of the year is Apple. I think you were disappointed with Apple last year. We got to go back.

5:22Let's look.

5:24Jon Krohn:Oh. I think, yeah, because Apple, did they announce Apple Intelligence last year? Or was that a this year thing? I don't know. Time is weird in the AI world. I think you're right. I think they announced Apple Intelligence in the autumn, Northern Hemisphere autumn of last year. And it was disappointing. But I mean, I guess it's kind of just like Google. It's like - And it's still disappointing. It's still, because it's like a year later. And what can I do additionally with AI in my phone that I actually use? Not much. I once made an emoji and then didn't do it again. you know just like used it built in gen ai to send an emoji over iMessage but i don't know i don't i don't really need to do that like there's there's yeah it seems like things like just having siri be be able to understand what i'm saying to it in the same way that opening i whisper can.

6:24I mean, you know, who's actually even done better than that is. So I will actually, I will use Grok in my Tesla. And so when I'm coming home from work and want to like brainstorm something or just learn a new subject, like I just push a button and it works seamlessly and the chat mode goes back and forth. Right. And I'm like, okay, if I can do that in my car, why can't I do it on my phone? That's seamlessly. So I hope that like next year, our comeback of the year is Apple because, you know, they've been in the disappointment quarter for far too long.

6:55Jon Krohn:Yeah, and they have a lot of potential because of how ubiquitous their devices are. There's a big opportunity for Apple if they can get it right. And now we're jumping straight into some of the most exciting frontier tech for 2026. In episode number 961, Cisco's Dr. Vijoy Pandey talks about distributed artificial superintelligence and how that idea is deeply rooted in the development of human language. Awesome. So now that we have this definition of artificial superintelligence under our belts, Vijoy, tell us about this idea of distributed artificial superintelligence, or as I've seen you abbreviating it, D-A-S-I.

7:35Right. So if you think about human intelligence, because one thing that we have in common moving forward is the comparison bar for all of us is human intelligence. And whether you think about artificial intelligence, AGI, ASI, powerful AI, whichever definition you might have, we just talked about two definitions, one which is economic in nature, one is technical in nature. There seems to be always this comparison metric, which is let's compare it against humans, because that's the best thing that we know when it comes to intelligence. So if you think about humans and how we evolved and how our intelligence evolved.

8:13We actually evolved intelligence across two axes. So the first 300 ,000, 400 ,000 years, human intelligence was actually scaling up vertically. So we were getting smarter and smarter. We were inventing tools. We were inventing processes, but it was limited because we weren't communicating that intelligence. So the communication was a big missing piece. And so what ended up happening was whatever we invented and whatever processes that we came up with and the dangers that we were aware of and how we reacted to those dangers and the way we stitched our clothes together or whatever we wore. I mean, I'm not an expert there, but whatever we did was limited to the lifetime of either that individual or that process.

9:02And so we became more and more intelligent, but it was very, very limited. And so we were scaling intelligence vertically. But as we all know, every system, including intelligence, can be scaled on two axes, vertically as well as horizontally. And so there was this big evolutionary jump around 70 ,000 years ago. It's called the cognitive evolution in humans, where we discovered language, not just sounds, not just paintings and patterns, but language. How do you convey meaning, semantics between people? And then how do you convey that across humans, but across tribes and across generations? So what happened when that cognitive evolution happened was we invented three things.

9:54The first thing being shared intent. So as a human society, we started sharing a common intent. Let's go and build this, not just in the lifetime of me as a person, but in the lifetime of this tribe or this group of people. The shared intent and coordination as a result. The second thing that we invented was shared knowledge. So this is what's colloquially known as standing on the shoulders of giants. So I build a knowledge base, then you add to it, then you add to it or you modify it and you keep doing that. and that is cumulative human knowledge. So we invented shared knowledge. And then the third thing, because of the first two, is now we could do shared innovation.

10:41So innovation itself wasn't a singular pursuit or an individual pursuit, but it was a shared pursuit. So that's what happened when language got invented and semantics got invented. And you started scaling horizontally because now you're inventing as a collective instead of as an individual. And so what we're seeing and the big thesis here is that so far in intelligence, in artificial intelligence, in artificial superintelligence, we've been building bigger and bigger individual geniuses. And the framework and the infrastructure to do collective intelligence, to do distributed intelligence has been missing.

11:25And that's what we want to go after. So distributed intelligence to us is to enable, to build a framework that allows for shared intent, shared knowledge, and shared innovation to happen in this multi-agent human society.

11:41Jon Krohn:A multi-agent human society sounds great, so let's hear from the people who are helping us to reach that goal. Evaluative frameworks are one way to get clear about how effective our AI systems are, and O 'Reilly educator, many-time best-selling author and AI entrepreneur, Sinan Ozdemer, explains that finding a common language for those frameworks will be crucial to their success. Here's a clip from episode number 959. Speaking of experimentation and research, let's jump to that chapter three that you mentioned earlier, the fun one. I mean, there's lots of fun chapters, but chapter three is particularly a good one.

12:14Jon Krohn:And in it, you present your comprehensive framework for AI evaluation. So you emphasize that accuracy alone isn't enough. And so you introduce multiple metrics across different task types, retrieval, classification, generation, and you stress the importance of reproducible experiments. You conclude the chapter by noting from now on, we will be incorporating evaluation language into every case study. So then throughout the rest of the book, you have these fantastic detailed case studies that build on each other. And you use this common evaluation language that you introduce in chapter three throughout, which is brilliant.

12:51So when you organize this evaluation

12:56Jon Krohn:by task buckets, generation, multiple choice, embedding, classification, why is that separation so critical? Yeah. So the split is usually of the types of LM tasks, there's generative and understanding. And then under generative, there's multiple choice and free text. Meaning it's basically like auto-encoding versus auto-aggressive is how I tried to think about that analogously. Meaning if you're chalking to a chatbot or an agent, which is just a chatbot with tools, you're asking it either to produce a paragraph, a sentence, several paragraphs, whatever, free text, or you're asking it to pick from a set of options.

13:36Should I proceed? Yes or no? is this good enough to post on LinkedIn? Yes or no? That's multiple choice. I'm basically collapsing the entirety of this deep learning architecture into a binary classification task. So versus understanding tasks, which are embeddings and classifications, which are similar to multiple choice, but just with a different architecture. Each one of those has their own suite of metrics because how I evaluate a child's essay on a catcher in the rye is going to be different than how I evaluate the embeddings that this embedding model is producing. They're just not the same task.

14:19They're not built for the same thing. They're all in LLMs. Open AI embeddings are produced by LLMs. Classification models are run, for the most part, by LLMs. So the evaluation is less on the model. It's more on the task that you're trying to perform. And whether you're performing classification through any kind of architecture. My book from 10 years ago, Principles of Data Science, talked about accuracy, precision, recall, sensitivity versus specificity. I also talk about that in my book from two months ago, Identic AI. It's the same classification that I'm asking an agent to do. It's the same task.

14:57It's just a different model is now doing it. So evaluation is tricky. It's the longest video I ever wrote or made. It was like nine, 10 hours on the O 'Reilly platform was evaluations because there is no one size fits all. It's what are you doing? I'm now going to walk through 20 case studies that are all very different from each other, all with different metrics.

15:18Jon Krohn:Yeah. And so to dig into this a little bit more, you mentioned there are all these different kinds of metrics for evaluating performance. So I already said in a question a few minutes ago, how accuracy isn't enough. In your book, you emphasize how using precision recall and for something, you know, where you're trying to rank results, something like mean reciprocal rank, MRR, using those metrics together because each exposes different failure modes of a model. Do you want to tell us a bit more about that? Yeah. So precision recall is probably the more, I would say, usable metrics for most people, meaning I'll say it this way.

15:58If you ask an LLM, you give it a LinkedIn post and you say, is this going to get a lot of engagement on LinkedIn? And it says, yes. Okay, great. You post it. It doesn't get a lot of engagement. That model had a false positive. It told you yes, but really it was no. When you care about false positives a lot, when they are expensive to you, you care about precision. Precision is the measurement of all the times the model said yes, how often can you trust it? So when the model says, yes, go ahead, how often is it correct in saying yes? That's precision. So when you care about false positives, precision is your metric.

16:47Recall is kind of the opposite. Recall is of all the times it should have said yes, how many times did it? So if false negatives are expensive to you, recall is the metric you care about. Because if the thing says, this is a terrible LinkedIn post, but you post it anyways, and it gets a lot of engagement, that's a false negative. It didn't want you to post that. And recall is a measurement, among other things. A recall is effectively a measurement of how many false negatives that you're seeing out of the system. So, and that was a pretty dense explanation for two, honestly, one of the simplest metrics in machine learning.

17:25And it kind of goes to show that the conversation around evaluation is not always as simple as, here's the fraction that you care about. It's, no, before we get to math, what do you, the human, care about? What's expensive to you? If you say this factory part off the line is good, but it's not good, is a plane going down or is someone's light going to break? how expensive is a false positive to you? If it's expensive, precision is the thing you need to look at. Recall shouldn't matter as much. I'll happily throw a part away on accident. At least if I know everything off the line is going to be right, precision matters the most.

18:04So again, it always comes down to not just the task, but even the risks of failing that task.

18:11Jon Krohn:From risks, we return full circle to incentives with my final clip, which I'm taking from episode number 957. In this episode, Ashwin Rajiva, who is the co-founder and CTO of Excel Data, a Bay Area AI startup that's raised over$100 million in venture capital, talked to me about how to find and keep the best developers and data scientists in their jobs. The kinds of things I was going to ask about in my next question, which were because you've said in past interviews that a talented programmer is looking for meaning in their day-to-day that, you know, highly paid engineers still just complain about their jobs.

18:47Jon Krohn:It's funny to me to think that like somebody who gets a hundred million dollars signing bonus at Meta is then just like, oh man. Um, but I'm sure that happens, you know, it's, I don't know what, where the stat is at today, but, um, when I was doing my PhD, something like 15 years ago, I attended a lecture on the economics of happiness, just for fun. I just went to this lecture. And at that time, it was showing that in the US, if a household was making over something like 80 ,000 or 100 ,000 US dollars a year, that is the happiest you can be. Making more money beyond that point didn't make people happier.

19:24Jon Krohn:Now with inflation, the numbers are probably a little bit higher. But directionally, I think this kind of gives the idea that you're explaining, which is that beyond having your basic needs taken care of and, you know, knowing that you have security for you and your loved ones, the extra money beyond that could end up being a hassle. Yeah, it is. It is. And it's also interesting, right? I mean, there are studies on developer productivity, right? And you would see that the numbers are insane. I mean, people talk about how an engineer, a software engineer is productive, maybe four hours a day or three hours a day.

20:01And the rest of the time is spent in meeting, planning, whatever it is, right? Now, I feel that even if you take two hours for meeting, you still are leaving a lot of this time out. And if you think about it, what better privilege can someone have than to sit in a usually a great office on a laptop without having to move? Moving is optional. Right. And then get paid top dollar for it. And most engineers then leave their jobs. You know, it's not just, you know, people leave Accelator, but people leave all sorts of companies. and there has to be a reason for it is that most people would be happy because knowledge work is something you know which is which which is which has to do with creativity it's hard to sit in one place and realize that the work which was presented to you or asked of you could be done maybe in two hours and then you got to sit and find something to do and it's good for a few days and you spend some time but after a while you start getting this feeling that hey what what am i doing i'm I have supposed to do something better.

21:05Let me find a mission which resonates with me and my work and my, you know, philosophy of it. And so I think making sure that no matter what the business environment you're in as an executive, the engineers or the R &D teams believe that fundamentally we are in the business of innovation in this field. and there are very few fields, whether it's, you know, something data management, even something as boring as the enterprise content management. I'm sure that there is innovation that could be done, new ways of doing things. And people should believe that they have the freedom to do it. And they're not just, you know, dictated by quarterly plans and this.

21:48I think if we can provide an environment like that, then new ideas come in. And, you know, for us, it's worked out because for a company of our age and size, we have a lot of, let's say, capability that we have built over the years. Whether it's to do with ODP, ADOC, we have a pulse monitoring system, we have ADM, we are working on the next version of our platform, which will be released in May. And so that is what allows us to do it, is where people believe that, hey, in this field of what the company has chosen, data management, there is innovation that can be driven through pure engineering work.

Read the full transcript

22:29And that's what drives people.

22:31Jon Krohn:Nice, I like that. How do you say in an interview, do you think you have a way of telling whether somebody is going to be passionate about a technical infrastructure-heavy mission like data management at Excel Data versus somebody who's just coming to collect a paycheck? I think it's easy to tell in some sense. Of course, we've made mistakes there as well, like everybody else. but I feel once you start talking to people the and this is what I felt I mean I've always felt that management in the technical field can only be done by people who have been in the trenches to some sense and I'm sure there are models everywhere else which are different And people have seen managers who work extremely well without actually being on the field.

23:31And so the number one thing, at least when I look for potential hires, is to see if they can build things. And it doesn't have to be working on some data problem. The question is, can you build? If you are given a problem, and it could be any problem. It could be something like, hey, how do you design, let's say, an e-commerce warehouse? How do you design a logistic system or any other business problem? And then can you translate it to something which you have learned? You probably know Go or Python or Java. can you put something together which represents a real world problem i think if you can then those are the people who bring the most value who can actually look at like a business problem and then convert it down to what they know and that's what technology is a good and of course this is for slightly senior people i think for people who are just coming out of college it's purely based on potential saying hey you know some of it is your scores and your background and some of it, hey, how interested are you into doing this?

24:43And then you take a bet and maybe after a few months you decide. But for most senior people, I would recommend checking if they can build things.

24:53Jon Krohn:All right, that's it for today's In Case You Missed It episode. To be sure not to miss any of our exciting upcoming episodes, subscribe to this podcast if you haven't already. But most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon. Thank you.

From the publisher

In this first of the year ICYMI episode, Jon Krohn selects his favorite moments from January’s SuperDataScience interviews. Listen to why incentivizing workers is the best way to get them to disclose their use of AI tools and pave the way for an AI-forward future, how AI continues to mimic human development in its own evolution, the importance of evaluation in building AI systems, and how to keep your best employees (and also: how to know your value) with guests Sadie St. Lawrence, Ashwin Rajeeva, Sinan Ozdemir, Vijoy Pandey, and Ethan Mollick.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/964⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
964: In Case You Missed It in January 2026Super Data Science: ML & AI Podcast with Jon Krohn · 25 min
Listen in VO