948: In Case You Missed It in November 2025

12 Dec 2025 · 29 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

A monthly “best of” roundup from Super Data Science Podcast (episode 948, Nov 2025), covering (1) IBM Granite 4.0’s hybrid state space + transformer design for long-context LLMs, (2) “zero trust for agents” using Cisco’s T-back permissions model, (3) preventing metric/BI “sprawl” and AI-driven metric drift, and (4) shifting AI product goals from addictive “stickiness” to human flourishing and creative empowerment.

Guests and backgrounds

Tyler Cox (distinguished engineer; discusses Granite 4.0 and state space models). Dr. Vijoy Pandey (developing Cisco’s open-source internet-of-agents platform; focuses on agent permissions/security). Mark Dupuis (Fabi.ai co-founder; prior product manager at Trasada, Clari, Assembled; discusses metrics/semantic layers). Maya Ackerman (Santa Clara University professor; discusses human-centered AI and creativity).

Key claims + notable examples

Granite 4.0 uses a 9:1 Mamba (state space) to attention layer ratio, claims linear context scaling and reduced memory (example: ~15GB vs ~80GB for 3B model at 8 sessions of 128K context on a micro). T-back enables just-in-time, task-scoped permissions with token revocation; requires identity/authorization support, semantic parsing of agent-human/tool communication, and sandboxed runtime. For BI, data teams should curate core metrics (ARR/churn/retention) while AI explores hypotheses before adding them to dashboards; otherwise AI can redefine metrics. For creativity, Ackerman argues against “oracle”/expected-output optimization; WaveAI-style co-creative tools should broaden vocabularies/styles, and enterprises should prioritize empowerment over stickiness.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction of Guests and Topic

0:18 to 0:44

Discussion about inviting guests to discuss state space models.

“First off, I invited Dels, Shirish Gupta, and Tyler Cox back to the show for episode number 939.”

Exploring State Space Models

0:44 to 6:00

In-depth conversation on state space models, their history, and applications in AI.

“Now I want to dig into state space models and hybrid architectures, which you touched on, Tyler.”

Zero Trust for AI Agents

6:00 to 11:30

Discussion of the importance of permissions and trust in AI agents.

“The other thing that Granite scores really well on is the Berkeley for function calling leaderboard, specifically BFCLV3.”

Evaluating Trust in Agents

11:30 to 14:00

Exploration of how agents prove their trustworthiness and manage evaluations.

“One thing that I guess I'm still, that I still don't quite get logistically.”

Navigating AI Metrics and Team Collaboration

14:00 to 21:23

Explore effective management of metrics in AI projects and team dynamics.

“And that trust can then feed back into the directory's reputation score and say, yep, all good to go.”

Human Interests in AI

21:23 to 21:56

Discuss the importance of keeping human welfare at the center of AI development.

“And in my final clip from The Month That Was, Santa Clara University professor Maya Ackerman and I talk about the importance of keeping human interests and welfare at the center of the conversation.”

Shifting Focus from Stickiness to Empowerment

21:56 to 28:00

Learn about the need to evolve engagement metrics towards human empowerment.

“And so how can investors and companies reframe their perception here, so reframe engagement metrics away from stickiness and towards sustained creative empowerment, kind of like you and David have with WaveAI?”

The Role of Creativity in AI Development

28:00 to 28:38

Explore how creativity impacts the construction of AI systems.

“Inherently, from a technical standpoint, you have to take creativity into account from the beginning in the way that you construct this machine brain.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00This is episode number 948, our In Case You Missed It in November episode.

0:07Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. This is an In Case You Missed It episode that highlights the best parts of conversations we had on the show over the past month. First off, I invited Dels, Shirish Gupta, and Tyler Cox back to the show for episode number 939. In this clip, I ask Tyler, a distinguished engineer, what state space models are. To give you a little context, in this clip, we're specifically discussing IBM's open-source Granite 4.0 language models, which include Mamba state space layers that theoretically allow infinitely large context windows.

0:43Here's my conversation with Tyler.

0:45Jon Krohn:Now I want to dig into state space models and hybrid architectures, which you touched on, Tyler. um so you know the granite 4.0 release introduced a hybrid architecture that combines state space models with transformers and that sounds important too but maybe maybe i've gone too far maybe we should start with digging a bit more into what state space models are first yeah yeah it's a so So state space models are a really important family of models, even pre-AI usage, right? So if you go back to the 1960s, these are being used in spaceflight control, right? They've been used in population studies and economic modeling for a long time.

1:32The core construct is that you have an input signal that you map onto a hidden or latent state space with one set of equations. And then you have a second set of equations that translate that state into an output that's observable. So there's lots of different ways that you can construct a state space model to represent different problems.

2:02What's happened over the last 10 years or so is that states-based models were investigated as part of the deep learning kind of revolution, is how do you construct your matrices for states-based models to better model different tasks without as much kind of classical feature engineering approach to it, right? So that was one important thing. But then there's a couple of researchers at Carnegie Mellon and Princeton who really kind of drove this home over the last five years or so. So there's a great set of papers. I invite everybody listening who wants to learn more to look up the work here on Mamba and structured state space models and things like that.

2:55There's a series of papers from 2021 all the way up to 2025, still working on it, that introduced some really great optimizations inside of state space models to make them appropriate for sequence transformation and language modeling tasks. So you get this evolution of structured state space models, your S4 paper, and then you go into structured state space models with selection and computation by scanning your S6, which turns into, hey, that's a lot of S's, sounds like a snake. Now we've got Mamba and the Mamba architecture really optimizes states based models for the kind of compute profile that's needed to be relevant for the language model tasks that we're applying here.

3:51And so some great optimization work that happened there, some great mathematical insights into the matrix properties of those. And I won't be able to do those full justice here, but really, really interesting work. That kind of accumulated into the Mamba and Mamba 2 language model blocks that IBM pulled in to the Granite 4H series models. Right. So composition here, you've got a nine to one ratio of Mamba layers to attention layers in the Granite 4H family. So quite a bit of state space model in that hybrid. Those are, like I said earlier, those are linear context scaling. They are the Granite 4H models are no positional embedding.

4:53So they have in the data set out to 512K context represented in the samples. They're validated out to 128K. The IBM team says, theoretically, you should be able to push it past that, right? So some really great long context performance. I think one of the key measurement points in the release notes are if you take eight sessions at 128K context on a micro, so a 3 billion parameter model, you get about a 15 gigabyte of memory usage versus about 80 on a pure transformer architecture, right? So some really great context reduction, which means that on more constrained devices, on the edge, you can make use of more useful context in RAG workflows, in multi-term workflows, and things of that nature.

5:55Just to close here, Sharish had mentioned IF eval. That's a really great benchmark for instruction following, right, structured output tasks. The other thing that Granite scores really well on is the Berkeley for function calling leaderboard, specifically BFCLV3. It shows up in the top five as we sit here recording, among a bunch of other frontier models and hundreds of billion parameter models, even at, for the small, 32 billion parameter footprint. So really, really punching above weight class there. Benchmarks have reappeared with considerable frequency in my interviews this year, and I expect this to continue into 2026, given the number of LLMs coming onto the market with ever more impressive capabilities.

6:48Permissions have been another big topic on the show this past year. This is especially important now that we're often working on projects that involve teams of AI agents working together, where we need to be careful about information flows between these agents and the tools they use. Dr. Vijoy Pandey is developing Cisco's open source platform for the internet of agents, tackling exactly these kinds of problems. In episode 941, I asked him about privacy and security when using AI agents.

7:17Jon Krohn:When we're working with something like T-back, it sounds like it might be tricky. I can't wrap my head around exactly how permissions are granted just in time. when you need to grant permission to an agent to be able to do a particular task and then you need to revoke that afterward, that sounds like it could be complicated. How do you handle it? We have to bring in a whole bunch of infrastructure around this notion of T-back to enable the end goal, which is the zero trust for agents, which is I give you permissions to do something specific for that just in time, just in time and then for that duration of that task to our transaction.

8:01And then I revoke that token or revoke those permissions the moment you're done. So what else do you need? So in my head, the equation runs like this, which is zero trust agency is I don't trust any agent and I just trust it once it's proven that it's supposed to do X, Y, or Z for the duration of that X, Y, or Z. and then I revoke those permissions. So that is zero trust for agents, is a combination of the availability of trust, a task tool, transaction-based access control. So all identity providers, authorization servers need to support T-back. That's step one. We need to have a parsing entity for all communication that's taking place between agents and agents and humans.

8:55that can parse that communication, that parse that discourse, and figure out the tasks or the tool access or the transactions that are taking place between agents. So that's the second piece, because that will help us define what those tasks, tool access, and transactions are. So there's a semantic parsing element. And then your just-in-time comment basically implies that you get a token, you do that task in a very contained sandbox jailed environment, and then you're taken out the moment you're done. So the analogy I draw here is I want to access a safe, which has a lot of money, but I want to withdraw$10 from the safe.

9:45Now you can give me, Since I have vJoy, you can give me access as vJoy to go and open that safe. And that's role-based. But then you can say, you know what, there are other people's money in that safe. So I'm going to give vJoy just enough to withdraw cash for the next 10 seconds and then move out. So that's like a task-based access control. But then I need to parse the communication that's happening between myself and somebody else, where we are talking about withdrawing$10 and say, okay, Vijay is allowed only to withdraw$10 and that is the task he's doing. So let me just give you access for that$10 withdrawal and not sit around to withdraw all of the money from the safe.

10:32So that's the task-based parsing that needs to happen. And finally, I'll let you in into the sandbox environment, give you that authorization to withdraw$10, and then I'm going to shut the door because I don't trust you beyond that point. I will not let you linger around. That's safe. So that is a sandbox runtime environment that needs to happen. So is the hooks in identity providers to provide task-based access control? Is a semantic parsing of the discourse of the communication to figure out what that task is? And then a runtime sandbox environment to just do that task with that authority and then get out.

11:15So those are the three things that need to come together for zero trust for agents to have.

11:20Jon Krohn:Semantic parsing, ephemeral runtimes, and human in the loop approvals. And overall, you gave me a really clear picture now of what this all involves. One thing that I guess I'm still, that I still don't quite get logistically. You said that you won't trust the agent even for the particular task tool or a single transaction that you're going to approve it for until the agent has proven itself. How do agents prove themselves trustworthy in the first place? This is where the entire pipeline comes from the picture. So we are looking at identity, which is the first stumbling block that everybody is running into.

12:02And so one of the things that we're seeing is in the agency framework, the identity piece is the problem to solve first, even before you can start deploying agents at scale within the enterprise. But then, as you pointed out, there are other aspects of trust. There are other aspects of semantic parsing. So there are these other aspects of the entire pipeline that we need to solve for. So coming back to trust, the simplest way you can start with is saying, is there a directory somewhere that allows me to discover agents that are trusted? So I want to find a financial agent, not from a particular vendor, maybe, but also if there are 10 vendors, I want the best of breed financial agent from a vendor that is highly reputable and highly trusted.

12:58So the simplest version of the question, the answer to your question is, is there a directory which tells me which agents are trusted agents? But then the next question is, so that's where the directory comes in. So we have a directory where you can discover agents through capabilities and through things like reputation and trust. But then the next question would be, how is that trust enforced or attributed? Is it crowdsourced? Crowdsourced could be one thing. So 10 people have used it. is like the Apple App Store, which says five stars, trust, or trust pilot score. And it's like, yes, it's awesome.

13:40But if you want to be a little bit more mathematical and provide some rigor, then you go to the last pillar of that four pillar thing that I talked about earlier, discovery, identity, communication, and evaluation. You look towards evaluations. And you look towards evaluating agents and multi-agent workflows and saying, over time, I've built trust by evaluating this agent. And that trust can then feed back into the directory's reputation score and say, yep, all good to go. Till something perturbs in the system, because these things are constantly evolving. From multiple agents, we move to multiple metrics.

14:20Is more always better? In episode 937, I speak to Fabi.ai co-founder Mark Dupuis about how to manage and evaluate several variables in a way that our teams don't lose sight of their goals and crucially, that the AI they're using stays on the same track as well.

14:36Jon Krohn:You have a decade of experience prior to co-founding Fabi as a product manager for lots of great companies, Trasada, Clari, Assembled. I'm probably mispronouncing all of those company names. And so at those companies, you build platforms, APIs, self-service analytics, drawing from that experience, when you're managing metrics, platforms like Fabi would allow product managers, like if you think about yourself in those previous roles, if you had a tool like Fabi, you would be able to very quickly spin up complex engagement metrics for your platform or your APIs on the fly. Would you be worried then about, or how would you prevent like metric sprawl where all of a sudden you're, you're, you just have tons of different measures flying around.

15:29Jon Krohn:You're not really like maybe people on the team that aren't going to be clear about what we're really building towards. Um, does that question make sense? Do you, do you see that as a potential problem or, or yeah, it makes a lot of sense. Yeah. The question makes perfect sense. Um, and I, I do think that like, there's always kind of this risk and also fear. and I think like justified fear that if you give everyone AI, then the AI is going to go and sort of like reinvent the metric every single time you ask a question. And so we could talk about semantic layers. That's probably like - Oh yeah, that I hadn't even thought.

16:01Jon Krohn:That's a huge problem too. Yeah, exactly. You could end up, you could have each person in the organization say, you know, just ask their AI enhanced BI tool, you know, how are we doing on this metric? And every time it's coming up with a new way and pulling different data, yeah, oh my goodness. Yeah, that's an even bigger potential problem here. Yeah. So that's maybe another episode for us to talk about semantic layers. And that sort of ties back to the guardrails as well. We talked about Fabian, make sure the team's actually supervising and collaborating and working with the business stakeholder.

16:33But to go back to your maybe original question, if that's not 100 % what you had in mind, which is like, how do you make sure that... And actually, let me make sure I ask the question back. is the question about how we, actually, let me ask you for clarification of the questions, make sure I understand, because that's what I was thinking about when you asked the question.

16:53Jon Krohn:Yeah, I mean, we should definitely answer the question that you brought up. But what I was thinking about, you would end up with the problem that you just brought up in what I described, where if project managers, if people across the organization, executives, product leads on individual products, If everybody can be using automated BI tools to be creating metrics on usage, you could end up having a lot of, it could start to muddy the water around what the organization at a whole is building towards. But even if there was agreement, so what I think is more interesting about the question that you heard is that even if you have agreement across everyone and humans are being consistent with their definitions and the humans all kind of have a clear idea of what they're building towards, what the metrics that they're trying to optimize are with the product that they're building, the AI could be surprising them by recalculating things a different way.

18:00Jon Krohn:So does that help kind of clarify what I was originally asking? Yeah, I think it does. And when you ask that question, I actually have two things that come to mind for me. The first is that the role of the data team, I don't think changes in the sense that data team still needs to be there to help make sure there are consistent metrics and that We're all working towards the same North Star as a whole, as an organization, as a department or whatever. And that's not changing, right? What you don't want is you don't want AI. This is why we talk about dashboards. I don't think dashboards are disappearing in that sense.

18:35I think that dashboards are actually going to be much more powerful and useful because there's going to be fewer of them. But the ones that are actually built are going to be curated and managed by the data team because they are tracking the actual core metrics that matter for the business. It's going to be tracking your churn, your ARR, your retention, or whatever it is. And the data team is going to be spending a lot more time making sure that those are correct. Now, when I think about AI for the business, what I think about is like all the other questions that like haven't yet made your way into your your North Star metrics or your OKRs.

19:08Right. So as a product manager, maybe going back to your original question, as a product manager, I and I still do today as a founder, product manager, I'm constantly exploring new ways to ask to think about the data. So, you know, you're thinking about your user activation, for example. Maybe at some point, if you're a mature organization that's growing, like your activation is very well defined. OK, it's seven. You have to friend seven people on Facebook and then that's your activation metric or that's our start. That's set. We're good. A lot of organizations don't have that or it's evolving or you're interested in your product.

19:39And so you don't want to go and like model what you what's effectively a hypothesis into your data and like pull in the data, create these new tables and then feed that and go through the entire process. and feed that through your BI solution before you've actually taken the time to figure out if that's actually what you want. So that's where you can, thanks to AI, actually let someone, ideally a pair of a data scientist and a product manager or a data scientist and a CSM or whoever, work together to kind of explore that messy phase and see like, okay, does this metric actually make sense? Is this how we want to think about it?

20:11And a lot of times I think what you'll find when you do that is, again, I'll draw my own experience to the product manager. You'll kind of look at the data and be like, hmm, actually that was the wrong question I was asking. Let me rethink about what I'm really asking and I'll get back to you. And if as a product manager, I can start asking my own questions off the data that they signed security for me or off the raw data, I'm going to get to my own answer much faster. And then we can also just sort of experiment and sit with that metric for a minute, for a month or a quarter. And then if it's like, okay, guys, like, you know, we're guys and gals, like we're looking at this metric and it's been the same metric for the last, you know, three months.

20:48Now, maybe it's the time for us to go and build this into, you know, add this to a dashboard. That's the point where you can say, OK, well, how do we pipe this data into our data warehouse? And like, how do we, you know, add this to our data modeling and feed it all the way through? So I think you just need to like, as you know, carefully think about like, OK, is this a metric that like we know that we've established and that is set? And if so, and like, let's go through like the quote unquote, like proper channels that have the right guardrails. If not, like, let's actually take the room to explore that before we go and overly invest in the measurement.

21:22So it's important that we return to thinking about the human motivations behind every AI project. And in my final clip from The Month That Was, Santa Clara University professor Maya Ackerman and I talk about the importance of keeping human interests and welfare at the center of the conversation.

21:38Jon Krohn:So you've previously argued that the sci-fi dream of AI as, yeah, I mean, this is, we talked about earlier as an infallible answer machine is, is disempowering and it's misguided. And in, in chapter 10 of your book, you note that disempowerment doesn't scale, but human flourishing does. And so how can investors and companies reframe their perception here, so reframe engagement metrics away from stickiness and towards sustained creative empowerment, kind of like you and David have with WaveAI? Yeah, it's interesting that this addiction model persists, even when the biggest problem with Gen.AI products is this kind of one hit wonder phenomenon.

22:26When a person feels really like they're really unnecessary, let's say there's like a little app where I upload my photos and then it does something fun with them. I might enjoy it once or twice, but if I feel like I'm not really doing anything, I don't have any control over what comes out, even if it's kind of cool, I'm very unlikely to come back. So we had a whole bunch of one hit wonder sort of Gen.AI products. People want to feel that they're doing something. People want to express themselves. They want to realize their own ideas. and industry and investors in particular have been very, very, very, very slow to realize that.

23:08I believe that the reason ChatGPT is successful is because it, for those who want to give of themselves, for those who want to really collaborate, it makes that possible. And that's why it's so successful. So it's sort of like, we need to sort of snap out of these old, outdated, exploitative practices that a lot of us sort of believe in them as gospel and really explore what's meaningful and appropriate with this technology, which I believe from my experience in business and otherwise that the real potential is to really elevate humans in our capacity.

23:49Jon Krohn:Yeah, it's a great idea. And I agree with it wholeheartedly. It does seem like it's going to be hard to convince product managers, you know, Do you have any sense? I realize this is a really tricky question. I don't have a good answer to this, but you're a lot more creative than I am. How do we convince big enterprises, product managers to move away from just stickiness to empowering people? I think we need to explain to them why something like, why the products that are successful, why they are successful. Chachapiti has not been beat yet. It's the number one. It's one of the most successful products in history.

24:30And that's the case, not because you press a button and it replaces you, but because of power users. You always have to look at power users because those are the ones who are, the customers that love you show you where the product needs to go. And those are the users that go deep, that become better as a result of using Chachapiti. I've heard people say that their vocabulary expanded. The writing style got more diverse as a result of using Chachapiti when you use it with a certain kind of intention. And also to show examples of how this sort of replacive paradigm has often failed in a lot of smaller products, but also how it ultimately fails collectively.

25:10This whole like everyone optimizes for themselves really fails when we want to apply AI to replace human workers. Because if we don't have any more human workers, the entire economy collapses. And so it's really time to wake up from these incredibly greedy principles that are guiding our economy. and to think more holistically in this moment, not assume that the old principles are just going to keep working and somehow magically everything is going to work out.

Read the full transcript

25:41Jon Krohn:Yeah, perhaps Gen AI will kind of force this shift that you're hoping for in product managers and in enterprises. You mentioned in your most recent response this idea of diverse responses. And so in interviews, you have previously contrasted convergent systems that gravitate towards safe, average outputs with divergent co-creative tools, presumably like your Lyric Studio and Melody Studio, that widen the search space and, like you just said, broaden people's vocabularies, increase the diversity of their writing styles. So as AI scales to creating music, to creating video? How do we keep models from nudging creators toward averaged, risk-averse aesthetics where, you know, you kind of, you hear another pop song that sounds the same?

26:38Jon Krohn:Yeah. How do we, how do we instead get diversity? Yeah. It's really impressive what industry has done to Gen.ai. Gen.ai has always been about expanding possibilities. That's how it was born. Going to new spaces. You know, we think of creativity as this massive search of possibilities. The next line of lyrics, there are billions of options. How do we search, right? You're about to play the piano. John, you were talking about playing other people's pieces, like we're all told. Where do we even start? What could be my first note? What could be my second note, it's a search space problem. And machines are phenomenal at just exploring different unlikely possibilities.

27:23And even in trying to take you to places that are pretty good, but not very common. The problem is that science fiction has convinced investors and some entrepreneurs that it's better to create an all-knowing oracle that gives you the most expected, most safe, most you've seen it a million times already answer, which is maybe good for a fact lookup table, but it really limits what Gen AI can do. And it's never actually going to be very good at being an all-knowing oracle, regardless of what we do. And so when you look at ChatGPT and you're like, oh, AI is not creative, that's because ChatGPT is optimized for the opposite.

28:04Inherently, from a technical standpoint, you have to take creativity into account from the beginning in the way that you construct this machine brain. But in the grand scheme of things, in the grand scheme of things, it's not that hard. If we figured it out as a team of three, granted we have David, but still a team of three people originally at Wave AI, then, you know, I'm sure somebody like OpenAI, Microsoft and Google can figure it out. It's just not their goal. Their goal is not to open our minds. Their goal is to replace search. And I think that goal needs to be modified. All right, that's it for today's In Case You Missed It episode.

28:43To be sure not to miss any of our exciting upcoming episodes, subscribe to this podcast if you haven't already. But most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there. and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

From the publisher

In this November episode of “In Case You Missed It” series, Jon Krohn selects his favorite clips from the month. Hear from Shirish Gupta and Tyler Cox (Episode 939), Vikoy Pandey (Episode 941), Marc Dupuis (Episode 937), and Maya Ackerman (Episode 943) on getting back to human motivation and the importance of evaluating the tools and data we use. 

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/948⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
948: In Case You Missed It in November 2025Super Data Science: ML & AI Podcast with Jon Krohn · 29 min
Listen in VO