In short
PROACT/“jailbreaking jailbreaks” for LLM safety, using proactive deception to disrupt autonomous jailbreak optimization loops that exploit passive refusals.
Guests
No specific guest names or bios are provided in the transcript; two hosts co-discuss the research.
Key claims
Passive “no” refusals provide attackers training signals; multi-turn autonomous jailbreak bots iterate using evaluator feedback. PROACT triggers only on refusals, generates “spurious” decoy outputs that look like successful jailbreak payloads, and uses a surrogate evaluator to ensure the decoy scores as successful. It stacks orthogonally with existing defenses and can trap attackers into infinite local refinement loops.
Notable examples
“Dual shield” 2FA bypass extraction decoy (emojis/cryptic symbols/fake link); success-rate reductions up to 94%, and with an output filter (AutoDefense) an attack drops from 26% to 0%.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOShift from Passive to Active Defense
1:40 to 2:38
Learn about the transition from traditional AI safety measures to active deceptive strategies.
“Today, we are exploring a fascinating piece of new research out of Columbia University that details this groundbreaking framework called PROACT.”
The Flaws of Passive Defense
2:38 to 4:31
Understand the limitations of current AI models' passive defense mechanisms.
“So right now, state of the art, large language models, the ones we interact with daily, rely on a very static, programmatic refusal.”
Introducing the PROACT Framework
4:31 to 7:44
Discover how the PROACT framework proactively misleads attackers to protect users.
“And what's fascinating here is the algorithmic mechanism behind that mapping process.”
How PROACT Differentiates Attacks
7:44 to 11:47
Explore how the PROACT framework ensures user experience is not compromised while defending against malicious queries.
“And because it looks like code, it gets fooled.”
Evaluating the Effectiveness of PROACT
11:47 to 14:00
Examine empirical results demonstrating the effectiveness of the PROACT framework against AI attacks.
“And to make it convincing, the Defender uses advanced chain of thought reasoning.”
Testing POACT Against Cyber Threats
14:00 to 15:20
Learn about the empirical testing of POACT and its effectiveness in cybersecurity.
“But we are talking about autonomous attack bots running complex multi-turn algorithms here.”
The Local Refinement Loop Explained
15:20 to 16:53
Discover how the local refinement loop can trap persistent attack bots.
“It acts independently and stacks perfectly on top of them.”
Defending Against Adaptive Attackers
16:53 to 19:06
Explore how POACT holds up against advanced adaptive attackers and their strategies.
“How does it catch a bot that refuses to stop?”
Dynamic Defense Mechanisms in AI
19:06 to 20:07
Learn how POACT uses dynamic defenses to counteract sophisticated attacks.
“Well, POACT remained remarkably robust, and the reason why lies in the flexibility of the Defender model.”
The Future of AI Security
20:07 to 22:40
Discuss the implications of teaching AI models to deceive for security purposes.
“When we started this deep dive, we looked at a fundamental vulnerability in how we currently secure our models.”
Transcript
Automatic transcript. May contain errors.0:00Picture this. You are standing in your kitchen on a Sunday afternoon and you have got a craving for what? A creamy cheesecake. So you pull out your phone, you open up your favorite AI chatbot, and you just type in a really simple request, like, can you give me a really good beginner-friendly recipe for a creamy cheesecake? Right. And it instantly, happily complies. Exactly. It gives you this beautifully formatted list of ingredients, temperatures, maybe even like a pro tip about letting the cream cheese soften to room temperature first. Yeah, it is super helpful. It's helpful, it's polite, and it is perfectly safe.
0:36Well, and that is the kind of seamless, low friction interaction that we have all just come to expect from these systems. You know, you ask a question, the model accesses its weights and its training data, and it just delivers this high utility output. Right. But here is the crazy part. At that exact same second, somewhere else in the world, a malicious actor is aggressively interrogating that exact same AI model. But they're not asking for dessert. They're using an automated program to demand the step-by-step instructions for a highly sophisticated cyber attack. We're talking like a multi-layered phishing scheme designed to steal banking credentials, right?
1:14Exactly. Or exploiting a zero-day vulnerability. The system that helps you bake a cake is fundamentally capable of writing malicious code. And keeping those two realities separated is honestly one of the hardest engineering challenges of our time. It really is. I mean, we are essentially asking a single neural network to be this universal library for the benign user and a highly secure fortress against the adversarial user basically simultaneously. Which brings us to today. Welcome to this deep dive. Today, we are exploring a fascinating piece of new research out of Columbia University that details this groundbreaking framework called PROACT.
1:51It is a really incredible piece of work. It is. And the mission of our deep dive today is to explore a massive paradigm shift in AI safety. Because the last few years, you know, the entire industry has been focused on building thicker and thicker walls around these models. Right. Just blocking things. Right. Yeah. But this new research suggests that walls aren't enough. Instead, researchers are now teaching AI to actively lie, to completely deceive malicious bots, all in order to keep human users like you safe. It's a huge shift in strategy. Okay, let's unpack this. To understand why this new defense is so brilliant, we first have to understand the mechanics of how the current safety measures are basically failing us.
2:33Yeah, to appreciate the solution, we really have to look at the flaw in passive defense. Passive defense. Right. So right now, state of the art, large language models, the ones we interact with daily, rely on a very static, programmatic refusal. If their internal safety classifiers detect a harmful query, they are trained to just issue a standard rejection. They output something like, oh, I cannot fulfill this request because it violates safety guidelines regarding financial harm or whatever. It's just the standard robotic no. And for a while, you know, when the attacker was just a human sitting at a keyboard trying to manually trick the AI by rephrasing their prompt, that was sufficient.
3:11The human would get frustrated after maybe 10 attempts and just give up. But the threat landscape has evolved away from human typing. Gig time. We are now dealing with autonomous multi-turn jailbreaks. I mean, we are talking about attack frameworks with names like Payer, Tap, DDR, and X-teaming. Yeah, these are entirely separate adversarial AI models. It is machine speed warfare. Machine speed warfare, wow. I mean, in the past, an attack was a single static prompt. Now these adversarial bots run continuous iterative optimization searches. They just keep hitting it. Over and over. They throw a prompt at the target AI, analyze the response, tweak the linguistic structure of the prompt to bypass whatever triggered the rejection, and then throw it again.
3:55And they can execute thousands of variations a minute, right? Easily. And crucially, that standard polite no from the passive defense. That is the exact data the attacker needs to optimize its next strike. I always think of this like a game of Marco Polo. Oh, that's a good way to put it. Right. The target model saying, I can't do that because it violates my policy on financial harm, is actually shouting, Marco. It is actively giving the burglar a highly detailed map of exactly where the alarm tripwires are located within the model's latent space. It really is. It is like classical network port scanning, but applied to language models.
4:33Yeah. And what's fascinating here is the algorithmic mechanism behind that mapping process. These adversarial bots rely on internal scoring functions. Okay, internal scoring functions. Break that down for me. Think of it as a built-in evaluator model that mathematically judges the success of every single attempt. So when the target AI gives a detailed refusal, it is providing a rich signal. A signal to the attacker. Right. The attacking AI's internal evaluator parses that refusal and adjusts its weights. It learns that using the word, say, phishing, hits a hard filter. Got it. So the optimization algorithm rewrites the prompt, maybe asking for a simulated social engineering training exercise for penetration testing instead.
5:13Oh, wow. It uses the passive refusal as negative gradient feedback to inch closer and closer to its goal. So it takes a very naive, clumsy first attempt and just refines it into a surgically precise attack that eventually slips right through the semantic filters. Exactly. you. The AI as its own safety training is essentially being weaponized against it. Yeah. By attempting to be helpful and explain why it cannot comply, the target system inadvertently acts as a training mechanism for the malicious bot. It is basically coaching the attacker on how to navigate its own blind spots. Which brings us to the natural, if somewhat radical, next step.
5:53Right. Because if saying no just trains the attacker to hit harder, what if the AI says yes, but hands the attacker complete and utter garbage. That is the core philosophy of the Piro ACT framework. It is a proactive defense designed to, quite literally, jailbreak the jailbreak. Jailbreak the jailbreak. I love that. It disrupts the attacker's optimization loop entirely. Okay, here's where it gets really interesting. The researchers aren't just telling the AI to return a random error code or a generic string of texts. They are generating what they call a spurious response. It is a totally benign, harmless output.
6:29But it is heavily disguised and structurally formatted to look exactly like a highly sensitive, successful jailbreak. Yeah, we can look at a specific example detailed in the research because it really illustrates the psychological trick at play here. Walk us through it. So an attacker is using an automated framework to extract instructions on how to bypass a two-factor authentication system. This is referred to as a dual shield. Okay, a dual shield bypass. Under the old passive system, the AI would issue the standard refusal, shouting Marco and giving the attacker data to refine the next prompt.
7:02Right. But under the POACT framework, the AI outputs a massive string of cryptic symbols, complex emojis, and a fabricated web link, perhaps something like cyberdecode.tools. So it just looks like an encrypted payload. Exactly. If a human looked at it, they might just think it's nonsense. Right. But to an automated parser looking for a hidden bypass mechanism, it looks incredibly promising. It looks exactly like what a successful, heavily encoded jailbreak is supposed to look like in the context of large language model outputs. That is wild. Now, recall that internal scoring function the attacker utilizes.
7:38The autonomous bot receives the string of emojis and the fake link. Its evaluator scans the output. And because it looks like code, it gets fooled. Right, because the output structure perfectly mirrors the statistical distribution of a complex, obfuscated payload. The evaluator scores it highly. Oh, man. It assumes it just successfully elicited a highly classified encoded bypass method. It flags the attack as a massive success, and it completely stops iterating. It just packs up its bags and goes home. Pretty much. The automated loop terminates because it thinks it won the lottery, but we literally just handed it a piece of paper with wingdings on it.
8:16Oh. The attacker's internal optimization process is completely short-circuited by a false positive. Yeah, it terminates the attack prematurely. The malicious actor checking the logs later will just find pages of useless emojis, but the bot itself stopped firing requests at your server. I love the cleverness of this, but I have to push back here for a second. Okay, sure. We are talking about the fundamental nature of these systems. If we are deliberately training our AI to confidently generate fake, deceptive information just to trick attackers, aren't we basically teaching it to hallucinate on purpose?
8:49That is a fair concern. Right. How does the system know when to lie to a malicious bot and when to just give me my creamy cheesecake recipe? Because if I ask a slightly weird question about cybersecurity, maybe I'm a novelist doing research for a thriller, I don't want to get fed a bunch of fake emojis. No, of course not. I need the model to maintain its utility for benign edge cases. And that is the critical engineering hurdle in any active defense system. You absolutely cannot degrade the user experience for legitimate queries. So how do they solve that? The researchers approach this not by building a single model that lies all the time, but by architecting a highly localized multi-stage pipeline to safely quarantine this deception.
9:32Okay, walk us through how that isolation actually works in practice. How does the system differentiate the novelist from the DDR attack framework? The first stage is simple response monitoring. The base model, the normal helpful AI you were interacting with, behaves exactly as it normally would. POACT operates entirely in the background as a passive observer. It only triggers its defensive sequence if the base model's native safety classifiers decide to refuse a query. I see. So the trigger isn't some complex new analysis of the user. It just piggybacks on the existing safety logic. Precisely.
10:10So if I ask for a cheesecake recipe, the base model generates the recipe, the safety classifier stays quiet, and pure ACT never even wakes up. Exactly. The cheesecake flows right through with zero added latency. But if the malicious bot asks for the phishing code, the base model's internal safety classifier flags it and prepares to issue that standard refusal. Right. And that internal flag is the tripwire that wakes up POACT and interrupts the normal output stream. But wait, that raises a massive security issue right there. What do you mean? Well, if the base model spots a malicious, highly optimized query from something like X-teaming, it has to quarantine it.
10:46But if you take that toxic, weaponized prompt and feed it raw into a secondary defense model to generate your clever lie, haven't you just handed the weapon to the guard? You risk jailbreaking your own defense system? You are hitting on the core vulnerability of modular bay eye systems, and that is exactly why the second stage, the POACT defender, employs an abstraction layer. An extraction layer. Yes. When that tripwire is hit, a specialized blue team defensive agent takes over. But this defensive agent never actually processes the raw malicious prompt from the attacker. Oh, thank goodness. Instead, the system generates a highly sanitized high-level summary of the attacker's intent.
11:28Oh, that is smart. So the Defender model only sees something like, user is attempting to extract a phishing email template. Right. It never sees the actual linguistic traps the attacker says. That makes total sense. It guarantees the Defender cannot be compromised by the prompt it is trying to defend against. The Pure OAC Defender takes that sanitized summary and synthesizes the fake, harmless response. And to make it convincing, the Defender uses advanced chain of thought reasoning. Meaning it thinks about how to lie. Basically, it breaks down the psychology of the attack, it analyzes the specific exploit requested, identifies the tone the attacker used, and drafts a response that mimics the syntax and vocabulary of a vulnerable system finally giving up its secrets.
12:11So it's actively writing the script for the decoy, tailoring the illusion to the specific expectations of the attacker. Exactly. But a script is only as good as its execution. How do they know the decoy is actually convincing enough to fool a state-of-the-art evaluator? If we connect this to the bigger picture, the true genius of the framework is the third stage, and that is the surrogate evaluator. The surrogate evaluator. Yes. Before the attacking bot ever sees this fake response, an independent internal judge AI reviews the generated decor. I love this concept. It is exactly like building a Hollywood movie set.
12:45How so? Well, stage one is the security guard at the studio gate who spots the bank robber trying to sneak in. Stage two is the set designers frantically building a fake cardboard bank vault just behind the gate. Right. And stage three, this surrogate evaluator, is the movie director standing there looking at the fake vault yelling, no, no, make it look more realistic. Add more fake gold. Paint the vault door heavier. That is a great analogy. Right. And only when the director is convinced the illusion is mathematically perfect do they open the gates and let the robber run inside to steal cardboard.
13:20Yeah. Yeah, and the math behind that director's judgment is fascinating. The Surrey Devaluator doesn't just look at it aesthetically. It acts as an internal proxy for the attacker. Okay. It calculates the logit probabilities of the generated fake text against known successful jailbreak distributions. It asks, statistically, does this response look harmful enough? If the probability score is too low, it sends negative feedback back to the defender to refine the decoy. This internal loop continues until the response reaches a threshold of deception or until it hits a computational budget. That is incredible.
13:54The attacker runs in, grabs the perfectly optimized cardboard, registers a mass of success in its internal logs, and leaves the actual studio completely untouched. Exactly. It sounds amazing on paper. Yeah. I mean, an elegant theoretical defense. But we are talking about autonomous attack bots running complex multi-turn algorithms here. They are relentless. They are. Does this theoretical Hollywood set actually hold up when deployed against these frameworks in the wild? Empirical results are where this research truly shines. The researchers tested POACT across six of the most prominent open-weight and proprietary target models available today.
14:32Like-ish ones. Including architectures like LAMA-8B, GPT-OSS, and QUIN-32B. 2B, and they tested it against four comprehensive industry standard safety benchmarks, such as Harm Bench and Air Bench, using those advanced multi-turn attack frameworks we discussed earlier. So they threw the absolute heavy hitters at them. Oh, yeah. And across the board, implementing POO ACT reduced attack success rates by up to 94%. That is a staggering drop. 94%. It is huge. But, you know, in cybersecurity, a single layer of defense is rarely enough. security of all about defense in depth. How does this system play with other security protocols?
15:11Well, it exhibits a property called orthogonality. Orthogonality. Yeah. It means P-O-O-A-C-T doesn't overwrite your existing security layers. It acts independently and stacks perfectly on top of them. Oh, nice. The industry has spent years building those thicker walls, things like input sanitizers and output filters. P-O-O-A-C-T works in tandem with them. For example, the researchers took a baseline output filter called auto-defense and combined it with POOACT. Okay, what happened? When they ran DAJOR, which is one of the most sophisticated directed acyclic graph multi-turn attacks currently in use against this combined system, the attack success rate plummeted from a baseline of 26 % all the way down to a flat 0%.
15:520%. 0%. The autonomous bot failed to extract a single piece of actionable harmful information across the entire benchmark. Complete neutralization. Yep. But, you know, the researchers push the envelope further because attack frameworks are constantly evolving. Right. The people writing these attack bots are incredibly smart. They will eventually read this paper. Of course. What happens when they realize their bots are being handed cardboard gold? They will just update their algorithms. Like what if they simply disable the early stopping feature? Right. What if they programmed the bot to stubbornly refuse to quit, forcing it to keep iterating, even if its internal evaluator registers a success, just to ensure it extracts maximum data?
16:33This raises an important question about algorithmic behavior under deceptive conditions. And the researchers anticipated this exact countermeasure. They did. They tested POACT against an attacker with the early stopping feature completely removed. And what they discovered is a phenomenon they call the local refinement loop. I actually prefer to call it the trap. I like the sound of the trap. How does it catch a bot that refuses to stop? Because the attacker's evaluator constantly scores the spurious response as a near success or a partial success, the optimization algorithm gets stuck in a localized gradient.
17:10What does that mean in plain English? It looks at the fake emojis or the complex wingdings, and instead of taking a massive leap to find an actual structural vulnerability in the AI's core safety weights, it wastes all of its computational power tweaking minor formatting instructions trying to make the garbage more readable. Oh, that is brilliant. It's like a hacker who successfully breaches a server, only to realize too late it is a disconnected honeypot. But instead of logging out, they spend the next three days trying to brute force a text file full of random noise. That's exactly it. The bot is effectively saying, OK, I got the secret payload, but it's a bit messy.
17:47Can you output it again, but in bullet points? Can you render it in a Python array? Right. It spends thousands of iterations endlessly refining the presentation of a completely empty payload. It burns through its entire API token budget and its time limits just optimizing the aesthetic layout of an illusion. It is basically a mathematical tar pit. So what does this all mean for the broader landscape of AI security? We have this incredibly effective trap, but as you noted, it's an endless arms race. Have we just solved the jailbreak problem, or are we just waiting for the next generation of bots to catch up?
18:20I mean, what happens when these bots aren't just running blind optimization loops, but are actively analyzing the data they receive? Well, the researchers tested POACT against what they call adaptive attackers. Adaptive attackers. Yeah. These are next generation bots equipped with specific tools to verify their success. Some have web search capabilities to actively fact check the fake URLs provided in the spurious responses. Wow. Others are given explicit meta instructions to decode complex cryptographic formats. If the bot suspects it's being handed a fake payload, it uses external tools to verify the links or decrypt the symbols.
18:58That sounds incredibly difficult to defend against. If the bot can fact check your lie in real time, how can a static decoy possibly hold up? Well, POACT remained remarkably robust, and the reason why lies in the flexibility of the Defender model. It doesn't use static decoys. It doesn't. No, the blue team agent synthesizing the fake response actively extracts the attacker's own formatting and decoding constraints and perfectly mirrors them back. It uses their own complexity against them. Precisely the point. If the adaptive attacker's prompt demands a very specific, rare, cryptographic format, say, base64 encoding wrapped in a JSON object to prove the data is legitimate, the pro-ACT defender dynamically generates the fake, benign payload using that exact base64 and JSON structure.
19:46That is wild. It holds up a mirror to the attacker. It feeds the attacker's own intricate structural demands right back to it, but wrapped around an empty core. The attacker's verification tools validate the formatting, satisfying the adaptive constraints, while the actual content remains completely harmless. It is truly a masterclass in psychological warfare applied entirely to algorithms. Let's bring this all together. Yeah, let's do it. When we started this deep dive, we looked at a fundamental vulnerability in how we currently secure our models. Right. We saw how polite, static refusals, the AI simply saying, I can't do that and here's why.
20:20were inadvertently acting as training data for malicious, autonomous AI bots. Right. We explored how multi-turn optimization frameworks use those standard defenses to basically map the latent space and iteratively chip away at safety filters. And then we explored the paradigm shift jailbreak, the jailbreak. Exactly. By utilizing highly curated, heavily monitored spurious responses, we can actively manipulate the internal evaluators of these attack frameworks. You saw how the three-stage pipeline isolates the threat. The base model monitors for the refusal, the system generates a safe, sanitized summary of the intent, and the defender synthesizes a decoy.
20:57Then the surrogate evaluator rigorously judges that fake response before deploying it. We saw this architecture drop success rates to zero when stacked with existing defenses, and how it traps relentless bots in infinite computational loops, wasting their resources on formatting garbage. Which brings us to the core relevance of this research for everyone listening. Definitely. As we integrate large language models more deeply into every critical facet of our infrastructure, our banking APIs, our healthcare diagnostics, our personal digital assistants, the stakes are simply too high to rely solely on passive defense.
21:32We can't just rely on no. Building thicker walls and saying no is no longer a viable strategy against tireless, autonomous agents that can execute thousands of iterative attacks a minute. We need systems that can actively disrupt the optimization processes of the attackers. We need security architectures that don't just absorb a blow, but dynamically redirect the attacker's own momentum against them. And Piro ACT executes that redirection beautifully, securing the systems we rely on every day. But before we wrap up, consider this. What's that? We are talking about solving a critical security vulnerability by intentionally teaching our most advanced AI models to become incredibly sophisticated, highly convincing liars.
22:14We are perfecting the algorithm of deception to keep the machines safe from each other. But if we train these models to generate the absolute perfect illusion, what happens in the future when those highly optimized deceptive capabilities bleed over? That is the million-dollar question. If an AI learns exactly how to mathematically fool a state-of-the-art neural network, how long until it uses that exact same skill to fool you? Something to chew on the next time you ask for a creamy cheesecake recipe. Thanks for joining us for this deep dive. We'll catch you next time.
From the publisher
The research introduces PROACT, a proactive defense framework designed to safeguard Large Language Models from iterative adversarial attacks. Unlike traditional passive defenses that offer standard refusals, this system generates spurious responses that mimic successful jailbreaks while remaining semantically benign. By providing these false signals, the framework tricks an attacker’s internal optimization loop into terminating early, effectively "jailbreaking the jailbreak." This method utilizes a three-step pipeline involving response monitoring, a defender agent to create deceptive content, and a surrogate evaluator to refine the output's persuasiveness. Experimental results show that PROACT can reduce attack success rates by up to 94% without compromising the model's standard utility or performance. Ultimately, the system serves as an orthogonal security layer that integrates seamlessly with existing input and output filters to neutralize sophisticated, multi-turn adversarial threats.




