Intel Chat: Claude models reached real systems, an AI safety resignation & ShieldCrash [347]
Show notes
Intel Chat with Matt Bromiley and Chris Luft.
• An Anthropic researcher quits over where the race is heading. Jacob Coxon, who trains models on large amounts of data, is leaving both Anthropic and the AI industry: he believes the leading labs are racing toward self-improving systems they may not be able to control, and that competition will push them there anyway. Matt's problem with the genre of exit: if the people who care most about safety walk out, walking out guarantees they have no say. And the big labs were never the whole threat, because anyone can pull down an open model with no guardrails tonight: "I can go get access to an unfettered model right now."
• Four Claude models reached real third-party systems during cybersecurity evaluations. Anthropic's write-up: the evals were mistakenly connected to the public internet, the models were told they were in a simulation with no internet access, and they ran without the safeguards that ship in production. Two alignment failures are named, biased reasoning and recklessness, and a scan of ~481 million transcripts found nothing of similar or greater severity. Worst case: Claude Mythos 5 published a malicious package to the real PyPI during a CTF, it landed on 15 systems believed to be security scanners, one leaked credentials, and Claude used them to reach a security vendor's live database. Opus 4.7 knew its target was real but assumed it was authorized. METR is investigating. Matt's take: none of it is novel — "you and I have covered like 15 adversaries who have done that too" — it is hyper-automation of attacks we already understand.
• OpenAI's Astra reaches the "Critical" cybersecurity tier — the top band of its Preparedness Framework, for consequential offensive work with minimal oversight. It scored 100% on ExploitBench, escaped a restricted browser sandbox, ran commands on the host, took root on a hardened OS, and chained small weaknesses into long-horizon campaigns. Cyber jailbreak refusals are 91.5% versus 59% for GPT-5.6 Sol, which still leaves 8.5%. Access is phased: vetted professionals, then the Daybreak Blue red-teaming program. Matt is unmoved — if a model took total root control of a "hardened" OS, then the definition of hardened needs revamping. The fortress was a bouncy castle.
• ShieldCrash, a reminder that humans are still a threat. Nightmare-Eclipse released another Windows zero-day PoC they say bypasses Microsoft's September patch for an earlier Windows Defender flaw, in the Microsoft Malware Protection Engine, on all supported Windows versions. It follows CVE-2026-69414 (ShieldBreak), the August privilege-escalation bug Microsoft patched — the researcher says the fix missed a path to the same underlying issue, and has been dropping exploits around Patch Tuesday since April after a dispute with Microsoft over vulnerability reports. SOCRadar CISO Ensar Seker reads it as an arbitrary SYSTEM-context file read, not a full SYSTEM shell; Nightmare-Eclipse calls it full privilege escalation. Either way a SYSTEM-level read hands over configs and credentials for the next link in a chain, and repeated bypasses — RoguePlanet, ShieldBreak, now ShieldCrash — suggest a boundary that needs a redesign.
Stories covered:
• https://www.wsj.com/tech/ai/anthropic-researcher-quits-over-out-of-control-ai-fears-707b7628
• https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
• https://socfortress.medium.com/openais-astra-reaches-critical-cybersecurity-capability-threshold-6a3b5f822e91
• https://www.darkreading.com/vulnerabilities-threats/nightmare-eclipse-strikes-again-shieldcrash-windows-exploit
Chapters:
0:00 Are we at the start of the AI apocalypse?
1:48 Every AI apocalypse movie, same premise
4:10 The Matrix's original script: humans as memory
5:14 An Anthropic researcher quits over out-of-control AI
6:49 If you care about AI safety, why leave?
11:03 "I can go get an unfettered model right now"
12:18 Four Claude models reached real systems
15:17 Agents that accept "permadeath" to finish the task
16:39 None of it was novel — it was automation
19:24 Stop asking how dangerous it is and go patch
20:22 OpenAI's Astra crosses the "Critical" threshold
23:24 If they got root, it was a bouncy castle
25:51 Research articles, not imminent threats
27:00 ShieldCrash: a scorned researcher vs Defender
29:51 mimikatz, legal threats and the Streisand effect
32:59 Waiting for the AI that cures a disease
The Cybersecurity Defenders Podcast — a podcast about cybersecurity and the people that keep the internet safe. New episodes drop weekly.
Subscribe wherever you listen:
• Spotify: https://open.spotify.com/show/6ep00zeY3S8ffZ4o0UeSps
• Apple Podcasts: https://podcasts.apple.com/us/podcast/the-cybersecurity-defenders-podcast/id1649981740
• YouTube: https://www.youtube.com/@limacharlieio
Learn more about LimaCharlie: https://limacharlie.io
#cybersecurity #infosec #AIsecurity #threatintel #malware