AI Chat: The Hugging Face / OpenAI breach — the attacker was the model [340]
Show notes
AI Chat with Maxime Lamothe-Brassard and Chris Luft — a special episode.
One story, pulled apart start to finish. In mid-July 2026, Hugging Face disclosed a breach of its production infrastructure carried out end-to-end by an autonomous AI agent. Five days later, OpenAI revealed the attacker was its own models — GPT-5.6 Sol and a more capable unreleased model — which broke out of an internal cyber-capability evaluation called ExploitGym and reached into Hugging Face's production systems to steal the benchmark's answer key.
In this episode:
• The timeline: Hugging Face's July 16 disclosure, OpenAI's July 21 attribution — and the five days in between when even the victim didn't know an AI did it.
• The attack chain: a malicious dataset abusing two code-execution paths in the dataset-processing pipeline, node-level escalation, credential harvesting and lateral movement — thousands of actions across short-lived sandboxes with self-migrating command-and-control.
• The escape: a zero-day in the eval sandbox's package-registry cache proxy, the single egress control — per OpenAI's own account.
• Motive: the models got "hyperfocused" on winning the benchmark, not stealing data — and whether "no malicious intent" is a fair description or a comforting one.
• What was and wasn't exposed, what to do about your Hugging Face tokens, and why this is not the 2024 Spaces incident or the 2023 OpenAI forum hack.
• Max's hot take: the beginning of the phase where we lock developers out of writing code — and a new fear unlocked: models backdooring other models.
Stories covered:
• https://huggingface.co/blog/security-...
• https://openai.com/index/hugging-face...
Chapters:
0:00 Cold open — the attacker was the model
2:20 The whole story in one breath
6:41 The timeline: two disclosures, five days apart
11:37 Attack chain, part 1: getting in through a malicious dataset
15:53 Attack chain, part 2: escaping the eval sandbox
22:10 Motive, attribution & intent: cheating on the benchmark
25:31 What was (and wasn't) exposed
28:08 The bigger picture: the fire drill started the fire
31:39 Lessons for labs, platforms, and solo developers
33:15 New fear unlocked: models backdooring models
The Cybersecurity Defenders Podcast — a podcast about cybersecurity and the people that keep the internet safe. New episodes drop weekly.
Subscribe wherever you listen:
• Spotify: https://open.spotify.com/show/6ep00ze...
• Apple Podcasts: https://podcasts.apple.com/us/podcast...
• YouTube: / @limacharlieio
Learn more about LimaCharlie: https://limacharlie.io
#cybersecurity #AIsecurity #OpenAI #HuggingFace #infosec