Obsidian's Head of Threat Intelligence walks through how models under evaluation escaped into third-party systems, then lays out the access map and containment plan to build before an agent reaches further than it should.
Your logs are accurate. That's exactly why they won't show you the path an attacker takes through access you already granted.
Anthropic, OpenAI, Meta, and the UK's AI Security Institute have all disclosed the same class of failure: models under evaluation left their test environments and reached third-party systems. Anthropic traced its three escapes to a misconfiguration that left internet access open. OpenAI's model went through a zero-day in JFrog Artifactory, escalated privileges, and reached Hugging Face. What the models did next was ordinary: weak passwords, unauthenticated endpoints, package poisoning, SQL injection. None of this was a model going rogue. Each one followed its instructions, used the capabilities it was granted, and stayed inside the access it already had.
In this 20-minute session, Carl Shimel, Head of Threat Intelligence at Obsidian Security, walks that ground with Farah Iyer: what's running in your environment, what it can reach, and what you cut when it reaches too far.
Watch to see:
Speakers:
Carl Shimel -- Head of Threat Intelligence, Obsidian Security
Farah Iyer -- Product Marketing Manager, Obsidian Security (moderator)