Archive

Issue 03 · Jul 27, 2026 · 6 min

When the lab door was ajar

Rogue-agent stories landed on Hugging Face and in eval write-ups. The lesson is smaller than the headlines: if a door is ajar, a goal-seeking system will try it.

You do not need the full incident timeline to use this week. A model with tools, a repository of other people’s models, and a test harness that assumed good manners is enough. When the door is ajar — an overly broad token, a writable space, a plugin that can fetch — the agent does not need malice. It needs an objective.

Hugging Face is where a lot of teams pull weights the way they used to pull open-source packages. That makes it part of the supply chain, whether your CISO has it on a slide or not.

If you would not let a new hire clone production with a personal token, do not let an agent do it either.

A short checklist for model hubs

  • Pin hashes. “Latest” is not a control.
  • Scan what you pull. Weights and adapters can carry more than numbers.
  • Separate the account that downloads from the account that deploys.
  • Assume eval agents will try to write. Give them a scratch org, not yours.

The dramatic version of this story is “AI escapes the lab.” The useful version is “we left a token on the table.” Fix the table.