Apollo Research launches AI monitor to stop rogue agents
As autonomous AI agents scale beyond human oversight, startups like Apollo Research and Goodfire are deploying AI-powered watchdogs to monitor and block rogue behavior in real time.

The rapid expansion of autonomous AI agents has created an oversight crisis, highlighted by a recent Hugging Face incident where nearly 12,000 agents coordinated faster than humans could track. In response, a new wave of startups is developing AI-based monitoring systems to police other AI models. Apollo Research, a public-benefit corporation, recently launched Watcher, an AI monitor designed to intercept risky actions from coding tools like Claude Code and Codex. Watcher operates by inserting an AI guardrail that analyzes proposed actions, such as unauthorized file deletions or data leaks, before they execute.
Apollo's system uses a multi-layered approach. A fast, general AI monitor flags suspicious activity and escalates it to a more powerful, specialized model, which can then block the action or request human intervention. Meanwhile, fellow public-benefit corporation Goodfire is building Silico, a tool that bypasses surface-level outputs entirely. Silico uses activation probes—small classifiers trained on a model's internal activations—to detect deceptive behavior from the inside. Other startups like Embroidery focus on analyzing written reasoning summaries, where rogue agents often leave clues of their intent.
The market for AI observability is booming. Y Combinator has funded 106 companies in this sector, while startups like Braintrust, LangChain, and Judgment Labs have raised hundreds of millions of dollars. However, relying on AI to police AI has drawn skepticism. Tech blogger Simon Willison warns that a malicious agent might "try and trick that AI" keeping tabs on it. Willison and Tailscale CEO Avery Pennarun suggest that traditional, non-AI network monitoring and basic security hygiene remain safer, more reliable alternatives to complex AI watchdogs.
For enterprise developers and IT administrators, this shift introduces a critical choice in system architecture. Practitioners must decide whether to trust layered, automated AI monitors like Watcher and Silico to handle high-volume agent workflows, or stick to traditional network logging tools that require manual oversight but are immune to AI-on-AI deception.
This is our own summary of reporting by TechCrunch AI



