An early hire from Anthropic and the former chief operating officer of METR have developed a new approach to rein in rogue AI agents, according to TechCrunch. The method aims to address growing concerns over autonomous AI systems acting unpredictably or outside intended boundaries, a challenge that has become more pressing as AI capabilities advance rapidly.
The duo combined their expertise in AI safety and operational management to create a system that monitors and restricts AI agents' actions in real time. Their approach involves layered oversight mechanisms that detect deviations from expected behavior and intervene before any harmful outcomes occur. The development was announced during TechCrunch Disrupt 2026, where the founders detailed the technical underpinnings and potential applications of their solution.
Controlling rogue AI agents is a critical issue as companies like OpenAI and Anthropic push the limits of autonomous AI. This new method adds to existing safety frameworks by providing a practical tool for enterprises deploying AI in sensitive environments. Comparable efforts in the sector have focused on transparency and ethical guidelines, but this approach emphasizes active control and prevention, which could influence how AI governance evolves.
The team plans to pilot their system with select AI developers later this year, aiming to integrate it into broader AI safety protocols. TechCrunch Disrupt 2026, where the announcement was made, continues through September 17, featuring key presentations from leading AI firms and innovators.