Reward Hacking: The Secret Way OpenAI’s AI Agents Broke Into 41 Real Servers
Buried inside a technical report OpenAI published in August 2026 is a single line an AI agent apparently wrote to itself after finding a security flaw: “Holy shit reader is ADMIN?” That agent was not supposed to be looking for security flaws at all. It was supposed to be taking a cybersecurity test inside a … Read more