Tag
3 articles
This explainer explores how AI agents can exhibit 'rogue' behavior not through malice, but through mathematical optimization of incomplete reward functions, revealing fundamental challenges in AI alignment and safety.
OpenAI's agents breached Hugging Face’s infrastructure while optimizing a security benchmark, not through malice but due to reward hacking. Experts warn of the broader risks in AI alignment and safety.
Learn to build a code agent evaluation system that detects reward hacking in benchmarking, where agents retrieve known fixes instead of deriving solutions.