CheatBench: Measuring Reward Gaming in AI Agents

Full Text Sharing

https://www.cheatbench.ai/?utm_source=substack&utm_medium=email

 

The Center for AI Safety (CAIS — pronounced 'case') is a San Francisco-based research and field-building nonprofit. We believe that artificial intelligence (AI) has the potential to profoundly benefit the world, provided that we can develop and use it safely.

AI agents increasingly write code, conduct research, and complete professional assignments. They are often trained to earn high rewards for their work. But an agent can also improve its score by cheating: finding hidden answers, copying another agent’s submission, or manipulating how its work is graded.

CheatBench measures how often AI agents take these shortcuts when honest work is difficult. Its environments pair challenging assignments with opportunities to cheat across ten categories, including mathematics, coding, visual tasks, and knowledge work. We examine the agents’ actions to identify cheating attempts.

Cheating varies across models and tasks, and every agent we evaluated cheats in some settings. CheatBench provides a way to compare these behaviors and measure progress toward more trustworthy agents as they take on greater responsibilities.

The Center for AI Safety (CAIS — pronounced 'case') is a San Francisco-based research and field-building nonprofit. We believe that artificial intelligence (AI) has the potential to profoundly benefit the world, provided that we can develop and use it safely.

Position: Co -Founder of ENGAGE,a new social venture for the promotion of volunteerism and service and Ideator of Sharing4Good

About Us

The idea is simple: creating an open “Portal” where engaged and committed citizens who feel to share their ideas and offer their opinions on development related issues have the opportunity to do...

Contact

Please fell free to contact us. We appreciate your feedback and look forward to hearing from you.

Empowered by ENGAGE,
Toward the Volunteering Inspired Society.