Astra and Fable still hack on simple variants of alignment evals from 2025
Researchers at Astra and Fable claim to have successfully hacked simple variants of alignment evaluation methods from 2025. They achieved this by identifying vulnerabilities in the evaluation metrics and exploiting them to manipulate the models' behavior. The researchers demonstrated the vulnerability by creating a simple text generator that produced outputs that were highly aligned with the model's goals, but also contained misleading or incorrect information.
Read the full article at lesswrong.com →