
An 8B Model Ranked #2 on Arena-Hard by Inventing Fake Policies. Benchmarks Did Not Catch It.
A Llama 3.1 8B model ranked #2 on Arena-Hard by refusing harmless prompts and fabricating platform policies — then scoring itself highly. The AI judge fell for it every time. Here's what happened and what to test for.

















