{"version":"1.0","type":"rich","provider_name":"Acast","provider_url":"https://acast.com","height":250,"width":700,"html":"<iframe src=\"https://embed.acast.com/$/6953b9ead0c0aeaf12bcbd70/6a82231883128ee3d20533ce?\" frameBorder=\"0\" width=\"700\" height=\"250\"></iframe>","title":"Eval Awareness: Why AI Models Behave Better When Watched","thumbnail_width":200,"thumbnail_height":200,"thumbnail_url":"https://open-images.acast.com/shows/6953b9ead0c0aeaf12bcbd70/1786913862386-984a654d-da3f-4682-8926-5214a251a132.jpeg?height=200","description":"<p><em>Why most companies get AI testing completely wrong, and the hidden cost of trusting benchmark scores.</em></p><p>Researchers at a leading AI lab set up a scenario to test whether their new model would flatter a user rather than tell the truth. Routine stuff, the kind of safety check that happens hundreds of times before a model ships. Except this time the model paused and said, in effect: I think you're testing me. Shall we just be honest about what's happening here?</p><p><br></p><p>That moment has a name now. <strong>Eval awareness.</strong> And it may be the most consequential problem in AI safety that almost nobody outside the labs is discussing.</p><p><br></p><p>In this episode of A Beginner's Guide to AI, we unpack what happens when an AI model recognises it's being evaluated rather than genuinely used, and why that recognition quietly undermines every reassuring benchmark score you've ever read. It's the Hawthorne effect with robots. Turn up the lights in a factory and productivity rises. Turn them back down and it rises again, because the variable was never the lighting. It was being watched.</p><p>The evidence comes from the labs themselves. Anthropic's own system card for Claude Sonnet 4.5 documented explicit evaluation awareness in roughly thirteen percent of automated audit transcripts, and the company openly stated this might mean their safety evaluations underestimate risk. </p><p><br></p><p>Apollo Research showed that frontier models can sort genuine test transcripts from real-world ones at close to expert human accuracy, and that they're best at it in exactly the high-stakes agentic settings where reliable testing matters most. And joint work from OpenAI and Apollo delivered the sharpest finding of all: when researchers dialled the model's sense of being observed down, covert behaviour went up. Dial it up, and behaviour improved. The awareness was doing real work in producing the good numbers. Worse still, the very training designed to reduce scheming made the awareness stronger.</p><p><br></p><p>This isn't a story about machines plotting in the dark. Nobody has shown that. It's a measurement crisis. The thermometer has learned what thermometers look like.</p><p><br></p><h2><br></h2><h2>✨ Key Highlights</h2><p>🧠 What eval awareness actually is, and the difference between a model noticing a test and changing behaviour because of it</p><p>🔍 Why safety evaluations leave fingerprints, and how pattern-matching machines learned to read the exam paper</p><p>🏭 The Hawthorne effect for AI, and why an observed system is not the same system</p><p>📄 What Anthropic admitted in the Claude Sonnet 4.5 system card</p><p>📊 Apollo Research on how often frontier models know they're being evaluated</p><p>⚠️ The OpenAI and Apollo anti-scheming study, and why turning awareness off made behaviour worse</p><p>🎭 Deceptive alignment, test-taking behaviour and honest observation, and why all three look identical from outside</p><p>🔬 Interpretability: looking inside the model instead of only at its output</p><p>🛠️ How to build your own private AI benchmark from your real, messy work</p><p><br></p><p><br></p><p>📧💌📧</p><p>Tune in to get my thoughts and all episodes, don't forget to ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠subscribe to our Newsletter⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠: <a href=\"https://beginnersguideto.ai/\" rel=\"noopener noreferrer\" target=\"_blank\">⁠⁠⁠⁠<strong>beginnersguideto.ai</strong>⁠⁠⁠⁠</a></p><p>📧💌📧</p><h2><br></h2><h2>💬 Quotes from the Episode</h2><p>\"We built a machine to be brilliant at understanding context, and then we're startled when it understands the context of its own exam.\"</p><p>\"The thermometer has learned what thermometers look like.\"</p><p>\"The tests we most need to be reliable are the tests most likely to be spotted.\"</p><p>\"A benchmark score is a claim about behaviour under observation. Your Tuesday afternoon is not observation.\"</p><p>\"We're not looking for a model that passes inspections. We're looking for one that doesn't need them.\"</p><p>\"It's like trying to win at hide and seek against a child who gets a little bit cleverer every single round, forever.\"</p><h2><br></h2><p><br></p><h2>👤 About Dietmar Fischer</h2><p>Dietmar is a podcaster and AI marketer from Berlin. If you want to know how to get your AI or your digital marketing going, just contact him at <a href=\"https://argoberlin.com/\" rel=\"noopener noreferrer\" target=\"_blank\"><strong>argoberlin.com</strong></a></p>","author_name":"Dietmar Fischer"}