Most voice agent evaluation happens against a set of expected intents — the questions the team anticipated when they scoped the project. That testing approach catches real bugs, but it structurally can't catch the failure mode that matters most in production: the question nobody anticipated.
When a voice agent has no explicit boundary around what it knows, it doesn't fail by going silent — it fails by answering anyway. Language models are built to produce a plausible continuation, and a plausible-sounding wrong answer on a phone call is worse than no answer at all, because the caller has no way to verify it on the spot the way they might scan a wrong web page.
This is distinct from a model simply being "wrong." The failure is architectural: nothing in the pipeline was checking whether the question fell inside the system's actual scope before generating a response.