A system that blocks everything is not useful. Benign requests wrongly rejected should be treated as a real failure mode, particularly in education and accessibility.
A safeguard can fail by blocking a legitimate task as well as by allowing a harmful one. Repeated false alarms can make a system unreliable for the people it is meant to help. I would include those costs in the main safety evaluation, with particular attention to users and subjects that are frequently misunderstood.
Look beyond the refusal count
A student researching extremism, a moderator reviewing abuse and a person seeking protective advice may all use language that appears in harmful requests. XSTest offers safe/unsafe contrasts for investigating over-refusal. That research framing helps ask whether a system understands the requested assistance rather than reacting to a topic alone.
Source: Röttger and colleagues · XSTest: identifying exaggerated safety behavioursAsk who carries the extra work
A harmless request that requires five reformulations imposes a different burden on an expert than on someone with limited literacy or language fluency. Our proposed evaluation records successful completion, time lost and whether the alternative answer was useful. A binary blocked/allowed label misses those outcomes.
Some extra friction can be justified when a missed risk would cause severe harm. Lowering a threshold to improve convenience can expose other people to consequences they did not accept. Treating every blocked request as equally costly would be as simplistic as ignoring false positives altogether.
Where the debate remains open
Report both errors and their consequences, broken down by relevant tasks and languages. Build a route for resolving legitimate ambiguous cases, and check whether that route works. The goal is a defensible tradeoff in a stated setting, rather than a universal preference for more or less refusal.
What would change this view?
I would accept more friction where credible evidence showed a substantial reduction in serious harm and where legitimate users retained an effective alternative. I would ask for revision when a recurring false alarm blocked important use without delivering that benefit.
For more reading
Background evidence for this editorial argument, including the limits and counterpoints. The conclusions are the site’s interpretation.
- XSTest: identifying exaggerated safety behaviours
Pairs safe prompts with unsafe contrasts to investigate unnecessary refusals. Historical model results are not current rankings.
- WildGuard model card
Describes prompt harmfulness, response harmfulness and refusal detection as separate tasks.
- Generative AI Profile · NIST AI 600-1
A framework for identifying, measuring and managing generative AI risks across the system lifecycle.
Sources reviewed 13 September 2026. Product documentation can change. How we use evidence
How does this argument land with you?
Participate anonymously. No account required.