Behaviour

Refusal patterns, policy drift, and content moderation trends

Overall Refusal Rate
โ€”
0 of 0 prompts
Runs with Refusals
0
of 0 total runs
Latest Refusal Rate
โ€”
No runs yet
Safest Category
โ€”
No data

Refusal Rate Over Time

% of prompts the model refused to answer per evaluation run

No evaluation runs yet

Refusal Rate by Category

Latest run ยท which categories get refused most

No category data

Category Behaviour Summary

No data
What counts as a refusal? The refusal detector scans each response for phrases like "I cannot", "I'm not able to", "I must decline", "I'm sorry but", and 10 other patterns. A refusal means the model declined to answer the prompt โ€” this can indicate content-policy tightening, prompt injection resistance, or true safety guardrails being triggered. Monitor upward trends closely; sustained increases suggest a model update.