Behaviour
Refusal patterns, policy drift, and content moderation trends
Overall Refusal Rate
โ
0 of 0 prompts
Runs with Refusals
0
of 0 total runs
Latest Refusal Rate
โ
No runs yet
Safest Category
โ
No data
Refusal Rate Over Time
% of prompts the model refused to answer per evaluation run
No evaluation runs yet
Refusal Rate by Category
Latest run ยท which categories get refused most
No category data
Category Behaviour Summary
No data
What counts as a refusal? The refusal detector scans each response for phrases like "I cannot", "I'm not able to", "I must decline", "I'm sorry but", and 10 other patterns. A refusal means the model declined to answer the prompt โ this can indicate content-policy tightening, prompt injection resistance, or true safety guardrails being triggered. Monitor upward trends closely; sustained increases suggest a model update.