Accuracy
Similarity vs baseline · cosine distance via sentence embeddings · higher = better
Overall Score Trends
Blue = embedding similarity vs baseline · Purple = LLM judge score (gpt-4o-mini) · Dashed = 95% threshold
No evaluation runs yet — click Run Now on the Overview page
Similarity by Category
Scores per prompt category for the selected run
No category data for this run
All Evaluation Runs
No evaluation runs yet.