Accuracy

Similarity vs baseline · cosine distance via sentence embeddings · higher = better

Overall Score Trends

Blue = embedding similarity vs baseline · Purple = LLM judge score (gpt-4o-mini) · Dashed = 95% threshold

No evaluation runs yet — click Run Now on the Overview page

Similarity by Category

Scores per prompt category for the selected run

No category data for this run

All Evaluation Runs

No evaluation runs yet.