Overview

Loading...

All systems healthy
Avg Similarity
––
No data yet
Refusal Rate
––
No data yet
Avg Latency
––
No data yet
Last Run Cost
––
No runs yet

Prompt Set

Total prompts
β€”
Categories
Loading…

Upload a .json (array of strings or { text, category } objects) or .csv (columns: prompt, category). Min 3, max 200 prompts.

Auto-Evaluation Schedule

Next run in
β€”
Interval
β€”
Last run
Not run yet
Scheduled runs
β€”
since server start

Similarity Score Over Time

Higher is better Β· Baseline = 100%

Run more evaluations to see trends

Latency Over Time

Average response time per evaluation run

Run more evaluations to see trends

Similarity by Category

Latest run Β· scored vs baseline

Run an evaluation to see category breakdown

Active Alerts

βœ…No active alerts β€” model is stable

Evaluation Run History

No runs yet.