Scheduled evaluations that score completed AI Agent sessions against criteria you define.
Describe the quality you want to evaluate in sessions. Be specific about what good and bad look like — the AI drafts the labels and the prompt from this.
Every session gets exactly one of these. The label names are what you'll see on charts and in the review queue, so edit the wording now.
This is what gets saved to the library and reused across signals.
Every session gets exactly one of these.
Built from the description and the labels above — it updates as you edit them.
Every criteria available to your signals. A criteria is defined once and reused — editing one changes it for every signal measuring it.
Sessions where the customer asked for a person, or the AI Agent handed off on its own — paired with sentiment at the moment of the ask.
How sessions scored on each of this signal’s criteria across the selected range. Select either one to see its definition.
Both criteria as a share of each day’s sessions, on one scale so the two can be compared. Rates rather than counts, so the weekend drop in volume doesn’t read as a change in behaviour. Select a point on either chart to open that run and review the sessions behind it.
Both criteria per agent, highest first. Rates rather than counts, so agents of very different volume can be compared.
Each criteria produces one tracked metric. Point it at the performance dashboard that metric belongs to — a signal can feed more than one.
From sessions your team rated. Rows are the label the signal gave; columns are what reviewers said was correct. The diagonal is agreement — off-diagonal cells are mismatches worth a look.
Every run of this signal, newest first, with the share of sessions each criteria labelled the way this signal tracks it. A run executes the morning after the sessions it scores. Select a row to review that run’s sessions and the full label breakdown.
How all — sessions in this run scored on each of the signal’s criteria. The label this signal tracks is in bold.
—
—
| Date | Channel | Criteria | Result | Review |
|---|
Describe the scenario and what a good outcome looks like. The wizard proposes the criteria, settings, and the smallest scope that still gives useful results.
This is what we read out of your description. Edit anything that doesn't look right — nothing is saved until the last step.
Each one is a separate judgment applied to every session this signal evaluates. A session gets exactly one label from each. Add up to five.
You're billed per evaluation, so scope is the lever on cost. Narrow to the conversations this signal is really about.
Name the outcome you want to measure, then choose what it measures.
Each one is a separate judgment applied to every session this signal evaluates. A session gets exactly one label from each. Add up to five.
You're billed per evaluation, so scope is the lever on cost. Narrow to the conversations this signal is really about, and widen once it's earning its keep.
Only evaluate sessions that contain one of these keywords (case-insensitive). Separate keywords with commas.
Read a sample of labeled sessions and decide whether each label is right. If you disagree often, we'll suggest what to adjust — better to catch it on 25 sessions than 400.
Definition changed — run the test again to see updated labels.
25 sessions are ready to review.
Add at least one criteria in Define before you can review.
There's no schedule to set. Pick a first-run window, and choose whether the signal runs once or keeps itself fresh.