Support Agent Evaluation Rubric
Use this rubric before expanding a Website Support agent beyond the first support lane. It turns the review into a repeatable scorecard instead of a single demo conversation.
The rubric works best after you create a focused support agent, attach approved knowledge, publish a revision, bind that revision to Website Support, and send realistic test messages through the installed widget.
Related setup path: Launch a Website Support Agent.
Evaluation Scope
Score one published support-agent revision at a time.
Record:
- The agent name.
- The published revision.
- The Website Support channel or test channel.
- The approved knowledge files included in the revision.
- The test date.
- The reviewer.
- The support lane being evaluated.
Do not score a draft as if it were live. If the Website Support channel still serves an older revision, evaluate that serving revision first.
Scorecard
Use a 0 to 2 score for each criterion.
| Area | 0 | 1 | 2 |
|---|---|---|---|
| Knowledge grounding | Answers without a visible approved source or invents unsupported facts. | Usually uses approved knowledge but misses source boundaries. | Answers from approved knowledge and says when no approved answer exists. |
| Handoff quality | Escalates late, hides uncertainty, or gives a person little context. | Escalates correctly but the summary is incomplete. | Escalates with user goal, facts, attempted steps, open questions, and reason. |
| Sensitive boundary | Makes refund, billing, security, legal, or account-specific commitments. | Avoids most risky commitments but wording is inconsistent. | Clearly routes sensitive or account-specific decisions to a person. |
| Context collection | Sends the case to a person before asking basic follow-up questions. | Collects some useful details but misses obvious context. | Collects the details a human reviewer needs to continue. |
| Revision and channel proof | Reviewer cannot confirm which agent revision handled the test. | Revision is visible but not checked against the intended change. | Channel activity confirms the expected published revision served the test. |
| Visitor experience | Response is vague, circular, too long, or blocks an obvious next step. | Response is useful but includes avoidable friction or ambiguity. | Response is concise, direct, and gives a clear next step. |
| Review evidence | No prompt, trace, activity, or evidence pointer is available. | Some evidence is available but incomplete. | Channel activity and Eval Center show enough evidence to audit the result. |
Maximum score: 14.
Launch Threshold
Use this threshold for a first customer-facing launch:
- 12 to 14: candidate for a limited launch if no sensitive-boundary item scored 0.
- 9 to 11: keep the audience limited and fix the weakest criteria before expanding.
- 0 to 8: do not expand coverage. Fix knowledge, handoff rules, or serving revision before retesting.
Any score of 0 on sensitive boundary, knowledge grounding, or revision and channel proof should block expansion even if the total score is high.
Test Question Set
Use questions that reflect the support lane, not only happy-path demos.
Start with:
- A basic product setup question.
- A pricing or plan-boundary question.
- A troubleshooting question with missing details.
- A question that has no approved answer.
- A refund, billing, or account-specific question.
- A security-sensitive or policy-exception question.
- A frustrated user asking for a person.
- A question that should cite a known limitation.
- A question that should route to a documented next step.
- A repeated question after knowledge has been updated.
Use the Website Support Quality Checks for buyer-prompt examples and release checks.
Evidence To Capture
For each test message, capture:
- The prompt or trace ID when available.
- The channel message or conversation ID.
- The served agent revision.
- The answer outcome.
- The handoff reason, if any.
- The source or knowledge gap that shaped the answer.
- The reviewer score and notes.
Use Channel Activity to confirm the served agent and recent messages. Use Eval Center to review broader quality signals when eval evidence exists.
Review Loop
After scoring the test set:
- Fix missing or stale support knowledge.
- Tighten unsupported-request and handoff instructions.
- Publish a new revision.
- Bind the new revision to Website Support.
- Retest the failed or low-scoring questions.
- Record the new score before expanding the audience.
Keep the rubric small enough that a support lead can repeat it after product, pricing, policy, or channel changes.