Use Eval Center - Docs menu

Use Eval Center

Use Eval Center after you publish or deploy a support agent. It helps you review delivery health, behavioral eval status, feedback, and redacted runtime evidence before you trust a support-agent change.

Eval Center is a review surface. It does not replace a real widget test, channel activity, or human review for sensitive support behavior.

When To Use It

Open Evals in the dashboard when:

  • You published a new support-agent revision.
  • You changed agent knowledge, tools, model route, or channel behavior.
  • A visitor reports a wrong, missing, slow, or unsafe answer.
  • You want to see whether support-agent quality signals are improving.
  • You need trace, prompt, or evidence IDs for follow-up.

1. Open Eval Center

  1. Open the dashboard.
  2. Select Evals in the sidebar.
  3. Wait for the page to show Eval Center and the latest update time.

If the page stays empty or unavailable, the workspace may not have recorded eval runs, imported suites, or recent channel evidence yet.

2. Read The Business Scorecard

Start with Business Scorecard. It summarizes recent support outcomes and delivery health.

Use it to check:

SignalHow To Read It
Recent conversationsRecent support conversations seen by the channel ledger.
Delivery failuresMessages that failed, timed out, or stayed pending.
Helpful feedbackReview signals from visitors or operators when available.
Resolved outcomesOutcome labels when the workspace records resolution state.

Delivery health is transport evidence. It tells you whether messages moved through the system. It does not prove the answer was correct.

3. Review Behavioral Evals

Use Behavioral Evals to inspect recorded runs for support-agent scenarios.

Look for:

  • The latest run status.
  • Pass, fail, and critical-failure counts.
  • The agent, bundle, or channel binding the run evaluated.
  • Links to run details when a run is available.

If no required gate has run, treat the page as a signal that automated eval coverage has not been recorded yet. Continue with manual widget tests and channel activity.

4. Check Eval Suites

Use Eval Suites to understand which contracts and cases are available.

An eval suite is a curated set of support scenarios. A suite can cover setup, billing boundaries, escalation behavior, unsupported requests, and no-secret or no-hidden-reasoning rules.

If the suite list is empty, no active suite has been imported for the workspace yet. It does not mean the deployed agent is safe or unsafe by itself.

5. Inspect Runtime Evidence

Use Runtime Evidence and Recent Evidence when you need to trace what happened around a support answer.

Useful evidence includes:

  • Prompt IDs.
  • Trace IDs.
  • Runtime session or channel message IDs.
  • Prompt-composition hashes.
  • Redacted evidence pointers.

Eval Center stores summaries, IDs, hashes, and redacted pointers. It should not show raw private transcripts, hidden reasoning, secrets, raw tool arguments, or private workspace files.

6. Compare With Agent Studio

After checking Eval Center, open the agent in Agent Studio.

Review:

  • The published revision currently bound to the channel.
  • The Eval contract panel.
  • The prompt, knowledge, tools, model route, and policy.
  • Whether a new draft needs to be published.

If you changed knowledge or instructions, saving the draft is not enough. Publish a new revision, then bind that revision from the channel page.

7. Confirm With Channel Activity

For Website Support, finish with channel activity.

  1. Open Channels > Website Support.
  2. Send a real widget test from the target website.
  3. Confirm the message is accepted and delivered.
  4. Confirm the served agent and revision match the change you intended.
  5. Confirm the answer uses the expected knowledge and does not expose hidden reasoning, internal prompts, secrets, or private workspace state.

Channel Activity proves what happened for a specific message. Eval Center helps you aggregate and investigate quality signals across messages and eval runs.

What Eval Center Proves

Eval Center can help prove:

  • The workspace has recent channel evidence.
  • The system recorded prompt, trace, and redacted evidence pointers.
  • A behavioral eval run passed or failed when a run is present.
  • Delivery failures or feedback signals need follow-up.

What Eval Center Does Not Prove

Eval Center does not prove:

  • Every support topic is covered.
  • Every answer is correct.
  • Draft knowledge is live in the public widget.
  • Empty suites or empty runs mean the agent is safe.
  • A delivery success means the answer resolved the visitor's issue.

Use Eval Center with a small manual test set for the highest-risk support questions your agent must handle.

Troubleshooting

No Eval Runs Yet

If Behavioral Evals says no runs have been recorded, continue with manual widget tests and channel activity. Automated run results appear only after a service runner or operator records them.

No Active Eval Suites

If Eval Suites is empty, no active suite has been imported yet. Use the page as an operational gap, not as a pass or fail result.

Delivery Failures Increased

Open Website Support activity and inspect the recent failed messages. Check the embed key, allowed origin, serving agent revision, runtime health, and recent channel errors.

Evidence Is Missing

Missing evidence usually means the relevant message or run did not record the expected pointer. Use channel activity first, then escalate with the prompt ID, trace ID, runtime session, or exact timestamp that is available.

The Agent Still Answers With Old Knowledge

The channel may still be serving an older published revision. Publish a new revision, open Channels > Website Support, use Change serving agent, and serve the new revision.