Version Compare

Persist, inspect, and approve or block baseline-versus-candidate evaluation deltas.

Baseline
The current production agent version — the known-good reference point
Model
Prompt version
Rulebook version
Retrieval top K
Candidate
The new agent version being evaluated for release readiness
Model
Prompt version
Rulebook version
Retrieval top K
Suite Readiness
Number of human-approved test cases available for comparison runs. Both baseline and candidate run against this identical frozen suite.
Trusted test cases
0
Critical-risk cases
0
Completed eval runs
0
Persisted Comparison
Select two completed evaluation runs and save a permanent record of how the candidate version compares to the baseline across all trusted test cases.
Quality Gate
Automated release decision based on configured thresholds for critical recall, precision, regressions, and hallucination rate.
Persist a comparison to evaluate the quality gate.

No persisted comparison yet

Complete baseline and candidate runs, then persist a comparison here to unlock metrics and quality-gate evidence.

Default Quality Gate Thresholds
A candidate version is blocked from release if any of these thresholds are not met.
Critical finding recall>= 0.95
Finding precision>= 0.90
Audit citation precision>= 0.98
Rule-reference accuracy>= 0.98
Schema validity= 1.00
CAP completeness>= 0.95
Critical regressions= 0
Hallucinated finding rate<= 0.02