CI integration & regression gating¶
Wire an eval into continuous integration so a pull request that degrades your skill
fails the build. The mechanism is deliberately small: you declare per-judge
thresholds in eval.yaml, and the scorer
exits non-zero when a run falls below them — which is all a CI job needs to gate on.
This is an evolving area
The harness ships the primitives for CI gating (thresholds + a non-zero exit), and the GitHub Actions job below works today. Turnkey CI recipes, reusable actions, and caching patterns are still being fleshed out — see remaining work. Treat the example as a starting point to adapt, not a drop-in.
How gating works¶
Scoring aggregates each judge by value type — boolean judges become a pass_rate,
numeric judges become a mean, and a --baseline pairwise
comparison produces a win_rate. Regression detection compares those aggregates against
the thresholds block and, on any breach, returns exit code 1.
flowchart TD
A[score.py judges] --> B{thresholds<br/>configured?}
B -- no --> P[exit 0]
B -- yes --> C[detect_regressions<br/>aggregates vs thresholds]
C --> D{any breach?}
D -- no --> E["print REGRESSIONS: 0<br/>exit 0"]
D -- yes --> F["print REGRESSIONS: N detected<br/>exit 1 → CI fails"]
Thresholds¶
thresholds is a top-level map of judge name → gate. Each gate sets one or more of
three keys; a run regresses if the matching aggregate is below the floor.
| Key | Applies to | Aggregate compared | Passes when |
|---|---|---|---|
min_pass_rate |
boolean judges | pass_rate (fraction of cases that returned True) |
pass_rate >= min_pass_rate |
min_mean |
numeric judges (e.g. 1–5 LLM scores) | mean (average value across cases) |
mean >= min_mean |
min_win_rate |
a pairwise judge (needs --baseline) |
win_rate |
win_rate >= min_win_rate |
judges:
- name: has_content # boolean judge → pass_rate
check: |
content = outputs.get("main_content", "")
return (len(content.strip()) >= 100,
f"{len(content.strip())} chars")
- name: output_quality # numeric judge (1–5) → mean
prompt: "Score the output 1-5 for completeness, clarity, and accuracy."
thresholds:
has_content: { min_pass_rate: 1.0 } # every case must pass
output_quality: { min_mean: 3.5 } # average score must stay >= 3.5
An unavailable metric counts as a regression
If a threshold is set but its aggregate is None, that is reported as a regression,
not silently skipped. This is almost always a config mistake — the judge was skipped
for every case (its if: condition, or it errored), or the key targets the wrong
judge type (e.g. min_pass_rate on a numeric judge, whose pass_rate is always
None). Match the key to the judge's value type.
Exit-code behavior¶
Two entry points enforce thresholds; both sys.exit(1) on regression so a CI runner
fails the step automatically.
score.py judges runs the judges and, if thresholds is set, checks them at the end
of scoring — so a normal /eval-run already gates. It prints REGRESSIONS: 0 or
REGRESSIONS: N detected (one line per breach) before exiting.
Runs live under $AGENT_EVAL_RUNS_DIR/<eval-name>/<run-id>/ (default eval/runs/), so
pass a stable --run-id (e.g. the commit SHA) to find the same directory later in the job.
Comparing against a baseline¶
Beyond absolute floors, you can gate a run relative to a previous one with
--baseline <run-id>. The baseline must be a prior run under the same eval-name.
/eval-run --baseline <run-id>adds a position-swapped pairwise judge on top of the regular judges. Each case is judged both A/B and B/A, and only a consistent preference counts as a win;summary.yamlgains apairwisesection (wins_a,wins_b,ties). Gate it with amin_win_ratethreshold.score.py regression --baseline <run-id>also does a direct aggregate comparison: formeanandpass_rate, a current value more than0.5below the baseline is flagged asDegraded vs baseline— catching drift even where no absolute floor was crossed.
# Score this run with a pairwise comparison against last week's baseline,
# then gate both the absolute thresholds and the vs-baseline deltas.
python3 skills/eval-run/scripts/score.py pairwise \
--run-id "$RUN_ID" --baseline 2026-07-09-opus --config eval.yaml
python3 skills/eval-run/scripts/score.py regression \
--run-id "$RUN_ID" --baseline 2026-07-09-opus --config eval.yaml
Absolute vs relative gates
Use min_mean / min_pass_rate for a hard quality floor that must always hold, and
--baseline for catch-any-drift protection between a known-good run and the PR under
test. They compose — a run can pass its absolute floors yet still fail because it
dropped sharply versus the baseline.
A GitHub Actions job¶
This skeleton shows how to gate a PR on thresholds. The score.py regression
step is the real copy-pasteable gate (exit 1 fails the job). Producing the run
itself is not a bare shell /eval-run — slash commands need Claude Code
(headless) or another driver; see the note below.
name: skill-eval
on: [pull_request]
jobs:
eval:
runs-on: ubuntu-latest
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
AGENT_EVAL_RUNS_DIR: eval/runs
RUN_ID: ci-${{ github.sha }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install the harness
run: pip install -e .
# NOT a shell builtin — drive /eval-run via Claude Code headless
# (or your own wrapper). See the note + headless guide.
# Expected output: eval/runs/<eval-name>/$RUN_ID/summary.yaml
- name: Run the eval
run: |
echo "TODO: invoke /eval-run headlessly — see guides/headless.md"
exit 1
# Gate: exits 1 on any threshold breach, failing the job.
- name: Check for regressions
run: |
python3 skills/eval-run/scripts/score.py regression \
--run-id "$RUN_ID" --config eval.yaml
- name: Upload report
if: always()
uses: actions/upload-artifact@v4
with:
name: eval-report
path: eval/runs/**/report.html
Driving the run in CI
/eval-run is a Claude Code slash command, not a binary on PATH. In CI, run
it through Claude Code headless (or an equivalent driver that
prepares workspaces, executes cases, and scores). For heavier or containerized
CI, use the same eval.yaml on Harbor with --runner harbor —
the config is unchanged; only the substrate flag differs.
Where to go next¶
-
Threshold semantics
The full reference for
min_mean,min_pass_rate, andmin_win_rate. -
Pairwise & sampling
How
--baselinecomparisons and repeated-sample stability work. -
Run headless
Auto-answer questions and gate external services for unattended runs.