Skip to main content
The quality patterns (Self-Refine, CoVe, Reflexion) make ChatCLI check its own answer while it works. That helps the answer in front of you, but it leaves nothing to compare from one version to the next, and a model checking itself shares its own blind spots. Evals fill that gap. An eval suite is a fixed set of cases with known pass criteria, run the same way every time. It produces a score: pass rate, mean score, cost and latency. You can save that score as a baseline and gate CI on it. chatcli eval is that harness. Use it to prove a prompt, model or engine change made ChatCLI better rather than worse, and to compare models and providers on the same tasks.
Evals spend real tokens: every case is a real model call, and judge checks add more. Use --max-cost to cap a run and validate/list to check a suite without spending anything.

Quick start

A real run of the starter suite, with Claude Haiku 5.5 as the candidate and Claude Sonnet 5.5 as the judge, compared against an earlier run in which the Go bug-fix case had timed out. Each finished trial prints a line on stderr while the run works, then the summary goes to stdout:
A failing case lists the checks that failed under its line, for example ✗ tests pass: exited 1 (expected 0): --- FAIL: TestSum …. The repository ships a starter suite in evals/ (chat language, JSON format, arithmetic, refusing to invent a flag, fixing a Go bug, creating a file, answering without touching files).

How a trial runs

1

Sandbox

A temp directory is created. The case’s fixture is copied into it, its inline files are written, and its setup commands are run (for example git init). A failing setup marks the trial as an error, and the candidate never runs.
2

Coder policy

For coder cases a workspace-local coder_policy.json (merge: true, so it layers over your global policy) is written. By default it allows @coder, including exec, inside the throwaway workspace. Set policy: on a case to narrow it. Safety-immune operations still require approval, and an eval has nobody to approve them, so they are denied. No “approve everything” switch exists for this.
3

Candidate

The real binary runs chatcli -p "<prompt>" --raw --no-anim, with /coder in front for coder cases, with no stdin and under the case timeout. The run is hermetic by default: long-term memory, bootstrap, memory and session recall, session autosave, coder checkpoints and REPL history are off, so the outcome depends on the suite alone. --with-memory evaluates with your real memory and recall instead. Your hooks never fire in a candidate, with or without --with-memory: they are side effects on your machine, and a hook that runs an eval would recurse.
4

Record

Through CHATCLI_EVAL_RECORD, the one-shot reports its final answer, tool calls, turns, tokens, cost and transcript to a file outside the sandbox, which the candidate can neither read nor change.
5

Grading

Every check runs against the answer and the sandbox. Commands such as go test ./... run inside it. A trial passes only if every check passes, and its score is the weighted mean of the check scores.
6

Cleanup

The sandbox is deleted. --keep leaves it on disk and the report records its path.
With --with-memory the run can also write to your memory: a one-shot queues its turn for memory extraction like any other run. Keep hermetic runs for baselines.

Suite format

A suite is a YAML file. Passing a directory loads every *.yaml / *.yml file directly inside it. Subdirectories are not scanned, so fixtures can carry YAML of their own. Keys are strict: a misspelled key is an error, never a check that is silently ignored.
A fixture is often deliberately broken, so keep it away from tooling that scans the host repository. A fixture directory that is a Go module needs its own go.mod, which keeps it out of go test ./.... Small fixtures can go inline under files: instead, and a YAML anchor (files: &name … files: *name) shares them between cases. The starter suite does this.

Checks

Deterministic checks come first: they are cheap, reproducible and not open to argument. Use the judge only for what has no deterministic test. Every check accepts name: (shown in reports) and weight: (its weight in the trial score, default 1). Paths in checks and files must be relative and stay inside the sandbox: absolute paths and .. escapes are rejected when the suite loads.

The judge

The judge receives the task, the rubric, the optional reference answer, the tools the candidate called and the candidate’s answer. It does not learn which model produced the answer. It replies with {"score": 0..1, "reasoning": "..."}, and the reply is parsed leniently: a fenced block is accepted, and a score given on a 0–10 or 0–100 scale is normalized. A failed or unparsable judge call counts as a zero sample. It never counts as a silent pass.
  • Which model judges: --judge-provider/--judge-model, else the first judge: a suite declares, else your configured default.
  • Self-grading is flagged: when the judge is the same model as the candidate, the report marks the run self_judged and warns.
  • Judge spend is booked on your real cost tracker and reported separately from the candidate’s.

Trials, pass@k and pass^k

LLM output varies, so a case can run trials times. For each case the report gives: The case verdict follows pass_policy. all (the default, pass^k) is strict: a case that passes two times out of three is flaky and fails, and the report says so instead of hiding it. any and majority are available per case. A trial whose run crashed or timed out is an error, kept separate from a fail, so “ChatCLI broke” never reads as “ChatCLI answered wrong”.

Baselines and the CI gate

--out saves the JSON report. A later run with --baseline compares case by case:
  • Regressions: cases that passed before and do not pass now.
  • Fixes: cases that did not pass before and pass now.
  • Added and removed cases.
  • Pass rate, mean score and cost before and after, measured only over the cases graded in both runs. A filtered run compared with a full baseline therefore compares like with like.
--markdown writes the same result as a Markdown document (summary table, baseline delta, failing checks) ready for a PR comment or a CI job summary.

Command reference

Ctrl+C stops scheduling new trials and kills the ones running. The report marks them skipped.
CHATCLI_EVAL_RECORD is the harness’s private contract with the binary it drives: when set, a one-shot run writes its record to that path at exit. Unset, which covers every run not driven by chatcli eval, the one-shot path writes nothing.

Evals and the quality pipeline

The quality patterns change how ChatCLI answers. Evals show whether that change was worth it. To measure, for example, whether turning on CoVe improves a model on your tasks, run the same suite twice and compare:
The candidate inherits the environment, so any CHATCLI_* setting can be compared this way, as can a case’s own env:.