chatcli eval is that harness. Use it to prove a prompt, model or engine change made ChatCLI better rather than worse, and to compare models and providers on the same tasks.
Evals spend real tokens: every case is a real model call, and judge checks add more. Use
--max-cost to cap a run and validate/list to check a suite without spending anything.Quick start
✗ tests pass: exited 1 (expected 0): --- FAIL: TestSum ….
The repository ships a starter suite in evals/ (chat language, JSON format, arithmetic, refusing to invent a flag, fixing a Go bug, creating a file, answering without touching files).
How a trial runs
1
Sandbox
A temp directory is created. The case’s
fixture is copied into it, its inline files are written, and its setup commands are run (for example git init). A failing setup marks the trial as an error, and the candidate never runs.2
Coder policy
For
coder cases a workspace-local coder_policy.json (merge: true, so it layers over your global policy) is written. By default it allows @coder, including exec, inside the throwaway workspace. Set policy: on a case to narrow it. Safety-immune operations still require approval, and an eval has nobody to approve them, so they are denied. No “approve everything” switch exists for this.3
Candidate
The real binary runs
chatcli -p "<prompt>" --raw --no-anim, with /coder in front for coder cases, with no stdin and under the case timeout. The run is hermetic by default: long-term memory, bootstrap, memory and session recall, session autosave, coder checkpoints and REPL history are off, so the outcome depends on the suite alone. --with-memory evaluates with your real memory and recall instead. Your hooks never fire in a candidate, with or without --with-memory: they are side effects on your machine, and a hook that runs an eval would recurse.4
Record
Through
CHATCLI_EVAL_RECORD, the one-shot reports its final answer, tool calls, turns, tokens, cost and transcript to a file outside the sandbox, which the candidate can neither read nor change.5
Grading
Every check runs against the answer and the sandbox. Commands such as
go test ./... run inside it. A trial passes only if every check passes, and its score is the weighted mean of the check scores.6
Cleanup
The sandbox is deleted.
--keep leaves it on disk and the report records its path.Suite format
A suite is a YAML file. Passing a directory loads every*.yaml / *.yml file directly inside it. Subdirectories are not scanned, so fixtures can carry YAML of their own. Keys are strict: a misspelled key is an error, never a check that is silently ignored.
Checks
Deterministic checks come first: they are cheap, reproducible and not open to argument. Use the judge only for what has no deterministic test.
Every check accepts
name: (shown in reports) and weight: (its weight in the trial score, default 1). Paths in checks and files must be relative and stay inside the sandbox: absolute paths and .. escapes are rejected when the suite loads.
The judge
The judge receives the task, the rubric, the optionalreference answer, the tools the candidate called and the candidate’s answer. It does not learn which model produced the answer. It replies with {"score": 0..1, "reasoning": "..."}, and the reply is parsed leniently: a fenced block is accepted, and a score given on a 0–10 or 0–100 scale is normalized. A failed or unparsable judge call counts as a zero sample. It never counts as a silent pass.
- Which model judges:
--judge-provider/--judge-model, else the firstjudge:a suite declares, else your configured default. - Self-grading is flagged: when the judge is the same model as the candidate, the report marks the run
self_judgedand warns. - Judge spend is booked on your real cost tracker and reported separately from the candidate’s.
Trials, pass@k and pass^k
LLM output varies, so a case can runtrials times. For each case the report gives:
The case verdict follows
pass_policy. all (the default, pass^k) is strict: a case that passes two times out of three is flaky and fails, and the report says so instead of hiding it. any and majority are available per case. A trial whose run crashed or timed out is an error, kept separate from a fail, so “ChatCLI broke” never reads as “ChatCLI answered wrong”.
Baselines and the CI gate
--out saves the JSON report. A later run with --baseline compares case by case:
- Regressions: cases that passed before and do not pass now.
- Fixes: cases that did not pass before and pass now.
- Added and removed cases.
- Pass rate, mean score and cost before and after, measured only over the cases graded in both runs. A filtered run compared with a full baseline therefore compares like with like.
--markdown writes the same result as a Markdown document (summary table, baseline delta, failing checks) ready for a PR comment or a CI job summary.
Command reference
Ctrl+C stops scheduling new trials and kills the ones running. The report marks them skipped.
CHATCLI_EVAL_RECORD is the harness’s private contract with the binary it drives: when set, a one-shot run writes its record to that path at exit. Unset, which covers every run not driven by chatcli eval, the one-shot path writes nothing.Evals and the quality pipeline
The quality patterns change how ChatCLI answers. Evals show whether that change was worth it. To measure, for example, whether turning on CoVe improves a model on your tasks, run the same suite twice and compare:CHATCLI_* setting can be compared this way, as can a case’s own env:.