> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Evals: measuring ChatCLI

> chatcli eval runs a fixed suite of cases through the real chatcli binary in throwaway sandboxes, grades them with deterministic checks and an LLM judge, aggregates repeated trials, and fails CI when a change makes ChatCLI worse than a saved baseline.

The quality patterns (Self-Refine, CoVe, Reflexion) make ChatCLI check its own answer while it works. That helps the answer in front of you, but it leaves nothing to compare from one version to the next, and a model checking itself shares its own blind spots. **Evals** fill that gap. An eval suite is a fixed set of cases with known pass criteria, run the same way every time. It produces a score: pass rate, mean score, cost and latency. You can save that score as a baseline and gate CI on it.

`chatcli eval` is that harness. Use it to prove a prompt, model or engine change made ChatCLI better rather than worse, and to compare models and providers on the same tasks.

<Info>Evals spend real tokens: every case is a real model call, and judge checks add more. Use `--max-cost` to cap a run and `validate`/`list` to check a suite without spending anything.</Info>

***

## Quick start

```bash theme={"system"}
chatcli eval validate evals/                    # parse and validate only, no LLM and no cost
chatcli eval list evals/ --filter tag:coder     # what a run would execute

# Use a judge that is a different model from the candidate: self-grading is biased.
chatcli eval run evals/ --provider CLAUDEAI --model claude-sonnet-5-5 \
  --judge-provider OPENAI --judge-model gpt-6.1-sol \
  --trials 3 --max-cost 2 --out runs/sonnet.json --markdown runs/sonnet.md
```

A real run of the starter suite, with Claude Haiku 5.5 as the candidate and Claude Sonnet 5.5 as the judge, compared against an earlier run in which the Go bug-fix case had timed out. Each finished trial prints a line on stderr while the run works, then the summary goes to stdout:

```text theme={"system"}
[1/7] ✓ core/chat-arithmetic-exact #1  1.00 · 2.9s · $0.0000
[2/7] ✓ core/chat-json-only #1  1.00 · 3.3s · $0.0001
[3/7] ✓ core/chat-ptbr-language #1  1.00 · 3.7s · $0.0023
[4/7] ✓ core/chat-admits-unknown #1  0.97 · 3.9s · $0.0029
[5/7] ✓ core/coder-create-file #1  1.00 · 11.1s · $0.0040
[6/7] ✓ core/coder-fix-off-by-one #1  1.00 · 29.4s · $0.0062
[7/7] ✓ core/coder-read-only-question #1  1.00 · 17.2s · $0.0039

✓ core/chat-ptbr-language                  chat    1/1 trials  1.00  $0.0023
✓ core/chat-arithmetic-exact               chat    1/1 trials  1.00  $0.0000
✓ core/chat-json-only                      chat    1/1 trials  1.00  $0.0001
✓ core/chat-admits-unknown                 chat    1/1 trials  0.97  $0.0029
✓ core/coder-fix-off-by-one                coder   1/1 trials  1.00  $0.0062
✓ core/coder-create-file                   coder   1/1 trials  1.00  $0.0040
✓ core/coder-read-only-question            coder   1/1 trials  1.00  $0.0039

Passed 7 of 7 (100.0%) · failed 0 · errors 0 · skipped 0 · flaky 0
Mean score 1.00 · cost $0.0146 + judge $0.0049 · tokens 293835 in / 2540 out
Latency p50 3.9s · p95 29.4s · wall time 42.9s
Candidate: CLAUDEAI claude-haiku-5-5
Judge: CLAUDEAI claude-sonnet-5-5

Against the baseline:
  pass rate 85.7% → 100.0%
  score 0.85 → 1.00
  cost $0.0210 → $0.0195
  ✓ core/coder-fix-off-by-one now passes (error → pass)
```

A failing case lists the checks that failed under its line, for example `✗ tests pass: exited 1 (expected 0): --- FAIL: TestSum …`.

The repository ships a starter suite in `evals/` (chat language, JSON format, arithmetic, refusing to invent a flag, fixing a Go bug, creating a file, answering without touching files).

***

## How a trial runs

<Steps>
  <Step title="Sandbox">
    A temp directory is created. The case's `fixture` is copied into it, its inline `files` are written, and its `setup` commands are run (for example `git init`). A failing setup marks the trial as an **error**, and the candidate never runs.
  </Step>

  <Step title="Coder policy">
    For `coder` cases a workspace-local `coder_policy.json` (`merge: true`, so it layers over your global policy) is written. By default it allows `@coder`, including `exec`, inside the throwaway workspace. Set `policy:` on a case to narrow it. Safety-immune operations still require approval, and an eval has nobody to approve them, so they are denied. No "approve everything" switch exists for this.
  </Step>

  <Step title="Candidate">
    The real binary runs `chatcli -p "<prompt>" --raw --no-anim`, with `/coder ` in front for coder cases, with no stdin and under the case timeout. The run is **hermetic** by default: long-term memory, bootstrap, memory and session recall, session autosave, coder checkpoints and REPL history are off, so the outcome depends on the suite alone. `--with-memory` evaluates with your real memory and recall instead. Your [hooks](/extensions/hooks-system#turning-hooks-off) never fire in a candidate, with or without `--with-memory`: they are side effects on your machine, and a hook that runs an eval would recurse.
  </Step>

  <Step title="Record">
    Through `CHATCLI_EVAL_RECORD`, the one-shot reports its final answer, tool calls, turns, tokens, cost and transcript to a file outside the sandbox, which the candidate can neither read nor change.
  </Step>

  <Step title="Grading">
    Every check runs against the answer and the sandbox. Commands such as `go test ./...` run inside it. A trial passes only if **every** check passes, and its score is the weighted mean of the check scores.
  </Step>

  <Step title="Cleanup">
    The sandbox is deleted. `--keep` leaves it on disk and the report records its path.
  </Step>
</Steps>

<Warning>With `--with-memory` the run can also **write** to your memory: a one-shot queues its turn for memory extraction like any other run. Keep hermetic runs for baselines.</Warning>

***

## Suite format

A suite is a YAML file. Passing a directory loads every `*.yaml` / `*.yml` file directly inside it. Subdirectories are not scanned, so fixtures can carry YAML of their own. Keys are strict: a misspelled key is an error, never a check that is silently ignored.

```yaml theme={"system"}
name: core
judge: {provider: OPENAI, model: gpt-6.1-sol}   # default judge; flags override it
defaults:                                        # inherited by every case
  mode: chat            # chat | coder
  timeout: 4m
  trials: 1
  pass_policy: all      # all (pass^k, default) | any (pass@k) | majority
  env: {KEY: value}     # merged into each case's env
  setup: ["git init -q"]
  checks:               # appended to every case
    - not_regex: '(?i)\bas an ai (language )?model\b'

cases:
  - id: chat-ptbr-language            # letters, digits, . _ -
    tags: [chat, i18n]
    prompt: "Explique em no máximo duas frases o que é um mutex."
    checks:
      - regex: '(?i)(thread|goroutine|concorr|exclus)'
      - judge:
          rubric: "Written in Brazilian Portuguese, correct about mutual exclusion, at most two sentences."
          threshold: 0.7

  - id: coder-fix-off-by-one
    mode: coder
    fixture: fixtures/go-off-by-one   # a directory relative to the suite file (the starter suite inlines this module under files:)
    setup: ["git init -q && git add -A && git -c user.email=e@x -c user.name=e commit -qm fixture"]
    prompt: "The tests fail. Fix the bug in the code. Do not modify the tests."
    checks:
      - name: tests pass
        command: {run: "go test ./...", timeout: 3m}
      - name: tests untouched
        command: {run: "git diff --exit-code -- sum_test.go"}
      - tool_called: "@coder"
      - max_turns: 25
      - max_cost_usd: 0.25
```

| Case field | Meaning |
| - | - |
| `id` | Unique within the suite. Reports key cases as `<suite>/<id>` |
| `prompt` | What the candidate receives |
| `mode` | `chat` (plain one-shot) or `coder` (the full ReAct loop with tools) |
| `fixture` / `files` / `setup` | Sandbox contents: a directory to copy, inline files, shell commands to run first |
| `trials` / `pass_policy` / `timeout` / `env` | Per-case overrides of the defaults (`trials` 1–20) |
| `policy` | Coder policy rules for the sandbox, `{pattern, action: allow\|deny\|ask}` |
| `skip` | Keep the case in the suite without running it; the value is the reason |
| `checks` | At least one, after the defaults' checks are appended |

<Tip>A fixture is often deliberately broken, so keep it away from tooling that scans the host repository. A fixture directory that is a Go module needs its own `go.mod`, which keeps it out of `go test ./...`. Small fixtures can go inline under `files:` instead, and a YAML anchor (`files: &name` … `files: *name`) shares them between cases. The starter suite does this.</Tip>

***

## Checks

Deterministic checks come first: they are cheap, reproducible and not open to argument. Use the judge only for what has no deterministic test.

| Check | Passes when |
| - | - |
| `contains` / `not_contains` / `equals` | The final answer contains, does not contain, or equals the text (`ignore_case: true` available) |
| `regex` / `not_regex` | The answer matches, or does not match, the pattern |
| `json: {require_keys: [...]}` | The answer, or its first fenced or embedded JSON value, parses and carries the keys |
| `file_exists` / `file_absent` | A sandbox file exists, or does not |
| `file_contains` / `file_regex: {path, text}` | A sandbox file contains the text, or matches the pattern |
| `command: {run, expect_exit, contains, timeout}` | A shell command run in the sandbox after the candidate exits with `expect_exit` (default `0`), and its output contains `contains` when set. Default timeout `2m` |
| `tool_called` / `tool_not_called` | A tool was, or was not, called. Matches the name (`@coder`) or the name plus the start of its arguments (`@coder write`), in native and XML tool-call forms |
| `max_cost_usd` / `max_turns` / `max_duration` | The candidate stayed within the budget |
| `judge: {rubric, reference, threshold, samples}` | An LLM judge's median score over `samples` calls (default 1, max 5) is at least `threshold` (default `0.7`) |

Every check accepts `name:` (shown in reports) and `weight:` (its weight in the trial score, default 1). Paths in checks and `files` must be relative and stay inside the sandbox: absolute paths and `..` escapes are rejected when the suite loads.

### The judge

The judge receives the task, the rubric, the optional `reference` answer, the tools the candidate called and the candidate's answer. It does not learn which model produced the answer. It replies with `{"score": 0..1, "reasoning": "..."}`, and the reply is parsed leniently: a fenced block is accepted, and a score given on a 0–10 or 0–100 scale is normalized. A failed or unparsable judge call counts as a zero sample. It never counts as a silent pass.

* **Which model judges:** `--judge-provider`/`--judge-model`, else the first `judge:` a suite declares, else your configured default.
* **Self-grading is flagged:** when the judge is the same model as the candidate, the report marks the run `self_judged` and warns.
* **Judge spend** is booked on your real cost tracker and reported separately from the candidate's.

***

## Trials, pass\@k and pass^k

LLM output varies, so a case can run `trials` times. For each case the report gives:

| Metric | Meaning |
| - | - |
| `pass_rate` | Passed trials ÷ graded trials |
| `pass_at_k` | At least one trial passed |
| `pass_hat_k` | Every trial passed |
| `flaky` | Some trials passed and some did not |
| `mean_score` | Mean trial score |

The case verdict follows `pass_policy`. `all` (the default, pass^k) is strict: a case that passes two times out of three is flaky and fails, and the report says so instead of hiding it. `any` and `majority` are available per case. A trial whose run crashed or timed out is an **error**, kept separate from a **fail**, so "ChatCLI broke" never reads as "ChatCLI answered wrong".

***

## Baselines and the CI gate

`--out` saves the JSON report. A later run with `--baseline` compares case by case:

* **Regressions:** cases that passed before and do not pass now.
* **Fixes:** cases that did not pass before and pass now.
* **Added** and **removed** cases.
* Pass rate, mean score and cost before and after, measured only over the cases graded in both runs. A filtered run compared with a full baseline therefore compares like with like.

```bash theme={"system"}
chatcli eval run evals/ --out runs/main.json                           # on main
chatcli eval run evals/ --baseline runs/main.json --max-regressions 0  # on the branch
chatcli eval compare runs/main.json runs/branch.json                   # offline, no LLM
```

| Exit code | Meaning |
| - | - |
| `0` | Everything within the gates |
| `1` | Pass rate below `--min-pass-rate` (default `1.0`) |
| `2` | Usage error or invalid suite |
| `3` | Regression against `--baseline`: more case regressions than `--max-regressions`, or a pass-rate drop larger than `--tolerance` |

`--markdown` writes the same result as a Markdown document (summary table, baseline delta, failing checks) ready for a PR comment or a CI job summary.

***

## Command reference

| Form | Description |
| - | - |
| `chatcli eval run [path]` | Run, grade, report and gate. `path` is a suite file or directory, default `./evals` |
| `chatcli eval list [path] [--filter …]` | The cases a run would execute, with mode, trials, check count and tags |
| `chatcli eval validate [path]` | Parse and validate only. Boots no provider and spends nothing |
| `chatcli eval compare <baseline.json> <current.json>` | Compare two saved reports; same gate flags and exit codes as `run` |

| `run` flag | Description | Default |
| - | - | - |
| `--provider`, `--model` | Candidate provider and model | Your configured default |
| `--judge-provider`, `--judge-model` | Judge model | Suite `judge:`, else your default |
| `--trials <n>` | Override every case's trial count (1–20) | Per case |
| `--concurrency <n>` | Trials running at once | `2` |
| `--filter <list>` | Comma-separated case ids (globs such as `coder-*`) or `tag:<name>` | All cases |
| `--timeout <dur>` | Override every case's timeout | Per case (`5m`) |
| `--max-cost <usd>` | Stop starting trials once candidate + judge spend reaches it; remaining trials are reported as skipped | No cap |
| `--out <file>` / `--markdown <file>` | Write the JSON / Markdown report | — |
| `--baseline <file>` | Report to compare against | — |
| `--max-regressions <n>` / `--tolerance <0..1>` | Regression gate | `0` / `0` |
| `--min-pass-rate <0..1>` | Pass-rate gate | `1.0` |
| `--with-memory` | Evaluate with your real memory and recall | Off (hermetic) |
| `--keep` | Keep each trial's sandbox | Off |
| `--json` | Print the report as JSON on stdout (progress stays on stderr) | Off |
| `--quiet` | No per-trial progress | Off |
| `--bin <path>` | chatcli binary that runs the cases | This binary |

Ctrl+C stops scheduling new trials and kills the ones running. The report marks them skipped.

<Note>`CHATCLI_EVAL_RECORD` is the harness's private contract with the binary it drives: when set, a one-shot run writes its record to that path at exit. Unset, which covers every run not driven by `chatcli eval`, the one-shot path writes nothing.</Note>

***

## Evals and the quality pipeline

The quality patterns change how ChatCLI answers. Evals show whether that change was worth it. To measure, for example, whether turning on CoVe improves a model on your tasks, run the same suite twice and compare:

```bash theme={"system"}
chatcli eval run evals/ --out runs/cove-off.json
CHATCLI_QUALITY_VERIFY_ENABLED=true chatcli eval run evals/ --baseline runs/cove-off.json --min-pass-rate 0
```

The candidate inherits the environment, so any `CHATCLI_*` setting can be compared this way, as can a case's own `env:`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.