# Validating a finding

tj optimize --validate downsize replays a sample of your own recorded calls against both models with your own API key, spending real money, and reports a measured token and cost delta plus a quality check. downsize is the only finding it accepts today.

---

Every analyzer finding is an estimate. `--validate` turns one into a measurement, by replaying a small sample of your own recorded calls twice and comparing what came back.

```bash
tj optimize --validate downsize
```

This spends real money on your own API key. An interactive run shows the cost ceiling and waits for you to confirm before the first call goes out. `-y` and `--json` both skip that prompt, so a scripted run spends without asking.

## What it does

For each sampled call, tj replays the prompt it recorded against the baseline model and again against the downsize candidate. It sums the tokens and the cost of both runs, and compares the two outputs for a quality signal. The result reports the measured token delta, the measured cost delta, and how many of the N outputs matched.

`downsize` is the only finding `--validate` accepts today. The `--validate` option's own choice list is the check: if a finding name is not in it, the engine has no candidate builder for that finding yet. Trim and Summarize share the same measurement engine and need a per-finding candidate builder and a content-presence quality signal, which is not written. Cache is deterministic and out of scope by design.

## The three gates

`--validate` refuses outright instead of half-running. Each refusal names what to do.

**Prompt capture must be on.** The replay needs the recorded prompt, so `[capture] prompts` has to be `true`. Capture defaults on, so this only fires when it has been turned off. Turn it back on, let a few captured calls accumulate, and retry.

**A provider key must be present.** Set `TJ_ANTHROPIC_API_KEY` in your environment. Version one replays same-family Claude models, so the key is Anthropic's. tj makes the calls with your key and holds nothing.

**The daemon must not hold the database lock.** `--validate` reads raw captured attributes, which needs a direct DuckDB connection and not the serve shim. Stop the daemon with `tj stop` and run it again.

If no captured call in the window matches the downsize candidate shape, tj says so instead of sampling something else. Widen `--since`, or let more captured calls accumulate.

## Sample size and cost

```bash
tj optimize --validate downsize --samples 10
tj optimize --validate downsize -y          # skip the confirmation prompt
```

`--samples` defaults to 5 and is capped at 20. The cap is a cost bound. Each sample is replayed twice, once on the baseline and once on the candidate, so ten samples means twenty live calls.

Before spending anything, tj prints the sample count and an estimated cost ceiling and asks you to confirm. Declining spends nothing. `-y` skips that prompt, and `--json` implies it, so a scripted run does not hang on a confirmation nobody will see.

## Reading the result

The labels below are the ones `tj` prints; the figures are illustrative.

```
Measured on a sample of 5 of your recorded calls (finding: downsize)
  Tokens:  41.2k -> 38.9k (-6%)
  Cost:    $0.62 -> $0.19 (-69%)
  Quality: preserved 4/5 (exact-match on output)
```

The quality line is the point, and it is worth knowing exactly what the CLI's `preserved` counter measures. It counts outputs that matched the baseline exactly after whitespace and case are normalized. That is a strict signal and a narrow one. A candidate output that is correct, useful and worded differently counts as not preserved. Exact-match suits deterministic and tool-shaped work, which is the class the finding targets, and it is blunt on open-ended prose, so a low number is a prompt to read the outputs rather than a verdict.

`--json` emits the same result with per-call measurements, for a script or a report.

## What a measured delta does and does not settle

A measured delta on five calls is a measured delta on five calls. The caveat is carried as a dataclass default so no surface can drop it:

> Measured on a sample of your own recorded calls, not a guarantee. A larger sample and human review of the outputs remain worthwhile before switching models.

Validation does not retire the downsize caveat. The finding still says "candidate", still says "review before switching", and a validated finding is not a finding that is safe to apply. What changed is the basis: the token and dollar figures came from calls that actually ran instead of from a heuristic. The judgment about whether the candidate model does your work is still yours, and it is made by reading the outputs.

## See also

- [Downsize](/docs/optimize-downsize) — the finding `--validate` measures
- [How the analyzers work](/docs/optimize-overview) — where estimates come from
- [Configuration](/docs/configuration) — the `[capture]` toggles