Sign in Book a demo 134

Optimize: the analyzers

Fifteen analyzers read your real usage history, surface reviewable candidates, and never auto-apply. A run covers the ones that apply to your workload. The honesty posture, the registry order, and plan-tier-aware rendering.

tj optimize runs the fifteen analyzers in the registry over the telemetry already in your local DuckDB store, minus any that do not apply to your workload. Several are skipped for an interactive-coding-agent window and several others for an SDK window; when an analyzer does not run lists both, and says why a skipped analyzer is not the same as a finding of zero. Each analyzer reads your real usage history and surfaces cost-saving candidates. You review them. You decide what to change.

tj optimize                      # the last 30 days, minus the analyzers skipped for your workload
tj optimize downsize cache       # run only the named analyzers
tj optimize --since 7d           # narrow the window

The model

Every analyzer works the same way at heart:

  1. It reads spans from your DB. Nothing leaves your machine.
  2. It looks for a structural pattern in that data.
  3. It reports the pattern as a candidate for you to look at.

An analyzer never rewrites a prompt, never changes a model, never edits a config. The output is a list of things worth reviewing, with the evidence attached so you can spot-check before you act.

Honesty posture

This is the part that matters most, so it comes first.

An analyzer reports what it can measure. It does not claim what it cannot. A downsize candidate means “the shape of this work matches a class where a cheaper model is worth a look” — not “this would have worked on the cheaper model.” A reuse cluster reports “these plans look structurally identical” — not “these plans were interchangeable.” Every savings figure is labelled estimated recoverable, computed from a stated heuristic over the analyzed window. None of them is a promise.

You will never see “saves you,” “safe to downgrade,” “provably unchanged,” or “certified” in TokenJam output. Where an analyzer flags a candidate, it also prints the caveat that says why the flag is structural and where the limits are. Those caveats are baked into the code so they cannot be dropped by accident.

The analyzers

The registry order is the order tj optimize runs them in and the order the CLI reports them.

AnalyzerCLI nameWhat it looks for
DownsizedownsizeSessions whose shape matches a cheaper-model candidate
Budget Projectionbudget-projectionYour run rate against any configured provider spend ceiling
CachecacheCurrent cache usage, plus breakpoint suggestions for Anthropic prompts
Cache Recommendcache-recommendAnthropic-only cache_control breakpoint candidates from stable prefixes in your prompt history
ResendresendContext your agent already sent in an earlier turn, sent again
ScriptscriptTool sequences that repeat identically across many runs
ReusereuseSessions that re-plan the same work
TrimtrimLow-significance regions in captured prompts
SubagentsubagentPer-subagent cost breakdown and right-sizing candidates
SummarizesummarizePrompt files worth summarizing without breaking their structure
RelearnrelearnBlockers your agents keep silently re-hitting across sessions, never written into a durable fix
VerbosityverbosityAgent/model cohorts whose output runs long relative to a baseline
DeadweightdeadweightConfigured MCP servers and always-injected context you never actually use
Stream Usagestream-usageStreaming calls that closed without reporting token usage
ShippedshippedThe commits each session produced, at a labelled confidence, and the measured cost of the sessions that left none or had them all reverted

placement is a sixteenth name the command accepts. It is not a registered analyzer: asking for it runs downsize and surfaces the batch-placement card from that run. For an interactive-coding-agent window the card is left out, so an absent card there is not a finding of no candidates.

For the spend-facing side, where your tokens actually went, see Cost visibility.

Plan-tier-aware rendering

What each finding shows depends on how you pay. API users see dollar figures. Subscription users (Claude Pro / Max, ChatGPT Plus / Team / Enterprise) see token-share framing instead, because a flat monthly fee has no per-call dollar spend to reclaim. Local (Ollama) users see token counts only. When TokenJam can’t tell your plan, dollar figures are suppressed and it points you at tj init --reconfigure.

The rendering rules live in one place (core/framing.py) so the CLI, the REST API, and the Lens dashboard all agree.