Work within a spend limit
A hard dollar ceiling, and what stella does as it approaches one — where the run stops and what the scope gate stops before you spend anything.
You have a number in mind, and you don't want to find out you went over it afterward. Several mechanisms hold that ceiling. One of them fires before you spend anything at all.
The flag
stella --spend-limit 2.00 run "port the utils module to the new API"--spend-limit is a global flag, so it can
go on either side of the subcommand. Putting it first is a good habit,
since it reads as a property of the whole command, which is what it is.
The default. Spend is tracked for the end-of-run cost summary and the local store, but it never stops a run. Use this to find out what a kind of task actually costs before you pick a ceiling for it.
A hard cap on the whole run or session — not per turn, not per round. Must be a positive, finite number.
Set it once for a whole shell session or CI job instead:
export STELLA_SPEND_LIMIT=5Where the run stops
This part is worth trusting, because it decides whether a stopped run leaves you with usable work or a mess.
Spend is checked between steps and stages, never in the middle of a tool call. So stopping a run never leaves a half-finished edit, a partly written file, or a tool whose changes happened but never got recorded. Work already saved to disk stays there. Nothing gets rolled back.
The exit code is 1, the same as any other failure. If you need to tell
the spend limit stopped this apart from this just failed, read the
JSON output:
stella --spend-limit 2.00 run --output-format json "…" | jq '{status, cost_usd, files_touched}'--spend-limit is measured against published catalog prices. Local and
offline models have no catalog price, so their cost always reads zero, and
a spend limit has nothing to measure against for them. Instead, bound those
runs with a verification plugin's --test-command check and goal-mode
round limits — see
Local and air-gapped.
The check that runs before you spend
This check is not built into stella itself. A plain stella run does
not have it, and the agent_engine_config.headless_scope_bypass setting
does nothing either way. The section below describes what a scope review is
for, because that behavior is something a verification plugin can add on
its own side of the connection. Treat it as something to look for in a
plugin, not as behavior this program has by default.
A spend limit stops a run partway through. A scope review stops it before the first edit, which is usually what you actually wanted.
When a plan's size crosses any of three limits, stella pauses and shows you the plan:
max_stepsMore than this many steps in the plan.
Default 5
max_filesEstimated number of files touched above this.
Default 8
max_cost_usdEstimated cost above this.
Default 1.00
Every check is a strict "greater than," so a plan that lands exactly on a
limit passes without pausing. When it pauses, a approves the plan, t
trims it down to the largest version that would not have triggered the
pause, and x cancels it. Anything longer than one of those letters is
read as what you want changed, and sends the plan back with your note
attached.
A gate like this needs one rule to be useful in an unattended run: it must never approve itself automatically. With nobody there to answer, a plan that's over the limits should stop the run, not wave itself through — because a scope check that approves itself in CI isn't really a check. That's something to verify a plugin actually does. It isn't a setting this program controls, and no flag here turns it on or off.
What the money is actually spent on
Before you tune anything, know the shape of the bill. Most attempts to save money target the wrong stage.
| Stage | Model calls | Cost share |
|---|---|---|
| Triage | 1 | negligible |
| Plan | 1 (+1 repair, rarely) | small |
| Witness | ~1 engine turn | only when you pass no --test-command |
| Execute | 5–40+ steps | most of the bill |
| Verify | 0 | free — the deterministic checks |
| Verifier | 0–1 per verification | cents |
Three things follow from that, and they're most of what you need to know:
- The worker's model price dominates the bill. Switching the worker to a cheaper model changes the total by several times over. Switching the verifier barely moves it at all.
- A top-tier verifier is cheap insurance. One verifier call is a few thousand input tokens and under a thousand output tokens. Keep it on a different model family than the worker.
--test-commandis the strongest cost-saving flag on a wrapped run. It arms the plugin's built-in fail-then-pass check, so most runs finish on that evidence alone, skipping the verifier entirely — and it also skips the witness-writing call. It's not accepted on the raw loop, which never makes those calls in the first place.
# Cheapest correct shape when the result must carry evidence:
# the plugin's oracle armed, a hard ceiling on top.
stella --spend-limit 1.50 run --pipeline my-verifier \
--test-command "cargo test -p api" "add cursor pagination to /orders"The multipliers
A ceiling only helps if you know what's pushing against it.
Goal modeRepeats checked rounds until the goal is met, up to 8. Each round is
a full working turn. stella monitor runs on goal mode, so it costs the
same way.
FleetsOne engine run per task. Worker spend counts against the parent
spend limit, so --spend-limit is one shared ceiling across the whole
group of tasks, not a separate allowance per worker — and the whole run
stops early once it's used up.
Best-of-NRuns N different attempts and keeps the best one by verification score, at N times the normal cost. This is strictly something you turn on, for exactly this reason.
RevisionsA failed verification sends the worker back with the evidence attached, limited to 2 revision attempts by default.
When the conversation outgrows the window
A long run's cost isn't only what it does. It's also how much conversation it has to carry while doing it. When the conversation gets close to the compaction budget (150,000 tokens by default), stella shortens it and keeps going, instead of failing or silently cutting content.
Shortening happens in four steps, from least to most lossy:
- Remove duplicates. If the exact same tool output appears more than once, only the earliest copy is kept — later copies just point back to it.
- Remove outdated results. When the same call ran twice with the same input, only the newest result reflects the current state, so the older ones are stripped out as stale.
- Trim old content. If it's still over budget, large old outputs get
cut down to their beginning and end. The beginning keeps the tool's
framing (a
PASSED/FAILEDline, file headers), and the end keeps the errors — so what's left is the part worth reading. - Drop content entirely. Only after that are the oldest large outputs replaced with a placeholder — and a result that's still the most recent one for its call is never dropped.
The system message and your latest message are never touched.
Deduplication keeps the earliest copy. Identical content doesn't depend on where it sits, so shortening the later copy instead would put the change in the newest part of the conversation and break the provider's cached prefix. Doing it the other way would save a few tokens and then lose much more money to a cold cache than it saved.
If a conversation is still over budget after all four steps, an overflow
summarizer runs — a real, extra model call, recorded as --call-seq 1 for
that step. Seeing these in stella inspect is
the sign that a run's conversation, not its actual work, is what got
expensive.
Picking the number
Set your spend limit from what you've actually seen, not a guess. Run the kind of task once with no limit, check what it cost, then set a ceiling with some room to spare:
stella run "the representative task" # observed mode
stella stats --format text # what it actually cost
stella --spend-limit 2.00 run "the next one" # now with a ceilingThe number worth tracking over time is cost per task solved, not the sticker price — a cheap model that needs three tries isn't actually cheap. The Observatory tracks this per model, and Examples & recipes has full setups at four different price points.
Next
The reference for observed versus enforced spend, and what the local store records.
Complete settings files from very cheap to maximum quality, with cost estimates.
Accounting for a run after the fact, down to the individual model call.
Every stage, what it costs, and when it gets skipped.
Make a change in an unfamiliar codebase
A repository you have never read, and a change that has to land in it. Orient with the code graph, find the blast radius, and ship the change verified.
Understand what a run cost, and why
Follow one expensive run from the dollar figure down to the exact bytes the model was sent — stats, the Observatory, stella inspect. Entirely local.