Same coding agents. Smaller bill.

Probe0 runs under Claude Code, Codex and Cursor on your machine. It cuts what they cost and shows you what every call spent.

All traffic · this machine
Today / cap$18.44 / $25.00
  • claude-code · 41822sonnet-4.6 · $0.0732Upstream
  • codex · 41903qwen3:8b · local · $0.0000Local
  • claude-code · 42011opus-4.7 · blockedCapped

1proxy slot

Your machine has one, and every tool wants it. So you pick an inspector, a cache or a router, and go without the rest.

93%of the 1-hour cache

Claude Code pays for an hour of cache it then uses for five minutes. Probe0 buys the five minutes.

92%of the bill is cache

51% reads, 41% writes, across 85,569 real turns. The money goes on buying the cache, not on your prompt.

Four ways it makes the same work cost less

It runs under the tools you already use, and any step can be turned off. On 85,569 real turns it cut the bill 12.8% with no model swaps, and 24% with swaps on.

  1. 01

    Intercept

    Install it once. Every agent you run goes through it — Claude Code, Codex, Cursor.

  2. 02

    Route

    Easy work goes to a model on your own machine. Hard work still goes to the best one. A weak local answer gets sent up automatically.

  3. 03

    Reuse

    Ask the same thing twice and the second answer is free. Two agents asking at once only pay once.

  4. 04

    Account

    Every call gets written down: model, tokens, cost, time, which agent. A spend cap pauses a run before it gets expensive.

Turn on only the parts you want

Ten switches. Each reports what it saved, so nothing stays on out of faith.

Send less

Strip Compression
Shrinks tool output before you pay for it.
Caveman
Optional terse mode. No articles, no hedges.

Send it somewhere cheaper

Local Routing
Whole classes of work run on your own model.
Cloud Routing
Easy turns go to a cheaper model in the same family.
Effort step-down
Same model, less thinking, on easy turns.

Don't send it at all

Exact Cache
Repeat questions never hit the network again.
Semantic Cache
Near-identical questions too, held to a strict match floor.
Duplicate Collapse
Two agents ask at once, you pay once.

Know what it cost

Spend Guard
A hard cap per run or per day. Warns, then pauses.
Recording
Every call, the agent behind it, cost you can check.

Pay for less output, keep the cache

Rewrite a prompt and you lose the copy your provider is already holding. Probe0 only shrinks what your tools send back, so that copy keeps paying off.

Strip Compressionon by default
Bash output comes back about 30% smaller, the turn it arrives. Old file reads shrink later, once the provider has stopped holding them. Worth about three points of the bill.
Cavemanoptional
A terse mode. Articles, linking verbs and hedges go, word by word. No saving is claimed for it.

One slot, and anyone can build in it

Long term, this is meant to be the layer other people build on.

Only one thing can hold the slot
Your machine has one proxy position. An inspector, or a cache, or a router. Not all three.
So the slot should be programmable
Probe0 takes the position and then opens it up. A module sees a request and decides what happens to it.
A module asks for what it needs
Read the details, change the body, reach the network — and nothing else. You approve that list when you install it.
Everything published is reviewed
Asking up front is what makes review possible: only what a module asked for has to be checked.

The cheapest call never leaves your machine

Plenty of agent work doesn't need the best model on the market.

Your models, your hardware
Point Probe0 at a model already running in Ollama or LM Studio. It becomes another tier.
Or a cheaper cloud model
Step down to a cheaper model in the same family instead. Requests carrying tool calls take the safe step only.
You set the share, not a rule
Both are a setting: how much of your traffic takes the cheaper path. Left at the default, about 11% of turns.
Nothing leaves your machine
The proxy, the cache and the record of every call all run here. There is no server to phone home to.

The month you didn't need that plan

Probe0 knows what you actually used. So it can tell you when you're paying for a tier above the one you need. Usually worth more than any cache will save you.

It works the other way too. Hitting your limit every week? It says so.

Last 30 days · this machine
Billed to provider
$96.20
Answered from cache
2,317
Duplicate calls collapsed
1,842
Cache writes avoided
3,106

You could drop to the tier below. You used 48% of a Max 20x plan over these 30 days.

Try it on your own agents

Private beta on macOS. Install it, point your shell at it, and it starts counting on the next run. Nothing gets uploaded.

macOS · Apple silicon and Intel.

No spam. One email when the beta opens.

Ask an AI instead

Opens a chat that already knows Probe0, rough edges included.