Every dollar figure derives from this table.
Prices are per million tokens at published list rates. Substitute your negotiated rates and the ratios hold; only the absolute numbers move.
| Endpoint | Input | Cache read | Output | Used for |
|---|---|---|---|---|
| Frontier baseline | $3.00 | $0.30 | $15.00 | The if-sent-direct column throughout |
| Small routed model | $0.05 | - | $0.40 | Trivial and simple buckets |
| Local handler or cache hit | $0.00 | $0.00 | $0.00 | Interception and cache exits |
Cache read at one tenth of input is the structural fact behind cache protection: a broken prefix costs about ten times what compressing that prefix saved.
The four demo requests, worked out.
The platform demo replays four requests. The fields in its wire log match the proxy session log schema: in_tokens, ratio, prospective_bucket, selected_provider, and result. The raw trace is shown below, and the platform display scales the dollar amounts by 24.27x so the examples match the $1,000/week heavy-usage workload.
| Request | Direct | Billed | Kept |
|---|---|---|---|
| 01 - long session | $1.69 | $0.64 | 61.9% |
| 02 - one-liner | $0.17 | $0.004 | 97.9% |
| 03 - cache hit | $0.36 | $0.00 | 100% |
| 04 - learned workflow | $2.79 | $0.00 | 100% |
| Sample of four | $5.01 | $0.65 | 87.1% |
This four-request sample is not a forecast. It has two free exits out of four requests, which is far above a normal mature registry. The 54-60% platform band comes from the benchmark below.
The headline band comes from a fixed traffic mix.
The 54-60% figure is produced by a 200-request suite with a 70% trivial, 20% simple, 10% complex traffic mix, priced at the rate card above and run through the full pipeline.
| Lever | Effect | Range | Basis |
|---|---|---|---|
| Compression | Reduction of prior-turn input tokens | 15-40% | Benchmark |
| Cache protection | Preserves the 90% discount on roughly 27% of baseline spend | - | Benchmark |
| Routing | Reduction of routable input cost | 47-95% | Benchmark |
| Combined | Reduction of the total bill | 54-60% | Benchmark |
| Learning loop | Reduction on intercepted requests only | 100% | Observed |
Traffic mix is the largest source of variance. A team doing mostly complex multi-tool agentic work will land below the band; a team doing many quick lookups will land above it. A pilot measures this first.
What the reference deployment shows.
These signals come from a machine running the proxy in daily use. They prove the mechanisms work end to end in real traffic; they are existence proof, not a clean production savings measurement.
request events logged
session log files
distinct models selected live
learning interception results
semantic cache exits
unit tests across Nexus optimizer and memory services
Caveat: that log set includes a developer's own traffic plus test and synthetic traffic, so aggregate compression ratio is not quoted as a production savings number.
The platform meter is scaled to one engineer-week.
A single session is too small to budget against, so the model uses one engineer doing heavy agentic coding for a 40-hour week.
| Horizon | Direct | After all four levers | Kept |
|---|---|---|---|
| Week | $1,000 | $377 | $623 |
| Month | $4,330 | $1,631 | $2,699 |
| Working year | $46,000 | $17,330 | $28,670 |
The working year uses 46 working weeks to allow for holiday and leave. The $4,330/month calculator default is the same $1,000/week model multiplied by 4.33.
Nexus reduces spend by shortening the conversation.
Every agent turn resends previous turns, so input cost grows with the square of the session length. Durable memory lets teams work in short sessions without losing state.
| Mechanism | Weekly | Derivation |
|---|---|---|
| No rediscovery | 9.5M | Heavy-use equivalent of 13,020 tokens saved per cold start x 6 cold starts/day x 5 days |
| Shorter prefixes | 14.8M | Heavy-use equivalent of about 7,800 fewer prefix tokens per turn x 78 affected turns/week |
| No re-litigating | 8.3M | Heavy-use equivalent of about 4 re-explanation exchanges/day at deep-session context depth |
| Survives compaction | 6.8M | Heavy-use equivalent of about 2 compaction-and-recover cycles/week |
| Total | 39.3M | 11.1% of a 354.85M-token heavy week |
The calculator does not double-count savings.
The levers apply in sequence, each to what the previous one left.
| Default | Value | Why |
|---|---|---|
h - Nexus | 11% | The four-mechanism table above against a 354.85M-token heavy week |
r - levers 1-3 | 54% | The bottom of the 54-60% benchmark band |
m - learned workflows | 8% | A mature registry after observation; day one is 0% |
What is not being claimed.
Not seat-plan savings
Savings apply to metered API spend. Flat-rate subscriptions need a different value model.
Learning starts at zero
A workflow must recur before it can be handled locally, and write-like workflows require approval.
Compression has tradeoffs
Cache-safe mode compresses less on purpose when a live cache prefix is more valuable.
Routing is guarded substitution
Cheap models handle cheaper buckets. Circuit breakers roll back if quality/fallback thresholds are breached.
Nexus is modelled
The arithmetic is exact, but cold-start and compaction frequencies vary by team.
Engineer time is excluded
Recovered time is likely larger than the token saving, but it is not counted in the calculator.
Where your data goes.
Both products run on your own hardware. Traffic goes from your machine to the provider you already configured; Tempell is not in the request path.
| Component | Where it runs | What it stores |
|---|---|---|
| Nexus local optimizer | Localhost, launchd, or systemd | Session logs and cost records on local disk |
| Semantic cache | Localhost | Workspace-hashed request/response bodies with TTL |
| Handler registry | Localhost | Generated handlers plus golden tests |
| Nexus context server | Your server, VM, or Pi | SQLite memories, tasks, graph rows, and embeddings |
| Provider keys | Your environment | Never transmitted anywhere except to the provider |
Run the pilot instead of trusting the page.
Two weeks, observe-only first. The proxy records what every request would have cost direct versus what it actually cost, per session, model, and engineer.
Install
Point one team's clients at Nexus.
Observe
Keep routing off for week one and record the traffic mix.
Enable
Turn routing on in week two and compare weeks.
Read
Use the dashboard at localhost to inspect deltas.
Decide
If measured savings do not clear the seat price, they do not clear it.