T
Tempell AI
Evidence & methodology

Where the numbers come from.

Every figure on the platform page is labelled as observed, benchmarked, or modelled. This page shows the rate card, the worked arithmetic, the sequencing rules, and the limits of the claim.

Claim provenance
Observed

Signals from a running deployment.

Benchmarked

A defined synthetic suite and fixed rate card.

Modelled

Arithmetic from stated assumptions or your inputs.

Rate card

Every dollar figure derives from this table.

Prices are per million tokens at published list rates. Substitute your negotiated rates and the ratios hold; only the absolute numbers move.

EndpointInputCache readOutputUsed for
Frontier baseline$3.00$0.30$15.00The if-sent-direct column throughout
Small routed model$0.05-$0.40Trivial and simple buckets
Local handler or cache hit$0.00$0.00$0.00Interception and cache exits

Cache read at one tenth of input is the structural fact behind cache protection: a broken prefix costs about ten times what compressing that prefix saved.

Demo arithmetic

The four demo requests, worked out.

The platform demo replays four requests. The fields in its wire log match the proxy session log schema: in_tokens, ratio, prospective_bucket, selected_provider, and result. The raw trace is shown below, and the platform display scales the dollar amounts by 24.27x so the examples match the $1,000/week heavy-usage workload.

01 - deep in a long session direct 18,942 in x $3.00/M = $0.056826 + 850 out x $15.00/M = $0.012750 = $0.069576 proxy 12,100 cache read = $0.003630 + 3,370 fresh in = $0.010110 + 850 out = $0.012750 = $0.026490 saved 61.9%
02 - a one-line question direct 1,240 in x $3.00/M = $0.003720 + 210 out x $15.00/M = $0.003150 = $0.006870 proxy 1,240 in x $0.05/M = $0.000062 + 210 out x $0.40/M = $0.000084 = $0.000146 saved 97.9%
03 - asked before direct 1,910 in x $3.00/M = $0.005730 + 620 out x $15.00/M = $0.009300 = $0.015030 proxy semantic cache hit = $0.000000 saved 100%
04 - learned workflow direct 4 agent turns 22,900 in total = $0.068700 + 3,090 out x $15/M = $0.046350 = $0.115050 proxy handler executed locally = $0.000000 saved 100%
RequestDirectBilledKept
01 - long session$1.69$0.6461.9%
02 - one-liner$0.17$0.00497.9%
03 - cache hit$0.36$0.00100%
04 - learned workflow$2.79$0.00100%
Sample of four$5.01$0.6587.1%

This four-request sample is not a forecast. It has two free exits out of four requests, which is far above a normal mature registry. The 54-60% platform band comes from the benchmark below.

Benchmark

The headline band comes from a fixed traffic mix.

The 54-60% figure is produced by a 200-request suite with a 70% trivial, 20% simple, 10% complex traffic mix, priced at the rate card above and run through the full pipeline.

LeverEffectRangeBasis
CompressionReduction of prior-turn input tokens15-40%Benchmark
Cache protectionPreserves the 90% discount on roughly 27% of baseline spend-Benchmark
RoutingReduction of routable input cost47-95%Benchmark
CombinedReduction of the total bill54-60%Benchmark
Learning loopReduction on intercepted requests only100%Observed

Traffic mix is the largest source of variance. A team doing mostly complex multi-tool agentic work will land below the band; a team doing many quick lookups will land above it. A pilot measures this first.

Observed

What the reference deployment shows.

These signals come from a machine running the proxy in daily use. They prove the mechanisms work end to end in real traffic; they are existence proof, not a clean production savings measurement.

70,929

request events logged

954

session log files

10+

distinct models selected live

135

learning interception results

4

semantic cache exits

1,923

unit tests across Nexus optimizer and memory services

Caveat: that log set includes a developer's own traffic plus test and synthetic traffic, so aggregate compression ratio is not quoted as a production savings number.

40-hour week

The platform meter is scaled to one engineer-week.

A single session is too small to budget against, so the model uses one engineer doing heavy agentic coding for a 40-hour week.

one engineer - 40 hours - heavy Claude/API spend - sent direct 1,280 high-context request groups 216.02M input x $3.00/M = $648.06 117.72M cache r x $0.30/M = $ 35.32 21.11M output x $15.00/M = $316.62 ------------------------------------------------ 354.85M tokens = $1,000 / week
HorizonDirectAfter all four leversKept
Week$1,000$377$623
Month$4,330$1,631$2,699
Working year$46,000$17,330$28,670

The working year uses 46 working weeks to allow for holiday and leave. The $4,330/month calculator default is the same $1,000/week model multiplied by 4.33.

Nexus model

Nexus reduces spend by shortening the conversation.

Every agent turn resends previous turns, so input cost grows with the square of the session length. Durable memory lets teams work in short sessions without losing state.

one long session - 60 turns b = 12,000 cold prefix g = 900 tokens added per turn 60 x 12,000 + 900 x (60 x 61 / 2) = 720,000 + 1,647,000 = 2,367,000 input tokens
six bootstrapped sessions - 10 turns each b = 4,200 small prefix + bootstrap g = 900 tokens added per turn per session: 10 x 4,200 + 900 x 55 = 91,500 x 6 sessions = 549,000 saved 76.8%
MechanismWeeklyDerivation
No rediscovery9.5MHeavy-use equivalent of 13,020 tokens saved per cold start x 6 cold starts/day x 5 days
Shorter prefixes14.8MHeavy-use equivalent of about 7,800 fewer prefix tokens per turn x 78 affected turns/week
No re-litigating8.3MHeavy-use equivalent of about 4 re-explanation exchanges/day at deep-session context depth
Survives compaction6.8MHeavy-use equivalent of about 2 compaction-and-recover cycles/week
Total39.3M11.1% of a 354.85M-token heavy week
Sequencing

The calculator does not double-count savings.

The levers apply in sequence, each to what the previous one left.

base = engineers x monthly spend - Nexus memory base x h = volume removed before optimization - Nexus optimization (base - memory) x r = compression, cache, routing - learned workflows remainder x m = local interceptions = net spend
DefaultValueWhy
h - Nexus11%The four-mechanism table above against a 354.85M-token heavy week
r - levers 1-354%The bottom of the 54-60% benchmark band
m - learned workflows8%A mature registry after observation; day one is 0%
Limits

What is not being claimed.

Not seat-plan savings

Savings apply to metered API spend. Flat-rate subscriptions need a different value model.

Learning starts at zero

A workflow must recur before it can be handled locally, and write-like workflows require approval.

Compression has tradeoffs

Cache-safe mode compresses less on purpose when a live cache prefix is more valuable.

Routing is guarded substitution

Cheap models handle cheaper buckets. Circuit breakers roll back if quality/fallback thresholds are breached.

Nexus is modelled

The arithmetic is exact, but cold-start and compaction frequencies vary by team.

Engineer time is excluded

Recovered time is likely larger than the token saving, but it is not counted in the calculator.

Deployment

Where your data goes.

Both products run on your own hardware. Traffic goes from your machine to the provider you already configured; Tempell is not in the request path.

ComponentWhere it runsWhat it stores
Nexus local optimizerLocalhost, launchd, or systemdSession logs and cost records on local disk
Semantic cacheLocalhostWorkspace-hashed request/response bodies with TTL
Handler registryLocalhostGenerated handlers plus golden tests
Nexus context serverYour server, VM, or PiSQLite memories, tasks, graph rows, and embeddings
Provider keysYour environmentNever transmitted anywhere except to the provider
Pilot

Run the pilot instead of trusting the page.

Two weeks, observe-only first. The proxy records what every request would have cost direct versus what it actually cost, per session, model, and engineer.

1

Install

Point one team's clients at Nexus.

2

Observe

Keep routing off for week one and record the traffic mix.

3

Enable

Turn routing on in week two and compare weeks.

4

Read

Use the dashboard at localhost to inspect deltas.

5

Decide

If measured savings do not clear the seat price, they do not clear it.

Want this measured against your own traffic?

Book a consultation and we will scope a two-week observe-only pilot on one of your teams.

Book a Consultation