Bonsai
Impact · Methodology

How we count the carbon

Bonsai claims to remove more carbon than it emits. This is the whole working: every constant, where it came from, and how confident we are in it.

What we actually promise

For every gram of CO₂e our infrastructure emits, Bonsai funds the removal of 10× that on Pro and 2× on the free plan. Both are net negative; only the size of the gift differs. The free plan is carried by us.

Two things about that sentence are deliberate.

"Our infrastructure", not "your prompts." We tried to account for inference alone and could not do it honestly — our provider reports a footprint for the whole organisation, across every product, and it cannot be decomposed into an inference-only slice. So the promise is defined against the whole figure. It is a simpler claim and an impossible one to under-deliver on.

Measured, not estimated. The token counts come from the provider's own usage reporting, never from your browser. What we do with them is the model below.

The part that makes the promise safe The model below is an estimate, and we say plainly how uncertain it is. The promise does not rest on it. At the end of each month we redeem the larger of our model's total and the provider's reported total. If our arithmetic is under, the provider's figure governs and we pay against that instead. An error in this page costs us accuracy in a chart; it cannot cost you removal.

The headline figures

A representative request — 1,500 prompt tokens and 300 generated, which is a typical tutoring exchange with context — costs and removes this much:

0.263 g emitted, one Everyday request
2.63 g removed for it on Pro
1.12 g emitted, same request on Deep Focus

Nearly all of the Everyday figure — 0.2562 g of it — is not the tokens at all. It is this month's share of fixed infrastructure that runs whether anyone prompts or not. The token work itself is under a hundredth of a gram. That ratio is the single most surprising thing in this model, and it is why a bigger model costs only about four times more per prompt today rather than the seven-odd times its marginal cost would suggest.

The chain

From a token count to a gram, one step at a time. Each step names the constant it applies; every constant is in the table below with its source.

  1. Tokens → GPU-seconds. Prompt and generated tokens are priced separately, because they cost genuinely different amounts of energy: prefill is compute-bound, generation is memory-bandwidth-bound. Each is divided by that model's measured aggregate throughput.
  2. GPU-seconds → watt-hours at the wall. GPU draw is scaled up by 1.4× to approximate whole-node power, because the host CPU, RAM, NICs and fans draw current too and the grid bills for all of it.
  3. Busy-hour → monthly average. Throughput was measured with the replica pushed toward saturation. Dividing by a busy-hour rate assumes a 100% duty cycle, so the figure is divided by a load factor of 0.35.
  4. Add embodied carbon. Manufacturing the GPUs, hosts and network gear emits before the hardware serves a single token. That debt amortises over its service life, added here as 1.2×.
  5. Watt-hours → grams CO₂e. Multiplied by the grid intensity of the region we serve from, 65 g CO₂e/kWh.
  6. Add the fixed baseline. Per request, not per token — idle infrastructure does not care how long your prompt was.
  7. Multiply by the promise. 10× on Pro, 2× on free. That is the removal we fund.
grams = BASELINE_PER_PROMPT × requests ← fixed infrastructure + prompt_tokens × Wh_per_prompt_token + completion_tokens × Wh_per_completion_token Wh_per_token = (GPU_W × GPUs × NODE_OVERHEAD) ÷ tokens_per_second ÷ 3600 × (EMBODIED_UPLIFT ÷ LOAD_FACTOR) removed = grams × 10 (Pro) = grams × 2 (free)

Deep Focus is charged a flat 4.25× the Everyday rate — the ratio of the two models' active parameters (17B against 4B). We use the ratio rather than the per-token physical model because the ratio is anchored to a published parameter count, and the absolute level of the physical model for that machine is the weakest number we have.

Every constant, and where it came from

ConstantValueSourceConfidence
Grid intensity 65 g/kWh Scaleway's published figure for the Paris region, derived from datacentre PUE and the French grid mix. PUE is already included — we do not apply it twice. Published
Node overhead 1.4× Accelerated nodes typically run 65–75% of node power through the GPUs. Mid-band. Industry band
Embodied uplift 1.2× Published lifecycle analyses put embodied carbon at 10–30% of lifetime operational emissions. We take the low end. Industry band
Load factor 0.35 Our best guess at the replica's time-averaged duty cycle. Published utilisation for inference fleets clusters at 10–40%; we sit at the optimistic end, which keeps the estimate honest rather than flattering. Not observable
Fixed baseline 2.44 g/day Derived from a complete provider-reported month (July 2026: 77.3 g over 295 prompts) minus what the token model accounts for. A residual, so it absorbs every error in the marginal model. Residual
Everyday · prefill 21,255 tok/s Measured 2026-08-07 against the live endpoint, from the extra time-to-first-token between a 26-token and a 10,036-token prompt, which cancels fixed network and queue overhead. About 38% of the card's compute ceiling — a normal sustained utilisation. Measured
Everyday · generation 7,425 tok/s Measured, best of several sustained runs at 80 concurrent streams. A second run of the same sweep peaked at 3,895 — the endpoint is shared, so what we measure is our slice of it at that moment. Measured, noisy
Deep Focus · generation 3,097 tok/s Measured, and unlike Everyday it reproduced: 3,097 tok/s at 64 streams and 3,088 at 48, two runs 0.3% apart. Measured
Deep Focus · prefill 40,000 tok/s Derived, not measured — the differential run was too noisy to use, so we applied Everyday's 38% derate to this machine's own compute ceiling. Derived
Deep Focus · GPUs 8 397B parameters at fp8 needs roughly 8×H100 to hold weights plus KV cache. The largest structural assumption left: if it is served on four, this model's Deep Focus figures halve. Assumed
Removal multiple 10× / 2× Not a measurement. This is the product promise — the only number here that is a choice. Our choice
A note on how small these numbers look We serve from Paris, and the French grid is unusually clean — roughly seven times below the world average, because it is heavily nuclear. A per-prompt figure from Bonsai will therefore look small next to published numbers for US-hosted assistants. That gap is real and it is in our favour, but it is a fact about French electricity, not evidence that we are more efficient than anyone else.

The water figure

The app also shows millilitres of water saved, and it is worth being plain that this number is not built like the carbon one. It is a flat 10.5 mL per prompt, multiplied by how many prompts you have sent this month.

It is not measured. Our provider reports tokens and a carbon footprint; it does not report water, so there is nothing to meter against. The constant descends from a published ~45 mL-per-response figure for a large hosted assistant, which we have now halved twice — a figure we choose to publish under rather than a quantity we derived. Real datacentre water use varies with the site, the season and the cooling design, and a saving measured against "other leading chatbots" depends entirely on which one you would otherwise have used.

How to read it As an order-of-magnitude comparison, not a meter. The carbon figure above is measured from your real token usage and backed by a promise we redeem monthly; the water figure is an illustration of the same choice, and we would rather say so here than let a precise-looking number imply a precision it does not have.

Check it yourself

Everything above comes out of one file, lib/carbon.mjs, which carries the full derivation in comments — including the arithmetic for the correction, the anchors we rejected and why, and a sensitivity table for the baseline. Nothing else in the app hardcodes a carbon number.

Two scripts are worth knowing about. scripts/check-carbon.mjs is the physics guard above. scripts/measure-energy.mjs is what produced the throughput figures, and will produce yours if you point it at your own endpoint.

If you find an error in this, we would rather hear it than not: [email protected]. A correction published is worth more to us than a number defended.

And this page is only the model. What was actually bought against it, month by month with the invoices, is at Removal receipts.

Start free Sign in Free forever — no card, ever