How we count the carbon
Bonsai claims to remove more carbon than it emits. This is the whole working: every constant, where it came from, and how confident we are in it.
What we actually promise
For every gram of CO₂e our infrastructure emits, Bonsai funds the removal of 10× that on Pro and 2× on the free plan. Both are net negative; only the size of the gift differs. The free plan is carried by us.
Two things about that sentence are deliberate.
"Our infrastructure", not "your prompts." We tried to account for inference alone and could not do it honestly — our provider reports a footprint for the whole organisation, across every product, and it cannot be decomposed into an inference-only slice. So the promise is defined against the whole figure. It is a simpler claim and an impossible one to under-deliver on.
Measured, not estimated. The token counts come from the provider's own usage reporting, never from your browser. What we do with them is the model below.
The headline figures
A representative request — 1,500 prompt tokens and 300 generated, which is a typical tutoring exchange with context — costs and removes this much:
Nearly all of the Everyday figure — 0.2562 g of it — is not the tokens at all. It is this month's share of fixed infrastructure that runs whether anyone prompts or not. The token work itself is under a hundredth of a gram. That ratio is the single most surprising thing in this model, and it is why a bigger model costs only about four times more per prompt today rather than the seven-odd times its marginal cost would suggest.
The chain
From a token count to a gram, one step at a time. Each step names the constant it applies; every constant is in the table below with its source.
- Tokens → GPU-seconds. Prompt and generated tokens are priced separately, because they cost genuinely different amounts of energy: prefill is compute-bound, generation is memory-bandwidth-bound. Each is divided by that model's measured aggregate throughput.
- GPU-seconds → watt-hours at the wall. GPU draw is scaled up by 1.4× to approximate whole-node power, because the host CPU, RAM, NICs and fans draw current too and the grid bills for all of it.
- Busy-hour → monthly average. Throughput was measured with the replica pushed toward saturation. Dividing by a busy-hour rate assumes a 100% duty cycle, so the figure is divided by a load factor of 0.35.
- Add embodied carbon. Manufacturing the GPUs, hosts and network gear emits before the hardware serves a single token. That debt amortises over its service life, added here as 1.2×.
- Watt-hours → grams CO₂e. Multiplied by the grid intensity of the region we serve from, 65 g CO₂e/kWh.
- Add the fixed baseline. Per request, not per token — idle infrastructure does not care how long your prompt was.
- Multiply by the promise. 10× on Pro, 2× on free. That is the removal we fund.
Deep Focus is charged a flat 4.25× the Everyday rate — the ratio of the two models' active parameters (17B against 4B). We use the ratio rather than the per-token physical model because the ratio is anchored to a published parameter count, and the absolute level of the physical model for that machine is the weakest number we have.
Every constant, and where it came from
| Constant | Value | Source | Confidence |
|---|---|---|---|
| Grid intensity | 65 g/kWh | Scaleway's published figure for the Paris region, derived from datacentre PUE and the French grid mix. PUE is already included — we do not apply it twice. | Published |
| Node overhead | 1.4× | Accelerated nodes typically run 65–75% of node power through the GPUs. Mid-band. | Industry band |
| Embodied uplift | 1.2× | Published lifecycle analyses put embodied carbon at 10–30% of lifetime operational emissions. We take the low end. | Industry band |
| Load factor | 0.35 | Our best guess at the replica's time-averaged duty cycle. Published utilisation for inference fleets clusters at 10–40%; we sit at the optimistic end, which keeps the estimate honest rather than flattering. | Not observable |
| Fixed baseline | 2.44 g/day | Derived from a complete provider-reported month (July 2026: 77.3 g over 295 prompts) minus what the token model accounts for. A residual, so it absorbs every error in the marginal model. | Residual |
| Everyday · prefill | 21,255 tok/s | Measured 2026-08-07 against the live endpoint, from the extra time-to-first-token between a 26-token and a 10,036-token prompt, which cancels fixed network and queue overhead. About 38% of the card's compute ceiling — a normal sustained utilisation. | Measured |
| Everyday · generation | 7,425 tok/s | Measured, best of several sustained runs at 80 concurrent streams. A second run of the same sweep peaked at 3,895 — the endpoint is shared, so what we measure is our slice of it at that moment. | Measured, noisy |
| Deep Focus · generation | 3,097 tok/s | Measured, and unlike Everyday it reproduced: 3,097 tok/s at 64 streams and 3,088 at 48, two runs 0.3% apart. | Measured |
| Deep Focus · prefill | 40,000 tok/s | Derived, not measured — the differential run was too noisy to use, so we applied Everyday's 38% derate to this machine's own compute ceiling. | Derived |
| Deep Focus · GPUs | 8 | 397B parameters at fp8 needs roughly 8×H100 to hold weights plus KV cache. The largest structural assumption left: if it is served on four, this model's Deep Focus figures halve. | Assumed |
| Removal multiple | 10× / 2× | Not a measurement. This is the product promise — the only number here that is a choice. | Our choice |
The water figure
The app also shows millilitres of water saved, and it is worth being plain that this number is not built like the carbon one. It is a flat 10.5 mL per prompt, multiplied by how many prompts you have sent this month.
It is not measured. Our provider reports tokens and a carbon footprint; it does not report water, so there is nothing to meter against. The constant descends from a published ~45 mL-per-response figure for a large hosted assistant, which we have now halved twice — a figure we choose to publish under rather than a quantity we derived. Real datacentre water use varies with the site, the season and the cooling design, and a saving measured against "other leading chatbots" depends entirely on which one you would otherwise have used.
Check it yourself
Everything above comes out of one file, lib/carbon.mjs, which carries the full derivation in comments — including the arithmetic for the correction, the anchors we rejected and why, and a sensitivity table for the baseline. Nothing else in the app hardcodes a carbon number.
Two scripts are worth knowing about. scripts/check-carbon.mjs is the physics guard above. scripts/measure-energy.mjs is what produced the throughput figures, and will produce yours if you point it at your own endpoint.
If you find an error in this, we would rather hear it than not: [email protected]. A correction published is worth more to us than a number defended.
And this page is only the model. What was actually bought against it, month by month with the invoices, is at Removal receipts.
Bonsai