measured — read from a real file on a real machine. A measurement, not a guess.
modelled — projected or estimated. Tap the chip for the basis and the assumptions behind it.
~N% measured — a total that mixes both. The chip states the exact split; the product refuses to present a mixed total as a pure measurement.
This is the product's third principle, applied to its own marketing.
RUNS LOCALLY · NO CODE LEAVES THE MACHINE · EVERY FIGURE TRACEABLE
GitHub now bills your organisation by the token — and stops working when the budget runs out. It tells you what you spent. It has never told you what you bought.
TokenLens reads the usage records Copilot already writes to your engineers' machines and turns an unexplained total into an itemised bill.
1 SEPTEMBER 2026
Copilot's launch allowances were promotional. On 1 September they dropped — the same subscription fee, substantially less included capacity. Everything beyond the allowance is billed as overage. When the pooled budget is exhausted, Copilot stops.
There is no cheaper fallback model. There is no degraded mode. It stops.
COPILOT BUSINESS · CREDITS / USER / MONTH
COPILOT ENTERPRISE · CREDITS / USER / MONTH
1 AI credit = $0.01 · Credits are pooled across the organisation, do not carry over, and are forfeited monthly.
THE FINDING NOBODY HAS PRICED
Modern AI assistants can use tools — a database connector, a ticketing integration, a documentation searcher. For the AI to use a tool, that tool must be described to it in full.
Not once. On every single request.
Install six integrations, use two, and you pay for six — on every request, from every developer, for as long as they remain installed.
THE TOLL GATE — DRAG THE SLIDER
This cost scales with what you have installed — not with what you use. It is invisible, it is large, and it is the most trivially reducible cost in the entire bill.
The model has no memory between steps. When the AI reads a file, then runs a command, then decides what to do next, everything it has gathered so far is re-sent at every step so it can “remember”.
Across real sessions, every token of genuinely new content was billed roughly ten times over.
Some of this is unavoidable — it is how the technology works. A meaningful share is not, and that is what we go after.
Because the whole conversation is re-sent every turn, a chat gets more expensive with every message. By the thirty-third message, each new message costs 61% more than it would in a fresh chat.
The fix is “start a new chat when you start a new task.”
It costs nothing. It requires no tooling. Almost nobody does it — because nobody can see the meter running.
COST PER MESSAGE · MESSAGE NUMBER
This is what an absent cost signal looks like. Nobody is being careless — the information simply does not exist at the moment the decision is made.
THE ENCOURAGING PART
Spend is extremely concentrated. The most expensive single session cost $32.70; the typical session cost $3.61.
This means the intervention is surgical, not cultural. Fix the worst tenth and you capture most of the value — without a training programme, a policy memo, or a single all-hands.
57 REAL SESSIONS · SIZED BY SPEND · SORTED DESCENDING
top session 8,644.6 credits
THE LARGEST NINE · 50% OF ALL SPEND
There is a whole industry of AI monitoring tools. None of them can see Copilot — for a structural reason, not a lack of effort.
Those tools work by sitting in the network path: you route your AI traffic through them and they inspect it as it passes. That works when your own application calls an AI provider.
Copilot does not work that way. The traffic goes from the editor to GitHub over a channel you do not control and cannot intercept. There is no point in the middle to stand.
The only place this information exists in readable form is a file on the developer's own machine — in an undocumented format that changes with every VS Code release. Reading it reliably is genuinely difficult.
That difficulty is exactly why the category is empty. It is also our moat.
HOW MONITORING TOOLS WORK
A place to stand.
HOW COPILOT WORKS
The local journal on the developer's own machine.
We did not find a better place to stand. We found that the receipt was already on the floor.
THE PRODUCT
TokenLens reads the usage records already on your engineers' machines, shows you exactly what every credit bought, names what was wasted and why, and hands your IT team a configuration change that removes a large part of it — without asking a single developer to work differently.
Read the records Copilot already writes. Produce an exact credit ledger. Every figure traceable to the byte it came from.
Available todayAssign wasted spend to fourteen named, individually detectable causes. Waste stops being a total and becomes a list of specific, fixable things.
Available todayReplay real recorded sessions under a proposed change and report what it would have saved — before anything is deployed.
On the roadmapEmit the exact configuration your IT team pushes through the device-management tooling you already own. No developer action.
On the roadmapRun it as a randomised trial against a holdout group on your own fleet. Report the actual saving, including where we were wrong.
On the roadmapSteps 01–02 are working software running on real data today. Steps 03–05 are specified and sequenced — see the roadmap below.
There is no server to run, no proxy to configure, no agent to install, and no data pipeline to build.
An engineer runs a single command on their own machine. It reads the local Copilot records and prints the ledger. Nothing is uploaded. Nothing is installed permanently. This takes minutes, not a project.
A local dashboard opens in the browser — running on that machine, reachable only from that machine. It shows the itemised bill: where the credits went, which sessions, which tools, which behaviours.
Repeat across a representative group of engineers — twenty is plenty. Only aggregate counts are combined. Source code never moves.
TokenLens produces the exact settings file. Your IT team pushes it through the same tooling used for every other setting. Developers notice nothing except that things stop running out.
Only aggregate counts and one-way hashes. Offline by default.
Files are opened for reading. Nothing is ever written, moved, or altered.
No proxy, no server, no agent, no pipeline.
One command prints the file and byte position behind any figure.
RUNNING TODAY, ON REAL DATA
Everything below is the tool's real dashboard — reconstructed from its actual source (packages/core/src/dashboard) — showing the live figures it reports on the developer machine it was built on: 713 requests across 57 chat sessions over 61 active days. Every figure carries its own receipt.
Time-series shapes are illustrative; every figure is real. Switch views with the sidebar — exactly as in the running tool.
| Cost centre | Tokens | Share |
|---|---|---|
| ■ Conversation history | 8,270,722 | 49% measured |
| ■ Tool results | 3,832,965 | 23% measured |
| ■ Tool descriptions | 2,601,022 | 16% measured |
| ■ Attached files | 1,231,513 | 7% measured |
| ■ System instructions | 789,442 | 5% measured |
| Model | Credits | Requests | Share | cr / 1k | Rate samples |
|---|
| Plan | enterprise |
| Monthly allowance | 3,900 credits |
| Month-to-date | — |
| Projected month-end | 9,922.0 modelled |
TokenLens makes no network calls, so it cannot read your organisation's actual limit. Set it with --allowance, TOKENLENS_MONTHLY_ALLOWANCE, or "monthlyAllowance" in .tokenlens/config.json.
Projected to exceed the allowance by 6,022.0 credits. modelled
Illustrated with the measured monthly run-rate of 9,922 credits.
Pick a custom or bounded range to compare it against the preceding window.
Change: remove integrations that were never invoked; keep the rest. Confidence high · fixed by: platform team, one settings change. Measured: 17,016 tokens per request, 21% of the median prompt. measured
Measured: 42.0% of file reads were repeats; one file was read 72 times in a single session. measured
In 115 days, automatic summarisation ran 84 times at a median of 92 seconds each — about two hours of engineers watching a progress indicator. measured
…plus W3–W8 and W10–W14 — the full fourteen are listed below.
“We found no waste” and “we could not look” are opposite claims — the dashboard keeps them visibly apart.
Ten views — Overview, Budget, Sessions, Models, Tools, Waste, Day, Projects, Data, Settings. Collapsible with Ctrl+B.
Three themes (system / light / dark — system follows the OS) and a compact density toggle. The tool respects the machine's existing choice.
Every view is driven by two filters: which workspaces count (scope) and which dates count (range). An empty scope says so loudly — a blank dashboard must never read as “you spent nothing”.
Six headline figures: credits, per-active-day, month-to-date, models in use, measured share, flagged days. Each carries its provenance chip.
Line chart with anomaly flags — days far enough from the median to be worth a look, detected by median absolute deviation, not σ.
The five-way cost-centre donut with the token table beneath. The same five colours, everywhere, forever.
Stacked credits per day. Answers “is the premium share growing” — the question a daily total cannot.
Only measured days are drawn. Click a day to open its detail view.
The provenance legend lives in the footer of every screen; Ctrl+K jumps to any view. The chips are the product's honesty principle, rendered as UI.
tokenlens dashboard serves on 127.0.0.1 behind a token — reachable only from that machine. There is no cloud half of this product.
Only 24% of these requests had a credit figure published by GitHub. The rest are estimated from rates measured on the ones that did — and every one of those is labelled as an estimate, never as a fact.
TAKE IT WITH YOU
Three exporters, one rule: an exported file is the artefact most likely to be forwarded to someone other than its subject, so it defaults to the share-safe scope. Per-session detail is withheld unless you explicitly ask for it.
The file the finance team receives. Same figures, same chips, zero dependencies.
Plan enterprise · monthly allowance 3,900 credits (plan default) · projected month-end 9,922.0 credits — projected to exceed the allowance by 6,022.0 credits.
| Cost centre | Tokens | Share | Credits |
|---|---|---|---|
| Conversation history | 8,270,722 | 49% | 8,242.6 |
| Tool results | 3,832,965 | 23% | 4,325.0 |
| Tool descriptions | 2,601,022 | 16% | 1,930.0 |
| Attached files | 1,231,513 | 7% | 1,021.0 |
| System instructions | 789,442 | 5% | 577.6 |
11 models in use. 60 requests on the most expensive model cost 17,872 credits; 212 requests on the cheapest cost 2,782 — three and a half times fewer requests, six and a half times more money.
Re-transmission multiplier 9.8× · 33rd-message penalty +61% · duplicate file-read rate 42.0%.
Fourteen named causes, each with its evidence and remedy — see the Waste board. Findings may overlap; they do not sum to the ledger.
Listed separately, with what data was missing and what would unblock each one. “We found no waste” and “we could not look” are opposite claims.
Installed integrations and their per-request description cost — the Tool-Definition Tax, itemised per tool.
tokenlens report writes the month as Markdown — structured for handing to an AI assistant or pasting into a doc. Same section spine: how to read this, the month in figures, allowance and forecast, where the tokens went, model mix, conversation economics, daily pattern, what the detectors found, priced model substitution, tool surface.
A report exists to be sent somewhere, so the safe scope is the default and the permissive one has to be asked for. --self keeps per-session detail — and the report tells you, in its header, exactly what was withheld and why.
“You are wasting money” is not actionable. TokenLens attributes wasted spend to fourteen specific, individually detectable causes — each with a way to detect it and a named remedy.
Paying to describe installed tools on every request, including tools never invoked.
The AI re-reads a file it already has in front of it.
A single command returns an enormous result that then rides along in every later step.
A chat left open across unrelated tasks, re-sending irrelevant history forever.
The most expensive model used for work a cheaper one handles identically.
Very long agent loops that consume heavily and produce no surviving change.
Credits spent on work that never reached the codebase.
Fifty engineers independently asking the same thing about the same internal system.
Automatic summarisation triggered by unmanaged context growth.
Standing instruction files that grew past their usefulness.
Paying for previews of files that were never opened.
Titles, summaries and commit messages generated on the most expensive model available.
Deep-reasoning settings applied to work that does not need them.
Expensive exploration done in the main conversation instead of an isolated cheaper one.
In 115 days, automatic summarisation ran 84 times and took a median of 92 seconds each — about two hours of engineers watching a progress indicator. That is not a billing problem. That is a productivity problem you were also not being shown.
THE PART THAT SURPRISES PEOPLE
The instinctive assumption is that reducing AI spend means asking engineers to use it less. That would be slow, unpopular, and would cost you more in lost productivity than it saved.
That is not the intervention.
The largest savings come from settings — which model handles which kind of work, how large a single tool result may be, which integrations stay installed. These are configuration values. They are deployed the same way every other setting in your organisation is deployed, through tooling you already own.
No training. No policy memo. No behaviour change. Your engineers keep working exactly as they do now.
TIER A — CONFIGURATION ONLY
Settings deployed through your existing device management.
Developer effort: None. They will not notice.
Availability: roadmap phase D5
TIER B — LIVE GUARDRAILS
Automatic prevention of specific wasteful actions as they happen.
Developer effort: None. Prevention is silent.
Availability: roadmap phase D6
Both figures are modelled — projected from measured cost shares, not yet observed on a deployed policy. That is precisely why the product's final phase is a randomised trial on your own fleet, with a holdout group, so the real figure replaces the projected one. We will report the gap between them, including if it is unflattering.
This product was built in a deliberate sequence: measure before attributing, attribute before recommending, recommend before enforcing, and prove before claiming. Each phase is only useful because the one before it is trustworthy.
The honest-measurement guarantees, enforced in code.
Exact, traceable accounting of every credit.
The itemised bill, on screen and exportable.
The fourteen named causes, detected and ranked.
“This change would have saved N last quarter.”
The exact settings file your IT team deploys.
Prevention at the moment of spend.
Live cost visibility for the developer.
Measured savings with confidence intervals.
Fleet-wide view and self-tuning policy.
Phases 1–3 are working software running on real data today. Everything shown in this presentation as measured came out of them. Everything shown as modelled is waiting on phase 9 to become measured.
WHAT THIS LOOKS LIKE FROM THE INSIDE
Priya is a senior engineer. She is good at her job. On this particular Tuesday she does nothing wrong, nothing unusual, and nothing anybody would flag in a review.
Watch the meter.
This day is constructed from measured medians — a typical request, a typical session penalty, a typical compaction. It is an illustration of measured behaviour, not a recording of one real Tuesday.
PRIYA'S TUESDAY, DECOMPOSED
Of Priya's Tuesday, roughly a third was avoidable — re-reads, a stale window, an oversized result, and a model heavier than the task needed.
Priya did nothing wrong. She was never shown a number. Every decision that created this was made in the absence of information that already existed on her own laptop.
Now multiply Priya by five thousand.
ILLUSTRATIVE EXAMPLE · MODELLED FROM MEASURED RATES
5,000 engineers · GitHub Copilot Enterprise · 3,900 credits included per seat per month
YOUR NUMBER, NOT OURS
This is a projection, not a quote.
Only aggregate counts and one-way hashes. Network access is denied by default.
One command prints the file and byte position behind any figure on any report.
Enforced in the type system — the code physically cannot print an unlabelled estimate. It is why every number on this page carries a chip.
Files opened for reading only. Nothing written, moved, or altered.
If something is not understood, it is reported and counted — never silently reported as zero.
No proxy, no server, no agent, no pipeline. One command on one laptop.
Notice that every figure on this page carries a green or amber chip. That is not a design flourish — it is the product's third principle, applied to our own marketing. If we would not label a number honestly here, you should not trust us to label one honestly in your reports.
We are asking for two weeks and twenty volunteers.
They run one command. Nothing is installed, nothing is uploaded, nothing changes about how they work. At the end, you get your own itemised bill — your requests, your tools, your sessions, your money — with every number traceable to the file it came from.
Then you decide whether the number is big enough to act on.
Where every credit went, decomposed into the five cost centres, for your own engineers.
Exactly what your installed integrations cost per request, and which have never once been invoked.
Modelled from your own measured shares — clearly labelled as modelled, with the assumptions stated.
You have been paying this bill for months.
This is the first time anyone has offered to read it to you.
FOR ENGINEERS
One Node.js/TypeScript binary (tokenlens) that runs as a CLI, a local dashboard server, a runtime hook target, an MCP server, and a policy compiler — one codebase, one build, one version. Not a SaaS, a proxy, a daemon, or an agent. The core makes zero network calls and zero model calls.
| TOKENLENS_HOME | where to store data (default ~/.tokenlens) |
| TOKENLENS_MONTHLY_ALLOWANCE | override the monthly budget from the command line |
| TOKENLENS_LOG_LEVEL | debug for verbose output |
| Command | What it does | Status |
|---|---|---|
| tokenlens ledger | Credits by day, model, session and cost centre | today |
| tokenlens sessions --top 10 | Where the spend went, ranked | today |
| tokenlens verify <requestId> | Every figure back to a file and byte offset | today |
| tokenlens budget | Burn-down against the plan allowance | today |
| tokenlens dashboard | Local web UI on 127.0.0.1, token-guarded | today |
| tokenlens dashboard --html / --json | Self-contained offline report / raw view model | today |
| tokenlens report --month YYYY-MM | The month as share-safe Markdown | today |
| tokenlens waste | Ranked causes, each with a named fix | today |
| tokenlens waste --explain W1 | The full evidence chain behind one cause | today |
| tokenlens mcp-roi | Per-server invocation ROI | today |
| tokenlens advise | Prescriptive advice from the measured findings | today |
| tokenlens hook install / status / disable | Runtime guards: install, inspect, or stand down one guard | roadmap |
| tokenlens simulate --all | Replay history under a derived policy and price it | roadmap |
| tokenlens policy detect / emit / verify | Detect the channel, emit the settings file, verify it landed | roadmap |
| tokenlens outcomes … | Survival, displacement and causal effect of AI-assisted work | roadmap |
| tokenlens holdout … | Randomised trials with a holdout group | roadmap |
| tokenlens org … | Fleet rollup, drift alerts, aggregate-only bundles | roadmap |
Plus the VS Code extension — a status-bar HUD that shows cost estimates in real time while you use Copilot. Install the .vsix from a release; it is a thin client over the same binary.