TokenLens on GitHub

RUNS LOCALLY · NO CODE LEAVES THE MACHINE · EVERY FIGURE TRACEABLE

Your Copilot invoice has no line items.
So we wrote them.

GitHub now bills your organisation by the token — and stops working when the budget runs out. It tells you what you spent. It has never told you what you bought.

TokenLens reads the usage records Copilot already writes to your engineers' machines and turns an unexplained total into an itemised bill.

terminal — quick start
$ npm install --global @tokenslens/core
$ tokenlens ledger # credits by day, model, session, cost centre
$ tokenlens dashboard # local UI → http://localhost:7331

1 SEPTEMBER 2026

The included allowance just fell by 37%.

Copilot's launch allowances were promotional. On 1 September they dropped — the same subscription fee, substantially less included capacity. Everything beyond the allowance is billed as overage. When the pooled budget is exhausted, Copilot stops.

There is no cheaper fallback model. There is no degraded mode. It stops.

COPILOT BUSINESS · CREDITS / USER / MONTH

3,000
BEFORE
1,900
AFTER · NOW BILLED AS OVERAGE

COPILOT ENTERPRISE · CREDITS / USER / MONTH

7,000
BEFORE
3,900
AFTER · NOW BILLED AS OVERAGE

1 AI credit = $0.01 · Credits are pooled across the organisation, do not carry over, and are forfeited monthly.

“Can you add validation to the checkout form?”

12 WORDS · 14 TOKENS

A developer asks Copilot a question. Watch what actually gets sent — and billed.

Where the money actually goes.

01 · 49%

Conversation history

Everything already said in this chat, re-sent in full with every new message. The model has no memory — so you buy it again each turn.

02 · 23%

Tool results

What the AI's tools handed back — file contents, command output, search results. Once returned, it rides along in every later step.

03 · 16%

Tool descriptions

Instructions describing every installed tool, sent in full on every request. Billed whether or not the tool is ever used.

04 · 7%

Attached files

Files deliberately attached to the question.

05 · 5%

System instructions

The fixed rules that tell Copilot how to behave.

Conversation history
8,270,722 tokens · 8,242.6 cr
Tool results
3,832,965 tokens · 4,325.0 cr
Tool descriptions
2,601,022 tokens · 1,930.0 cr
Files
System

Rows 3 and 5 are fixed overhead — identical on every request, unrelated to the question asked. Together: 21% of every credit.

THE FINDING NOBODY HAS PRICED

You are paying to describe tools you may never use.

Modern AI assistants can use tools — a database connector, a ticketing integration, a documentation searcher. For the AI to use a tool, that tool must be described to it in full.

Not once. On every single request.

Install six integrations, use two, and you pay for six — on every request, from every developer, for as long as they remain installed.

0
TOKENS SPENT DESCRIBING TOOLS, ON THE MEDIAN REQUEST
0%
OF A TYPICAL PROMPT, BEFORE THE QUESTION IS EVEN READ
0%
THE HIGHEST SHARE OBSERVED ON A SINGLE REAL REQUEST

THE TOLL GATE — DRAG THE SLIDER

INSTALLED TOOLS 6
012
~9,600
TOKENS PER REQUEST
12%
SHARE OF PROMPT
$9.60
COST PER DEVELOPER PER MONTH
$48,000
ACROSS 5,000 SEATS · PER MONTH

This cost scales with what you have installed — not with what you use. It is invisible, it is large, and it is the most trivially reducible cost in the entire bill.

You pay for the same words about ten times.

The model has no memory between steps. When the AI reads a file, then runs a command, then decides what to do next, everything it has gathered so far is re-sent at every step so it can “remember”.

Across real sessions, every token of genuinely new content was billed roughly ten times over.

0×
RE-TRANSMISSION MULTIPLIER · MEASURED ACROSS 790 REQUESTS

Some of this is unavoidable — it is how the technology works. A meaningful share is not, and that is what we go after.

The most expensive habit in your engineering org is leaving a chat window open.

Because the whole conversation is re-sent every turn, a chat gets more expensive with every message. By the thirty-third message, each new message costs 61% more than it would in a fresh chat.

The fix is “start a new chat when you start a new task.”

It costs nothing. It requires no tooling. Almost nobody does it — because nobody can see the meter running.

COST PER MESSAGE · MESSAGE NUMBER

1 33
31.2
A FRESH CHAT
56.9
THE 33rd MESSAGE · +61%

This is what an absent cost signal looks like. Nobody is being careless — the information simply does not exist at the moment the decision is made.

THE ENCOURAGING PART

You do not have to change how everyone works.

0%
OF CHAT SESSIONS
0%
OF ALL MONEY SPENT

Spend is extremely concentrated. The most expensive single session cost $32.70; the typical session cost $3.61.

This means the intervention is surgical, not cultural. Fix the worst tenth and you capture most of the value — without a training programme, a policy memo, or a single all-hands.

57 REAL SESSIONS · SIZED BY SPEND · SORTED DESCENDING

top session 8,644.6 credits

⎴

THE LARGEST NINE · 50% OF ALL SPEND

Why has nobody shown you this before?

There is a whole industry of AI monitoring tools. None of them can see Copilot — for a structural reason, not a lack of effort.

Those tools work by sitting in the network path: you route your AI traffic through them and they inspect it as it passes. That works when your own application calls an AI provider.

Copilot does not work that way. The traffic goes from the editor to GitHub over a channel you do not control and cannot intercept. There is no point in the middle to stand.

The only place this information exists in readable form is a file on the developer's own machine — in an undocumented format that changes with every VS Code release. Reading it reliably is genuinely difficult.

That difficulty is exactly why the category is empty. It is also our moat.

HOW MONITORING TOOLS WORK

YOUR APP ──▶ MONITORING TOOL ──▶ AI PROVIDER

A place to stand.

HOW COPILOT WORKS

EDITOR ━━━━━━▶ GITHUB
✕ NOWHERE TO STAND
◑ THE RECORD IS ALREADY HERE ──▶ TOKENLENS

The local journal on the developer's own machine.

We did not find a better place to stand. We found that the receipt was already on the floor.

THE PRODUCT

An itemised bill for AI-assisted engineering.

TokenLens reads the usage records already on your engineers' machines, shows you exactly what every credit bought, names what was wasted and why, and hands your IT team a configuration change that removes a large part of it — without asking a single developer to work differently.

01

Measure

Read the records Copilot already writes. Produce an exact credit ledger. Every figure traceable to the byte it came from.

Available today
02

Attribute

Assign wasted spend to fourteen named, individually detectable causes. Waste stops being a total and becomes a list of specific, fixable things.

Available today
03

Simulate

Replay real recorded sessions under a proposed change and report what it would have saved — before anything is deployed.

On the roadmap
04

Enforce

Emit the exact configuration your IT team pushes through the device-management tooling you already own. No developer action.

On the roadmap
05

Prove

Run it as a randomised trial against a holdout group on your own fleet. Report the actual saving, including where we were wrong.

On the roadmap

Steps 01–02 are working software running on real data today. Steps 03–05 are specified and sequenced — see the roadmap below.

What adopting this actually involves.

There is no server to run, no proxy to configure, no agent to install, and no data pipeline to build.

1

One command, one laptop.

An engineer runs a single command on their own machine. It reads the local Copilot records and prints the ledger. Nothing is uploaded. Nothing is installed permanently. This takes minutes, not a project.

$ tokenlens ledger
# credits by day, model, session and cost centre
2

Look at the report.

A local dashboard opens in the browser — running on that machine, reachable only from that machine. It shows the itemised bill: where the credits went, which sessions, which tools, which behaviours.

$ tokenlens dashboard
# local web UI on 127.0.0.1, token-guarded → http://localhost:7331
3

Widen to a sample.

Repeat across a representative group of engineers — twenty is plenty. Only aggregate counts are combined. Source code never moves.

$ tokenlens dashboard --html team-report.html
# one self-contained file — email it to finance, opens offline
4

Deploy the configuration.

TokenLens produces the exact settings file. Your IT team pushes it through the same tooling used for every other setting. Developers notice nothing except that things stop running out.

$ tokenlens policy emit --out ./out
# .reg / .mobileconfig / settings.json + rollbacks

Nothing leaves the machine

Only aggregate counts and one-way hashes. Offline by default.

Read-only, always

Files are opened for reading. Nothing is ever written, moved, or altered.

No infrastructure

No proxy, no server, no agent, no pipeline.

Every number traceable

One command prints the file and byte position behind any figure.

$ tokenlens verify <requestId>

RUNNING TODAY, ON REAL DATA

This is not a mock-up.

Everything below is the tool's real dashboard — reconstructed from its actual source (packages/core/src/dashboard) — showing the live figures it reports on the developer machine it was built on: 713 requests across 57 chat sessions over 61 active days. Every figure carries its own receipt.

Time-series shapes are illustrative; every figure is real. Switch views with the sidebar — exactly as in the running tool.

3

Overview

Scope: All projects · 61 active days · UTC days
4
Credits
66,008.2~24% measured
713 requests over 61 active day(s)
Per active day
1,082.1
61 active days measured
Month to date
—
allowance: 3,900 credits/mo
Models in use
11
Top: claude-sonnet-4-6
Measured share
24%
the rest is a rate-card estimate
Flagged days
—
median absolute deviation, not σ
5

Daily spend

Cumulative credits across the 61 active days. Hover for the exact figure on any day.
66,008.2
6

Where the prompt goes

The five-way split VS Code records on every request.
16,076 credits
Cost centreTokensShare
■ Conversation history8,270,72249% measured
■ Tool results3,832,96523% measured
■ Tool descriptions2,601,02216% measured
■ Attached files1,231,5137% measured
■ System instructions789,4425% measured
7

Model mix over time

Stacked credits per day, biggest model at the bottom. The chart that answers “is the premium share growing”.
Illustrative shape — the live tool stacks your real models. 60 requests on the dearest model cost 17,872 credits; 212 on the cheapest cost 2,782. measured
8

When the spend happens

Click a day to open its detail. Only measured days are drawn — a padded-out calendar would render “before you started” identically to “spent nothing”, which are not the same claim.
Illustrative pattern shown — your real days render here.
1
Sidebar navigation

Ten views — Overview, Budget, Sessions, Models, Tools, Waste, Day, Projects, Data, Settings. Collapsible with Ctrl+B.

2
Theme & density

Three themes (system / light / dark — system follows the OS) and a compact density toggle. The tool respects the machine's existing choice.

3
Scope & range filters

Every view is driven by two filters: which workspaces count (scope) and which dates count (range). An empty scope says so loudly — a blank dashboard must never read as “you spent nothing”.

4
KPI row

Six headline figures: credits, per-active-day, month-to-date, models in use, measured share, flagged days. Each carries its provenance chip.

5
Daily spend

Line chart with anomaly flags — days far enough from the median to be worth a look, detected by median absolute deviation, not σ.

6
Where the prompt goes

The five-way cost-centre donut with the token table beneath. The same five colours, everywhere, forever.

7
Model mix over time

Stacked credits per day. Answers “is the premium share growing” — the question a daily total cannot.

8
Calendar heatmap

Only measured days are drawn. Click a day to open its detail view.

9
Footer & command palette

The provenance legend lives in the footer of every screen; Ctrl+K jumps to any view. The chips are the product's honesty principle, rendered as UI.

10
Token-guarded, local only

tokenlens dashboard serves on 127.0.0.1 behind a token — reachable only from that machine. There is no cloud half of this product.

Only 24% of these requests had a credit figure published by GitHub. The rest are estimated from rates measured on the ones that did — and every one of those is labelled as an estimate, never as a fact.

TAKE IT WITH YOU

The report, as it leaves the machine.

Three exporters, one rule: an exported file is the artefact most likely to be forwarded to someone other than its subject, so it defaults to the share-safe scope. Per-session detail is withheld unless you explicitly ask for it.

terminal — exports
$ tokenlens dashboard --html copilot-spend.html
✓ wrote copilot-spend.html — one self-contained file: inline CSS, no external assets, no fetch calls. Email it to finance; it opens fully offline.
$ tokenlens dashboard --json ledger.json
✓ wrote ledger.json — the raw provenance-annotated view model, for scripting and archival.
$ tokenlens report --month 2026-07 --out july.md
✓ wrote july.md — the month as Markdown, ready to hand to an AI assistant. Per-session detail withheld as share-safe.
$ tokenlens report --month 2026-07 --self --out july-full.md
! kept per-session detail — for your eyes only.

What the HTML export looks like

The file the finance team receives. Same figures, same chips, zero dependencies.

◑
GitHub Copilot spend — 2026-07
Generated 2026-07-30 · scope: shared — per-session detail withheld · UTC days

The month in figures

66,008.2
TOTAL CREDITS · = $660.08
713
REQUESTS · 57 SESSIONS · 61 ACTIVE DAYS

Allowance and forecast

Plan enterprise · monthly allowance 3,900 credits (plan default) · projected month-end 9,922.0 credits — projected to exceed the allowance by 6,022.0 credits.

Where the tokens went

Cost centreTokensShareCredits
Conversation history8,270,72249%8,242.6
Tool results3,832,96523%4,325.0
Tool descriptions2,601,02216%1,930.0
Attached files1,231,5137%1,021.0
System instructions789,4425%577.6

Model mix

11 models in use. 60 requests on the most expensive model cost 17,872 credits; 212 requests on the cheapest cost 2,782 — three and a half times fewer requests, six and a half times more money.

Conversation economics

Re-transmission multiplier 9.8× · 33rd-message penalty +61% · duplicate file-read rate 42.0%.

What the detectors found

Fourteen named causes, each with its evidence and remedy — see the Waste board. Findings may overlap; they do not sum to the ledger.

Classes that could not be assessed

Listed separately, with what data was missing and what would unblock each one. “We found no waste” and “we could not look” are opposite claims.

Tool surface

Installed integrations and their per-request description cost — the Tool-Definition Tax, itemised per tool.

And the Markdown, for the AI-assisted workflow

tokenlens report writes the month as Markdown — structured for handing to an AI assistant or pasting into a doc. Same section spine: how to read this, the month in figures, allowance and forecast, where the tokens went, model mix, conversation economics, daily pattern, what the detectors found, priced model substitution, tool surface.

july.md — sample
# GitHub Copilot spend — 2026-07

## How to read this
Figures tagged [measured] were read from your local journals. Figures tagged [modelled] are rate-card estimates — the chip on each says which.

## The month in figures
- Total: 66,008.2 credits ($660.08) [~24% measured]
- Requests: 713 · Sessions: 57 · Active days: 61

## Where the tokens went
- Conversation history — 49% (8,270,722 tokens) [measured]
- Tool results — 23% (3,832,965 tokens) [measured]
- Tool descriptions — 16% (2,601,022 tokens) [measured]

## Priced model substitution
What moving premium-model work to the cheapest sufficient model would have cost…

A report exists to be sent somewhere, so the safe scope is the default and the permissive one has to be asked for. --self keeps per-session detail — and the report tells you, in its header, exactly what was withheld and why.

Waste is not a number. It is a list.

“You are wasting money” is not actionable. TokenLens attributes wasted spend to fourteen specific, individually detectable causes — each with a way to detect it and a named remedy.

W1 · DETECTED TODAY OR NEXT PHASE

The tool-definition toll

Paying to describe installed tools on every request, including tools never invoked.

Measured: 17,016 tokens per request · 21% of the median prompt · 9.1 credits per request.
W2 · DETECTED TODAY OR NEXT PHASE

Reading the same file twice

The AI re-reads a file it already has in front of it.

Measured: 42% of file reads were repeats; one file was read 72 times in a single session.
W3 · DETECTED TODAY OR NEXT PHASE

Oversized tool results

A single command returns an enormous result that then rides along in every later step.

One observed result cost $0.90 on its own — then kept being re-sent.
W4 · DETECTED TODAY OR NEXT PHASE

Stale conversations

A chat left open across unrelated tasks, re-sending irrelevant history forever.

By the 33rd message each new message costs 61% more than in a fresh chat.
W5 · DETECTED TODAY OR NEXT PHASE

Over-powered model choice

The most expensive model used for work a cheaper one handles identically.

Model price spread, most vs least expensive: 15.1×.
W6 · DETECTED TODAY OR NEXT PHASE

Runaway loops

Very long agent loops that consume heavily and produce no surviving change.

Detected from loop length vs. surviving-change analysis in the outcomes module.
W7 · SPECIFIED

Abandoned work

Credits spent on work that never reached the codebase.

Detector specified for the waste-attribution phase.
W8 · SPECIFIED

The same question, fifty times

Fifty engineers independently asking the same thing about the same internal system.

Detector specified for the waste-attribution phase.
W9 · SPECIFIED

Context compaction

Automatic summarisation triggered by unmanaged context growth.

In 115 days it ran 84 times at a median of 92 seconds each — about two hours of engineers watching a progress indicator.
W10 · SPECIFIED

Instruction bloat

Standing instruction files that grew past their usefulness.

Detector specified for the waste-attribution phase.
W11 · SPECIFIED

Search-snippet leakage

Paying for previews of files that were never opened.

Detector specified for the waste-attribution phase.
W12 · SPECIFIED

Premium models on trivial work

Titles, summaries and commit messages generated on the most expensive model available.

Detector specified for the waste-attribution phase.
W13 · SPECIFIED

Over-provisioned reasoning

Deep-reasoning settings applied to work that does not need them.

Detector specified for the waste-attribution phase.
W14 · SPECIFIED

Missing context isolation

Expensive exploration done in the main conversation instead of an isolated cheaper one.

Detector specified for the waste-attribution phase.

W9 is worth pausing on.

In 115 days, automatic summarisation ran 84 times and took a median of 92 seconds each — about two hours of engineers watching a progress indicator. That is not a billing problem. That is a productivity problem you were also not being shown.

THE PART THAT SURPRISES PEOPLE

Most of the saving needs nobody to do anything differently.

The instinctive assumption is that reducing AI spend means asking engineers to use it less. That would be slow, unpopular, and would cost you more in lost productivity than it saved.

That is not the intervention.

The largest savings come from settings — which model handles which kind of work, how large a single tool result may be, which integrations stay installed. These are configuration values. They are deployed the same way every other setting in your organisation is deployed, through tooling you already own.

No training. No policy memo. No behaviour change. Your engineers keep working exactly as they do now.

TIER A — CONFIGURATION ONLY

Settings deployed through your existing device management.

Developer effort: None. They will not notice.

~35%
MODELLED REDUCTION IN SPEND

Availability: roadmap phase D5

TIER B — LIVE GUARDRAILS

Automatic prevention of specific wasteful actions as they happen.

Developer effort: None. Prevention is silent.

up to ~53%
COMBINED MODELLED REDUCTION

Availability: roadmap phase D6

Both figures are modelled — projected from measured cost shares, not yet observed on a deployed policy. That is precisely why the product's final phase is a randomised trial on your own fleet, with a holdout group, so the real figure replaces the projected one. We will report the gap between them, including if it is unflattering.

Built in order, with the measurement first.

This product was built in a deliberate sequence: measure before attributing, attribute before recommending, recommend before enforcing, and prove before claiming. Each phase is only useful because the one before it is trustworthy.

PHASE 1 · COMPLETE

Foundation

The honest-measurement guarantees, enforced in code.

PHASE 2 · COMPLETE

Credit ledger

Exact, traceable accounting of every credit.

PHASE 3 · COMPLETE

Dashboard & reports

The itemised bill, on screen and exportable.

PHASE 4 · NEXT

Waste attribution

The fourteen named causes, detected and ranked.

PHASE 5 · SPECIFIED

Simulation

“This change would have saved N last quarter.”

PHASE 6 · SPECIFIED

Policy compiler

The exact settings file your IT team deploys.

PHASE 7 · SPECIFIED

Runtime guardrails

Prevention at the moment of spend.

PHASE 8 · SPECIFIED

Editor integration

Live cost visibility for the developer.

PHASE 9 · SPECIFIED

Randomised proof

Measured savings with confidence intervals.

PHASE 10 · SPECIFIED

Organisation rollup

Fleet-wide view and self-tuning policy.

Phases 1–3 are working software running on real data today. Everything shown in this presentation as measured came out of them. Everything shown as modelled is waiting on phase 9 to become measured.

WHAT THIS LOOKS LIKE FROM THE INSIDE

An ordinary Tuesday.

Priya is a senior engineer. She is good at her job. On this particular Tuesday she does nothing wrong, nothing unusual, and nothing anybody would flag in a review.

Watch the meter.

CREDITS TODAY
0
$0.00
09:12
Opens yesterday's chat window to pick up where she left off. Asks a small question.
Yesterday's entire conversation is re-sent with it.
09:40
The assistant reads the same configuration file for the fourth time this session.
It already had it.
10:15
Asks a one-line question: “what does this function do?”
Costs almost the same as a full refactor. The floor, not the work.
11:30
A command returns a very large output. It stays in context for the rest of the day.
Every later step re-sends it.
12:05
Still the same chat window. Now on a completely different task.
The morning's unrelated discussion is still being paid for, every turn.
14:20
The assistant pauses for a minute and a half to summarise the conversation so far.
It ran out of room. She waits.
15:45
A complex refactor. Genuinely hard work, genuinely well done.
This one was worth it.
16:30
Same chat window. Fourth unrelated task of the day.
17:50
Closes the laptop.
She has no idea. There was never a number.

This day is constructed from measured medians — a typical request, a typical session penalty, a typical compaction. It is an illustration of measured behaviour, not a recording of one real Tuesday.

PRIYA'S TUESDAY, DECOMPOSED

Of Priya's Tuesday, roughly a third was avoidable — re-reads, a stale window, an oversized result, and a model heavier than the task needed.

Priya did nothing wrong. She was never shown a number. Every decision that created this was made in the absence of information that already existed on her own laptop.

Now multiply Priya by five thousand.

ILLUSTRATIVE EXAMPLE · MODELLED FROM MEASURED RATES

MERIDIAN FINANCIAL GROUP

5,000 engineers · GitHub Copilot Enterprise · 3,900 credits included per seat per month

ACT I · SEPTEMBER — THE FIRST BILL AFTER THE ALLOWANCE DROPPED

0
ANNUAL OVERAGE · BEYOND THE INCLUDED ALLOWANCE
Annual consumption at measured rate$5,953,200
Included in subscription−$2,340,000
Annual overage$3,613,314

Nobody at Meridian can explain this number. It is not that they lack the will — the information does not exist. There is a total, a date, and an amount due.

Somewhere inside it is a large amount of money buying nothing at all. Nobody can point at it.

Real fleets have a long tail of lighter users, so a real total would likely be lower — but spend concentrates in heavy users, and this profile is a heavy user.

YOUR NUMBER, NOT OURS

What would this be worth to you?

5,000
5020,000

What this calculator assumes

  • Your engineers' usage resembles the measured profile of a heavy Copilot user.
  • Model prices and allowances stay as they are today.
  • The reduction percentages are modelled from measured cost shares — they have not yet been observed on a deployed policy.
  • Real organisations have a long tail of lighter users, so a real total is likely lower than the heavy-profile projection.
  • Light / Moderate tiers scale the measured heavy run-rate by 0.25 / 0.55 — illustrative, not measured.

This is a projection, not a quote.

$3,613,320
ANNUAL OVERAGE TODAY
$1,529,658
ANNUAL OVERAGE AFTER
BEFORE
AFTER
$2,083,662
ANNUAL SAVING
58%
OVERAGE BILL REDUCED BY

What we will not do.

No source code ever leaves the machine.

Only aggregate counts and one-way hashes. Network access is denied by default.

Every number is traceable.

One command prints the file and byte position behind any figure on any report.

We never present an estimate as a measurement.

Enforced in the type system — the code physically cannot print an unlabelled estimate. It is why every number on this page carries a chip.

Read-only on your data.

Files opened for reading only. Nothing written, moved, or altered.

We fail loudly.

If something is not understood, it is reported and counted — never silently reported as zero.

Nothing to adopt.

No proxy, no server, no agent, no pipeline. One command on one laptop.

Notice that every figure on this page carries a green or amber chip. That is not a design flourish — it is the product's third principle, applied to our own marketing. If we would not label a number honestly here, you should not trust us to label one honestly in your reports.

We are not asking you to buy anything.

We are asking for two weeks and twenty volunteers.

They run one command. Nothing is installed, nothing is uploaded, nothing changes about how they work. At the end, you get your own itemised bill — your requests, your tools, your sessions, your money — with every number traceable to the file it came from.

Then you decide whether the number is big enough to act on.

Your itemised bill

Where every credit went, decomposed into the five cost centres, for your own engineers.

Your tool-definition tax

Exactly what your installed integrations cost per request, and which have never once been invoked.

Your projected saving

Modelled from your own measured shares — clearly labelled as modelled, with the assumptions stated.

You have been paying this bill for months.
This is the first time anyone has offered to read it to you.

FOR ENGINEERS

One binary. Every surface.

One Node.js/TypeScript binary (tokenlens) that runs as a CLI, a local dashboard server, a runtime hook target, an MCP server, and a policy compiler — one codebase, one build, one version. Not a SaaS, a proxy, a daemon, or an agent. The core makes zero network calls and zero model calls.

Install

terminal
# prerequisites: Node.js ≥ 20, npm ≥ 10
$ npm install --global @tokenslens/core
$ tokenlens --version && tokenlens --help

# or build from source
$ git clone https://github.com/Ashutosh-Panda2004/TokensLens.git && cd TokensLens
$ npm ci && npm run build && npm run cli:link

CONFIGURATION · .tokenlens/config.json

{
  "plan": "pro",
  "monthlyAllowance": 300
}

ENVIRONMENT

TOKENLENS_HOMEwhere to store data (default ~/.tokenlens)
TOKENLENS_MONTHLY_ALLOWANCEoverride the monthly budget from the command line
TOKENLENS_LOG_LEVELdebug for verbose output

Command reference

CommandWhat it doesStatus
tokenlens ledgerCredits by day, model, session and cost centretoday
tokenlens sessions --top 10Where the spend went, rankedtoday
tokenlens verify <requestId>Every figure back to a file and byte offsettoday
tokenlens budgetBurn-down against the plan allowancetoday
tokenlens dashboardLocal web UI on 127.0.0.1, token-guardedtoday
tokenlens dashboard --html / --jsonSelf-contained offline report / raw view modeltoday
tokenlens report --month YYYY-MMThe month as share-safe Markdowntoday
tokenlens wasteRanked causes, each with a named fixtoday
tokenlens waste --explain W1The full evidence chain behind one causetoday
tokenlens mcp-roiPer-server invocation ROItoday
tokenlens advisePrescriptive advice from the measured findingstoday
tokenlens hook install / status / disableRuntime guards: install, inspect, or stand down one guardroadmap
tokenlens simulate --allReplay history under a derived policy and price itroadmap
tokenlens policy detect / emit / verifyDetect the channel, emit the settings file, verify it landedroadmap
tokenlens outcomes …Survival, displacement and causal effect of AI-assisted workroadmap
tokenlens holdout …Randomised trials with a holdout grouproadmap
tokenlens org …Fleet rollup, drift alerts, aggregate-only bundlesroadmap

Plus the VS Code extension — a status-bar HUD that shows cost estimates in real time while you use Copilot. Install the .vsix from a release; it is a thin client over the same binary.