THE TOOLROOM

Every tool onthe wall.

A toolroom is the room behind the workshop where each tool hangs in its own painted silhouette — so anyone can see at a glance what is there, what is in use, and what is missing. This is ours, and we leave the door open.

بجانب كل أداة.
bi-jānib kull adāhBESIDE EVERY TOOL
THE WORKSHOP · WHERE THE INSTRUMENTS ARE KEPT

PRICES READ AT THE PROVIDERS’ OWN PAGES 27 SEP 2026 · REGISTER: LAST FULL REVIEW 11 SEP 2026 · NEXT REVIEW 11 OCT 2026

Most firms treat their tools as the secret. We think the tools were never the secret — they are on sale to everyone, for twenty dollars a month.

What is scarce is knowing which instrument answers which question, what its answer is actually worth, and where it must not be trusted at all. A brilliant model pointed at the wrong problem produces confident nonsense faster than any human could. Half this page is about what each tool is for. The other half is about what it is not for — which is the more valuable half, and the one nobody publishes.

Machines draft. Code computes. A person decides, corrects, and signs — and that signature is not a formality at the end. It is the work.

EVERY CONCLUSION ON THIS SITE IS A HUMAN'S
HOW A DELIVERABLE IS ACTUALLY MADE
Four stages. The machine sits in the middle of the process — never at either end.
SOURCEdocuments, ledgers, published rulesMACHINEdrafts, extracts, checks, codeCODEarithmetic, reconciliation, datesHUMANjudgement, correction, signatureEVERY DELIVERABLE, EVERY TIMEOnly the last box can sign.
The middle two boxes are fast and cheap. The last one is neither, and it is the one you are paying for.

The letter from the toolroom. One worked governance workflow, named tools included, every fortnight.

OPENING SHORTLY · DOUBLE OPT-IN · NO TRACKING PIXELS · WRITE TO THE DESK AND WE WILL ADD YOU BY HAND

Eight jobs. Who actually does what.

This is the part most people mean when they ask “do you use AI?” — so here it is, one job at a time. The machine never holds a number and never holds the pen at the end. It does the reading, the drafting and the arguing, which is real work and a great deal of it.

Eight jobs, and for each one: what the machine does, what deterministic code does, what the person does — and which instruments on the wall below are reached for first. Choose a job and the bench answers.

The rack

Every tool hangs in its own painted silhouette, so anyone can see at a glance what is on the wall. Choose one to see what it is for, and what it is never given. We are these platforms' user, not their partner — named with credit, nothing implied.

Anyone can rent the same wall for the price of a working lunch. That has never been the scarce part.

Three things the leaderboards do not tell you

The window is not the window

Million-token windows are now common, not universal: the current Grok publishes 500,000, one Mistral model 256,000. And a published limit is not a usable one. Independent long-context testing in 2026 found retrieval and multi-step reasoning falling away well before the limit, at a point that differs by model and by task. If you hand a model a data room and assume it read all of it, you have made an assumption nobody sold you.

A model can be switched off

On 12 June 2026 Anthropic suspended two of its models under a U.S. export-control direction; the controls were lifted on 30 June and access returned on 1 July. Whatever you build on, build it so that a supplier disappearing for a fortnight is an inconvenience rather than an outage.

The frontier is crowded and close

Three laboratories shipped their current top model inside three weeks in September: OpenAI on the 3rd at $10 and $50 per million tokens, xAI on the 21st, Anthropic on the 22nd at $4 and $20. When the leaders are separated by a point or two, the right question stops being which is smartest and becomes which fits the job, the budget and the rules you have written.

The research desk

A bench like this only stays honest if somebody keeps testing it. That work is continuous here, it is written down, and where it is worth reading by anyone else it is published rather than kept.

STANDING

The monthly re-read

Every model, price and claim on this page is opened at its provider’s own page on a fixed date, not when we remember. Three research engines returning the same answer is correlation, not verification.

STANDING

The correction log

What we got wrong and fixed is published, dated, on the story page. A real product called fabricated, a superseded model called current, a stale price — all three were ours, and all three are on the record.

OPEN

Context that isn’t there

Measuring how much of an advertised context window a model reliably uses on multi-document finance work. Independent testing says the usable window ends before the published one; where it ends is the question. If that holds on a data room, a great deal of diligence practice is built on sand.

OPEN

Where the boundary should sit

Which finance tasks may approach a model at all, tested against the containment register rather than against capability. The interesting answers are the ones where the model could do it and still should not.

PUBLISHED

Research desk №001

The counsellor’s game — what a twelfth-century chess set says about governance. Read it →

INTENDED

Academic and practitioner writing

Where a finding is genuinely new, it belongs in a journal or a professional body’s hands rather than a marketing page. Nothing is claimed here that has not been submitted; this line exists so the intent is on the record and can be held against us.

This is teachable, and we teach it. The register, the map and the rules above are the same instruments we build for finance teams that want their own. Learning & development →
WHERE A FRONTIER CHAT MODEL IS STRONG — AND WHERE IT ISN'T
The shape most people get wrong. Indicative, from our own use.
REASONINGLONG DOCUMENTSCODEWEB RESEARCHARITHMETIC
The amber bar is the one that matters: a language model is a poor calculator, however fluent it sounds. That is why our numbers live in code.

Five ways good tools get used badly

Not a criticism of anyone — these are the habits we had to unlearn ourselves.

COMMON HABITAsking a chat model to add up a column, reconcile a ledger, or work out a filing deadline — because it answers instantly and sounds certain.
WHAT WE DOArithmetic never touches a model. It runs in code, where the same input gives the same answer today and in a year, and every step can be shown.
COMMON HABITUsing a research tool for analysis — taking the summary it produces as the finding, because it arrived with citations attached.
WHAT WE DOResearch tools fetch; they do not conclude. Every citation gets opened at source. Twice in September 2026 a cited claim did not survive that test.
COMMON HABITAsking a reasoning model about this morning's news or this year's regulation, when its knowledge stops months earlier.
WHAT WE DOAnything time-sensitive is verified against the primary source — the regulator's page, the provider's own documentation — before it enters a deliverable.
COMMON HABITAsking the same model the same question twice and treating the agreement as confirmation.
WHAT WE DOA second family, not a second attempt. Models are more likely to repeat their own reasoning than to catch it.
COMMON HABITChoosing the cheapest tier available — without reading what the provider does with what you send it.
WHAT WE DOWe read the terms at source before a tool is used on client-adjacent work. One major provider currently offers a tier roughly twelve times cheaper that trains on your inputs. That is not a discount; it is a price paid in confidentiality.
PUBLISHED API PRICE · PER MILLION TOKENS, STANDARD RATE
Each rate says where it was read: at the provider's own page, through the research run, or not yet. A million tokens is roughly 750,000 words. The next review is 11 October 2026.
ModelInputOutputRead
Gemini 3.8 FlashGoogleResearch run of 25 Sep; introductory rate to 31 Dec$0.75$3.75Research run of 25 Sep; introductory rate to 31 Dec
Haiku 4.5AnthropicAnthropic's page, 27 Sep$1$5Anthropic's page, 27 Sep
Sonnet 5AnthropicAnthropic's page, 27 Sep$2$10Anthropic's page, 27 Sep
GPT-6 SolOpenAIOpenAI's page, 27 Sep$2$10OpenAI's page, 27 Sep
Grok 4.7xAINot yet read at the provider——Not yet read at the provider
Opus 5.5AnthropicAnthropic's page, 27 Sep$4$20Anthropic's page, 27 Sep
GPT-6 AstraOpenAIOpenAI's page, 27 Sep$10$50OpenAI's page, 27 Sep
Fable 5.1AnthropicAnthropic's page, 27 Sep$10$50Anthropic's page, 27 Sep
Output costs five times input on every rate read here. That ratio, not the headline rate, decides whether a tool is worth pointing at a job. A frontier model earns its price when it removes an hour of rework, and only then.

What we will not do with any of it

No client data enters a general-purpose tool unredacted, and lightly redacted is not redacted. No model owns a number we sign. No tool is used on client work until its terms have been read at source. Credentials are never prompt material — not in a consumer tool, not in an enterprise one, not ever. The full register names every tool, what it may see, what it is never given, and the date we last checked.

Workshops for client teams run on this bench — real workflows, named tools, the limits included.The containment register →

This page is dated on purpose. The field moves monthly; a toolroom that is never re-hung is a museum. Anthropic's and OpenAI's rates were read at their own pages on 27 September 2026; the next review is 11 October 2026. Platform names are used with credit, as their user — not their partner. No affiliation is implied.