Every tool onthe wall.
A toolroom is the room behind the workshop where each tool hangs in its own painted silhouette — so anyone can see at a glance what is there, what is in use, and what is missing. This is ours, and we leave the door open.
PRICES READ AT THE PROVIDERS’ OWN PAGES 27 SEP 2026 · REGISTER: LAST FULL REVIEW 11 SEP 2026 · NEXT REVIEW 11 OCT 2026
Most firms treat their tools as the secret. We think the tools were never the secret — they are on sale to everyone, for twenty dollars a month.
What is scarce is knowing which instrument answers which question, what its answer is actually worth, and where it must not be trusted at all. A brilliant model pointed at the wrong problem produces confident nonsense faster than any human could. Half this page is about what each tool is for. The other half is about what it is not for — which is the more valuable half, and the one nobody publishes.
Machines draft. Code computes. A person decides, corrects, and signs — and that signature is not a formality at the end. It is the work.
The letter from the toolroom. One worked governance workflow, named tools included, every fortnight.
OPENING SHORTLY · DOUBLE OPT-IN · NO TRACKING PIXELS · WRITE TO THE DESK AND WE WILL ADD YOU BY HAND
Eight jobs. Who actually does what.
This is the part most people mean when they ask “do you use AI?” — so here it is, one job at a time. The machine never holds a number and never holds the pen at the end. It does the reading, the drafting and the arguing, which is real work and a great deal of it.
Eight jobs, and for each one: what the machine does, what deterministic code does, what the person does — and which instruments on the wall below are reached for first. Choose a job and the bench answers.
The rack
Every tool hangs in its own painted silhouette, so anyone can see at a glance what is on the wall. Choose one to see what it is for, and what it is never given. We are these platforms' user, not their partner — named with credit, nothing implied.
Anyone can rent the same wall for the price of a working lunch. That has never been the scarce part.
Three things the leaderboards do not tell you
The window is not the window
Million-token windows are now common, not universal: the current Grok publishes 500,000, one Mistral model 256,000. And a published limit is not a usable one. Independent long-context testing in 2026 found retrieval and multi-step reasoning falling away well before the limit, at a point that differs by model and by task. If you hand a model a data room and assume it read all of it, you have made an assumption nobody sold you.
A model can be switched off
On 12 June 2026 Anthropic suspended two of its models under a U.S. export-control direction; the controls were lifted on 30 June and access returned on 1 July. Whatever you build on, build it so that a supplier disappearing for a fortnight is an inconvenience rather than an outage.
The frontier is crowded and close
Three laboratories shipped their current top model inside three weeks in September: OpenAI on the 3rd at $10 and $50 per million tokens, xAI on the 21st, Anthropic on the 22nd at $4 and $20. When the leaders are separated by a point or two, the right question stops being which is smartest and becomes which fits the job, the budget and the rules you have written.
The research desk
A bench like this only stays honest if somebody keeps testing it. That work is continuous here, it is written down, and where it is worth reading by anyone else it is published rather than kept.
The monthly re-read
Every model, price and claim on this page is opened at its provider’s own page on a fixed date, not when we remember. Three research engines returning the same answer is correlation, not verification.
The correction log
What we got wrong and fixed is published, dated, on the story page. A real product called fabricated, a superseded model called current, a stale price — all three were ours, and all three are on the record.
Context that isn’t there
Measuring how much of an advertised context window a model reliably uses on multi-document finance work. Independent testing says the usable window ends before the published one; where it ends is the question. If that holds on a data room, a great deal of diligence practice is built on sand.
Where the boundary should sit
Which finance tasks may approach a model at all, tested against the containment register rather than against capability. The interesting answers are the ones where the model could do it and still should not.
Research desk №001
The counsellor’s game — what a twelfth-century chess set says about governance. Read it →
Academic and practitioner writing
Where a finding is genuinely new, it belongs in a journal or a professional body’s hands rather than a marketing page. Nothing is claimed here that has not been submitted; this line exists so the intent is on the record and can be held against us.
Five ways good tools get used badly
Not a criticism of anyone — these are the habits we had to unlearn ourselves.
| Model | Input | Output | Read |
|---|---|---|---|
| Gemini 3.8 FlashGoogleResearch run of 25 Sep; introductory rate to 31 Dec | $0.75 | $3.75 | Research run of 25 Sep; introductory rate to 31 Dec |
| Haiku 4.5AnthropicAnthropic's page, 27 Sep | $1 | $5 | Anthropic's page, 27 Sep |
| Sonnet 5AnthropicAnthropic's page, 27 Sep | $2 | $10 | Anthropic's page, 27 Sep |
| GPT-6 SolOpenAIOpenAI's page, 27 Sep | $2 | $10 | OpenAI's page, 27 Sep |
| Grok 4.7xAINot yet read at the provider | — | — | Not yet read at the provider |
| Opus 5.5AnthropicAnthropic's page, 27 Sep | $4 | $20 | Anthropic's page, 27 Sep |
| GPT-6 AstraOpenAIOpenAI's page, 27 Sep | $10 | $50 | OpenAI's page, 27 Sep |
| Fable 5.1AnthropicAnthropic's page, 27 Sep | $10 | $50 | Anthropic's page, 27 Sep |
What we will not do with any of it
No client data enters a general-purpose tool unredacted, and lightly redacted is not redacted. No model owns a number we sign. No tool is used on client work until its terms have been read at source. Credentials are never prompt material — not in a consumer tool, not in an enterprise one, not ever. The full register names every tool, what it may see, what it is never given, and the date we last checked.
This page is dated on purpose. The field moves monthly; a toolroom that is never re-hung is a museum. Anthropic's and OpenAI's rates were read at their own pages on 27 September 2026; the next review is 11 October 2026. Platform names are used with credit, as their user — not their partner. No affiliation is implied.