AA-Briefcase rates Thinking Machines Lab's Inkling on decks and memos, not trivia. Route models on cost-per-deliverable.

Via Artificial Analysis: How Thinking Machines Lab's Inkling performs on agentic knowledge work
Artificial Analysis scored Thinking Machines Lab's Inkling at 836 Elo on AA-Briefcase - a benchmark built around multi-file knowledge work: spreadsheets, presentations, reports, mock-ups. That is closer to how owner-led firms actually burn tokens than trivia quizzes.
Inkling lands ahead of some peers (e.g. DeepSeek V4 Flash in AA's write-up) but below leading open-weights on the same board. Split scores matter: about 863 Elo on presentation quality versus 764 on analytical quality, with high average output tokens per task (AA cites roughly 52K output tokens per task). Pretty decks can still be thin analysis - and expensive ones.
Thinking Machines Lab's Inkling showing up on AA-Briefcase matters because the benchmark grades finished knowledge work, not parlor tricks. For SMEs drowning in decks and spreadsheets, that is the right scoreboard - as long as you also count dollars per finished pack.
AA-Briefcase combines rubric checks with pairwise grading on analysis and presentation. Models chew through thousands of input files across multi-week-style projects. That is the agentic knowledge-work pattern SMEs are quietly adopting: "make the board pack," "build the model," "draft the memo."
If you pick an open model because it is fashionable, you may buy presentation polish and pay in tokens for mediocre analysis - or the reverse. Cost-per-deliverable beats leaderboard vibes.
Inkling's split - stronger on presentation (~863 Elo) than analytical quality (~764 Elo), with high average output tokens - is the cautionary chart. Pretty slides can hide thin reasoning. High token counts can erase the "open model is cheaper" story if the agent loops forever polishing bullets.
For owner-led firms, the useful habit is to copy AA's grading instinct without copying their entire harness: score finished artifacts on accuracy, usefulness, and cost. A model that wins a public Elo and loses money on your recurring Friday pack is not a win.
AA also surfaces a practical warning about format breadth: models can look multimodal in marketing and still stumble on messy real-world files. That maps to owner-led offices where the critical spreadsheet is a decade of duct tape. Benchmark wins on clean expert projects do not automatically transfer.
Use Inkling's result as a category signal - agentic knowledge work is measurable - then insist on your own cost-quality sheet for the packs you actually ship.
Treat AA-Briefcase as a reminder that knowledge work is multi-file and multi-week, not a chat prompt. Your internal scoreboard should look more like that and less like a trivia leaderboard screenshot in Slack.
Inkling scores an Elo of 836 on AA-Briefcase - higher on presentation (863) than analytical quality (764). - Artificial Analysis
Lead with Model Selection and Continuity Planning when the question is which model for which deliverable. Lead with a Workflow ROI Audit when the question is whether agent loops are saving money or burning margin on the wrong open-weights pick. Both beat collecting Elo screenshots.
We help firms define the jobs, pick the routes, and measure cost per finished pack - then keep the shortlist current when the next Inkling-class model shows up. Selection without measurement is shopping. Measurement without ownership is a spreadsheet nobody opens.
Inkling's AA-Briefcase result is useful because it grades finished knowledge work. Use that habit internally: cost and quality per memo, not vanity benchmarks. Assessment when you want routing rules that stick.
Pick one recurring deliverable. Track quality and token spend for two models over two weeks. Promote the winner for that job only. Repeat. That is model selection without the hype hangover.
This article summarizes publicly reported information and is for general informational purposes only. It does not constitute legal, tax, financial, investment, security, or compliance advice. AgentsROI.ai is not a law firm, accounting firm, or registered investment adviser. Facts, pricing, statistics, and product capabilities cited here reflect the sources listed at the time of writing and may change. Readers should verify current information independently and consult qualified professionals regarding obligations specific to their industry, jurisdiction, and circumstances - including applicable New York State and New York City requirements. AgentsROI.ai may have commercial relationships with vendors mentioned; where material, such relationships are disclosed. Nothing in this article is an endorsement of any specific AI product, model, or provider.