Kimi K3 sits near the top of AA-Briefcase, then bills like a premium hour. Score cost and latency before you crown a default model.

Via Artificial Analysis: Kimi K3: second only to Fable 5 on AA-Briefcase
Artificial Analysis reports that Moonshot AI Kimi K3 scores an AA-Briefcase Elo of 1543, second only to Claude Fable 5 at 1574, after a +727 jump over Kimi K2.6. That is a serious agentic knowledge-work result on a private dataset of realistic tasks across thousands of complex input files.
It is also expensive and slow on that bench: about USD 10.57 per task and 56.4 minutes average time per task, with roughly 83 turns and 120k output tokens per task. Capability without a cost and latency card is how SME pilots burn calendar and cash.
The business decision is not who won the leaderboard. It is whether your model route prices correctness, presentation quality, turn count, and wall-clock time for the jobs you actually run.
Artificial Analysis says AA-Briefcase combines correctness, analytical quality, and presentation into one Elo. Kimi K3 posts a 51 percent rubric pass rate (second to Fable 5 at 56 percent) and analytical quality Elo 1754, near Fable 5 at 1744. Presentation Elo 1471 lags GPT-5.6 Sol max at 1660 and Claude Opus 4.8 max at 1492.
Cost and latency tell the other half. Kimi K3 is priced at USD 3 / 15 per 1M input / output tokens with a 90 percent cached-token discount, yet averages among the highest AA-Briefcase task costs and times - about 2.5x Fable 5 time and about 3.8x Grok 4.5 high time, per Artificial Analysis. High turn use (83 vs 67 for Fable 5 and 50 for GPT-5.6 Sol max) drives both bill and wait.
Kimi K3 also scores 57 on the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5 class models. Independent benches help. They do not replace a job map for a 10-50 person firm that cannot afford hour-long tasks for every deliverable.
Use the Artificial Analysis numbers as planning inputs, not as a mandate to switch overnight.
Leaderboards move weekly. Your operating card should survive the next release.
Kimi K3 averages about USD 10.57 and 56.4 minutes per AA-Briefcase task while ranking second only to Fable 5 (Artificial Analysis).
This story maps first to Model Selection and Continuity Planning: right model for cost, capability, privacy, and latency, plus a replacement path when the economics shift.
It also maps to a Workflow ROI Audit when you need to prove which agentic knowledge tasks repay USD 10-class run costs and which should stay human or use a cheaper route.
AgentsROI is vendor-neutral. Kimi K3, Fable 5, GPT-5.6 Sol, and the rest can all be correct answers for different jobs. The service is judgment and continuity - not cheering a single lab.
If your team is shopping frontier agentic models on Elo alone, book a Model Selection conversation before the next pilot turns a 56-minute average task time into an unpaid queue.
This article summarizes publicly reported information and is for general informational purposes only. It does not constitute legal, tax, financial, investment, security, or compliance advice. AgentsROI.ai is not a law firm, accounting firm, or registered investment adviser. Facts, pricing, statistics, and product capabilities cited here reflect the sources listed at the time of writing and may change. Readers should verify current information independently and consult qualified professionals regarding obligations specific to their industry, jurisdiction, and circumstances—including applicable New York State and New York City requirements. AgentsROI.ai may have commercial relationships with vendors mentioned; where material, such relationships are disclosed. Nothing in this article is an endorsement of any specific AI product, model, or provider.