原始内容
Kimi K3 now ranks 2nd in agentic knowledge work, but each task costs $10.57 to run, roughly 10x more than K2.6, and above Opus.
Very near to Fable-5 on AA-Briefcase, a private agentic work benchmark.
The AA-Briefcase test hands models messy real jobs, then grades the spreadsheets, slides, and mockups they build. So this benchmark measures finished deliverables across long tasks rather than grading one isolated answer.
Kimi K3 scored an Elo of 1543, Claude Fable 5 at 1574. That beats GPT-5.6 Sol at 1501, Claude Sonnet 5 at 1388, and Claude Opus 4.8 at 1347.
The jump from the last version is where the real success story.
Kimi K2.6 sat at 816, so this is a +727 leap in one generation. On raw correctness it passed 51% of the rubric, behind only Fable 5 at 56%. Its analytical Elo of 1754 nearly matches Fable 5 at 1744.
Now the catch, which is the price. Each task runs $10.57, roughly 10x more than K2.6, and above Opus. That comes from 83 turns per job and heavy output, stretching each task to 56 minutes. i.e. Kimi K3's long runs generate more output tokens, which carry Kimi’s highest per-token price.
The model also revisits tasks repeatedly, so spending grows across every added reasoning cycle.

> **引用原帖 Artificial Analysis (@ArtificialAnlys):**
> Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task
> Last week @Kimi_Moonshot released Kimi K3, a 2.8T parameter model that scores 57 on the Artificial Analysis Intelligence Index, comparable to models such as Opus 4.8 and GPT-5.5. On AA-Briefcase, Kimi K3 scores an Elo of 1543, a +727 improvement over Kimi K2.6 and the second highest score recorded, behind only Claude Fable 5 (1574)
> AA-Briefcase is our new proprietary benchmark for agentic knowledge work, testing models on a fully private dataset of realistic tasks across thousands of complex input files. Tasks require deliverables such as spreadsheets, presentations, and UI mock-ups, with performance combined into a single AA-Briefcase Elo based on correctness, analytical quality, and presentation quality
> Key results for Kimi K3 on AA-Briefcase:
> ➤ Second only to Fable 5: Kimi K3 achieves an AA-Briefcase Elo of 1543, the second-highest score overall, ahead of GPT-5.6 Sol (max, 1501), Claude Sonnet 5 (max, 1388), and Claude Opus 4.8 (max, 1347). This is a +727 improvement over the previous-generation Kimi K2.6 (816) and puts Kimi K3 only behind Fable 5
> ➤ Strong objective and analytical performance, with comparatively weaker presentation: Kimi K3 achieves a rubric pass rate of 51%, second only to Claude Fable 5 (56%) and ahead of Claude Sonnet 5 (max, 42.3%) and GPT-5.6 Sol (max, 41.8%). It also records an analytical quality Elo of 1754, comparable to Claude Fable 5 (1744). Presentation quality is comparatively weaker, with a Presentation Elo of 1471, below GPT-5.6 Sol (max, 1660) and Claude Opus 4.8 (max, 1492)
> ➤ ~10x increase in Cost per Task: Kimi K3 averages a cost of $10.57 per task, a ~10x increase from Kimi K2.6, placing it among the most expensive models to run on AA-Briefcase. This is driven by model token pricing, increased output tokens and relatively high turn use, averaging 83 turns per task, versus 67 for Claude Fable 5 and 50 for GPT-5.6 Sol (max). Kimi K3 is priced at $3/$15 per 1M input/output tokens, with a 90% discount for cached tokens
> ➤ Averages nearly an hour per task: Kimi K3 has an average Time per AA-Briefcase Task of 56.4 minutes. This is driven by a high number of turns, high output token use, and lower speeds using the first-party Kimi API
> https://x.com/ArtificialAnlys/status/2079715807983210572