Signal Brief

Atomic Agent GAIA基准胜Hermes

Atomic Agent在GAIA Level 1基准测试中以69.8%正确率超越Hermes的58.5%,速度快1.6倍(3h12m vs 5h10m),使用相同4-bit Qwen-3.6-35B模型和Apple M4 Max硬件,开源MIT许可。

twitter关注列表 Chubby♨️ (@kimmonismus) 发布 2026-07-24 收录 2026-07-24 观察

一句话判断

该结果展示了通过KV-cache复用和工具调用压缩等运行时优化而非模型微调提升性能的方法,值得关注开源Agent工具的进步。

核心信息

Atomic Agent在GAIA Level 1基准测试中以69.8%正确率超越Hermes的58.5%,速度快1.6倍(3h12m vs 5h10m),使用相同4-bit Qwen-3.6-35B模型和Apple M4 Max硬件,开源MIT许可。

原始内容

Atomic Agent beat Hermes on GAIA: 69.8% vs. 58.5%, and finished 1.6× faster! They ran both agents through all 53 GAIA Level 1 tasks using the same 4-bit Qwen-3.6-35B on the same Apple M4 Max. - Atomic: 37/53 solved in 3h 12m - Hermes: 31/53 solved in 5h 10m Atomic solved six more tasks and finished nearly two hours sooner. It’s an Open-Source-Agent-Runtime model under MIT license. Open Source got a huge push! Love to see! https://video.twimg.com/amplify_video/2080752105984331776/vid/avc1/1920x1080/SopLIylYg-Ylckpf.mp4?tag=29 > **引用原帖 Atomic Agent (@atomicagent_io):** > Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! > We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. > Results: > ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m > ✦ Hermes Agent: 31 of 53 solved, took 5h 10m > Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. > Where it showed: > ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a blank. > ✦ Vietnamese specimens, which city they ended up in: Atomic pulled it from the first source and normalized the answer in 33s. Hermes spent 7.3 min and never answered. > ✦ The dinosaur featured-article nominator: Atomic walked the Wikipedia chain to "FunkMonk" in 57s. Hermes guessed a wrong name after 11 min. > Atomic keeps a byte-stable prompt prefix, so llama-server reuses the KV-cache instead of re-encoding the whole context every turn, and it emits one JSON array of tool calls per inference, then compresses results back instead of pasting them in full, so the context never balloons and a small model stays sharp deep into a task. On top of that a no-progress guard vetoes repeated identical tool calls (warn at 3, hard veto at 5) and forces a reply, so Atomic never sinks 15 minutes into re-scanning one page the way Hermes did. Both agents missed some of the same questions, and on a few Hermes got there and Atomic did not, usually format slips where Atomic computed the right number but printed the working instead of the bare value. But on identical hardware and identical weights, the runtime that reuses its cache and refuses to spin came out ahead on accuracy and speed. Getting this from the runtime alone is wild. > Run the same 53 GAIA tasks on Atomic Agent! > https://x.com/atomicagent_io/status/2080751328452604127

相关动态

02

Opus 5 性能大幅提升

在发布仅两个月内,Opus 5模型在ARC-AGI 3基准上从Opus 4.8的低于5%提升至超过30%,并在Artificial Analysis基准上以比Fable 5低26%的成本提供相当智能。

twitter关注列表2026-07-24#模型发布#技术突破
观察
03

OpenCode爆炸增长数据

Y Combinator 在播客中披露,OpenCode(开源替代 Claude Code 和 Codex 的 AI 编码工具)自年初起增长至 460 万周活用户、1300 万月活用户,年化收入约 4000 万美元,处理了 7 万亿 tokens,CEO Jay V 谈到 Anthropic 争议和 20 倍增长驱动因素。

twitter关注列表2026-07-24#行业动态#AI#开源
观察