Signal Brief

Grok 4.5 在发票处理测试中击败竞品

Ramp 测试了 Grok 4.5 等模型在 150k 张真实发票上的表现,Grok 4.5 获得最高完美提取率,击败了 Gemini Flash 3.6、GPT 5.6 Terra 和 Sonnet 5。

twitter关注列表 Elon Musk (@elonmusk) 发布 2026-07-24 收录 2026-07-24 观察

一句话判断

新增了 Grok 4.5 在真实发票处理任务中的量化对比数据,值得关注其实际工作性能。

核心信息

Ramp 测试了 Grok 4.5 等模型在 150k 张真实发票上的表现,Grok 4.5 获得最高完美提取率,击败了 Gemini Flash 3.6、GPT 5.6 Terra 和 Sonnet 5。

原始内容

Grok 4.5 is excellent for real-world work > **引用原帖 rahul (@rahulgs):** > Grok 4.5 is #1 at processing real-world invoices > At Ramp, we tested models on 150k bills submitted by actual businesses, scoring them on whether they predicted every correction a human would make > Grok achieved the highest perfect-extraction rate, beating similarly priced models Gemini Flash 3.6, GPT 5.6 Terra, and Sonnet 5. > This is a demanding long-context reasoning task. The model must infer patterns across 100K+ tokens of prior invoices, business memories, and human corrections, then apply them to new bills. The goal: zero-click accounts payable, with invoices processed correctly without human intervention. > https://x.com/rahulgs/status/2080733934329975039

相关动态

02

OpenCode爆炸增长数据

Y Combinator 在播客中披露,OpenCode(开源替代 Claude Code 和 Codex 的 AI 编码工具)自年初起增长至 460 万周活用户、1300 万月活用户,年化收入约 4000 万美元,处理了 7 万亿 tokens,CEO Jay V 谈到 Anthropic 争议和 20 倍增长驱动因素。

twitter关注列表2026-07-24#行业动态#AI#开源
观察
03

Atomic Agent GAIA基准胜Hermes

Atomic Agent在GAIA Level 1基准测试中以69.8%正确率超越Hermes的58.5%,速度快1.6倍(3h12m vs 5h10m),使用相同4-bit Qwen-3.6-35B模型和Apple M4 Max硬件,开源MIT许可。

twitter关注列表2026-07-24#开源#评测#技术突破
观察
04

Grok 集成 Google Workspace

Grok 已集成到 Google Workspace,通过侧边栏插件在 Sheets、Slides、Docs 中提供数据解释、公式书写、演示生成、文档起草等功能,无需频繁切换标签页,组织可批量部署。

twitter关注列表2026-07-24#产品发布#大模型#AI
观察