Signal Brief

AI 解决 19 个 Erdős 问题宣称

Przemek Chojecki 声称使用 GPT 模型(主要是 GPT-5.5 Pro)在四月最后两周内解决了 19 个 Erdős 问题,其中 3 个新宣称由 GPT-5.6 Sol Ultra 在七月完成,但部分解决方案尚未经过正式验证。

twitter关注列表 Marc Andreessen 🇺🇸 (@pmarca) 发布 2026-07-15 收录 2026-07-16 观察

一句话判断

该推文具体列出了不同 GPT 版本(5.2, 5.4, 5.5, 5.6)在数学问题上的表现差异,并指出大多数解决方案仅需少量提示,值得关注其实际效果。

核心信息

Przemek Chojecki 声称使用 GPT 模型(主要是 GPT-5.5 Pro)在四月最后两周内解决了 19 个 Erdős 问题,其中 3 个新宣称由 GPT-5.6 Sol Ultra 在七月完成,但部分解决方案尚未经过正式验证。

原始内容

Marc Andreessen 🇺🇸 (@pmarca) 转发了 Przemek Chojecki | PC (@prz_chojecki) 的帖子: 19 Erdős problems I claimed with AI Just to update the list for all the solution claims I've got with GPTs - 5.5 Pro solved the most, but each 5.2, 5.4, 5.6 solved something in the end. 3 New Claims with GPT-5.6 Sol Ultra in July: #421, #793, #415 (loose ends closed from previous attempt by GPT 5.4) 16 Previous Solution Claims: #258, #522, #603, #610, #750, #856, #858, #888, #896, #953, #956, #1092, #1133, #1148, #1151, #1190 5 almost/partial: #514 (finished with another person), #689 (finished with another person, bookkeeping is tricky), #741 (cleaning what's already proven from OpenAI/Deepmind), #906 (was completed by someone else first, a bit different), #1201 (partial, but close). Most of the Solution Claims I've got with the help of GPT-5.5 Pro in the last two weeks of April 2026. What a marathon that was! I'm writing solution claims not solutions, because not all of the solutions were verified - either formally or by other mathematicians. However each was run through an adversial LLM check at least twice and read by me. Since May I wasn't as active pursuing these problems, but it makes sense to revisit some now with the release of GPT-5.6. ![photo](https://pbs.twimg.com/media/HNQICuLXUAA2WTS.jpg) > **引用原帖 Przemek Chojecki | PC (@prz_chojecki):** > 12 Erdos Problems I claimed with AI. > They still need to be verified properly for the final score, though most look fine. > Here's the full summary of my AI+math marathon: > 12 Solution Claims: #522, #603, #610, #856, #888, #896, #953, #956, #1092, #1133, #1151, #1190 > 4 almost/partial: #514 (finished with another person), #741 (cleaning what's already proven from OpenAI/Deepmind), #906 (was completed by someone else first, a bit different), #1201 (partial, but close) > On top of that I had 3 solutions verified before: #258, #858, #1148 > Also I've got MANY partial results in other problems, that I haven't posted. It was usually a small increment like optimizing a constant in an estimate that didn't really give any new math or insight into a full solution. > One thing to add is I'm pretty humbled by the whole experience. > Was it really me that solved them? > Except for maybe 3-4 problems where my LLM guidance was very explicit with methods to be used, most of these didn't require much guidance besides crafting a good prompt with references, potential approaches, computations. Couple of them were proved basically in one prompt (prompt -> solution). > A general remark is that except for maybe 2 problems, I've got solutions in max 4 prompts. GPT-5.5 Pro is way more efficient than GPT-5.4 Pro, but also if it doesn't hit quick, it makes little sense to push it. > https://x.com/prz_chojecki/status/2050160288108892407

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
05

开源模型长期网络能力差距缩小至4-7个月

长期网络能力方面,领先开源模型落后闭源前沿从2025年大部分时间的6-10个月缩短至4-7个月。GLM-5.2匹配了约4个月前发布的闭源模型。AI安全研究所的32步“The Last Ones”任务中,GPT-5.6 Sol在10次尝试中完成7次,每次预算1亿token,且性能随推理token增加而提升。

twitter关注列表2026-07-17#模型#技术突破#AI安全
观察