Signal Brief

GPT-5.6 Sol解决6个Erdős问题

一位研究员使用GPT-5.6 Sol在5天内解决了6个开放Erdős问题(尝试13个,成功率46%)。他采用合同式提示,明确证明要求、排除弱结果,并指定搜索策略,包括同时探索多种路径和对抗性检查。

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-23 收录 2026-07-23 观察

一句话判断

这种合同式提示和搜索策略对解决复杂数学问题有效,可尝试复现。

核心信息

一位研究员使用GPT-5.6 Sol在5天内解决了6个开放Erdős问题(尝试13个,成功率46%)。他采用合同式提示,明确证明要求、排除弱结果,并指定搜索策略,包括同时探索多种路径和对抗性检查。

原始内容

. > **引用原帖 Rohan Paul (@rohanpaul_ai):** > A single researcher solved 6 open Erdős problems in 5 days using GPT-5.6 Sol. > He attempted 13 problems in total, so the success rate sits near 46%. > He wrote the prompt almost as a contract, not a question. > It restated the problem, then spelled out what a finished proof must actually establish. listed the weaker results that would not count, so near-misses could not sneak through. > Known traps and edge cases went in before the model ever started thinking. Then he told it how to search, which is where most people stop. > Chase many approaches at once, keep incompatible ones alive, and never commit early. > Hunt for counterexamples to your own lemmas, and kill a route that only leads to another open problem. > Separate adversarial agents then tried to break every draft that survived. > None of this works without the right target, so he picked problems mathematicians already argue about, avoiding anything welded to a famous conjecture. > Codex held the whole search in memory for hours while that loop ran itself. > https://x.com/rohanpaul_ai/status/2080143115885916189

相关动态

01

Dreamina Seedance 2.0 4K 更新

BytePlus 旗下 Dreamina 发布 Seedance 2.0 4K 视频生成模型,原生 4K 分辨率,支持文本、图像、视频、音频参考输入及视频编辑扩展,通过 BytePlus ModelArk API 运行。导演 Neill Blomkamp 使用该模型创作了 13 分钟科幻恐怖短片 Nightborne,并计划用于长片制作。

twitter关注列表2026-07-23#技术突破#产品更新#AI
观察
02

红点笔记模型获 IMO 完美分数

RedNote 的 dots‑note‑3.0 AI 模型在国际数学奥林匹克赛中得满分 42/42,超越去年 DeepMind 和 OpenAI 的 35/42,且仅有 7 名 666 名人类选手匹配。该模型仍在 beta,为 dots3 系列中最小的变体,公司承诺日后开源。

twitter关注列表2026-07-23#模型#技术
值得跟进
04

RedNote AI 模型获 IMO 满分 42/42

RedNote(小红书母公司)旗下 AI 模型 dots-note-3.0 在国际数学奥林匹克竞赛(IMO)中获得 42/42 满分,成为首个达成此成就的 AI 系统。该模型直接阅读原始竞赛文档,通过结合自然语言推理与 Python 代码执行的代理循环,经草稿测试、破坏、修复及自我审查后提交答案,全程无人工改写或提示。对比之下,Googl...

twitter关注列表2026-07-23#AI#技术突破#大模型
值得跟进
05

RedNote AI模型IMO满分42分

RedNote(小红书)的AI模型dots-note-3.0在国际数学奥林匹克竞赛中获得满分42分,而Google DeepMind和OpenAI去年仅得35分。该模型直接阅读原题,使用自然语言推理与自执行Python代码的智能体循环,并具备自我审查机制,最终在禁止外部帮助的条件下取得满分。模型为dots3系列最小版本,承诺未来开源。

twitter关注列表2026-07-23#技术突破#大模型#AI模型
观察