Signal Brief

LLM研究想法范围比人类更窄

耶鲁大学和芝加哥大学研究者通过分析11683篇真实论文发现,人类研究者仅12.1%的创意是连接已有工作,而LLM的47.1%至64.2%为此类,使用频率是人类的4至5倍,表明LLM研究想法范围更窄。

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-04 收录 2026-07-06 观察

一句话判断

该研究提供了LLM与人类研究想法差距的量化证据,值得关注其方法论是否适用于评估AI科研创意。

核心信息

耶鲁大学和芝加哥大学研究者通过分析11683篇真实论文发现,人类研究者仅12.1%的创意是连接已有工作,而LLM的47.1%至64.2%为此类,使用频率是人类的4至5倍,表明LLM研究想法范围更窄。

原始内容

This is the prompt Yale and Univ of Chicago researchers used when asking LLMs for new research ideas. Feed LLMs prior work, ask for ideas, then measure how repetitive the ideas get. The surprising finding is that LLMs often treat research ideation as connecting what already exists, while humans use a wider set of problem-finding moves. LLM-generated ideas reveal a bias toward safe bridge-and-combine proposals. ![photo](https://pbs.twimg.com/media/HMaahLdaIAAkHaH.jpg) > **引用原帖 Rohan Paul (@rohanpaul_ai):** > This Yale + University of Chicago paper shows that real gap between LLM generated research ideas vs humans is not idea quality, but idea range: LLMs think narrower than human researchers. > The researchers built a controlled test from 11,683 real papers, using each paper’s nearby prior work as the shared starting point. > They asked models to propose a new motivation and method from those same prior papers, then compared those ideas with the real human paper ideas. > Instead of asking whether 1 idea looked novel, they labeled each idea by what gap it noticed and what kind of contribution it made. > Human ideas spread across many patterns, such as explaining mechanisms, testing failures, measuring evidence, building systems, and improving efficiency. > Only 12.1% of human ideas were mainly about connecting separate work, but 47.1% to 64.2% of LLM ideas did that, meaning models used this move about 4 to 5 times more often. > Even extra reasoning made this pattern stronger, suggesting models often polish a familiar recipe instead of finding more varied research moves. > --- > – arxiv. org/abs/2607.01233 > Title: "Measuring the Gap Between Human and LLM Research Ideas" > https://x.com/rohanpaul_ai/status/2073502907920703553

相关动态

01

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
02

端到端开源的VideoChat3视频多态大模型(Zeros-R RTX 3090/L40 无显存瓶颈)

Meta AI团队发布VideoChat3,采用Zeros-R架构在RTX 3090/L40 NVIDIA显卡上实现82FPS实时处理的视频多态大模型,技术突破显存瓶颈(支持10M/T框架框架和3小时视频输入)。模型包含20亿参数的基本配置和更大版本,首次在与Meta AudioGen、Microsoft VALL-E的基准测试中展示出在...

twitter关注列表2026-07-17#技术突破#大模型#视频理解
值得跟进
03

Sakana AI 提出无需反向传播且遵守 Dale 原理的神经网络训练方法

Sakana AI 发表论文《Diffusing Blame》,提出一种训练神经网络的方法,该网络严格遵循 Dale 原理(每个神经元要么兴奋要么抑制),无需反向传播。该方法基于 Error Diffusion 扩展,通过模数错误路由在图像分类和强化学习任务(Ant、Humanoid、HalfCheetah、Craftax)上取得有竞争力...

twitter关注列表2026-07-17#技术突破#研究#AI
观察
04

Kimi K3发布:2.8万亿参数模型

Moonshot AI 发布 Kimi K3 模型,2.8 万亿参数,百万 tokens 上下文,原生多模态;采用 Kimi Delta Attention 实现百万 tokens 解码速度 6.3 倍提升,Attention Residuals 以小于 2% 额外成本提升训练效率约 25%;已上线多个平台,开源权重承诺 2026 年 7...

twitter关注列表2026-07-17#模型发布#大模型#技术突破
值得跟进
05

AI说服力远超人类:三倍效果

在一系列包含20000次对话和7000名参与者的实验中,AI模型在说服力上超越了人类冠军辩手和专业募捐者,即使人类获得准备、培训和高达1000英镑的奖金也无法赶上。AI的优势来自速度和信息量,当限制到人类水平时差距消失;无限制时AI说服捐赠货币的效果是专业募捐者的近三倍。

twitter关注列表2026-07-17#AI#研究#技术
观察