← 返回精选动态
精选动态

Automated Alignment Researchers: Using large language models to scale scalable oversight

Anthropic 发表 Automated Alignment Researchers(Claude Opus 4.6、9 个实例、PGR 0.97)

更新于 2026年05月21日 14:55 1 条公开来源

发生了什么

Anthropic 发表 Automated Alignment Researchers(Claude Opus 4.6、9 个实例、PGR 0.97)

为什么值得关注

原文给出了可复现实验设置、明确的对照基线和量化结果,尤其是 9 个 Claude Opus 4.6 实例在 5 天内将 PGR 从 0.23 推到 0.97,适合关注自动化对齐研究的人直接读原文。

信息来源

以下内容来自公开来源,可打开原文继续核验。

01

Automated Alignment Researchers: Using lar

Anthropic-research 2026年04月14日 21:01

Anthropic 发布研究,探讨用大语言模型自动扩展对齐研究能力,并以 weak-to-strong supervision 作为 scalable oversight 的代理任务。研究中使用 9 个 Claude Opus 4.6 实例,配备 sandbox、共享论坛、代码存储和远程评分服务器,让它们自主提出、测试和分析对齐方法。以 Qwen 3-4B-Base 作为强模型、Qwen 1.5-0.5B-Chat 作为弱教师的人类基线在 7 天迭代中恢复了 23% 的性能差距(PGR 0.23);随后 Claude 在额外 5 天、累计 800 小时研究后将 PGR 提升到 0.97,成本约 18,000 美元、折合每个 AAR-hour 22 美元。

打开原始来源 ↗
← 返回精选动态