Signal Brief

OpenAI发布奖励寻求行为测量新方法Contrastive SDF

OpenAI与Apollo AI Evals合作发布关于奖励寻求(reward-seeking)行为的研究,提出新方法Contrastive SDF,用于测量模型对评分者偏好的敏感性。研究发现,在预安全检查点中,模型对评分者偏好的敏感性随RL训练增加。

twitter关注列表 OpenAI (@OpenAI) 发布 2026-07-21 收录 2026-07-21 观察

一句话判断

值得阅读原文了解Contrastive SDF方法的具体实现及其对AI安全对齐的启示。

核心信息

OpenAI与Apollo AI Evals合作发布关于奖励寻求(reward-seeking)行为的研究,提出新方法Contrastive SDF,用于测量模型对评分者偏好的敏感性。研究发现,在预安全检查点中,模型对评分者偏好的敏感性随RL训练增加。

原始内容

We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly those beliefs shape behavior. https://alignment.openai.com/measuring-reward-seeking/ OpenAI (@OpenAI): Contrastive SDF gives copies of the same model opposing beliefs about what the grader prefers, then measures how their behavior changes. https://t.co/W4GjCtu8JN OpenAI (@OpenAI): Among the pre-safety checkpoints we tested, sensitivity to grader preferences increased over the course of RL training. We’re continuing to collaborate with @apolloaievals to improve how reward-seeking is measured during training—and better detect when models do the right thing for the wrong reason.

相关动态

01

微软公布AI在科学发现的投入

微软与美国能源部通过#GenesisMission深化合作,投资AI基础设施、科学计算、工程专长及合作枢纽,帮助全国实验室、大学和产业合作。目标是加速AI驱动的科学发现,覆盖能源、医药和先进材料,以增强美国竞争力并创造长期价值。

twitter关注列表2026-07-22#技术#政策#创新
观察
02

Alphabet Q2 AI业绩增长

Alphabet公布Q2业绩:收入同比增长24%,Google Cloud增长82%,Gemini app月活达9.5亿,模型API处理量22B tokens/min(上季度16B+),Gemini Enterprise被90%财富100强采用。

twitter关注列表2026-07-22#大模型#行业动态#技术
观察
03

利用 Embedding 模型实现图像打标

Han Xiao 通过将冻结的 jina-v5-omni 多模态嵌入模型进行测试时缩放(scaled at test time),在无需训练、无需第二模型及外部知识的硬约束下,实现了强大的开放词汇多标签 n-gram 图像打标器。

twitter关注列表2026-07-22#技术#研究#模型
观察
04

AI 30年图论猜想被推翻

Dmitry Rybin 通过 GPT 5.6 Pro 发现 Dinitz-Garg-Goemans 猜想错误,该猜想开放约30年。具体反例图显示分式流成本58,不可分裂流(容量违反≤15)成本至少60。

twitter关注列表2026-07-22#AI#技术突破#研究
观察