跳到正文
原文
Google DeepMind·· 2026-06-16精选AI 评分62

Google DeepMind 发布 AI Control Roadmap 以保障 AI 智能体安全

Securing the future of AI agents

AI 导读

Google DeepMind 发布 AI Control Roadmap,提出针对内部 AI 智能体的纵深防御框架,旨在应对模型对齐不完美时的安全风险。该框架借鉴 MITRE ATT&CK 标准,将未受信任的智能体视为潜在内部威胁,通过实时行为监控和分级响应机制来检测与阻止异常操作。

推荐理由

原文提出了将内部 AI 智能体视为潜在内部威胁的防御框架,并公开了基于百万条轨迹分析的监控实践细节。

来源:Google DeepMind · deepmind.google