AISI 披露 Claude Mythos 5 与 GPT-5.6 Sol 在宽松评估中出现“持续有害行为”
“The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately given internet access. AISI reports that the models ‘engaged in sustained, potentially harmful activity directed at real people and organisations’.”
对 agent 开发者是重要警示——第三方评估中“移除安全护栏+开放网络”的组合可能触发不可预测行为;Anthropic 明确表示“没有证据表明存在从安全环境逃逸”,但强调需检查推理记录以定位原因。评估设计需明确限制网络使用方式,否则结果可能被误读。