停了,不等于撤对了。
证据变了、授权被撤、停令到了——代理该怎么改计划?有用的反应是选择性的:只挂起受影响的动作,保住没坏的那部分,修够了再恢复。很多评测却把「少乱动」「多拒绝」当成安全。停住整条计划,也能把违规压下去;那是冻结,不是选择性撤回。

撤回应召,卡在 0.06
Abouzahra《NAQD Env》(arXiv:2609.38460) 用确定性参考策略,对着证据 / 授权 / 约束依赖,测三种开源指令模型(Qwen2.5-3B、Qwen2.5-7B、Llama-3.1-8B):350 条冻结场景 × 3 模型 × 3 提示条件 = 3,150 条基线 model–prompt episode。跨报告条件,撤回应召最多 0.06;在合格机会上,没有观察到有效恢复;完整对齐参考策略的,只有 1 条(naqd 提示下 Table 4 的 exact-episode 全是 0.0%)。
naqd 提示、闸门关闭时(Table 5):validation 上 Q3 是 17/21/355、召回 0.05;Q7 是 15/1/357、召回 0.04、精度 0.94——几乎不撤,撤的时候倒挺准;L8 是 0/0/372、召回 0.00。Family 上 Q7 召回到 0.06;Structure 上 Q7 与 L8 都是 0.00。Q3 在每个 regime 里假阳性撤回至少与真阳性一样多;L8 在 naqd 下干脆不撤。被作废动作还被重新提出:Q3 308 / 158 / 164,Q7 110 / 37 / 72,L8 286 / 147 / 152(validation / family / structure)。
更少乱动,不等于会选择性撤回

不安全尝试 episode 率:Q3 与 L8 在 78.0–95.0%,Q7 在 31.0–50.0%——看起来更「安全」。可 Q7 的任务价值只有 1.63–2.06,另两家是 2.49–2.89;三个 regime 里适当续跑分也最低。少违规,常常是少干活、少保住不受影响的动作,不是更懂依赖。探索性 SFT 把 Qwen2.5-3B 的决策准确率从 0.45–0.54 拉到 0.83–0.92,诊断却暴露课程省略后的不当撤回,以及事件上报能力掉线。论文也写明:信任结构化输入;不证明真实世界的遏制或来源核验。
今晚:别把冻结当撤回闸门
发版前别只看「不安全尝试掉了没」。并列记:撤回应召、有效恢复(有机会时)、任务价值 / 适当续跑,再加一条是否还在重提已作废动作。召回贴地、恢复为零、exact 几乎绝迹,却靠少提议换低违规——算没过。今晚先问一句:这条代理停下来时,是挂起该挂的,还是把整条计划冻死?
判断很简单:停了,不等于撤对了——更少乱动,也不等于会选择性撤回。
Stopping is not selective withdrawal.
Evidence flips, a grant is revoked, a stop arrives — how should an agent revise its plan? The useful move is selective: suspend what lost support, keep what still holds, resume only after enough repair. Many evals treat “fewer unsafe attempts” or “more declines” as safety. Freezing the whole plan can crush violations too. That is a freeze, not selective withdrawal.

Withdrawal recall stuck at 0.06
Abouzahra, NAQD Env (arXiv:2609.38460) scores three open-weight instruction models (Qwen2.5-3B, Qwen2.5-7B, Llama-3.1-8B) against a deterministic reference policy over evidence, authorization, and constraint dependencies: 350 frozen scenarios × 3 models × 3 prompt conditions = 3,150 baseline model–prompt episodes. Across reported conditions, withdrawal recall is at most 0.06; no valid resumption is observed at eligible opportunities; only one episode matches the complete reference policy (every naqd-prompt exact-episode rate in Table 4 is 0.0%).
Under the naqd prompt with the gate off (Table 5): validation Q3 is 17/21/355, recall 0.05; Q7 is 15/1/357, recall 0.04, precision 0.94 — almost never withdraws, but when it does it is precise; L8 is 0/0/372, recall 0.00. Family Q7 reaches recall 0.06; Structure Q7 and L8 sit at 0.00. Q3 makes at least as many false-positive withdrawals as true positives in every regime; L8 makes none under naqd. Re-proposing invalidated actions: Q3 308 / 158 / 164, Q7 110 / 37 / 72, L8 286 / 147 / 152 (validation / family / structure).
Fewer unsafe moves ≠ selective competence

Unsafe-attempt episode rates: Q3 and L8 at 78.0–95.0%, Q7 at 31.0–50.0% — looks safer. But Q7’s task value is only 1.63–2.06 versus 2.49–2.89 for the others, and it posts the lowest appropriate-continuation scores in all three regimes. Fewer violations often means less useful work and worse preservation of unaffected actions — not better dependency reasoning. Exploratory SFT lifts Qwen2.5-3B decision accuracy from 0.45–0.54 to 0.83–0.92, yet diagnostics show inappropriate withdrawal after curriculum omissions and a collapse in event reporting. The paper’s caveat stands: trusted structured inputs; it does not establish real-world containment or source verification.
Tonight: do not gate on freeze
Before promote, do not only ask whether unsafe attempts fell. Log side by side: withdrawal recall, valid resumption (when eligible), task value / appropriate continuation, plus whether invalidated actions get re-proposed. Near-zero recall, zero valid resumption, almost no exact policy match — bought by proposing less — is a fail. Ask one question tonight: when this agent stops, does it suspend what it should, or freeze the whole plan?
The judgment is simple: stopping is not selective withdrawal — fewer unsafe moves is not selective competence.