AI

满意了,不等于办成了(中英切换)

满意了,不等于办成了。

团队越来越爱用便宜的离线闸门:LLM 用户模拟器聊完,再让 LLM 当裁判打满意度,高分就晋升。问题不在「裁判像不像人」,而在 人验证过的锚点本身可能钉错了——满意,几乎不告诉你任务有没有办成。

满意了,不等于办成了

满意过关,失败率几乎不动

Bodhwani、Tran、Wei《GAUGE》(arXiv:2609.12191) 在 τ²-bench 上抽 150 条对话,做盲评三人板:评分为满意(≥5/7)的对话里,57.5% 仍未办成客户任务。满意度与可验证成功的 Spearman ρ = -0.147,AUC 0.44。该样本基线失败率是 57.3%——条件在「满意」上,失败率几乎不降。这不是「弱相关」,是 无信息。

模式跨五类评分人群(约 47.6–59.5%)、两个基板都成立;SimulatorArena 数学辅导里,人类打 ≥8/10 的对话仍有 38.7% 答错。尊重、清晰、有帮助、会再来——每一维主观维度与成功的 |ρ| 都 ≤0.17。体验好与办成事,是两条轴。

总排序很稳,发版那一刀最晃

跨 25 个代理、六家提供商,闸门相对可验证回报的排序 Spearman ρ 约 0.94——粗筛与回归检测够用。但真要在 近乎势均力敌 的强代理里二选一(|ΔR|<0.1)时,闸门在 31% 的对子上推了回报更低的那一方;宽差距对子只有 <1%。发版决策恰恰发生在近平等区间。

看见工具与目标的策略感知闸门,能把接受集里的失败风险大约砍半(20% vs 基线 40.2%);过程盲的满意度代理做不到。证据通路决定可靠性,不是模型面子。

今晚:发版认证办成

别再用满意度单独当 promote/kill。发版闸门并列挂 可验证结果检查(DB 状态 / 动作 oracle / 能映射到真值的完成位);满意度只当体验遥测。便宜的完成位可拦截断回归;周期性可验证审计用来校准廉价闸门——先校准,再信任。

判断很简单:满意了,不等于办成了——闸门认证办成,满意只做旁路仪表。

Satisfied is not succeeded.

Teams increasingly ship agent variants through a cheap offline gate: an LLM user-simulator chats, an LLM-as-a-judge scores satisfaction, and the higher score gets promoted. The failure is not “does the judge sound human.” It is that a human-validated anchor can still be mis-anchored — satisfaction carries essentially no information about whether the task succeeded.

Satisfied is not succeeded.

Satisfied clears; failure barely moves

Bodhwani, Tran, and Wei, GAUGE (arXiv:2609.12191) take 150 τ²-bench transcripts through a blind three-person panel: among conversations rated satisfied (≥5/7), 57.5% still failed the customer’s task. Spearman ρ(satisfaction, success) = -0.147, AUC 0.44. Base failure in that sample is 57.3% — conditioning on “satisfied” does not lower failure. That is not a weak correlation. That is uninformative.

The pattern holds across five rater populations (~47.6–59.5%), both substrates; on SimulatorArena math tutoring, 38.7% of conversations humans rated ≥8/10 were still wrong. Respect, clarity, helpfulness, would-return — every subjective dimension stays at |ρ| ≤0.17 against success. Feeling good and getting it done are separate axes.

Ranking holds; the release cut wobbles

Across 25 agents and six providers, gate-vs-verifiable-reward ranking Spearman ρ is about 0.94 — fine for coarse quality and regression. But among near-equal strong agents (|ΔR|<0.1), the gate promotes the lower-reward agent on 31% of pairs versus <1% on wide pairs. Real release decisions live in that near-equal band.

A policy-aware gate that sees tools and goal roughly halves fail risk among accepts (20% vs 40.2% base); a process-blind satisfaction proxy does not. Evidence access decides reliability — not model pedigree.

Tonight: certify success at the gate

Stop promoting or killing on satisfaction alone. Hang a verifiable outcome check (DB-state / action oracle / a completion bit that maps to ground truth) beside any satisfaction score; treat satisfaction as experience telemetry. A free completion bit catches truncation regressions; periodic verifiable audits calibrate the cheap gate — calibrate, then trust.

The judgment is simple: satisfied is not succeeded — the release gate certifies success; satisfaction is a side gauge.