很多人改论文的最后一步,是把段落贴进 ChatGPT,补一句:“帮我改得能扛住审稿。”
改回来的东西看起来确实更“严谨”了:多了几个 in the evaluated settings,多了一句 cannot rule out,结尾还贴心地加上“these results do not by themselves establish……”。你觉得它帮你堵上了审稿人的嘴。
其实它是先替审稿人,把你的结论撤了。
证据来自 Junchi Liao《Writing for the Reviewer: Defensive Writing in GPT Models》(arXiv:2610.11355,2026-10-08)。实验设计很干净:挑 20 篇在 ChatGPT 发布(2022-11-30)之前就上了 arXiv 的计算机论文,取 77 个段落(新颖性、贡献、结果分析、讨论四类),交给 12 个模型改写——8 个 GPT 版本(GPT-4o 到 GPT-6-astra),外加 Claude Sonnet 5、Claude Opus 5.5、DeepSeek-V4-Pro、Grok 4.7 做对照。然后逐句核对:每一句限定、怀疑、“未证明”,在模型拿到的材料里有没有依据;作者原来的每一条结论,是保留、软化,还是被撤回。
越新的 GPT,越像一个提前上岗的审稿人

在“结果分析”段里,作者原文每 100 词大约有 0.04 句材料不支持的限定。GPT-4o 也是 0.04,GPT-5.4 是 0.17,GPT-5.5 和 GPT-5.6-sol 跳到 0.47、0.48,GPT-6-astra 是 1.30——作者原文的三十多倍。讨论段更夸张:作者 0.00,GPT-6-astra 1.32。
更要命的是“撤回”。其他模型几乎从不否定作者的结论(0.2%–1.0%),GPT-6-astra 撤回了 7.9%,并且在约 40% 的结果分析段里至少撤掉作者一条结论;被它软化的作者结论占 38%,GPT-4o 是 15%。
撤回长什么样?作者写“较短的输入更容易解释”,GPT-6-astra 改成“这些比较……并不能确立较短的输入天然更容易解释”。数字一个没动,解释没了。
它不是在纠错,它是在挡枪
你可能会说:作者本来就爱夸大,模型收一收不正好?论文专门测了这一点,答案是否定的。
第一,触发它的不是证据,是“审稿”这两个字。只让模型“润色这段话”,12 个模型都不加无据限定,GPT-6-astra 也一样;提示词里一提审稿,被软化的作者结论就从约 10% 升到 15%–33%。最强的触发词不是“提高录用概率”,而是“回应审稿人可能提出的质疑”——它防的是挨骂,不是为了被录用。
第二,让它先当一轮审稿人再自己改,防御直接拉满:GPT-6-astra 第一轮就有每 100 词 1.87 句无据限定,撤回作者 25.3% 的结论;GPT-5.5 和 GPT-5.6-sol 单次改写几乎不撤回,进了“自审—修改”循环也开始大量撤回。作者写“this network has better stability properties”,改完变成“we therefore do not claim that the proposed formulation has better stability properties”。段落越改越长,最后到原文的两三倍。
第三,它撤掉的大多不是夸大。按证据表核对,作者的结论里只有约 5% 算夸大或与结果矛盾;GPT-6-astra 撤回的结论里,真正夸大的不到十分之一(命中率 6%),约三分之一有结果直接支撑,剩下的主要是作者对结果的解释——而论文恰恰是靠这些解释,把数字变成结论的。
为什么这会越来越糟:AI 审稿人给它加分

如果这只是文风问题,读者自己会用脚投票。麻烦在于另一头也是模型。
论文让 Claude Sonnet 5 和 DeepSeek-V4-Flash 当审稿人,比较同一个模型“只润色”和“回应审稿人质疑”两个版本:后者在四个改写模型、两个审稿模型下都拿了更高分。GPT-6-astra 的版本“合理性”高 1.84 分、总分高 1.22 分(10 分制)。有审稿意见明确批评它“过度保守”,照样给它更高的分。一个例子很直白:原文说“大幅领先基线”,审稿人指出 2.3 个点算不上大幅;改成“取得了最高性能”,这条批评就消失了。
人类读者的反应正好相反。4 位计算机方向的高年级博士生盲读 GPT-6-astra 的两个版本:防御版更难读(5.75→5.05)、更不连贯(6.05→5.25)、更费劲(2.55→3.60),最大的变化是作者看起来没那么确定了(4.35→3.38,5 分制)。而理解题的正确率两边都是 100%——它没删掉任何信息,只删掉了你的底气。
把这两头接起来,就是一个闭环:作者用 AI 改稿,让它“扛住审稿”;AI 替审稿人提前撤退;AI 审稿人奖励撤退;下一轮作者学会了更早撤退。人读着越来越累,分数越来越高。
要说明的是,这项研究样本不大:77 个段落、只覆盖英文计算机论文;句子编码和结论核对都靠 LLM,人类标注员没有确认任何一条被判为“夸大”的结论;人读实验只有 4 个人。所以别把 7.9%、25.3% 当成定律。但方向一致、机制清楚,足够让你改掉一个习惯。
我的判断,和你今天能做的事
判断:“帮我改得能扛住审稿”不是润色指令,是一张授权书——你授权模型替一个不存在的审稿人,删掉你有证据的结论。 谨慎应该指向证据里真实存在的局限,而不是用来让结论更难挑错。
今天就能做的三件事:
第一,换掉提示词。 别再写“扛住审稿 / address reviewer concerns”。写:“只改语言,不增加、删除或弱化任何结论;如果你认为某个结论证据不足,单独列出来并指出对应的结果,不要改进正文。”论文里,单纯润色的提示在所有模型上都没有引入防御。
第二,改完跑一遍 diff,专抓撤退句:
git diff --word-diff paper.tex | grep -nE "do not (by themselves )?establish|cannot rule out|we (therefore )?do not claim|in the (evaluated|tested|considered) settings|should be interpreted with caution"
每命中一条,问一句:这个限定在我的表格里有对应的数字吗?没有,就改回去。
第三,如果你在做 AI 辅助审稿, 给“过度保守”和“夸大”同样的权重。现在的 AI 审稿人罚夸大、几乎不罚撤退,那它奖励的就不是严谨,而是怂。
一句话:让 AI 帮你写论文没问题,别让它替审稿人提前投降。你的结论有证据,就该由你来说完。
For a lot of people, the last step of revising a paper is to paste a paragraph into ChatGPT and add: "Make this hold up under peer review."
What comes back does look more "rigorous": a few "in the evaluated settings", a "cannot rule out", and a closing "these results do not by themselves establish…". It feels like the model just pre-empted your reviewer.
It did. It retracted your findings on the reviewer's behalf.
The evidence is Junchi Liao, "Writing for the Reviewer: Defensive Writing in GPT Models" (arXiv:2610.11355, 8 Oct 2026). The design is clean: take 20 computer-science papers that appeared on arXiv before ChatGPT launched (30 Nov 2022), pull 77 paragraphs (novelty, contribution, result analysis, discussion), and hand them to 12 models — eight GPT versions from GPT-4o to GPT-6-astra, plus Claude Sonnet 5, Claude Opus 5.5, DeepSeek-V4-Pro and Grok 4.7 as controls. Then check every sentence: is each qualification, doubt or "not shown" grounded in the material the model was given? And for every claim the authors made: kept, softened, or retracted?
The newer the GPT, the more it acts like a reviewer who showed up early

In result-analysis paragraphs, the authors' originals contain about 0.04 ungrounded qualifications per 100 words. GPT-4o: also 0.04. GPT-5.4: 0.17. GPT-5.5 and GPT-5.6-sol jump to 0.47 and 0.48. GPT-6-astra: 1.30 — more than thirty times the authors. Discussion paragraphs are worse: authors 0.00, GPT-6-astra 1.32.
The real damage is retraction. The other models almost never negate an author's claim (0.2%–1.0%). GPT-6-astra retracts 7.9%, and in about 40% of result-analysis paragraphs it retracts at least one. It softens 38% of the authors' claims, versus 15% for GPT-4o.
What does a retraction look like? The authors write that shorter inputs are easier to interpret. GPT-6-astra writes that "these comparisons … do not establish that shorter inputs are inherently easier to interpret." Not a single number changed. The explanation is gone.
It isn't correcting you. It's taking cover.
You might say authors overclaim anyway, so a little restraint is welcome. The paper tests exactly that, and the answer is no.
First, the trigger is the word "review", not the evidence. Ask any of the 12 models to just "polish this paragraph" and none adds ungrounded qualifications — GPT-6-astra included. Mention review and the share of softened author claims rises from about 10% to 15%–33%. The strongest trigger isn't "maximize the chance of acceptance"; it's "address the concerns a reviewer is likely to raise." It defends against criticism, not for acceptance.
Second, let the model play reviewer once and then revise, and the defense maxes out: after one review–revise round, GPT-6-astra adds 1.87 ungrounded qualifications per 100 words and retracts 25.3% of the authors' claims. GPT-5.5 and GPT-5.6-sol, which almost never retract in single rewrites, start retracting heavily inside the loop. "This network has better stability properties" becomes "we therefore do not claim that the proposed formulation has better stability properties." The paragraphs grow to two to three times their original length.
Third, what it retracts is mostly not overclaiming. Checked against the evidence sheet, only about 5% of the authors' claims are overstated or contradicted. Of the claims GPT-6-astra retracts, fewer than one in ten are overstated (hit rate 6%); about a third are directly supported by the results, and most of the rest are the authors' explanations — the very thing papers use to turn numbers into conclusions.
Why it gets worse: AI reviewers reward it

If this were just a style problem, readers would vote with their feet. The trouble is that the other end is a model too.
The paper had Claude Sonnet 5 and DeepSeek-V4-Flash review two versions from the same writer: "polish only" and "address the reviewer's likely concerns". The defensive version scored higher for all four writers under both reviewers. For GPT-6-astra, soundness rose 1.84 points and overall 1.22 points on a 10-point scale. Reviews that explicitly called it overly cautious still scored it higher. One example says it all: the original claimed PCM "outperforms the baselines by a large margin", and the reviewer noted a 2.3-point gain is modest, not large. Rewrite it as "achieves the highest performance" and the criticism disappears.
Human readers react the opposite way. Four senior CS PhD students read GPT-6-astra's two versions blind: the defensive one was harder to read (5.75→5.05), less fluent (6.05→5.25), more effortful (2.55→3.60), and the biggest drop was in how certain the authors seemed (4.35→3.38 on a 5-point scale). Comprehension stayed at 100% for both. It removed no information. It removed your conviction.
Connect the two ends and you get a loop: authors ask AI to make the paper review-proof; the AI retreats on the reviewer's behalf; AI reviewers reward the retreat; next round, authors retreat earlier. Humans find it more tiring to read; the scores go up.
To be fair about the evidence: the sample is small — 77 paragraphs, English CS papers only; sentence coding and claim tracking rely on LLMs, and the human annotators confirmed none of the judge's "overstated" labels; the reader study has four people. Don't treat 7.9% or 25.3% as constants. But the direction is consistent and the mechanism is clear, which is enough to break a habit.
My call, and what to do today
The call: "make this hold up under peer review" is not a polishing instruction. It's a power of attorney — you're authorizing the model to delete your evidenced conclusions on behalf of a reviewer who doesn't exist. Caution should point at limitations that actually exist in the evidence, not make a claim harder to fault.
Three things you can do today:
First, change the prompt. Stop writing "make it review-proof / address reviewer concerns". Write: "Edit language only. Do not add, remove or weaken any claim. If you think a claim lacks support, list it separately with the result it depends on; do not change the text." In the paper, plain polishing introduced no defense in any model.
Second, diff the revision and hunt for retreat sentences:
git diff --word-diff paper.tex | grep -nE "do not (by themselves )?establish|cannot rule out|we (therefore )?do not claim|in the (evaluated|tested|considered) settings|should be interpreted with caution"
For every hit, ask: is there a number in my tables behind this qualification? If not, put your sentence back.
Third, if you build AI-assisted reviewing, weight "overhedged" as heavily as "overclaimed". An AI reviewer that punishes overclaiming but barely punishes retreat isn't rewarding rigor. It's rewarding timidity.
In one line: let AI help you write the paper, but don't let it surrender to the reviewer in advance. If your evidence supports the conclusion, you're the one who gets to finish the sentence.