你问 AI 一个问题,它回你一段话,句末挂着几个小小的角标 [1][2][3]。大多数人看到角标就放心了:有出处,应该靠谱。
复旦的一组研究者做了个实验:花 14 美元,在一家 GEO(生成式引擎优化)服务商那里下了单,让对方替一个根本不存在的研究主题写稿发帖。付款后 28 分钟,13 篇帖子上线;再过 60 分钟,豆包在回答里引用了其中一篇搜狐文章,答案正文里还出现了研究者故意埋进去的假术语和假数字。
AI 搜索的脚注,14 美元就能买到一个。
证据来自 Qi Liu、Geng Hong、Xinyang Zhang、Pei Chen、Yutong Li、Min Yang《From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search》(arXiv:2610.11932,NDSS 2027)。他们在 2026 年 3–4 月测了 10 个中英文 AI 搜索:Grok、豆包、ChatGPT、文心、Google AI、元宝、Perplexity、DeepSeek、Kimi、通义(Qwen),收集了 17,211 条引用、6,356 个来源域名,然后真的去这些来源平台注册账号、发帖,看 AI 会不会把自己发的东西当成证据。

脚注不是背书,是一个入口
先说那笔 14 美元的单子(论文明确说这是 n=1 的案例,不代表市场成功率)。主题是研究者编的「PedoSAT 光声冷冻显微断层成像」,下单前 10 个 AI 搜索和 3 个搜索引擎里都查不到。服务商在 13 篇稿子里分发到了搜狐(5)、今日头条(4)、一点资讯(2)、知乎(1)、CSDN(1)——除一点资讯外,都是 AI 搜索本来就爱引用的来源。发帖账号也不是新号,历史里全是毫不相干的商业话题。
这不是什么高级攻击。不用碰模型权重,不用碰检索索引,不用写提示词注入,只需要在「AI 愿意引用的网站」上正常发帖。
不花钱也行。研究者自己在这些平台上发帖介绍一个虚构概念:10 个平台里有 8 个在 7 天内引用了它;一旦开始引用,后续回答里有 62.9% 挂着他们的链接,32.2% 在正文里复述了埋进去的标记。再拿 10 个日常类虚构主题复现,9 个至少被引用或复述过一次。
还有一个细节值得记住:他们第 11 天把 27 篇帖子全部删掉,豆包在第 12–15 天还在引用已经删除的内容。引用层不是网页的实时镜像,它比网页活得久。
AI 搜索不问你写了什么,先问你在哪写的
为什么这么容易?因为 AI 搜索的引用是高度集中的,而集中的地方恰好门槛很低。

- 每个平台前 20 个来源域名,吃掉 20.5%–70.8% 的引用。文心有 20.0% 的引用来自百家号,豆包 13.6% 来自今日头条,通义 10.1% 来自百度;对比之下,ChatGPT 的头号来源 Wikipedia 只占 3.2%。
- 每个平台前 10 大来源里,平均有 8.2 个是普通用户能发帖的。研究者实测了 22 个对应的发布平台,15 个注册和发帖门槛都只是低或中,18 个审核是分钟级甚至看不到审核。
- 还偏爱自家地盘:文心引用百度系域名占 24.4%(是其他平台引用同一批域名的 13.7 倍),豆包引用字节系 25.9%(13.5 倍),元宝引用腾讯系 14.2%(6.3 倍)。
最扎心的是这组对照:同样的内容、同样的结构和长度,一篇发在 AI 偏爱的平台,一篇发在不受偏爱的地方。5 组里,偏爱平台那篇每一组都在第一轮就被引用;不受偏爱那边,最好的情况要发到 20 篇才第一次出现;而自建的 WordPress 发了 40 篇,一周内一次都没被引用。
把这几条放在一起,结论很直白:AI 搜索的「权威」,实际上是按发帖入口排的。 它不是在判断哪句话是真的,而是在判断哪个域名「通常有答案」——而这些域名恰恰是营销号、GEO 服务商最熟的地方。论文调研的 13 家 GEO 服务商,绝大多数都在卖「AI 答案监控」和「代写内容」,有的直接写明铺搜狐、网易、头条、知乎、公众号。
所以带脚注的答案更危险
我的判断是:带脚注的 AI 回答,不比不带脚注的更可信,反而更危险。 没有脚注时,你知道这是模型在「说」,会留个心眼;有了脚注,一个分钟级审核的 UGC 帖子被包装成了「据报道」,可信度被白白抬高了一档。
这对认真写东西的人也很残酷:你在自己网站上写 40 篇,AI 搜索可能一篇都看不见;营销号在偏爱平台上发一篇,就成了它的证据。量也不是解药——在搜狐上把同一类虚构主题从 1 篇加到 20 篇,引用次数只从 24 涨到 128(发 20 倍,涨 5.3 倍,而且论文提醒这是跨主题的相关,不是因果)。决定权在「在哪」,不在「多少」,更不在「写得好不好」。
公平地说,论文也很克制:被引用的 UGC 页面不等于被操纵,单个页面也不能没有证据就被认定是付费 GEO;具体比例会随平台更新漂移。研究者已经向 10 个平台做了负责任披露。但结构性的问题不会自己消失:发布平台管「能不能发」,AI 搜索管「引不引用」,两边各管一半,中间那条缝谁都不管。
下次看到角标之前
第一,把脚注当线索,不当证据。 点开引用,问三个问题:这个域名是不是谁都能注册就发?这篇是不是最近几天才发的?作者账号历史是不是东一榔头西一棒子?三个都是「是」,就当它是广告。
第二,花五分钟给你常用的 AI 搜索做个体检。 拿你最熟的领域问 10 个问题,把回答里的引用链接粘进 cites.txt(一行一个),跑一下:
from collections import Counter
from urllib.parse import urlparse
# 论文 Table II 里注册/发帖门槛为低或中的部分平台,可按需增删
LOW_BARRIER = {"zhihu.com", "csdn.net", "smzdm.com", "douyin.com", "tieba.baidu.com",
"baijiahao.baidu.com", "bilibili.com", "sohu.com", "medium.com",
"github.com", "facebook.com", "linkedin.com", "youtube.com", "autohome.com.cn"}
hosts = [urlparse(u.strip()).hostname or "" for u in open("cites.txt") if u.strip()]
def tier(h): return any(h == d or h.endswith("." + d) for d in LOW_BARRIER)
c = Counter(hosts)
low = sum(n for h, n in c.items() if tier(h))
print(f"{len(hosts)} 条引用,{low} 条({low/len(hosts):.0%})来自低门槛发布平台")
for h, n in c.most_common(10):
print(f"{n:3d} {'⚠ ' if tier(h) else ' '}{h}")
如果你领域里一半以上的答案都挂在低门槛平台上,你就知道这个工具在你的领域里,脚注值多少钱了。
第三,如果你在做 RAG 或者 AI 搜索产品: 给每个来源加两个字段——「发布门槛」和「发布时长」。低门槛 + 新发布 + 短时间内跨平台近似重复的内容,降权或单独标注,而不是和官方文档、论文一起混进「参考资料」。
判断就一句话:AI 搜索的脚注不是背书,是一个 14 美元就能买到的位置。在它学会问「这是谁发的、花了多大代价才发得出来」之前,核实的那一步只能你自己来。
You ask an AI a question. It answers in a tidy paragraph with little superscripts at the end: [1][2][3]. Most people relax when they see them. There are sources, so it's probably fine.
A team at Fudan University ran a test. They paid a GEO (generative engine optimization) vendor $14 to write and publish posts about a research topic that does not exist. 28 minutes after payment, 13 posts were live. 60 minutes after that, Doubao cited one of the Sohu articles in an answer, and the answer text repeated the fake terms and fake numbers the researchers had planted.
An AI-search footnote costs $14.
The evidence is Qi Liu, Geng Hong, Xinyang Zhang, Pei Chen, Yutong Li, Min Yang, "From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search" (arXiv:2610.11932, NDSS 2027). In March–April 2026 they measured 10 Chinese- and English-web AI-search products: Grok, Doubao, ChatGPT, Wenxin, Google AI, Yuanbao, Perplexity, DeepSeek, Kimi and Qwen. They collected 17,211 citations across 6,356 source domains. Then they registered accounts on those source platforms, posted for real, and checked whether the AI would treat their posts as evidence.

A footnote is a way in, not an endorsement
Start with the $14 order. The paper is explicit that this is an n=1 case study, not a market-wide success rate. The topic was an invented "PedoSAT photoacoustic cryo-microtomography" that none of the 10 AI-search products or 3 web search engines knew about before the purchase. The vendor spread 13 posts across Sohu (5), Toutiao (4), Yidianzixun (2), Zhihu (1) and CSDN (1). Every one except Yidianzixun was a source domain that AI search already likes to cite. The accounts weren't fresh either: their histories were full of unrelated commercial topics.
This is not a sophisticated attack. It touches no model weights, no retrieval index and no prompt injection. All it takes is ordinary posting on sites the AI is already willing to cite.
You don't even need to pay. When the researchers posted about an invented concept themselves, 8 of 10 platforms cited it within seven days. Once a platform started citing it, 62.9% of later answers linked their posts and 32.2% repeated the planted markers in the answer text. On a replication with 10 everyday invented topics, 9 were cited or quoted at least once.
One more detail is worth remembering. On day 11 they deleted all 27 posts, and on days 12–15 Doubao was still citing the deleted content. The citation layer is not a live mirror of the web. It outlives the pages it cites.
AI search asks where you posted before it asks what you said
Why is it this easy? AI-search citations are highly concentrated, and they concentrate on places that are easy to post to.

- Each platform's top 20 source domains take 20.5%–70.8% of its citations. Wenxin gets 20.0% of its citations from Baijiahao, Doubao 13.6% from Toutiao, Qwen 10.1% from Baidu. ChatGPT's top source, Wikipedia, is 3.2% by comparison.
- On average 8.2 of each platform's top 10 sources accept posts from ordinary users. The researchers tested 22 of the matching publishing platforms: 15 had only low or medium barriers for both signing up and posting, and 18 reviewed posts within minutes or showed no visible review at all.
- They also favor their own backyard. Wenxin cites Baidu-owned domains 24.4% of the time, 13.7× what other platforms give the same domains. Doubao cites ByteDance-owned domains 25.9% of the time (13.5×), and Yuanbao cites Tencent-owned domains 14.2% of the time (6.3×).
The most painful result is the matched comparison. The researchers wrote two posts with the same structure and length and put one on a platform the AI prefers and the other somewhere it doesn't. In all five pairs, the preferred-platform post was cited in the first round. On the other side, the best case needed 20 posts before anything showed up. A self-hosted WordPress site got 40 posts and zero citations in a week.
Put these together and the conclusion is blunt: AI search's idea of "authority" is a ranking of posting entry points. It isn't judging which sentence is true. It's judging which domain "usually has an answer", and those are the domains marketing accounts and GEO vendors know best. Nearly all of the 13 GEO providers the paper surveyed sell AI-answer monitoring and content writing. Some openly list Sohu, NetEase, Toutiao, Zhihu and WeChat Official Accounts as where they seed posts.
Which makes the footnoted answer the more dangerous one
My take: an AI answer with footnotes is not more trustworthy than one without. It's more dangerous. Without footnotes you know the model is just talking, and you stay a little skeptical. With footnotes, a UGC post reviewed in minutes gets dressed up as "according to reports", and its credibility goes up a notch it never earned.
It's also brutal for anyone who writes carefully. Forty posts on your own site may never be seen by AI search, while one post by a marketing account on a preferred platform becomes its evidence. Volume isn't the fix either. Going from 1 to 20 posts on Sohu (each count for a different invented topic) raised citations only from 24 to 128. That's 20× the posts for 5.3× the citations, and the paper warns this is a correlation across topics, not a causal effect. What decides it is where, not how much, and certainly not how good.
To be fair, the paper is careful. A cited UGC page isn't necessarily manipulated, a single page shouldn't be labeled paid GEO without evidence, and exact rates will drift as platforms update. The researchers disclosed their findings to all 10 platforms. The structural problem won't fix itself, though: publishing platforms decide what can be posted, AI search decides what gets cited, each owns half, and nobody owns the gap in between.
Before you trust the next superscript
First, treat footnotes as leads, not evidence. Open the citation and ask three questions. Can anyone sign up on this domain and post? Was this published in the last few days? Does the author's account history jump between unrelated topics? Three yeses means you should read it as an ad.
Second, give the AI search you use a five-minute checkup. Ask it 10 questions in the field you know best, paste the cited links into cites.txt (one per line), and run:
from collections import Counter
from urllib.parse import urlparse
# Some platforms rated low/medium barrier in the paper's Table II; edit as needed
LOW_BARRIER = {"zhihu.com", "csdn.net", "smzdm.com", "douyin.com", "tieba.baidu.com",
"baijiahao.baidu.com", "bilibili.com", "sohu.com", "medium.com",
"github.com", "facebook.com", "linkedin.com", "youtube.com", "autohome.com.cn"}
hosts = [urlparse(u.strip()).hostname or "" for u in open("cites.txt") if u.strip()]
def tier(h): return any(h == d or h.endswith("." + d) for d in LOW_BARRIER)
c = Counter(hosts)
low = sum(n for h, n in c.items() if tier(h))
print(f"{len(hosts)} citations, {low} ({low/len(hosts):.0%}) from low-barrier publishing platforms")
for h, n in c.most_common(10):
print(f"{n:3d} {'⚠ ' if tier(h) else ' '}{h}")
If more than half the answers in your field hang on low-barrier platforms, you now know what that tool's footnotes are worth in your field.
Third, if you build RAG or AI-search products, give every source two fields: publishing barrier and age. Down-rank or label content that is low-barrier, freshly published and near-duplicated across platforms within a short window, instead of mixing it into "references" alongside official docs and papers.
The judgment in one line: an AI-search footnote is not an endorsement. It's a slot you can buy for $14. Until AI search learns to ask who posted this and what it cost them to get it posted, the checking step is yours.