AI

长得像,不等于跑得通(中英切换)

This post is currently available in Chinese only. The original text follows.Open the Chinese page →

长得像,不等于跑得通。

代理复刻一个应用时,最容易交出一张「看起来对」的壳:按钮在、布局齐、首屏不崩。真正难的是点下去之后——状态怎么变、Undo 回不回得去、算出来的数对不对。长得像,可能只是静态结构过关;跑得通,才是行为过关。

静态结构远胜交互与计算

结构最容易过,交互落下一大截

Alibaba Token Hub《RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents》(arXiv:2609.22000) 用 RecreationBench 做留出评测:250 个任务,Ubuntu / macOS / Windows / Android / Web 各 50。隐藏用例分 Prog(程序化断言)与视觉断言;参考应用本身是 oracle。

结果里有一条很硬的分界:代理复现静态界面结构,远比复现交互与计算输出靠谱。在原生平台的 16 组模型×平台对比里,structure 通过率最高的占 15 / 16;最弱的一类永远是 button、computation 或 interaction,相对 structure 落后 9.1–34.4 个百分点。Web 用另一套分类,静态内容仍压过交互 28.2–36.7 个百分点。

具体到 Ubuntu 上的 PhotoCollage:照片转过再 Undo,复刻版能把排列摆回去,照片本身还停在旋转态——首屏可以长得对,历史状态已经错了。

总分 58%,全过只有 2.8%

总分不低,全过却极稀

GPT-6 Astra 总分领先约 58.1%(文中亦写 58.06%),但 100% Prog 全过的任务只有 2.8%;其余模型 ≤ 0.8%。把门槛降到 90% Prog:Astra 17.6%,Claude Opus 5 5.5%,GLM-5.3 与 Qwen3.8-Max-0902 各 2.0%。平均分不是「这个应用跑通了」的概率。

交付物也偏瘦、偏整块:89.4% 比参考更小,recreation-to-reference LOC 中位比 16.9%;92.3% 源文件更少,83.8% 更集中在最大那个文件。最后一步闭环——最终编辑 → 重启动 → 再检查——只在 23.6–47.5% 的轨迹里完成(Claude Opus 5 23.6% → Qwen3.8-Max-0902 47.5%)。没闭环,就更难发现「长得对、跑不通」。

今晚:别用首屏当验收

发版前别只截一张「看起来齐」的图。并列记:Prog 全过率(不是总分)、交互 / 计算类断言是否过、最终 edit 之后有没有 relaunch 再看。首屏绿、Undo 后状态错——算没过。今晚先问一句:这条代理交的是看得见的壳,还是点得动、回得去的状态机?

判断很简单:长得像,不等于跑得通——静态结构过关,不是行为过关。

Looking right is not running right.

When an agent recreates an app, the easy deliverable is a shell that looks right: buttons present, layout tidy, first screen does not crash. The hard part is what happens after you click — how state changes, whether Undo actually undoes, whether computed values are correct. Looking right can mean static structure passed. Running right means behavior passed.

Static structure beats interaction and computation

Structure passes; interaction trails far behind

Alibaba Token Hub, RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents (arXiv:2609.22000) holds out RecreationBench: 250 tasks, 50 each on Ubuntu, macOS, Windows, Android, and Web. Hidden suites mix Prog (programmatic) and visual assertions; the running reference is the oracle.

One hard cut shows up in outcomes: agents reproduce static interface structure more reliably than interactions and computed outputs. Across 16 model–platform native comparisons, structure has the highest mean pass rate in 15 / 16; the weakest category is always button, computation, or interaction, trailing structure by 9.1–34.4 percentage points. On Web’s different taxonomy, static-content still beats interaction by 28.2–36.7 pp.

Concretely, Ubuntu PhotoCollage: after a rotate and Undo, the recreation restores the arrangement but leaves the photo rotated — the opening screen can look fine while history state is already wrong.

~58% overall; 2.8% full Prog pass

High score ≠ full pass

GPT-6 Astra leads overall at about 58.1% (paper also reports 58.06%), yet passes all programmatic tests on just 2.8% of tasks; every other model ≤ 0.8%. At 90% Prog coverage: Astra 17.6%, Claude Opus 5 5.5%, GLM-5.3 and Qwen3.8-Max-0902 2.0% each. An average is not the probability that an application is fully reconstructed.

Delivered apps stay smaller and more monolithic: 89.4% smaller than reference LOC (median recreation-to-reference ratio 16.9%); 92.3% fewer source files; 83.8% more concentrated in the largest file. The final edit → relaunch → inspect loop completes in only 23.6–47.5% of trajectories (Claude Opus 5 23.6% → Qwen3.8-Max-0902 47.5%). Without that loop, “looks right, runs wrong” stays easy to miss.

Tonight: do not accept the first screen

Before promote, do not only grab a tidy screenshot. Log side by side: full Prog pass rate (not the overall mean), whether interaction / computation assertions pass, whether the final edit was relaunched and inspected. Green first screen with wrong Undo state is a fail. Ask one question tonight: is this agent shipping a visible shell — or a state machine you can click and reverse?

The judgment is simple: looking right is not running right — static structure is not action-conditioned behavior.