b001AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction–claim gap is especially consequential in agentic model discovery, where language-model agents generate and revise predictors using predominantly score-based feedback. We introduce CellAudit, which audits registered input-use claims through three distinct questions: whether the input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed responses. On a paired morphology–transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor achieves a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 yet remains exactly invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key–value attention; the invariance persists after refitting with physically disjoint control wells for inputs and target references. In a stratified audit of 48 generated candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining predictive gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback shows higher mean held-out predictive performance and larger mean compound and dose contributions across five paired trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort further shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not persist. CellAudit therefore adds a falsification layer to agentic model discovery, moving from generate–score–revise toward discover–falsify–revise. Project page: https://limengran98.github.io/CellAudit/.
1. 这段在干什么
这是全文开篇:提出"预测好 ≠ 真用了输入"这一核心 gap,并引出 CellAudit 的三问审计框架和主要实验结果。
2. 需要解释的地方
3. 值得留意
b003AI virtual cells aim to predict cellular responses to chemical and genetic interventions and ultimately support intervention design (Bunne et al., 2024). Yet a model can predict well by exploiting cellular context, control profiles, or other systematic structure while making little or no use of the supplied perturbation information (Ahlmann-Eltze et al., 2025; Viñas Torné et al., 2026). Predictive performance alone therefore does not establish perturbation use (D’Amour et al., 2022; Geirhos et al., 2020; Lapuschkin et al., 2019; DeGrave et al., 2021). We call a mismatch between a registered claim about input use and the model’s implementation or fitted behavior the prediction–claim gap.
这段是 Introduction 的问题提出:先给 AI 虚拟细胞的目标(预测干预反应、辅助设计),再指出一个隐患,最后命名全文核心概念。
关键概念
值得留意
性能指标无法证明用了扰动,这是后文要用代码审计 + 贡献归因去"证伪"的动机;注意它是"对不上"而非"造假"。
b004This gap poses an additional challenge for agentic model discovery (Huang et al., 2024; Li et al., 2024; Jiang et al., 2025; Lu et al., 2026). Systems such as CellScientist (Li et al., 2026), CellForge (Tang et al., 2025), and HarmonyCell (Huang et al., 2026) generate executable predictors, train them, observe validation feedback, and iteratively revise or select candidates. When feedback primarily rewards predictive performance, search can favor high-scoring predictors without testing whether they use the inputs specified by their designs. As generated models become more numerous and diverse, manual verification becomes increasingly difficult. Agentic discovery therefore needs scalable tests of the claims attached to its selected models, not only mechanisms for proposing better predictors.
承接上文提出的「预测–声明缺口」,指出它在 *agentic model discovery* 场景下更难查,从而论证为什么需要可扩展的检验机制,而不只是更好的预测器。
b005We introduce CellAudit, a framework that links registered input-use claims to evidence from source code and fitted models. It asks three non-equivalent questions: can the input enter the cited computation, do fitted predictions depend on it, and does that dependence improve prediction of the observed response? These correspond to source consumption, fitted dependence, and target-relevant predictive contribution. A cited pathway can be mathematically inactive; an executable pathway may have little effect after fitting; and input dependence need not provide predictive benefit. To distinguish these cases, CellAudit combines source inspection with matched input replacement (Fisher et al., 2019; Chamma et al., 2023) on frozen checkpoints, testing both whole-model dependence and target-relevant contribution. During development, the same diagnostics can guide revision; final claims are evaluated only after the selected source, checkpoint, and audit rules are frozen.
这段提出 CellAudit 框架:把"注册的输入用途声明"和源代码、已拟合模型里的证据连起来,作为对上一段"需要可扩展检验"的回应。
关键概念:它问三个不等价的问题——输入能否进入所指计算(源代码层面)、拟合后的预测是否依赖它(依赖层面)、这种依赖是否改善对观测响应的预测(目标贡献层面)。领域常识:三者可以脱节,比如某通路数学上不激活、可执行但拟合后影响很小、依赖却无预测收益。
值得留意:方法是在冻结的 checkpoint 上做 matched input replacement;诊断在开发期可指导修改,但最终声明只在源码、checkpoint、审计规则都冻结后才评估。
b006CellAudit reveals these distinctions in agent-generated cellular-response predictors. On a paired morphology–transcriptomics perturbation task built from BBBC047 (Bray et al., 2017; Haghighi et al., 2022), an agent-selected model reaches a mean held-out Global PCC of 0.3153, compared with 0.3142 for a control-profile-only predictor, while remaining exactly invariant to compound replacement. Source inspection identifies a singleton key–value attention pathway whose query cannot transmit compound information. The model remains invariant after refitting with physically disjoint control wells for inputs and target references. A broader audit separates fitted dependence from evidence of predictive contribution. In a stratified sample of 48 generated candidates across the two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 have registered target-loss scaffold intervals that lie entirely above zero on both folds.
这段在干什么:用具体例子展示 CellAudit 如何区分「模型拟合出的依赖」与「真正有预测贡献的证据」——一个反例加一次更广的审计。
关键概念:
值得留意:关键区分是「换了输入预测就变」≠「该输入有真实预测贡献」。48 个候选里 47 个对化合物替换敏感,但只有 20 个的目标损失区间在两折上都整体大于零——即多数只是拟合了关联,而非稳健贡献。
b007These diagnostics also distinguish what a revision needs to address. Making a perturbation pathway executable does not by itself ensure a target-relevant contribution. On BBBC047, falsification-guided revisions learn a residual over a control-profile-only predictor, recovering positive compound contributions together with predictive gains. In matched sci-Plex searches (Srivatsan et al., 2020), audit-enriched feedback produces higher mean held-out predictive performance and larger mean compound and dose contributions than score-only feedback. However, paired confidence intervals for these differences across five trajectories include zero, so the evidence for improved discovery remains suggestive rather than conclusive.
1. 这段在干什么
承接上段的“诊断结果”,说明诊断也能指出修订该改哪儿:能跑通≠有贡献,并对比两种反馈策略的效果。
2. 需要解释的地方
3. 值得留意
作者自己留了退路:BBBC047 上看似有增益,但 sci-Plex 上跨五条轨迹的配对置信区间含零,结论只是“suggestive rather than conclusive”——这是领域常识层面的谨慎,不是论文额外下的定论。
b008Finally, we distinguish predictive generalization from claim generalization. The latter asks whether a registered input-use claim remains supported when a fixed model design is refit on an independently acquired cohort. When LINCS-selected designs are fixed and refit on LKCP (Keenan et al., 2018; Subramanian et al., 2017; Weisbart et al., 2024), predictive gains and dose contribution persist, whereas compound-identity contribution changes substantially. A separate CRISPRa evaluation applies the same auditing principle to unseen genetic combinations (Norman et al., 2019). Together, these studies trace input-use claims from source code through fitted dependence to predictive contribution, use the resulting evidence to guide revision, and re-evaluate those claims on new data.
1. 这段在干什么
作为引言的收尾,把讨论从「预测泛化」推进到更难的「论断泛化」:换一套独立数据重训后,原来的输入使用论断还站得住吗?并总结全文的研究路线。
2. 需要解释的地方
3. 值得留意
关键不对称——预测增益和剂量贡献能保持,但化合物身份贡献变化很大。这说明「有用」和「靠什么有用」是两回事,正是审计要抓的点。
b010Cellular perturbation models use latent transfer, compositional representations, neural transport, genetic graphs, and foundation models to predict chemical or genetic responses (Lotfollahi et al., 2019; Lotfollahi et al., 2023; Bunne et al., 2023; Roohani et al., 2024; Cui et al., 2024). Benchmarks standardize evaluation across splits and baselines (Wu et al., 2025; Wenteler et al., 2025), while recent studies examine linear references, shared response variation, calibrated metrics, and distribution shifts (Ahlmann-Eltze et al., 2025; Viñas Torné et al., 2026; Miller et al., 2025; Mao et al., 2026). These works ask how well perturbation responses can be predicted and evaluated; CellAudit asks whether registered input-use claims are supported by a model’s source code and fitted behavior.
1. 这段在干什么
交代相关工作的版图,最后一句划清界限:别人问「预测得准不准」,CellAudit 问「模型声称用了的输入,代码和拟合行为真的支持吗」。
2. 需要解释的地方
3. 值得留意
前半句全在罗列方法和基准,读起来像文献堆砌,但它是为最后一句的反差做铺垫——前人的关注点是「预测性能」,本文的关注点是「输入使用声明是否属实」,这个对照才是这段的真正落点。另外,代码与拟合行为是两个独立证据源,缺一不可。
b011Scientific agents generate and revise executable hypotheses, programs, and models (Boiko et al., 2023; Romera-Paredes et al., 2024; Jiang et al., 2025). Permutation reliance, conditional replacement, and behavioral testing probe fitted-model behavior under targeted input interventions (Fisher et al., 2019; Chamma et al., 2023; Ribeiro et al., 2020), while ConceptSMILE studies explanation reliability and POPPER automates hypothesis falsification (Mollapour et al., 2026; Huang et al., 2025). Unlike reliance tests that begin from prediction behavior, CellAudit starts with a registered input-use claim and its cited computation, then follows it through source consumption, fitted dependence, and target-relevant predictive contribution. This evidence can also guide model revision. Appendix K provides additional context.
这段在干什么:摆在相关工作里做定位对比——先列出科学智能体和依赖测试两条线,再一句"Unlike…"标出 CellAudit 的差异:从注册的输入使用声明出发,而非从预测行为出发。
需要解释的地方:
值得留意:CellAudit 走的是"声明→引用代码→源码消费→拟合依赖→预测贡献"这条链,且证据还能反过来指导模型修改;末句把细节推给附录 K。
b013Each example contains biological context cc, perturbation representation pp, optional attributes aa such as dose, and response y=[y(1),…,y(M)]y=[y^{(1)},\ldots,y^{(M)}]. Source ss and fitted parameters θ\theta define fs(c,p,a,θ)f_{s}(c,p,a;\theta); outputs may include imaging, molecular, or other readouts. Before search, the task registers input coordinates and admissible interventions. Each audit records whether an input-use claim comes from a candidate or the task contract and, after selection, fixes its cited computation and test mapping. A prediction–claim gap occurs when the cited source contradicts a registered claim or fitted predictions are invariant to input replacement. Predictive performance affects the importance of the case, not the definition; target-relevant contribution is separate.
定义后续审计的形式化框架:每个样本的输入/输出是什么、模型由谁定义、claim 从哪来、什么算“预测—claim 差距”。
作者划了两条线:性能只影响案例权重,不影响 gap 的定义;“目标相关贡献”是另一回事,别混为一谈。
b014Data processing, budgets, evaluation, and selection rules are fixed before search. Folds 1–2 fit candidates; Fold 3 provides feedback and selects checkpoints. Selected checkpoints are evaluated on Folds 4 and 5, except prediction-score discovery also selects a source across policies on Fold 4 and refits it for Fold 5. Norman instead selects sources before Fold 4 and refits the same sources on Fold 5 without reselection. Appendices B and N.1 specify study-specific refitting and reuse rules.
这段交代评估协议:数据处理、预算、评估与选择规则在搜索前就固定,属于防数据泄漏和防调参作弊的常规设计(领域常识)。
留意:这段是"规则先定死",与上段"性能不影响定义"呼应——性能只动案例重要性,不动定义。
b017The existing CellScientist policy (Li et al., 2026) starts from a common predictor h0h_{0}. It receives the current model, development diagnostics, execution status, and compact history, and proposes a hypothesis and executable revision within the registered permissions. Shared preflight and training code provide up to three repairs per failed proposal; persistent failures consume a slot. Each executable candidate trains once, and a fixed Fold-3 rule selects the endpoint. Records preserve parent links, emitted hypotheses and declarations, source, diagnostics, repairs, checkpoint, and costs. The language model proposes revisions; input-use tests are deterministic.
1. 这段在干什么
描述 CellScientist 这个 agent 搜索策略的具体运行机制:给定固定数据和评估,它如何从初始模型出发提出并执行修改。
2. 需要解释的地方
3. 值得留意
分工很关键——语言模型只负责"提修改",确定性测试负责"验输入使用"。这一句点出了全文核心:把"发现、证伪、修正"的验证环节与生成环节分开,避免全靠 LLM 判断。
b019Given registered inputs and cited source locations, CellAudit checks the computation attached to each input-use statement. Its abstract-syntax-tree checker returns implemented, implementation-contradicted, or unresolved. Coverage is the fraction of checked candidates whose cited locations are all resolved; unresolved code remains eligible for behavioral testing. With one key–value pair, attention has constant normalized weight and cannot transmit compound information through its query (Vaswani et al., 2017). This localizes a recognized implementation defect. Whole-model replacement then establishes whether fitted predictions depend on the input, including through other routes; target-loss contrasts test whether that dependence helps prediction.
讲 CellAudit 怎么先静态检查源码里的输入使用声明,再对没查清的做行为测试。
b021With checkpoint and targets fixed, CellAudit replaces one registered input xqx_{q} while retaining the others. Prediction distance measures target-blind dependence; zero identifies invariance on the tested replacements. To test predictive benefit, it computes the decrease in a higher-is-better score mm and the increase in loss:
这段在交代评测协议:固定其余输入,只替换一个已注册输入,看预测怎么变。是方法过渡到结果的一步。
解释
留意
b022Positive effects indicate benefit from the observed input relative to its registered replacements. Replacement distributions and their observed-support coverage are specified by task. We distinguish single-coordinate contrasts Δmq\Delta m_{q} from factorial effects EqE_{q}, which average over the other input’s two states in the 2×22\times 2 compound–context audit. Loss effects EqℒE^{\mathcal{L}}_{q} use the same averaging. The transfer analysis separates compound and dose allocations (Lundberg and Lee, 2017). Appendix L gives the full contrasts.
这段在干什么:承接上文「计算分数下降与损失上升」,定义怎么比对:把观测输入和它的「注册替代品」比,正效应即观测输入更有益。
需要解释的地方:
值得留意:两种对比(单坐标 vs 因子平均)不是一回事,别混用;具体替换分布和覆盖范围「由任务指定」,本文没给,细节都在附录 L。
b023Maps are repeated measurements within a checkpoint, averaged before across-model inference. Continuous contributions have paired training-seed or trajectory intervals; fixed-model source-group resampling tests variation across biological groups. Known-truth checks distinguish fixed-dataset and population uncertainty (Appendix F.1). BBBC, LINCS, and LKCP separately report whether the effect quantile interval lies above, below, or across a reference threshold. A positive mean contribution can occur in any category. Appendix B defines the threshold, quantiles, and original category labels; Norman uses a separate loss-based criterion (Appendix D).
好,我们看这段。它紧接上一段“分离化合物与剂量分配”,开始讲怎么做统计检验。
1. 这段在干什么:交代本节的检验方法——先平均重复测量,再做跨模型推断,并逐项说明各数据集的报告方式。
2. 需要解释的地方:
3. 值得留意:作者特意强调“正的平均贡献可以出现在任何一类”——别把正负号直接当成“高于阈值”。阈值、分位数、类别标签的定义在附录 B,Norman 用的是另一套基于损失的判据(附录 D),这段只点名字、没展开。
b025The audit distinguishes two reasons to revise a model. A contradicted source path motivates an executable route; weak target contribution motivates learning an increment over context. Path-constrained discovery addresses the first by compiling structured design cards into models with an explicit compound route. BBBC falsification-guided discovery addresses the second through residual modeling and input-use feedback:
1. 这段在干什么
承上启下:把"修订模型"分成两类原因,并预告两种对应方法。
2. 需要解释的地方
3. 值得留意
两种"修订理由"与两种方法是一一对应的(first/second),别读成两件独立的事。
b026where g(c)g(c) is fitted on Folds 1–2 and fixed. The compiler enforces centered perturbation/dose branches, context gates, and bias-free readouts, giving zero residual at their joint reference input. Training and Fold-3 feedback use prediction error and replacement loss changes. BBBC047 selects by predictive score among qualifying candidates; BBBC036 prioritizes compound loss gain among predictively noninferior candidates. Each trajectory evaluates the anchor plus nine proposals; the harness selects scale α\alpha and checkpoint on Fold 3 and freezes that model for Folds 4 and 5. These formulations test the joint design changes; Appendix B specifies their rules.
交代 BBBC 两条线(047/036)的具体选择规则与训练/冻结流程,说明假证据如何落到下一轮搜索。
两条线排序准则不同:047 看预测分,036 在"预测不劣"前提下看复合损失增益——作者未细说阈值,规则在 Appendix B。
b027The sci-Plex comparison instead shares a context-plus-dose anchor g(c,a)g(c,a) and permissions for direct predictors or anchor residuals. Score feedback supplies metrics and learning curves; Audit adds grouped errors, source checks, and development-set compound/dose replacements. After each fit, CellAudit computes these diagnostics and inserts them into the agent’s next prompt, linking input-path findings and measured contributions to model revision. Loss, optimizer, training limits, and PCC-first selection stay fixed. Paired outcomes evaluate the complete audit-enriched feedback package; selected sources and checkpoints are frozen before both held-out evaluations.
这段在干什么:说明 sci-Plex 对比实验里用什么反馈驱动“下一轮模型修正”——反馈分 Score 和 Audit 两层,并交代评估怎么保持公平。
需要解释的地方:
值得留意:Loss、优化器、训练上限、PCC-first 选择全部固定,来源与 checkpoint 在两次留出评估前冻结——作者刻意只改反馈,排除其他变量干扰。
b030Development uses linked tasks from one paired cellular-profile release: BBBC036, the known-bioactive subset, and its parent BBBC047 (Haghighi et al., 2022; Bray et al., 2017; Subramanian et al., 2017). Each predicts joint morphology–transcription responses from a same-plate control profile, compound representation, and dose. Response-blind Bemis–Murcko scaffold folds (Bemis and Murcko, 1996) hold out compound families. Global PCC on concatenated responses selects models; MSE and modality-specific metrics diagnose gains. Appendix A gives provenance, transformations, fold inventories, and a complementary plate-family split.
交代实验的数据与评测设计:用同一批配对细胞数据里的两个任务,规定输入、划分方式、选模指标,为下一段结果做铺垫。
任务标注为"response-blind"——划分时不看响应值;控制来自同板;细节都推给附录 A。
b031Prediction-score discovery compares CellScientist, AIDE, CellForge, and HarmonyCell policies (Li et al., 2026; Jiang et al., 2025; Tang et al., 2025; Huang et al., 2026) under a matched initial model, executor, feedback, and ten-candidate budget, with ten trajectories per task and five paired refits per selected model. Both follow-up formulations retain the folds and budget and carry all ten selected designs through Folds 4 and 5; failures, costs, and repeated designs are retained (Appendix C).
1. 这段在干什么
说明"预测分数发现"实验的对照设置:四个方法在相同初始模型、执行器、反馈和十候选预算下比较,并交代重复次数与后续两种变体的延续方式。
2. 需要解释的地方
3. 值得留意
"匹配"(matched)的是初始模型、执行器、反馈和预算——公平比较的前提;但十轨迹、五配对 refit 属于领域常见设定,具体理由这段未提。
b032For BBBC047, each plate’s control wells are split into equal-size banks A/B. With the target reference fixed, shared and disjoint arms take the input profile from the same or the other bank, so their contrast isolates physical control-well overlap. Both orientations refit the score-selected source, a control-only source, and one fixed guided design under five seeds; all 60 checkpoints are fixed before evaluation on Folds 4 and 5 (Appendix A.1).
这段在干什么:交代 BBBC047 的对照设计——用同/异控制孔库(shared/disjoint arms)来隔离控制孔重叠带来的混杂,并说明评估前的固定流程。
需要解释的地方:
值得留意:作者特意强调"所有 60 个 checkpoint 在评估前就冻结",这是在防"事后调整"的嫌疑。
b033We predict 2,000-gene well-level responses in sci-Plex3 (Srivatsan et al., 2020), holding out compound–cell-line combinations, with whole treatment wells in one partition, while both marginals occur in training; gene selection and scaling use training wells only. Both conditions use the same DeepSeek serving model, five paired trajectory seeds, five new candidate slots per trajectory, and a shared anchor. All ten selected checkpoints are evaluated on two held-out combination partitions of this source. Compound replacements match dose and experimental context; dose replacements retain compound identity (Appendix N.1).
交代所选 checkpoint 的评测数据集与评测设置:在 sci-Plex3 的留出组合划分上做评测。
所有 10 个 checkpoint 在评测前已固定;化合物替换与剂量替换的对照条件不同。
b035The fixed falsification-guided protocol first runs on LINCS–Pilot1 (Keenan et al., 2018; Subramanian et al., 2017). Its ten selected designs and sources are then frozen and refit on independently acquired LKCP Batch 2 (Weisbart et al., 2024), with response- and fold-blind interface adaptation fixed before training. Compound permutations retain dose; dose permutations stay within compound (Appendix E). The Norman CRISPRa task predicts 5,045-gene responses for unseen double-perturbation combinations (Norman et al., 2019). Its registered input bundles multi-hot pair identity with the matching mean single-perturbation anchor. GEARS supplies a task-native reference (Roohani et al., 2024).
1. 这段在干什么
交代实验怎么落地:先在 LINCS–Pilot1 上跑固定协议、冻结设计,再原样搬到 LKCP Batch 2 上重训验证;同时说明两个扰动任务(Norman、GEARS)的设置和参照。
2. 需要解释的地方
3. 值得留意
b038CellAudit’s automated audit identifies the BBBC047 source selected on Fold 4 as compound-invariant. Its source-selection score is 0.30350.3035; mean Global PCC reaches 0.31530.3153 on held-out replication Fold 5. Compound shuffling changes PCC by 0.00000.0000 in every refit, whereas control replacement lowers PCC by 0.19320.1932 on Fold 4 and 0.20040.2004 on Fold 5 (Table 1). Control-only reaches 0.31420.3142 on Fold 5; Full−-control-only is +0.0011+0.0011 (95% CI [−0.0005,+0.0027][-0.0005,+0.0027]). Prediction-space distance is exactly zero under compound replacement and large under control replacement (Appendix I).
这段在干什么:用一组具体数字,给出 CellAudit 对高分预测器的一条审计结论——该模型是「化合物不变的」。
需要解释的地方:
值得留意:打乱化合物 PCC 变化恰为 0、预测空间距离恰为 0,等于模型完全没依赖化合物;真正起作用的是对照。
b039The source checker identifies a concrete computation to revise: the cited compound-conditioned cross-attention has one key and one value, making its normalized weight constant. The behavioral test establishes complete-model invariance; the source result localizes an inactive cited pathway. A compound-only positive control produces a shuffle effect of 0.0111±0.00100.0111\pm 0.0010, with all five refits above the reference threshold.
这段在干什么:承上给出定位结果——源检查指出该被引通路存在一处具体可修正的计算,行为测试则确认整个模型对该化合物替换完全不变。
需要解释的地方:cross-attention 的 key/value(领域常识:注意力里用来算权重的两路向量);只有一个 key 和一个 value,归一化权重就成了常数——即这一路对化合物条件实际不生效;阳性对照是验证实验本身有效的参照;shuffle effect 指打乱后造成的效应量。
值得留意:源检查与行为测试结论一致,一个定位、一个确证;0.0111 的效应量连同五次重拟合全部过阈值,说明阳性对照确实生效,反衬出化合物条件那一路是真的没起作用。
b040A source-stratified sample of 48 real candidates separates these outcomes further (Figure 2a,b). An implemented multi-tower source is exactly invariant; 47 candidates show fitted dependence on both folds, but only 20 have positive target-loss scaffold intervals on both. Source consumption and fitted dependence thus leave predictive contribution unresolved for many candidates (Appendix M.3). CellScientist supplies a competitive discovery setting: mean Best@10 and frontier AUC lead both linked tasks, with intervals versus HarmonyCell crossing zero (Table 19).
这段在干什么:用一批真实候选样本,进一步验证“源码里用了”和“拟合上依赖”都不能直接等同于“真有预测贡献”。
需要解释的地方:
值得留意:48个里只有20个两折都正向——说明“被用上”“拟合依赖”与“实际贡献”之间存在大缺口。
b041Control: control-only predictor. Reference: effect quantile interval above/below/across the reference threshold (Appendix B); distinct from a positive mean contribution. aOne-key attention cannot transmit compound identity. bCompiler-enforced compound path. cExplicit compound-residual path.
这是表格的脚注说明:定义「Control」「Reference」两个对照基线,并解释表内 a/b/c 三个角标各自的含义。
「distinct from a positive mean contribution」是在提醒:落在区间外 ≠ 有正向平均贡献,两者别混为一谈。这段不是正文论述,读表时容易直接跳过。
b043We test whether the mismatch persists when control inputs and CP target references use disjoint physical wells. Newly reconstructed shared and disjoint arms have identical targets within each bank orientation (Figure 2c,d). In the disjoint arm, the score-selected source reaches Global PCC 0.30850.3085 and 0.31930.3193, while control-only reaches 0.30820.3082 and 0.31850.3185 on Folds 4 and 5. The score-selected source remains exactly compound-invariant in every refit and evaluation. All six shared–disjoint PCC comparisons have five-seed intervals crossing zero. The detected mismatch therefore persists without physical control-well reuse at this reference-construction layer.
这段在干什么:它检验前面的错配是不是由「对照组与目标组共用同一批物理孔」造成的——换成互不重叠的孔后,错配依然存在。
需要解释的地方:
值得留意:两组数值极接近(0.3085 vs 0.3082、0.3193 vs 0.3185),且「五种子区间均跨零」说明差异不显著——作者据此强调错配不是孔复用造成的假象。
b044The fixed guided design retains both predictive and compound-use increments under the same separation (Table 6). Relative to disjoint control-only, its PCC gain is +0.0028+0.0028 (95%95\% paired-seed CI [+0.0017,+0.0038][+0.0017,+0.0038]) on Fold 4 and +0.0025+0.0025 ([+0.0012,+0.0037][+0.0012,+0.0037]) on Fold 5. Compound replacement lowers PCC by 0.00430.0043 and 0.00330.0033 and increases target loss by 1.5106×10−41.5106\times 10^{-4} and 1.1939×10−41.1939\times 10^{-4}; each mean effect has a positive five-seed interval. Effects are slightly smaller than in the shared arm. Same-plate context and the original scaffold folds are retained (Appendix A.1).
这段在干什么:验证固定引导设计在分离对照下依然保持预测增益与复合使用增益,回应上段"错配持续存在"的结论。
需要解释的地方:
值得留意:所有区间都为正,效果虽比 shared 组小,但方向一致;同板上下文与原始 scaffold 折被保留,说明结论不是在无控制条件下取得的。
b046Path-constrained discovery addresses the source defect by compiling an explicit compound route: all 200 candidate evaluations pass source and interface checks without candidate-code repair. On BBBC047, mean factorial compound effects are 0.00830.0083 and 0.00740.0074, with positive mean loss effects on Folds 4 and 5. Relative to the reference threshold, Fold 4 has six below-threshold and four overlapping cases; Fold 5 has seven and three. Control-profile effects exceed the threshold throughout (Table 1). An executable route therefore admits compound dependence without ensuring a consistently above-reference contribution in each model.
它检验「显式复合路径」这个修复方案:200 个候选全部通过源码与接口检查,说明源码缺陷被绕过了,但拟合出的复合效应在 Fold 4、5 上只是「部分」超过参考阈值。
b048Falsification-guided discovery next asks the perturbation residual to improve prediction beyond a fixed control-profile-only model. On BBBC047, the full predictor exceeds that baseline by +0.0037+0.0037 on Fold 4 and +0.0028+0.0028 on Fold 5. Mean factorial compound effects are 0.00530.0053 and 0.00480.0048, with positive loss effects (Table 1); joint model–scaffold intervals for these increments are above zero. The explicit-route models have larger compound effects but lower predictive scores. Guided discovery combines higher scores with positive target-relevant effects, although none of its ten models exceeds the reference-relative criterion.
论证「假说引导的残差发现」在 BBBC047 上既提升了预测分数(超过仅用对照的基线),又带来正向的目标相关化合物效应——但十模型都未过参考相对标准。
分数增益(+0.0037/+0.0028)和效应量(0.0053/0.0048)是不同量纲的两类指标,别当成一回事。另外"joint intervals above zero"只说明这两个增量不为零,并不等于超过参考标准——末句正是这个转折。
b049For the original guided checkpoints, observed-support replacements retain positive contributions on BBBC047 while holding context and dose fixed. They cover 3341/43803341/4380 and 2927/38782927/3878 rows on Folds 4 and 5: compound PCC drops are 0.005340.00534 and 0.004040.00404, and loss gains are 1.84×10−41.84\times 10^{-4} and 1.44×10−41.44\times 10^{-4}, all with positive source-scaffold intervals. The score-selected model remains exactly invariant on these same eligible rows (Table 57).
这段在干什么:承接上段,报告引导式检查点在观测支持替换下的稳健性——贡献为正、性能只小幅下降。
需要解释的地方:
值得留意:合格行占比极小(3341/4380、2927/3878),贡献结论只覆盖少数行,作者没明说其代表性有限。
b050Fixed residual learners clarify the trade-off: Ridge has strong compound sensitivity but low joint-response scores, and on BBBC047 the residual MLP has larger compound effects but lower scores than the guided models, which lead both references in mean Global PCC at all four BBBC evaluations. On BBBC036 Fold 5, their gain over the context anchor is −0.0020-0.0020, with model–scaffold intervals for predictive and compound increments crossing zero (Table 16).
这段拿固定残差学习器做对照。作用:论证 guided 模型更好——它在四个 BBBC 评估上平均 Global PCC 都领先两个参照物。
关键概念:残差学习器=在已有预测基础上只学"补差"的模型(领域常识);Global PCC 是整体预测相关性指标;模型–scaffold 区间是评估波动范围,跨过零意味着效应不稳健。
留意:BBBC036 Fold 5 上它们只比上下文锚点提升 −0.0020,且两类增量区间都跨零——作者用这个"没提升"的例子反衬 guided 的增益。
b051Known-function experiments separate invariance, sensitivity without benefit, benefit, and harm across 6,400 independent simulated datasets. On the sensitive-without-benefit null, map-only intervals attain near-nominal coverage for the fixed-dataset conditional effect but severely undercover the population effect. Source-group intervals improve population coverage, with finite-group undercoverage remaining. Holding compound benefit fixed while changing reference-context variation can also change qualification. These checks support reporting continuous contribution, across-model uncertainty, and reference-relative status as distinct results (Appendix F.1).
这段在干什么:总结"已知函数"模拟实验的四种判别情形,并据此提出报告建议——贡献度、跨模型不确定性和参照相对状态应作为三种独立结果分别汇报。
需要解释的地方:
值得留意:benefit 固定不变时,仅改动参照上下文的变化就能改变"是否合格"的判定,说明结论对基准选择高度敏感;6,400 这个规模也提示结论来自大量独立模拟,而非单次实验。
b053Do audit measurements help when returned to an agent? Figure 3a retains all five paired sci-Plex search frontiers. The context-plus-dose anchor already predicts much of the whole-expression profile, making incremental error reduction informative. Audit feedback yields higher mean PCC and lower mean MSE on both later partitions (Figure 3b): MSE is 9.58%9.58\% lower than Score on Fold 4 and 6.10%6.10\% lower on Fold 5. Both improve on the anchor.
1. 这段在干什么
论证把审计测量结果反馈给 agent 后,预测效果确实变好了——用 Figure 3 的 PCC、MSE 数字支撑这一点。
2. 需要解释的地方
3. 值得留意
b054The higher mean predictive performance of Audit is accompanied by higher mean compound and dose contributions (Figure 3c). Compound target-loss gain rises from 0.01860.0186 to 0.02200.0220 on Fold 4 and from 0.00640.0064 to 0.00790.0079 on Fold 5. Paired tt intervals for these Audit–Score contribution contrasts include zero at both folds (Table 68). Correct inputs help prediction in all five endpoints of both conditions. Audit–Score PCC is +0.0024+0.0024 (paired 95%95\% CI [−0.0013,+0.0060][-0.0013,+0.0060]) on Fold 4 and +0.0015+0.0015 ([−0.0005,+0.0034][-0.0005,+0.0034]) on Fold 5. Both means, and the favorable MSE and compound loss-gain directions, persist after removing any one trajectory pair (Figure 3d), whereas prediction-RMS differences can reverse after one omission. Percentile-bootstrap predictive intervals exclude zero even when resampling only trajectories, so uncertainty depends on interval construction at five pairs (Appendix N.2). Both conditions complete all 25 new candidates; Audit uses 38 versus 36 logical model calls and 592,347 versus at least 286,995 reported tokens (Table 65).
这段在干什么:在上一段报告主预测误差(MSE)更低的基线上,这段补上贡献度与相关性证据,说明 Audit 的优势不只在单一指标,而是多角度一致。
需要解释的地方:
值得留意:贡献增益方向一致,但 t 区间跨零、PCC 区间也跨零——作者没宣称显著。真正排除零的是百分位自助区间,而这是否成立取决于区间构造方式(仅 5 对)。
b055In the first registered Audit trajectory, CellAudit returns source checks, compound/dose replacement effects, grouped errors, and learning curves for the bilinear incumbent and a non-improving challenger. The agent returns to that incumbent and proposes simpler encoders with dropout, retaining dose-conditioned FiLM. Parameters decrease while selection PCC and compound loss gain increase (Figure 1D); search continues to a distinct final endpoint (Appendix O).
这段在干什么:描述第一次登记审计的实际过程——审计工具给出诊断,agent 据此回到原模型并改出更简单的编码器,得到更好的结果。
需要解释的地方:
值得留意:challenger 是"non-improving"——它没赢,agent 因此退回原模型继续改,而不是接受挑战者。这是"审计→修正"闭环的关键一步,但原文没点明。
b057To test claim generalization, the frozen LINCS-selected designs are refit on LKCP without source revision, model reselection, or audit-rule changes. They improve over the control-profile-only baseline by +0.0104+0.0104 (95% CI [+0.0103,+0.0106][+0.0103,+0.0106]) on Fold 4 and +0.0211+0.0211 (95% CI [+0.0205,+0.0217][+0.0205,+0.0217]) on Fold 5. Figure 4a,b separates compound identity from dose. On Folds 4 and 5, the dose-replacement Shapley allocation is +0.0486+0.0486 and +0.0467+0.0467 Global PCC on LINCS and +0.0237+0.0237 and +0.0341+0.0341 on LKCP, whereas the compound-identity allocation is +0.0175+0.0175 and +0.0126+0.0126 on LINCS but −0.000070-0.000070 and +0.0033+0.0033 on LKCP. All 50 trajectory-endpoint refits pass the dose-use criterion at every cohort–fold boundary, but none passes the compound-identity criterion at either LKCP boundary. LINCS shows why average effects and per-model decisions must remain separate: its compound-identity criterion is passed by 35/50 refits on Fold 4 but 10/50 on Fold 5 (mean lower effect quantile 0.01390.0139 to 0.00900.0090), although average compound contribution remains positive.
测试前文冻结的设计在独立数据集 LKCP 上能否复现:预测增益能泛化,但“输入用途”是否成立要分开看。
平均效应为正,不代表每个模型都通过判据:LINCS 上通过化合物身份判据的重训模型从 35/50 掉到 10/50,作者明确提醒二者要分开看。
b059On Norman, all five CellScientist refits pass the registered pair-input-bundle criterion on both held-out folds; shuffling the multi-hot pair identity together with its matching single-perturbation anchor reduces PCC by 0.55780.5578 and 0.62800.6280. Mean PCC exceeds h0h_{0} by +0.0060+0.0060 and +0.0045+0.0045 and the observed-single additive baseline by +0.0209+0.0209 and +0.0114+0.0114 (Table 26; Appendix D).
这段在给「pair bundle 审计」下结论:在 Norman 数据上,五个 CellScientist 重拟合都通过了成对输入捆绑的检验。
关键概念:multi-hot pair identity 指把「哪一对扰动」编码成向量;shuffle 是把它和对应的单扰动锚点一起打乱,看性能掉多少——掉得越多,说明模型越依赖这对输入。PCC 是预测与真实的相关系数。h₀ 是零假设基线,这里指随机打乱后的水平。
值得留意:打乱后 PCC 掉 0.56 和 0.63,是很大的跌幅,说明模型确实靠这对身份信息;而它相对加性基线只高 0.006 和 0.0045,增益其实很小。这两组数字一对比,才是这节真正的张力。
b061CellAudit shows that predictive success and claimed input use should be evaluated separately. Source consumption, fitted dependence, and target-relevant contribution can diverge: on BBBC047, a high-scoring source remains compound-invariant after disjoint control-reference refitting, while falsification-guided revisions recover positive compound contributions with predictive gains. The broader candidate audit shows this is not specific to singleton attention. In sci-Plex, audit-enriched feedback improves mean prediction and input contributions, but paired intervals over five trajectories span zero. Independent-acquisition refits further show that predictive generalization need not imply claim generalization.
这段在干什么:讨论段收尾——总结本文核心论点:预测好≠用了所声称的输入,并用多个数据集佐证这不是个例。
需要解释的地方:
值得留意:作者主动承认 sci-Plex 的配对区间跨零(即提升不显著),这是诚实而非示弱;末句「预测泛化不蕴含主张泛化」是本段最重的结论。
b062These conclusions are conditional on the fitted model, evaluation setting, and registered replacement distribution. Positive contribution is distinct from passing the reference-relative criterion, and support for one fit should not be inherited by later refits. Observed-support replacements avoid unobserved combinations but do not establish biological exchangeability, mechanism, or causal effects. The sci-Plex study evaluates the complete audit-feedback package rather than individual diagnostics. CellAudit therefore adds a falsification layer to agentic model discovery, tracing input-use claims from cited computation to fitted behavior and predictive contribution.
收束全文:把前面结论限定条件摆明,定位 CellAudit 的贡献是给智能体模型发现加了一层"证伪"。
b065BBBC036 is the known-bioactive subset of its parent BBBC047 screen; the two tasks are therefore linked cohorts from one paired cellular-profile release rather than independent acquisitions. Their source collections in cpg0003-rosetta are CDRPBIO-BBBC036-Bray and CDRP-BBBC047-Bray, and the abbreviated task names retain the formal BBBC accession mapping (Ljosa et al., 2012). Inputs come from released replicate-level Cell Painting and L1000 profile tables (Weisbart et al., 2024; Haghighi et al., 2022). The morphology screen originates from the CDRP/BBBC047 Cell Painting resource (Bray et al., 2017), and the transcriptional profiles use the L1000 platform (Subramanian et al., 2017). Both assays apply matched compound–dose conditions on parallel plates and are joined at the condition level.
1. 这段在干什么
交代两个数据集 BBBC036 和 BBBC047 的来龙去脉,说明它们是同源配对,不是各自独立采集的。
2. 需要解释的地方
3. 值得留意
两个任务共用同一次细胞画像发布,属于"linking cohorts"。这一点很关键——后续审计输入使用声明时,"任务相关"可能来自同源而非独立,容易被忽略。
b066Table 2 reports both the released profile rows and the matched samples used for modeling. BBBC036 begins with 21,122 CP and 6,929 L1000 replicate profiles and retains 1,916 matched compound–dose conditions; BBBC047 begins with 153,386 CP and 68,120 L1000 profiles and retains 20,081 matched conditions. The matched subsets span 47 CP/22 L1000 treatment plates for BBBC036 and 273 CP/360 L1000 treatment plates for BBBC047. Before aggregation, 143 BBBC047 L1000 treatment profiles without a same-plate control are removed.
这段交代两个数据集从原始 profile 到建模样本的筛选过程。
关键概念:CP 和 L1000 是两种不同的检测平台,同一 compound–dose 条件在两平台上各测一遍,所以能按 condition 配对;treatment plate 是实验批次单位。
值得留意:CP 与 L1000 的样本量差距悬殊(如 BBBC047 15 万 vs 6.8 万),最终保留的 matched conditions 远少于任一平台原始数;末尾一句说明 L1000 里缺同板对照的 profile 会被剔除,但只提了 BBBC047。
b067For each condition, cc is the featurewise median of same-plate control-only CP profiles. Compound identity pp is represented by a 2,048-bit radius-2 Morgan fingerprint (Rogers and Hahn, 2010), while the separate attribute aa is the log-transformed observed dose. Before treatment filtering and condition aggregation, numeric profile columns are converted to float32 and any column containing a non-finite value in the released profile table is removed. CP treatment responses undergo the empirical same-plate-control mid-CDF transform, are centered by subtracting 0.50.5, and are aggregated by condition median; L1000 responses are centered by the same-plate control median and likewise aggregated by condition median. The resulting joint targets contain 591 CP and 977 L1000 outputs for BBBC036, and 701 CP and 977 L1000 outputs for BBBC047, yielding 1,568 and 1,678 dimensions. For training, each target coordinate is centered and scaled using its Folds 1–2 mean and population standard deviation; scales at or below 10−810^{-8} are replaced by one. Reported metrics invert this standardization, retaining the control-transformed response scale. Bemis–Murcko groups (Bemis and Murcko, 1996) are assigned without response values and remain disjoint across all five folds. Table 3 lists the complete fold inventory and experimental roles.
这段在干什么:交代数据预处理流水线——控制基线怎么定、化合物怎么表示、响应怎么变换聚合,最后给出两套数据集的目标维度。
需要解释的地方:
值得留意:标准化参数只用 Folds 1–2 估计,且报告指标时被反变换回控制变换后的尺度;骨架分组"无响应值"且跨五折不相交,是为防泄漏。
b068For the complementary plate-family/context transfer, connected components are formed from released CP treatment-plate identifiers, their same-plate control contexts, linked L1000 treatment plates, and matched-condition counts, without reading responses. The construction yields six components for BBBC036 and 64 for BBBC047, assigning every CP/context and linked L1000 plate to exactly one role. Fit/selection/audit/replication condition counts are 640/320/637/319 for BBBC036 and 8,264/3,868/3,826/4,123 for BBBC047. BBBC036 therefore provides a six-family transfer case, while BBBC047 supplies broader 64-family coverage.
交代“互补的板系/情境迁移”数据是怎么组装的,并给出两个数据集划分出的组件数与各角色条件数。
作者特别强调“without reading responses”,这是防止信息泄漏的关键措辞,容易读漏。
b069Cohort CP source profiles L1000 source profiles A. Released source profiles BBBC036 21,122 (17,594/3,528) 6,929 (3,451/3,478) BBBC047 153,386 (126,814/26,572) 68,120 (64,642/3,478)
这是附录里的一张数据表片段,列出 CP 与 L1000 两种来源画像下、BBBC036 与 BBBC047 两个队列的样本量。
BBBC047 的 L1000 分母两个数都是 3,478,与上一段提到的家族覆盖数(六族 vs 64 族)呼应,但具体对应关系原文未点明。
b070Parentheses: treatment/control profile rows.
这段在干什么:这是表格的脚注说明——括号里的数字表示 treatment(处理组)/ control(对照组)两类的样本行数。
需要解释的地方:「profile」在这里指表达谱样本;「treatment/control」是实验分组,即受药物等干预的组 vs 未干预的对照组。这是生物实验的领域常识,具体定义这段没展开。
值得留意:上一段那些数字(如 17,594/3,528)加起来正好等于括号外的总数,读表时可自行核对;两个数据集 BBBC036 和 BBBC047 的 control 数竟然都是 3,478,是巧合还是共用对照,这段没说。
b071Cohort Matched conditions Unique SMILES Murcko groups CP/context dim. L1000 dim. Joint target dim. B. Matched samples and model dimensions BBBC036 1,916 1,916 1,196 591 977 1,568 BBBC047 20,081 20,068 4,540 701 977 1,678
这是附录里的数据统计表:列出两个数据集(BBBC036、BBBC047)的样本配对情况、分子多样性指标和各模块的维度数。起补充数据支撑作用。
BBBC047 中 Matched conditions(20,081)比 Unique SMILES(20,068)略多——配对后可复用片段,这与上段"括号内为处理/对照行"的说明呼应。原文只给数字,未解释差异。
b072Statistic Fold 1 Fold 2 Fold 3 Fold 4 Fold 5 BBBC036 Matched conditions 285 509 356 411 355 Murcko groups 206 244 258 265 223 CP replicate profiles 2,222 3,995 2,779 3,213 2,803 L1000 replicate profiles 519 919 626 739 648 BBBC047 Matched conditions 3,788 3,921 4,114 4,380 3,878 Murcko groups 853 912 941 913 921 CP replicate profiles 15,891 17,060 17,263 18,390 16,429 L1000 replicate profiles 10,722 10,692 11,531 12,252 10,815
这段在干什么:用表格列出两个数据集(BBBC036、BBBC047)在5个交叉验证折上的四种样本计数。
需要解释的地方:
值得留意:BBBC047 各指标远大于 BBBC036,说明两数据集规模差异大;每折数值不等,说明划分并非等分。
b073Frozen source Audit Global PCC Replication Global PCC Audit CP PCC Audit L1000 PCC Parameters BBBC036 h0h_{0} -0.0030 [-0.0400, 0.0339] 0.0484 [0.0223, 0.0746] 0.0255 -0.0130 1,144,864 CellScientist 0.1921 [0.1903, 0.1939] 0.1567 [0.1517, 0.1617] 0.2758 0.1568 3,999,776 AIDE 0.1746 [0.1495, 0.1998] 0.1497 [0.1193, 0.1800] 0.2520 0.1400 4,419,872 CellForge -0.0091 [-0.0397, 0.0215] 0.0640 [0.0187, 0.1093] 0.0821 -0.0435 3,824,672 HarmonyCell 0.0657 [0.0378, 0.0936] 0.0102 [-0.0426, 0.0629] 0.0486 0.0719 2,061,343 BBBC047 h0h_{0} 0.1805 [0.1769, 0.1840] 0.2014 [0.1941, 0.2087] 0.2688 0.0494 1,201,294 CellScientist 0.1781 [0.1716, 0.1846] 0.2036 [0.1940, 0.2132] 0.2615 0.0469 1,777,486 AIDE 0.1805 [0.1769, 0.1840] 0.2014 [0.1941, 0.2087] 0.2688 0.0494 1,201,294 CellForge 0.1771 [0.1716, 0.1826] 0.2011 [0.1963, 0.2059] 0.2652 0.0444 2,531,982 HarmonyCell 0.1808 [0.1699, 0.1916] 0.1963 [0.1921, 0.2006] 0.2689 0.0484 2,284,875
这段在干什么:这是附录里的一张结果表,逐行列出五种模型(h₀ 到 HarmonyCell)在两个数据集(BBBC036、BBBC047)上的四项 PCC 指标和参数量。
需要解释的地方:PCC 是皮尔逊相关系数(领域常识,衡量预测与真实的线性相关)。列名里 "Frozen source" 是冻结源模型,其余几列是不同层面的相关性检查;方括号是置信区间。
值得留意:表中 h₀ 在 BBBC047 一行的数值与 AIDE 完全相同,参数量也一样,值得核对是否为复制粘贴错误。另外 CellForge 在 BBBC036 出现负值,与其余模型方向不同。
b074Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.
这段在干什么:这是表格的图注,规定表里加粗和下划线分别代表同一任务/折中排名第一、第二的预测均值。
需要解释的地方:加粗=该任务最好的;下划线=第二好的,同列中"distinct"即去掉并列后分别取两档。平局并列共享同一排名;区间上下界和诊断/资源类列不参与排名,故不加标记。
值得留意:排名只在"同一任务或同一折内"比较,跨任务不横比;平局共享名次意味着可能出现没有下划线的列。
b075The plate-family audit crosses the observed or permuted compound representation with the observed or an alternative control profile from the same plate family. Each selected model is refit under five fixed seeds, and each refit averages 32 permutations constructed without response values. Compounds are permuted within the observed control-profile group whenever possible, covering every held-out BBBC036 row and 98.8%98.8\% of BBBC047 rows; each remaining row receives a fixed donor with a different compound from the same held-out partition. Table 5 separates compound, control-profile, and interaction effects on Folds 4 and 5.
1. 这段在干什么
交代「plate-family audit」的实验设计:怎么组合表征与控制 profile、怎么重复采样,并说明 Table 5 用这些数据拆解三类效应。
2. 需要解释的地方
3. 值得留意
b076Model Global PCC Compound effect Control effect Interaction BBBC036 / Audit h0h_{0} -0.0030 [-0.0400, 0.0339] 0.0025 [0.0020, 0.0031] 0.0012 [-0.0126, 0.0150] -0.0005 [-0.0008, -0.0002] CellScientist 0.1921 [0.1903, 0.1939] 0.0000 [0.0000, 0.0000] 0.0146 [0.0114, 0.0178] 0.0000 [0.0000, 0.0000] AIDE 0.1746 [0.1495, 0.1998] 0.0024 [0.0019, 0.0028] 0.0260 [0.0076, 0.0445] -0.0008 [-0.0012, -0.0003] CellForge -0.0091 [-0.0397, 0.0215] 0.0018 [0.0011, 0.0025] -0.0086 [-0.0169, -0.0004] -0.0005 [-0.0010, 0.0001] HarmonyCell 0.0657 [0.0378, 0.0936] 0.0019 [0.0002, 0.0037] 0.0768 [0.0632, 0.0905] -0.0028 [-0.0044, -0.0013] BBBC047 / Audit h0h_{0} 0.1805 [0.1769, 0.1840] 0.0046 [0.0039, 0.0053] 0.0249 [0.0228, 0.0271] 0.0008 [0.0005, 0.0012] CellScientist 0.1781 [0.1716, 0.1846] 0.0000 [0.0000, 0.0000] 0.0295 [0.0243, 0.0346] 0.0000 [0.0000, 0.0000] AIDE 0.1805 [0.1769, 0.1840] 0.0046 [0.0039, 0.0053] 0.0249 [0.0228, 0.0271] 0.0008 [0.0005, 0.0012] CellForge 0.1771 [0.1716, 0.1826] 0.0050 [0.0039, 0.0061] 0.0208 [0.0163, 0.0252] 0.0014 [0.0004, 0.0023] HarmonyCell 0.1808 [0.1699, 0.1916] ×10−96.892\!\times\!10^{-9} [−1.390,2.768]×10−8[-1.390,2.768]\!\times\!10^{-8} 0.0365 [0.0245, 0.0484] ×10−81.211\!\times\!10^{-8} [−2.216,4.637]×10−8[-2.216,4.637]\!\times\!10^{-8} BBBC036 / Held-out replication h0h_{0} 0.0484 [0.0223, 0.0746] 0.0026 [×10−6,0.0053][9.873\!\times\!10^{-6},0.0053] 0.0043 [-0.0117, 0.0203] 0.0002 [-0.0006, 0.0009] CellScientist 0.1567 [0.1517, 0.1617] 0.0000 [0.0000, 0.0000] -0.0070 [-0.0088, -0.0053] 0.0000 [0.0000, 0.0000] AIDE 0.1497 [0.1193, 0.1800] 0.0013 [-0.0004, 0.0031] -0.0016 [-0.0168, 0.0136] -0.0001 [-0.0004, 0.0002] CellForge 0.0640 [0.0187, 0.1093] 0.0034 [0.0026, 0.0043] 0.0018 [-0.0087, 0.0123] 0.0001 [-0.0006, 0.0009] HarmonyCell 0.0102 [-0.0426, 0.0629] -0.0063 [-0.0100, -0.0025] -0.0038 [-0.0350, 0.0274] -0.0037 [-0.0080, 0.0005] BBBC047 / Held-out replication h0h_{0} 0.2014 [0.1941, 0.2087] 0.0030 [0.0023, 0.0038] 0.0136 [0.0077, 0.0194] 0.0002 [-0.0010, 0.0013] CellScientist 0.2036 [0.1940, 0.2132] 0.0000 [0.0000, 0.0000] 0.0231 [0.0168, 0.0293] 0.0000 [0.0000, 0.0000] AIDE 0.2014 [0.1941, 0.2087] 0.0030 [0.0023, 0.0038] 0.0136 [0.0077, 0.0194] 0.0002 [-0.0010, 0.0013] CellForge 0.2011 [0.1963, 0.2059] 0.0037 [0.0030, 0.0044] 0.0088 [0.0029, 0.0148] -0.0004 [-0.0010, 0.0002] HarmonyCell 0.1963 [0.1921, 0.2006] ×10−93.073\!\times\!10^{-9} [−5.460,11.61]×10−9[-5.460,11.61]\!\times\!10^{-9} 0.0205 [0.0142, 0.0268] ×10−92.608\!\times\!10^{-9} [−4.632,9.848]×10−9[-4.632,9.848]\!\times\!10^{-9}
这段是附录里的一张完整数值表,紧接上一段"把效应拆成三块"的说法,把每个模型在四个数据集上的拆分结果全摆出来。
这段在干什么:用 Table 5 给出各模型($h_0$、CellScientist、AIDE、CellForge、HarmonyCell)的 Global PCC,及其分解出的 compound、control、interaction 三项效应和区间。
需要解释的地方:Global PCC 是预测与真值的整体相关性(领域常识);方括号是置信区间;中间三列把整体表现拆成"化合物本身""对照谱""两者交互"三部分贡献。
值得留意:CellScientist 的 control 效应和交互几乎恒为 0.0000,AIDE 在 BBBC047 两行与 $h_0$ 数值完全相同——这两点容易一扫而过,但可能正对应正文的审计结论。
b077Fixed residual baselines isolate a perturbation-specific function class while keeping the same task. They retain the selected control-profile predictor g(c)g(c) and learn one joint residual h(p,a)h(p,a) using either multi-output Ridge or a shallow multi-output MLP. Both baselines use Fold 3 for selection and are refit under the same five seeds. Their Fold-4 and Fold-5 predictive scores, compound effects, and target-loss gains are reported together in Table 48.
这段在干什么:介绍两个“固定残差基线”,用来对照扰动本身的效应:保留同一个预测器 g(c),只学一个联合残差 h(p,a)。
需要解释的地方:
值得留意:两个基线只在 h 的模型类上不同(Ridge vs. MLP),g 和任务保持一致——这正是“隔离出扰动专属函数类”的意思;Fold 4/5 才报告分数,说明 3 折用于选择、后两折用于评估。
b078Two additional small-molecule cohorts test the final procedure beyond BBBC development. We first apply it to paired morphology–transcription LINCS–Pilot1, then transfer its ten trajectory-selected design instances to independently acquired Cell Painting LKCP Batch 2 under the same compound, dose, and control-profile inputs. Appendix E records data provenance, the order in which models and rules were fixed, the input permutations, and the compound–dose decomposition.
这章在交代数据来源:用两个额外的小分子队列,检验前面在 BBBC 上定下来的流程能不能推广出去。
几个词:小分子队列指一批用化合物处理的实验样本。LINCS–Pilot1 是配对好的「形态+转录」数据(领域常识:形态指细胞成像,转录指基因表达)。LKCP Batch 2 是另一批独立采集的 Cell Painting 图像数据。
值得留意:「ten trajectory-selected design instances」是从最开始那份数据里挑出来、再原样搬到新数据上的十个设计实例,化合物、剂量、对照输入都保持不变——这正是「迁移检验」的关键,读的时候别把它当成又跑了一遍新实验。
b080Model Fold Shared PCC Disjoint PCC Shared PCC drop Disjoint PCC drop Shared Loss gain Disjoint Loss gain Score F4 0.30810.3081 0.30850.3085 00 00 00 00 Control only F4 0.30840.3084 0.30820.3082 00 00 00 00 Guided F4 0.31110.3111 0.31100.3110 0.00470.0047 0.00430.0043 ×10−41.65\!\times\!10^{-4} ×10−41.51\!\times\!10^{-4} Score F5 0.31910.3191 0.31930.3193 00 00 00 00 Control only F5 0.31870.3187 0.31850.3185 00 00 00 00 Guided F5 0.32140.3214 0.32100.3210 0.00360.0036 0.00330.0033 ×10−41.31\!\times\!10^{-4} ×10−41.19\!\times\!10^{-4}
这是附录里的一张对照表,报告各模型在 Shared / Disjoint 两种划分下的 PCC、PCC 降幅和 Loss 增益。
Shared 与 Disjoint 两列数值几乎相同,说明划分方式几乎不影响结果。三组 Score/Control only/Guided 中,只有 Guided 的 drop 与 gain 非零,Score 与 Control only 全为 0。
b081Means of five seeds, averaging directions A/B within seed; compound replacement uses 32 fixed maps. PCC drop and loss gain are correct-minus-replacement PCC and replacement-minus-correct MSE. Guided minus control-only disjoint-arm PCC: F4 0.00280.0028 [0.0017, 0.0038][0.0017,\,0.0038], F5 0.00250.0025 [0.0012, 0.0037][0.0012,\,0.0037] (paired-seed 95% tt CIs).
这段在干什么:报告引导组与控制组之间 PCC 差值的汇总统计,用五个种子的均值和配对置信区间来支撑“控制输入与目标引用可分离”这一主张。
需要解释的地方:
值得留意:表格里 F5 出现两次(0.3214 / 0.3210),差值 0.0036 恰好是上一段末尾的数;而这里的 0.0025 是“引导减仅控制”的分离臂口径,两者不是同一个量,别混。
b082The shared and disjoint arms are newly constructed BBBC047 comparisons; the historical shared-control result is not their absolute baseline. Physical control wells are hash-sorted within each plate and well-row stratum and allocated to balanced, disjoint A/B banks (seed 2026091401). Within direction A, the target uses bank A as reference: shared context uses A, disjoint context uses B; direction B reverses these roles. CP targets are treated-well empirical mid-CDF values relative to the direction’s reference bank, centered by 0.50.5 and aggregated by condition medians. The 701 CP features are fixed from the historical schema; no new feature selection is performed. Shared/disjoint arms have byte-identical targets and row identities within direction, and L1000 targets are unchanged. A/B are averaged within each seed, not counted as independent replicas. This intervention separates physical wells at the reconstructed CP reference layer; it does not establish that all upstream preprocessing or biological dependence has been removed. No historical qualification label is changed.
1. 这段在干什么
交代 A/B 对照臂的构造细节,说明这只是一次「物理隔离」干预,不改变历史标签。
2. 需要解释的地方
3. 值得留意
作者主动划界:只说隔离了物理孔,不声称清除了上游预处理或生物学依赖;A/B 取平均后不算独立重复。
b083Quantity Value Interpretation Predictor refits / anchors 60 / 20 80 training stages; 120 fold evaluations Seeds / directions 5 / 2 A/B are paired inside each seed F4 retained support 3341 / 4380 76.28% of evaluation rows F5 retained support 2927 / 3878 75.48% of evaluation rows F4 / F5 source scaffolds 913 / 921 Bootstrap over source-scaffold labels Maps / scaffold draws 32 / 2000 All refits and donor maps held fixed Donor kernel total variation 0 (both folds) Every eligible compound contributes one row Guided coordinate signs 40 / 40 positive RMS, loss gain and PCC drop, each record Score/control coordinates 80 / 80 exact zero All three compound coordinates Physical controls, A / B 9144 / 9144 Unique plate/well identities; banks disjoint CP plates / retained conditions 273 / 20081 0 conditions excluded
1. 这段在干什么
这是一张审计配置清单:用表格列出评估中固定了哪些量、保留了多少样本,作用是证明控制输入与目标引用在物理上确实被分开了。
2. 需要解释的地方
3. 值得留意
表里「Maps / scaffold draws = 32 / 2000,所有 refit 与 donor map 保持固定」意味着打乱只发生在标签层面,配置本身没变——这正是"物理分离"主张的关键。
b084Only row-uniform maps were evaluated. An independent metadata calculation proves equality of row-uniform and compound-balanced donor probabilities on all 6268 retained source pools; pre-generated compound-balanced maps are not an additional evaluated sensitivity analysis. Kernel equality is not biological exchangeability.
1. 这段在干什么
这是一段方法上的“澄清与边界声明”,回应读者可能提的质疑:为什么只评估 row-uniform 地图、没有评估 compound-balanced 地图。
2. 需要解释的地方
3. 值得留意
两层关键声明:一、compound-balanced 地图并非“漏做的分析”,因为概率已被证明相等,预先生成的版本不算额外的敏感性分析;二、作者特意划清界限——核(kernel)相等不等于生物学可交换性,即数学上的等价不能直接推到生物意义上的可互换。这句话很短,但很可能是作者在防过度解读。
b085Separately from the direction-specific bank targets, a full-bank CP mid-CDF reconstruction was compared with historical normalized CP ranks. It recorded a maximum absolute discrepancy of 0.191406250.19140625 and mean absolute discrepancy ×10−62.76\!\times\!10^{-6}, with 99.9799%99.9799\% agreement at 10−610^{-6}. This discrepancy was retained, not repaired by choosing a favorable target transform. The portable metadata retains the complete check.
1. 这段在干什么
报告一次"全库重建 vs 历史排名"的对齐检查结果:两者差异极小(最大绝对差 0.191,平均约 2.76×10⁻⁶,99.97% 一致),并声明没为了好看去挑变换把差异抹平。
2. 需要解释的地方
3. 值得留意
作者强调"差异被保留、未被修复",这是在主动承担不完美,以反驳"调参数凑结果"的质疑;末尾"metadata 保留完整检查"意味着可复现。这是承上一段"kernel 相等≠生物可换"的同一防守姿态——但具体机制这段没展开。
b086Model Fold Dir. Arm Global PCC MSE CP PCC CP MSE L1000 PCC L1000 MSE Fold 4 Score F4 A Shared 0.30820.3082 0.04600.0460 0.33310.3331 0.05320.0532 0.28020.2802 0.04080.0408 Score F4 A Disjoint 0.30870.3087 0.04600.0460 0.33340.3334 0.05320.0532 0.28070.2807 0.04080.0408 Score F4 B Shared 0.30790.3079 0.04600.0460 0.33170.3317 0.05320.0532 0.28100.2810 0.04080.0408 Score F4 B Disjoint 0.30840.3084 0.04600.0460 0.33170.3317 0.05320.0532 0.28180.2818 0.04080.0408 Control only F4 A Shared 0.30930.3093 0.04600.0460 0.33270.3327 0.05320.0532 0.28280.2828 0.04080.0408 Control only F4 A Disjoint 0.30860.3086 0.04600.0460 0.33220.3322 0.05330.0533 0.28180.2818 0.04080.0408 Control only F4 B Shared 0.30740.3074 0.04600.0460 0.33030.3303 0.05330.0533 0.28130.2813 0.04080.0408 Control only F4 B Disjoint 0.30770.3077 0.04600.0460 0.33040.3304 0.05330.0533 0.28180.2818 0.04080.0408 Guided F4 A Shared 0.31220.3122 0.04590.0459 0.33840.3384 0.05300.0530 0.28230.2823 0.04080.0408 Guided F4 A Disjoint 0.31150.3115 0.04590.0459 0.33730.3373 0.05310.0531 0.28210.2821 0.04080.0408 Guided F4 B Shared 0.31000.3100 0.04590.0459 0.33500.3350 0.05310.0531 0.28140.2814 0.04080.0408 Guided F4 B Disjoint 0.31050.3105 0.04590.0459 0.33560.3356 0.05310.0531 0.28170.2817 0.04080.0408 Fold 5 Score F5 A Shared 0.32190.3219 0.04540.0454 0.34920.3492 0.05360.0536 0.28950.2895 0.03960.0396 Score F5 A Disjoint 0.32220.3222 0.04540.0454 0.35000.3500 0.05350.0535 0.28890.2889 0.03960.0396 Score F5 B Shared 0.31630.3163 0.04550.0455 0.33890.3389 0.05370.0537 0.28940.2894 0.03960.0396 Score F5 B Disjoint 0.31650.3165 0.04550.0455 0.33860.3386 0.05370.0537 0.29030.2903 0.03960.0396 Control only F5 A Shared 0.32170.3217 0.04540.0454 0.34800.3480 0.05360.0536 0.29040.2904 0.03960.0396 Control only F5 A Disjoint 0.32180.3218 0.04550.0455 0.34930.3493 0.05360.0536 0.28890.2889 0.03960.0396 Control only F5 B Shared 0.31560.3156 0.04550.0455 0.33730.3373 0.05380.0538 0.28970.2897 0.03960.0396 Control only F5 B Disjoint 0.31530.3153 0.04560.0456 0.33640.3364 0.05390.0539 0.29010.2901 0.03960.0396 Guided F5 A Shared 0.32490.3249 0.04540.0454 0.35360.3536 0.05340.0534 0.29050.2905 0.03960.0396 Guided F5 A Disjoint 0.32360.3236 0.04540.0454 0.35250.3525 0.05350.0535 0.28920.2892 0.03960.0396 Guided F5 B Shared 0.31780.3178 0.04550.0455 0.34160.3416 0.05370.0537 0.28950.2895 0.03960.0396 Guided F5 B Disjoint 0.31840.3184 0.04550.0455 0.34200.3420 0.05360.0536 0.29030.2903 0.03960.0396
这段在干什么:用一张大表汇报 Shared 与 Disjoint 两种输入模式下、各模型架构在两个 Fold 上的六项指标,为「物理分离控制输入与目标引用」提供证据。
需要解释的地方:
值得留意:Shared 与 Disjoint 的数值几乎一致,差异都在小数点后第三四位;L1000 MSE 列基本恒为 0.0408/0.0396,几乎不区分任何条件——这两点是全表的重点,但作者在本段没有点明。
b087Model Fold Dir. Arm Compound RMS Compound loss gain Compound PCC drop Fold 4 Score F4 A Shared 00 00 00 Score F4 A Disjoint 00 00 00 Score F4 B Shared 00 00 00 Score F4 B Disjoint 00 00 00 Control only F4 A Shared 00 00 00 Control only F4 A Disjoint 00 00 00 Control only F4 B Shared 00 00 00 Control only F4 B Disjoint 00 00 00 Guided F4 A Shared 0.01260.0126 ×10−41.75\!\times\!10^{-4} 0.00500.0050 Guided F4 A Disjoint 0.01180.0118 ×10−41.48\!\times\!10^{-4} 0.00410.0041 Guided F4 B Shared 0.01180.0118 ×10−41.55\!\times\!10^{-4} 0.00450.0045 Guided F4 B Disjoint 0.01180.0118 ×10−41.54\!\times\!10^{-4} 0.00440.0044 Fold 5 Score F5 A Shared 00 00 00 Score F5 A Disjoint 00 00 00 Score F5 B Shared 00 00 00 Score F5 B Disjoint 00 00 00 Control only F5 A Shared 00 00 00 Control only F5 A Disjoint 00 00 00 Control only F5 B Shared 00 00 00 Control only F5 B Disjoint 00 00 00 Guided F5 A Shared 0.01250.0125 ×10−41.43\!\times\!10^{-4} 0.00390.0039 Guided F5 A Disjoint 0.01160.0116 ×10−41.17\!\times\!10^{-4} 0.00320.0032 Guided F5 B Shared 0.01170.0117 ×10−41.19\!\times\!10^{-4} 0.00330.0033 Guided F5 B Disjoint 0.01170.0117 ×10−41.21\!\times\!10^{-4} 0.00340.0034
这段是一张消融结果表,用来对比不同条件下模型的表现。
这段在干什么:按 Fold 4/5 分组,列出 Score、Control only、Guided 三类模型在 Shared/Disjoint 两种设置下的三项指标,用数字说明哪些输入真正有用。
需要解释的地方:
值得留意:Score 和 Control only 两组的数值全是 0,只有 Guided 行有非零值(如 0.0126、×10⁻⁴、0.0050),说明前两类模型在这段里没有可测得的贡献。
b088Contrast Fold Mean Seed 95% CI Scaffold 95% CI Global PCC (×103\times 10^{3}) Score: D−-S F4 0.4270.427 [−0.130, 0.984][-0.130,\,0.984] [−0.070, 0.945][-0.070,\,0.945] Score: D−-S F5 0.2430.243 [−0.583, 1.069][-0.583,\,1.069] [−0.158, 0.638][-0.158,\,0.638] Control only: D−-S F4 −0.186-0.186 [−0.769, 0.398][-0.769,\,0.398] [−0.794, 0.338][-0.794,\,0.338] Control only: D−-S F5 −0.136-0.136 [−1.679, 1.407][-1.679,\,1.407] [−0.687, 0.393][-0.687,\,0.393] Guided: D−-S F4 −0.099-0.099 [−0.972, 0.773][-0.972,\,0.773] [−0.442, 0.225][-0.442,\,0.225] Guided: D−-S F5 −0.327-0.327 [−0.963, 0.309][-0.963,\,0.309] [−0.640, 0.004][-0.640,\,0.004] Score−-control, S F4 −0.295-0.295 [−1.223, 0.633][-1.223,\,0.633] [−1.724, 1.066][-1.724,\,1.066] Score−-control, S F5 0.4280.428 [−0.662, 1.517][-0.662,\,1.517] [−0.938, 1.539][-0.938,\,1.539] Guided−-control, S F4 2.6882.688 [2.059, 3.317][2.059,\,3.317] [1.012, 4.541][1.012,\,4.541] Guided−-control, S F5 2.6742.674 [1.609, 3.739][1.609,\,3.739] [0.272, 5.761][0.272,\,5.761] Score−-control, D F4 0.3180.318 [−0.514, 1.149][-0.514,\,1.149] [−0.853, 1.491][-0.853,\,1.491] Score−-control, D F5 0.8070.807 [−0.037, 1.650][-0.037,\,1.650] [−0.554, 1.931][-0.554,\,1.931] Guided−-control, D F4 2.7752.775 [1.731, 3.818][1.731,\,3.818] [1.264, 4.468][1.264,\,4.468] Guided−-control, D F5 2.4832.483 [1.227, 3.739][1.227,\,3.739] [0.187, 5.548][0.187,\,5.548] MSE (×105\times 10^{5}) Score: D−-S F4 −1.079-1.079 [−2.884, 0.726][-2.884,\,0.726] [−2.847, 0.536][-2.847,\,0.536] Score: D−-S F5 −0.564-0.564 [−3.258, 2.131][-3.258,\,2.131] [−1.846, 0.775][-1.846,\,0.775] Control only: D−-S F4 1.9231.923 [−1.011, 4.858][-1.011,\,4.858] [0.232, 3.870][0.232,\,3.870] Control only: D−-S F5 1.4901.490 [−4.983, 7.964][-4.983,\,7.964] [−0.383, 3.538][-0.383,\,3.538] Guided: D−-S F4 1.0471.047 [−2.859, 4.954][-2.859,\,4.954] [−0.192, 2.309][-0.192,\,2.309] Guided: D−-S F5 1.7801.780 [−0.662, 4.222][-0.662,\,4.222] [0.555, 2.980][0.555,\,2.980] Score−-control, S F4 −0.598-0.598 [−3.662, 2.465][-3.662,\,2.465] [−4.858, 4.209][-4.858,\,4.209] Score−-control, S F5 −2.943-2.943 [−7.456, 1.569][-7.456,\,1.569] [−6.772, 1.784][-6.772,\,1.784] Guided−-control, S F4 −7.236-7.236 [−11.010,−3.463][-11.010,\,-3.463] [−13.811,−1.332][-13.811,\,-1.332] Guided−-control, S F5 −7.329-7.329 [−14.377,−0.282][-14.377,\,-0.282] [−18.153, 0.990][-18.153,\,0.990] Score−-control, D F4 −3.601-3.601 [−6.091,−1.110][-6.091,\,-1.110] [−7.476, 0.520][-7.476,\,0.520] Score−-control, D F5 −4.998-4.998 [−8.488,−1.507][-8.488,\,-1.507] [−8.982,−0.105][-8.982,\,-0.105] Guided−-control, D F4 −8.112-8.112 [−12.543,−3.681][-12.543,\,-3.681] [−14.199,−2.832][-14.199,\,-2.832] Guided−-control, D F5 −7.040-7.040 [−12.443,−1.636][-12.443,\,-1.636] [−17.585, 0.849][-17.585,\,0.849]
这段是一张数值表,列出各组对比的 PCC、MSE 及两种置信区间。
在干什么:用统一格式汇总消融实验的量化结果,支撑正文关于「控制输入/目标参考分离」的论证。
需解释:Fold 指数据划分折,Seed 指随机种子;Scaffold 是另一种划分方式(领域常识:按分子骨架划分以测泛化)。PCC 为相关,MSE 为误差。表头 D−S 是「去脚手架 vs 留脚手架」差异;Score−control、Guided−control 等为消融组相减。
值得留意:Guided−control 两行数值大(约 2.7)且置信区间不跨 0,而 Score/Control only 的区间多跨 0——但作者没在原文点明。
b089D−-S: disjoint minus shared wells. Within-arm contrasts subtract control only. Seed CIs use five A/B-averaged paired values (t4t_{4}). Scaffold CIs share each draw across all arms/models/directions and fix all five refits, 32 maps and the donor pool; 2000 draws, seed 2026091433+fold2026091433+\mathrm{fold}. PCC is recomputed after pooling sufficient statistics, never averaged across scaffold PCCs. These are source-scaffold, not compound-cluster, intervals; scaffold labels do not ensure fully independent biological groups.
这段在干什么:交代区间估计的构建方式——用 D−S(去共享井)做对照、用 bootstrap 重采样(2000 次)算置信区间,并说明为什么这些区间是 source-scaffold 级别而非真正独立的生物学分组。
需要解释的地方:
值得留意:作者主动声明这些区间不保证生物学独立性——这是重要的自我限制,别当成严格的统计推断。
b090Contrast Fold Mean Seed 95% CI Scaffold 95% CI Compound RMS (×103\times 10^{3}) Score: D−-S F4 00 [0, 0][0,\,0] [0, 0][0,\,0] Score: D−-S F5 00 [0, 0][0,\,0] [0, 0][0,\,0] Control only: D−-S F4 00 [0, 0][0,\,0] [0, 0][0,\,0] Control only: D−-S F5 00 [0, 0][0,\,0] [0, 0][0,\,0] Guided: D−-S F4 −0.414-0.414 [−1.567, 0.738][-1.567,\,0.738] [−0.495,−0.333][-0.495,\,-0.333] Guided: D−-S F5 −0.406-0.406 [−1.563, 0.750][-1.563,\,0.750] [−0.490,−0.327][-0.490,\,-0.327] Score−-control, S F4 00 [0, 0][0,\,0] [0, 0][0,\,0] Score−-control, S F5 00 [0, 0][0,\,0] [0, 0][0,\,0] Guided−-control, S F4 12.22012.220 [11.209, 13.231][11.209,\,13.231] [11.249, 13.186][11.249,\,13.186] Guided−-control, S F5 12.08512.085 [11.011, 13.158][11.011,\,13.158] [11.456, 12.679][11.456,\,12.679] Score−-control, D F4 00 [0, 0][0,\,0] [0, 0][0,\,0] Score−-control, D F5 00 [0, 0][0,\,0] [0, 0][0,\,0] Guided−-control, D F4 11.80511.805 [9.826, 13.784][9.826,\,13.784] [10.884, 12.725][10.884,\,12.725] Guided−-control, D F5 11.67811.678 [9.740, 13.616][9.740,\,13.616] [11.106, 12.221][11.106,\,12.221] Compound loss gain (×105\times 10^{5}) Score: D−-S F4 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.36\!\times\!10^{-12},\,2.64\!\times\!10^{-12}] Score: D−-S F5 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.71\!\times\!10^{-12},\,1.87\!\times\!10^{-12}] Control only: D−-S F4 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.91\!\times\!10^{-12},\,1.53\!\times\!10^{-12}] Control only: D−-S F5 00 [0, 0][0,\,0] [−×10−12,×10−12][-1.46\!\times\!10^{-12},\,3.05\!\times\!10^{-12}] Guided: D−-S F4 −1.403-1.403 [−4.179, 1.374][-4.179,\,1.374] [−2.134,−0.809][-2.134,\,-0.809] Guided: D−-S F5 −1.163-1.163 [−3.432, 1.107][-3.432,\,1.107] [−1.938,−0.376][-1.938,\,-0.376] Score−-control, S F4 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.43\!\times\!10^{-12},\,2.50\!\times\!10^{-12}] Score−-control, S F5 00 [0, 0][0,\,0] [−×10−12,×10−12][-1.39\!\times\!10^{-12},\,3.26\!\times\!10^{-12}] Guided−-control, S F4 16.50916.509 [13.377, 19.640][13.377,\,19.640] [10.136, 24.392][10.136,\,24.392] Guided−-control, S F5 13.10213.102 [10.729, 15.475][10.729,\,15.475] [5.985, 20.741][5.985,\,20.741] Score−-control, D F4 00 [0, 0][0,\,0] [−×10−12,×10−12][-1.32\!\times\!10^{-12},\,3.12\!\times\!10^{-12}] Score−-control, D F5 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.57\!\times\!10^{-12},\,2.01\!\times\!10^{-12}] Guided−-control, D F4 15.10615.106 [9.906, 20.305][9.906,\,20.305] [9.267, 22.336][9.267,\,22.336] Guided−-control, D F5 11.93911.939 [7.423, 16.455][7.423,\,16.455] [5.390, 18.963][5.390,\,18.963]
这段是两张消融表,直接支撑上段“scaffold 标签不保证独立性”的论点。
在干什么:把各类方法(Score、Control、Guided)在 Compound RMS 和 loss gain 两个指标上的 D−S 差值列出来,做证据。
需要解释:
值得留意:纯 Score 和 Control 行的 D−S 全是 0,置信区间覆盖零;只有 Guided 行才出现非零值(如 12.220、11.805)。说明差异只在方法引入引导后才显现。两列 CI(Seed/Scaffold)在 Guided 行常不重叠,值得对比。
b091D−-S: disjoint minus shared wells. Within-arm contrasts subtract control only. Seed CIs use five A/B-averaged paired values (t4t_{4}). Scaffold CIs share each draw across all arms/models/directions and fix all five refits, 32 maps and the donor pool; 2000 draws, seed 2026091433+fold2026091433+\mathrm{fold}. PCC is recomputed after pooling sufficient statistics, never averaged across scaffold PCCs. These are source-scaffold, not compound-cluster, intervals; scaffold labels do not ensure fully independent biological groups.
说明这套方法里对照组怎么切分(D−-S)以及置信区间怎么算——即结果可信度的计算口径。
b092Contrast Fold Mean Seed 95% CI Scaffold 95% CI Score: D−-S F4 00 [0, 0][0,\,0] [−×10−13,×10−13][-1.50\!\times\!10^{-13},\,1.83\!\times\!10^{-13}] Score: D−-S F5 00 [0, 0][0,\,0] [−×10−13,×10−13][-1.89\!\times\!10^{-13},\,1.78\!\times\!10^{-13}] Control only: D−-S F4 00 [0, 0][0,\,0] [−×10−13,×10−13][-1.61\!\times\!10^{-13},\,1.83\!\times\!10^{-13}] Control only: D−-S F5 00 [0, 0][0,\,0] [−×10−13,×10−13][-2.33\!\times\!10^{-13},\,1.44\!\times\!10^{-13}] Guided: D−-S F4 −0.424-0.424 [−1.191, 0.342][-1.191,\,0.342] [−0.627,−0.248][-0.627,\,-0.248] Guided: D−-S F5 −0.337-0.337 [−0.944, 0.270][-0.944,\,0.270] [−0.559,−0.115][-0.559,\,-0.115] Score−-control, S F4 00 [0, 0][0,\,0] [−×10−13,×10−13][-2.05\!\times\!10^{-13},\,1.11\!\times\!10^{-13}] Score−-control, S F5 00 [0, 0][0,\,0] [−×10−13,×10−13][-2.28\!\times\!10^{-13},\,1.33\!\times\!10^{-13}] Guided−-control, S F4 4.7084.708 [3.832, 5.583][3.832,\,5.583] [2.949, 6.840][2.949,\,6.840] Guided−-control, S F5 3.6253.625 [2.955, 4.296][2.955,\,4.296] [1.650, 5.741][1.650,\,5.741] Score−-control, D F4 00 [0, 0][0,\,0] [−×10−13,×10−13][-2.22\!\times\!10^{-13},\,1.33\!\times\!10^{-13}] Score−-control, D F5 00 [0, 0][0,\,0] [−×10−13,×10−13][-1.89\!\times\!10^{-13},\,1.67\!\times\!10^{-13}] Guided−-control, D F4 4.2834.283 [2.823, 5.744][2.823,\,5.744] [2.673, 6.238][2.673,\,6.238] Guided−-control, D F5 3.2883.288 [2.042, 4.535][2.042,\,4.535] [1.471, 5.215][1.471,\,5.215]
1. 这段在干什么
用一张对照表检验:把输入和目标做物理隔离后,各类预测分数是否还站得住。
2. 需要解释的地方
3. 值得留意
所有含 Score 的对比 Mean 都是 0,CI 里是 10⁻¹³ 级,即隔离后分数信号基本归零;而 Guided 相关对比仍显著偏离 0(如 4.708),差异不在一个量级。
b093D−-S: disjoint minus shared wells. Within-arm contrasts subtract control only. Seed CIs use five A/B-averaged paired values (t4t_{4}). Scaffold CIs share each draw across all arms/models/directions and fix all five refits, 32 maps and the donor pool; 2000 draws, seed 2026091433+fold2026091433+\mathrm{fold}. PCC is recomputed after pooling sufficient statistics, never averaged across scaffold PCCs. These are source-scaffold, not compound-cluster, intervals; scaffold labels do not ensure fully independent biological groups.
这段在干什么:交代统计推断的口径——(D−S) 是「减去共享孔」的差集,并按「臂内」和「scaffold 重采样」两套规则分别构造置信区间。
需要解释的地方:
值得留意:作者自认这是 source-scaffold 区间,不是 compound-cluster 区间;scaffold 标签不等于完全独立的生物分组。
b094The guided model has positive compound RMS, target-loss gain and PCC drop in all 40 evaluations, whereas Score and control only are exactly invariant. The guided disjoint-arm PCC advantage over control only remains positive under both reported interval methods. However, all guided disjoint-minus-shared coordinate differences include zero under the five-seed interval while their fixed-refit source-scaffold intervals are negative; the two uncertainty scopes do not establish an unchanged effect. The F5 guided-minus-control MSE contrast likewise crosses zero under the source-scaffold interval despite a negative seed interval. A positive prediction RMS indicates input use, not target benefit. RMS is the mean over maps of the square root of mean squared prediction distance, not the square root after averaging maps.
汇报引导模型在全部 40 次评估中都有正向效应,但两种区间方法结论不一致,故不能断言效应不变。
作者自陈:正向预测 RMS 只说明「用了输入」,不等于「对目标有好处」——这是防止误读的关键限定。
b096The fixed executor owns joint inputs (c,p,a)(c,p,a), the response target, data transformations, folds, evaluator, checkpoint rule, and training configuration; candidate code cannot alter data loading, targets, folds, or metrics. Discovery accesses Folds 1–3 only. Global PCC is one Pearson correlation after flattening all post-search rows and all dimensions of the inverse-standardized joint output. CP and L1000 PCC use the same operation within each response block. MSE is averaged over every element of the inverse-standardized joint output. At search step tt, the policy maps the current model, Fold-3 diagnostics, and compact search history to a proposed revision,
这段在干什么:交代实验的固定约束与评价口径——执行器锁死输入、目标、折与指标,发现阶段只能用 Fold 1–3,并逐条定义 PCC/MSE 怎么算。
需要解释的地方:
值得留意:发现阶段只碰 Fold 1–3,说明 Fold 4 之后是留给验证的——原文这里没明说,但"only"是刻意的。
b097where edits are limited to architecture, fusion, response readout, loss, optimizer, and scheduler. The experiment log stores the proposal’s parent, hypothesis, source or structured design, declared input use, execution and repair results, checkpoint, metrics, provider usage, and compute. Within each ten-slot prediction-score trajectory, the incumbent changes only under a strict improvement in Fold-3 Global PCC. After all ten trajectories finish, the first, within-policy selection chooses one source per search policy from that policy’s complete executable candidate pool by maximum Fold-3 Global PCC, with lower Fold-3 MSE, fewer parameters, and stable candidate identity as tie-breakers. Neither Fold 4 nor Fold 5 enters this within-policy selection. The four policy-selected agent sources and three distinct diagnostic sources are then refit under five paired seeds and evaluated on Fold 4; the h0h_{0} source is also displayed under its simple-concatenation role. A second, pre-specified across-policy score rule chooses the agent source with the highest mean Fold-4 PCC for transfer to Fold 5. Within source selection, Fold 4 enters only this disclosed across-policy choice, while Fold 5 never enters. The path-constrained and falsification-guided studies instead carry all ten Fold-3 winners through Folds 4 and 5 without Fold-4 reselection. Before Fold 4, each study freezes the candidate set and registers the source-check rules, input-intervention construction, map seeds, threshold-construction rule, status rule, and aggregation. The realized model-specific thresholds are computed from the 32 reference scores on the fold being audited. For the prediction-score panel, Fold 5 independently refits the frozen source on Folds 1–2, selects its checkpoint on Fold 3 under the same paired-seed schedule, and evaluates Fold 5 once. Thus Fold 5 does not reuse Fold-4 weights and never changes the selected source. LINCS and LKCP subsequently test the final procedure outside BBBC development.
紧接着上一段的搜索步骤,交代实验协议:编辑范围、日志记录、候选如何按 Fold-3 择优、如何在 Fold-4/5 上检验。
b098The compound intervention first seeks a different-fingerprint derangement within rows sharing the same observed control-profile vector. If that class derangement is impossible, a deterministic partition-wide donor with a different fingerprint is used and recorded. Only the compound representation changes: the source control profile and dose remain fixed. The resulting effect is measured under this control-stratified, dose-preserving replacement distribution. The control intervention maps each row to a different observed control-vector class; donor reuse is allowed to ensure that every control vector changes. Dose maps replace only the scalar dose, with observed-dose support checked against the source compound’s released BRD identity and available cell/time annotations. None of these maps uses responses, model predictions, or evaluation scores to choose donors.
1. 这段在干什么
这段在定义三种"干预"的精确操作规则——分别改化合物、改对照、改剂量,并强调选供体时不看响应或模型分数。
2. 需要解释的地方
"不同指纹的错排(derangement)"是领域常识:指每个元素都不留在原位的重新排列,这里要求供体行的"指纹"与原行不同,即不能照抄自己。"对照分层"指只在与源行共享同一对照向量的行里找供体;找不到才退而用全分区供体。"保存剂量的替换分布"指换对照时剂量不动,换剂量时只动标量剂量。
3. 值得留意
最后一句"不使用响应、预测或评分选供体"是作者在防一种质疑:如果选供体时偷看了结果,后面的归因就有循环论证之嫌。另外剂量替换要对照源化合物已发布的 BRD 身份和细胞/时间注释来核验——这是可复现性上的约束。
b099For the primary scaffold-fold compound maps, within-control-group derangements cover 411/411411/411 and 355/355355/355 rows on BBBC036, 4330/43804330/4380 and 3831/38783831/3878 on BBBC047, 847/847847/847 and 868/868868/868 on LINCS, and 378/405378/405 and 378/432378/432 on LKCP, respectively on Folds 4 and 5. The remaining rows use partition-wide donors. These counts specify control matching; the response contrasts retain the source dose, including when the donor compound was not observed at that dose. The plate-family analysis uses its separately specified maps in Appendix A.
这段在干什么:交代主分析中化合物映射的匹配完成度,说明哪些行用了组内错配、哪些改用分组宽供体,并声明对照方式。
需要解释的地方:
值得留意:作者强调对照只换供体、响应仍保留原始剂量(哪怕该剂量下没观测到供体化合物);具体映射规则在附录 A,不在本段。
b100For the 2×22\times 2 audit, the model-specific reference values are derived from the intervened predictions rather than from responses used to choose a permutation. For compound use they are bj,r(chem)=12(mp~c,r+mp~c~,r)b^{(\mathrm{chem})}_{j,r}=\tfrac{1}{2}(m_{\tilde{p}c,r}+m_{\tilde{p}\tilde{c},r}); for control-profile use, bj,r(context)=12(mpc~,r+mp~c~,r)b^{(\mathrm{context})}_{j,r}=\tfrac{1}{2}(m_{p\tilde{c},r}+m_{\tilde{p}\tilde{c},r}); and for their interaction, bj,r(int)=mp~c,r−mp~c~,rb^{(\mathrm{int})}_{j,r}=m_{\tilde{p}c,r}-m_{\tilde{p}\tilde{c},r}. Dose uses the prediction score under a within-compound dose permutation. In every case, τj,q\tau_{j,q} is the 95th percentile of absolute pairwise differences among the 32 fixed reference values. The lower and upper bounds used for the status are the empirical 2.5th and 97.5th percentiles of the 32 effect values; the interaction uses absolute effects. A model is qualified when the lower bound exceeds τj,q\tau_{j,q}, unsupported when the upper bound does not exceed τj,q\tau_{j,q}, and inconclusive otherwise. Inputs without enough legal interventions are reported as not identifiable. For the five-refit prediction-score panel, each refit receives its own status and an aggregate status requires agreement in at least four of five refits; if no category reaches four of five, the aggregate is inconclusive. Path-constrained and falsification-guided results report status counts across ten selected models without a vote. The 50-refit coordinate analysis is descriptive at the refit level and uses the ten trajectory-selected endpoint instances as the policy-distribution units. These instances contain six unique configurations; repeated configurations are retained because they are independent trajectory selections. For the BBBC, LINCS, and LKCP panels, the 32 permutations are repeated measurements within a trained model, not independent sample units. The sci-Plex feedback study uses 32 development maps and 128 final held-out maps, which are likewise repeated measurements; Norman uses a separate set of 32 complete derangements (Appendix D).
这是一段操作手册式的协议说明:交代 2×2 审计中各类"参考值"怎么算、状态怎么判定、以及各面板的统计单元是什么。
"not identifiable"(干预不足)和"inconclusive"是两种不同结果,别混。聚合状态需五次 refit 中至少四次一致。
b101The falsification-guided trainer evaluates α∈{0.05,0.10,0.20,0.35,0.50,0.75,1.00}\alpha\in\{0.05,0.10,0.20,0.35,0.50,0.75,1.00\} on Fold 3. Predictive noninferiority allows at most a 0.0030.003 decrease from the control-only predictor in Global PCC, a 0.0050.005 decrease in either response block, and a 0.00050.0005 increase in MSE. Search qualification uses eight fixed Fold-3 replacement maps. The empirical 2.5th percentile of their compound-related Δℒ\Delta\mathcal{L} values (increases in inverse-standardized joint-output MSE) must exceed both the 95th percentile of absolute pairwise differences among the replaced-input losses and 0.0010.001 times the reference loss. Here reference loss is the current candidate’s Fold-3 MSE with correct inputs. Dose uses the same rule when eligible, with reference loss restricted to eligible rows. These eight development maps are separate from the 32 held-out audit maps. The discovery code groups equal Morgan fingerprints for its replacement maps. Its dose gate requires at least 32 eligible rows and eight fingerprint groups in both the fit partition and Fold 3, with at least 0.10.1 log-dose contrast. For each candidate, epoch/scale selection maximizes Fold-3 Global PCC among qualifying states, breaking ties by lower MSE; if none qualifies, the best predictive state is retained with its qualification status recorded.
1. 这段在干什么
这是附录方法论部分:把「证伪引导训练器」的筛选规则(α 网格、非劣性容差、置换图显著性检验、剂量门槛、epoch/scale 选择)逐条写死,让协议可复现。
2. 需要解释的地方
3. 值得留意
b102Trajectory selection is distinct from this within-candidate checkpoint rule. BBBC047 selects the highest-PCC qualifying candidate, with lower MSE as tie-breaker, or reports no qualified endpoint. BBBC036 prefers predictively noninferior candidates and ranks them by compound-related loss increase, then Global PCC, lower MSE, and stable candidate identity. If none is noninferior, that ordering retains a predictive-fallback endpoint from the executable candidates. Strict qualification is reported separately for BBBC036. These cohort-specific rules select the ten sources before either post-search boundary; held-out results never reselect them.
1. 这段在干什么
承接上句的"候选内检查点规则",说明轨迹选择是另一套规则:不同队列(BBBC036/047)用不同排序挑出十个源。
2. 需要解释的地方
3. 值得留意
排序规则各队列不同,且都发生在"两个边界"之前——即held-out 结果不参与重选,这是防止数据泄漏的关键。
b103Held-out dose testing requires at least 64 eligible rows and coverage of at least 10% of the partition. The prediction-score and path-constrained tests accept any observed within-compound dose change; the falsification-guided test additionally requires eight compounds and an absolute log-dose difference of 0.10.1. Compound identity here is the BRD compound identifier, with sample/batch suffixes removed, rather than a fingerprint-equivalence class. On BBBC047, the resulting coverage is 0/43800/4380 rows on Fold 4 and 14/387814/3878 rows from seven compounds on Fold 5 (0.3610%0.3610\%); neither boundary meets the minimum, and neither contains a legal 0.10.1-log-dose contrast. Dose is therefore not identifiable in the BBBC047 panels (Table 50). Equal-fingerprint grouping alone would merge distinct released identities. On LINCS and LKCP, all replacement dose values in the 32 frozen maps are observed for the source BRD compound and its available cell/time annotations; the original input tensors and effect estimates are retained. The held-out PCC threshold is computed from fold-specific reference-score variation and is separate from the Fold-3 loss gate.
这段在干什么:交代 held-out dose 测试的准入门槛(≥64 行、覆盖 ≥10%),并报告 BBBC047 上 coverage 太低、剂量不可识别,因此不做该检验。
需要解释的地方:
值得留意:Fold 4 是 0 行,Fold 5 只有七个化合物、0.361%,两边都没过线;作者特意说用指纹分组会合并不同 release 身份。
b104The prediction-score discovery study contains 80 ten-slot trajectories: two tasks, four search methods, and ten independent trajectory seeds. Each trajectory’s budget includes the common initial predictor, and discovery intervals treat paired trajectory identifiers as the statistical unit because all policies share the same candidate-training seed schedule within an identifier. Selected-model comparisons use five paired fitting seeds on the relevant held-out fold. Path-constrained and falsification-guided studies retain all ten trajectory winners and report their distributions. All summaries are computed from fixed experiment logs that retain failed slots, repairs, and provider attempts.
这段交代实验的统计口径,是协议的操作性收尾。
术语:
值得留意:所有汇总都基于保留失败槽、修复和 provider 尝试的固定日志——即不清理失败样本,这是审计可复现的关键。
b105The accompanying package separates three reproducibility targets. Numerical source records and deterministic renderers regenerate reported tables. Frozen-source refitting trains new parameters under the recorded data and selection rules. Checkpoint replay takes the selected source, matching fitted weights, prepared data, and replacement specification to recompute predictions and input-use effects by inference. Its standalone feedback entry point verifies checkpoint and data hashes and writes fresh outputs. Biological matrices and fitted weights are supplied as external inputs; tabulated records include checkpoint identities and counterfactual sufficient statistics. The package’s docs/REPRODUCIBILITY.md gives the corresponding commands and required assets.
这段在干什么:交代配套代码包如何实现三种可复现目标,并说明数据、权重等资产从哪来。
需要解释的地方:
值得留意:生物矩阵和拟合权重是外部输入,并不打包在代码里;要完全复现得自己备齐这些资产。
b106Comparator Best@10 Δ\Delta [95% CI]; pp Frontier AUC Δ\Delta [95% CI]; pp BBBC036 AIDE +0.0084 [+0.0044, +0.0124] p=9.638×10−4p=9.638\times 10^{-4} +0.0054 [-0.0003, +0.0111] p=6.157×10−2p=6.157\times 10^{-2} CellForge +0.0151 [+0.0107, +0.0196] p=3.24×10−5p=3.24\times 10^{-5} +0.0106 [+0.0056, +0.0155] p=9.601×10−4p=9.601\times 10^{-4} HarmonyCell +0.0046 [-0.0007, +0.0100] p=8.064×10−2p=8.064\times 10^{-2} +0.0025 [-0.0016, +0.0066] p=1.962×10−1p=1.962\times 10^{-1} BBBC047 AIDE +0.0214 [+0.0105, +0.0324] p=1.636×10−3p=1.636\times 10^{-3} +0.0180 [+0.0078, +0.0282] p=3.18×10−3p=3.18\times 10^{-3} CellForge +0.0181 [+0.0057, +0.0304] p=9.051×10−3p=9.051\times 10^{-3} +0.0159 [+0.0046, +0.0271] p=1.086×10−2p=1.086\times 10^{-2} HarmonyCell +0.0028 [-0.0168, +0.0223] p=7.56×10−1p=7.56\times 10^{-1} +0.0094 [-0.0033, +0.0220] p=1.293×10−1p=1.293\times 10^{-1}
1. 这段在干什么
这是附录 B 的对比结果表:把 AIDE、CellForge、HarmonyCell 三个方法在 BBBC036 / BBBC047 两个数据集上的 Best@10 和 Frontier AUC 提升幅度、置信区间和 p 值列出来,用来支撑正文里"方法确实带来增益"的统计证据。
2. 需要解释的地方
3. 值得留意
HarmonyCell 的多行 Δ 区间都跨 0、p 值偏大(如 0.756、0.129),说明它的增益在这两个数据集上并不显著;CellForge 则普遍区间不含 0。这张表是"用统计说话",别只看 Δ 的正负号。
b107Comparator Δ\Delta Global PCC 95% CI pp nn BBBC036 HarmonyCell -0.0015 [-0.0037, +0.0007] 1.228×10−11.228\times 10^{-1} 5 AIDE +0.0117 [+0.0030, +0.0205] 2.027×10−22.027\times 10^{-2} 5 CellForge +0.0158 [+0.0095, +0.0222] 2.307×10−32.307\times 10^{-3} 5 TabM +0.0048 [+0.0016, +0.0081] 1.445×10−21.445\times 10^{-2} 5 RealMLP +0.0118 [-0.0103, +0.0338] 2.13×10−12.13\times 10^{-1} 5 Standard MLP +0.0193 [+0.0072, +0.0313] 1.128×10−21.128\times 10^{-2} 5 h0h_{0} +0.0207 [+0.0130, +0.0283] 1.733×10−31.733\times 10^{-3} 5 Joint TabR +0.0369 [+0.0343, +0.0395] 2.576×10−62.576\times 10^{-6} 5 Ridge +0.2324 [+0.2301, +0.2347] 1.033×10−91.033\times 10^{-9} 5 BBBC047 HarmonyCell +0.0012 [-0.0083, +0.0108] 7.394×10−17.394\times 10^{-1} 5 AIDE +0.0337 [+0.0283, +0.0390] 6.271×10−56.271\times 10^{-5} 5 CellForge +0.0293 [+0.0256, +0.0330] 2.57×10−52.57\times 10^{-5} 5 TabM +0.0181 [+0.0150, +0.0211] 7.867×10−57.867\times 10^{-5} 5 RealMLP +0.0202 [+0.0154, +0.0250] 3.062×10−43.062\times 10^{-4} 5 Standard MLP +0.0337 [+0.0275, +0.0398] 1.103×10−41.103\times 10^{-4} 5 h0h_{0} +0.0337 [+0.0283, +0.0390] 6.271×10−56.271\times 10^{-5} 5 Joint TabR +0.0500 [+0.0451, +0.0549] 9.339×10−69.339\times 10^{-6} 5 Ridge +0.2541 [+0.2530, +0.2552] 3.077×10−113.077\times 10^{-11} 5
这是两套数据集(BBBC036、BBBC047)上各模型的 ΔGlobal PCC 与显著性结果汇总表,接在上一段显著性数据之后。
在干什么:用表格列出各对比模型相对基线的整体相关变化、置信区间与 p 值,支撑“预测贡献”审计结论。
解释:ΔGlobal PCC 指相对基线整体皮尔逊相关系数的增量,正数代表更好;95% CI 是不确定性区间;p 值越小越可能非偶然;n=5 是重复次数。
留意:Ridge 的 Δ 最大(0.23、0.25)却并非主结论重点;HarmonyCell 在 BBBC047 上几乎为零;RealMLP 的 CI 跨零,接近不显著——别只看点估计。
b108For chemical-generalization contrasts, complete held-out Murcko-scaffold clusters are jointly resampled across all five paired refits, and Global PCC is recomputed from cluster sufficient statistics. This preserves within-scaffold response dependence and quantifies uncertainty over held-out chemical families. The Holm step-down adjustment applies to the three registered discovery-policy contrasts reported in Table 18; the broader frozen-predictor comparison is descriptive and reports its unadjusted paired intervals explicitly (Holm, 1979).
这一段在交代统计推断的决策规则:重采样怎么做、Holm 校正用在哪些对比上、哪些只是描述性报告。
需要解释的地方
值得留意
b109Contrast Evaluation Δ\Delta Global PCC [95% CI] BBBC036 CellScientist −- control-only Audit -0.0002 [-0.0020, +0.0015] CellScientist −- HarmonyCell Audit -0.0014 [-0.0032, +0.0003] score-selected −- h0h_{0} Replication +0.0344 [+0.0195, +0.0514] BBBC047 CellScientist −- control-only Audit +0.0016 [+0.0001, +0.0032] CellScientist −- HarmonyCell Audit +0.0052 [+0.0034, +0.0069] score-selected −- h0h_{0} Replication +0.0366 [+0.0308, +0.0424]
1. 这段在干什么
这是一张对比评估表,报告两个数据集(BBBC036/BBBC047)下三类对比的全局 PCC 及 95% 置信区间,是上一段"报告未校正配对区间"的具体数值落实。
2. 关键概念
3. 值得留意
Replication 行在两个数据集区间均明显为正(+0.0344、+0.0366)且不跨 0;而 Audit 各行区间几乎都贴着 0 甚至跨 0,说明前者才是稳健信号,后者的差异不宜过度解读——这正是上句"descriptive、不做多重校正"的含义。
b110For the central falsification-guided study, effects are first computed from all 32 fixed within-model input permutations. The conditional interval resamples held-out Murcko scaffolds while holding the ten selected models fixed; the hierarchical sensitivity interval jointly resamples selected models and scaffolds. Models and scaffolds are the resampling units, whereas the 32 permutations remain fixed repeated measurements within a model.
这段交代中心证伪研究的统计口径。
1. 这段在干什么:说明效应值与两类区间的计算方式,即“用什么单位重采样”。
2. 需要解释的地方:
3. 值得留意:两类区间差别就在“模型动不动”,这点最易读漏。
b111Quantity Estimate Scaffold 95% CI Model×\timesscaffold 95% CI BBBC036 / Fold 4 Full −- anchor PCC +0.0046 [-0.0042, +0.0165] [-0.0043, +0.0164] Compound PCC drop +0.0091 [+0.0001, +0.0211] [+1.78×10−5+1.78\times 10^{-5}, +0.0213] Compound target-loss gain +0.00039 [+0.00002, +0.00089] [+0.00002, +0.00090] BBBC036 / Fold 5 Full −- anchor PCC -0.0020 [-0.0104, +0.0052] [-0.0102, +0.0053] Compound PCC drop +0.0022 [-0.0050, +0.0090] [-0.0051, +0.0092] Compound target-loss gain +0.00009 [-0.00022, +0.00038] [-0.00022, +0.00039] BBBC047 / Fold 4 Full −- anchor PCC +0.0037 [+0.0018, +0.0060] [+0.0018, +0.0060] Compound PCC drop +0.0053 [+0.0033, +0.0078] [+0.0033, +0.0078] Compound target-loss gain +0.00018 [+0.00011, +0.00026] [+0.00011, +0.00026] BBBC047 / Fold 5 Full −- anchor PCC +0.0028 [+0.0003, +0.0061] [+0.0002, +0.0061] Compound PCC drop +0.0048 [+0.0024, +0.0079] [+0.0023, +0.0079] Compound target-loss gain +0.00017 [+0.00009, +0.00026] [+0.00009, +0.00026]
这段在干什么:这是一张数值明细表,逐条列出两个数据集(BBBC036/BBBC047)在两个 fold 下三种指标的点估计与两种 95% CI。
需要解释的地方:CI 是置信区间,即估计的不确定范围。表里给了两种:Scaffold 与 Model×scaffold——领域常识是,换一种重采样单位,区间会略变,用来检验结论对"统计单位"选择是否敏感。
值得留意:同一指标两列 CI 几乎相同,说明换单位影响很小;但 BBBC036/Fold 5 的 Compound PCC drop 等区间跨越 0,即该 fold 下效应不显著,与 BBBC047 四条几乎都远离 0 形成对比。
b112The larger BBBC047 cohort retains positive hierarchical intervals for all three quantities on both held-out folds: Full−-control-only PCC, compound-shuffle PCC drop, and compound target-loss gain. On BBBC036, Fold-4 compound and target-loss intervals remain positive, whereas Fold-5 intervals include zero. Scaffold resampling therefore shows which compound-family effects persist without selecting models or permutations again.
1. 这段在干什么
报告稳健性检验的结果:把化合物按骨架重新抽样(scaffold resampling)后,哪些效应仍站得住。
2. 需要解释的地方
3. 值得留意
BBBC047(大队列)三个量在两个 held-out fold 上全正;BBBC036 只有 Fold-4 正,Fold-5 跨零——同一数据不同折结论不一致,作者没直说,但这正是稳健性检验要暴露的边界。
b113Trajectory-level Best@10 Global PCC is the primary discovery outcome, with frontier AUC as the co-primary search-efficiency summary. Each contrast pairs the shared trajectory seed and reports its effect estimate, two-sided paired interval, and complete ten-trajectory distribution. Holm adjustment is applied to the three pre-specified policy contrasts within each task and outcome (Table 18). Response blocks, response-sensitive metrics, cost records, code-check outcomes, and input-shuffle effects retain their own estimands and statistical units rather than being pooled into one omnibus claim. Input permutations are repeated transformations within a trained refit and are never treated as independent biological replicates.
这段在干什么:在协议收尾处规定统计口径——主结局、配对方式、多重比较校正,并划清各类指标的统计单元边界。
需要解释的地方:
值得留意:Holm 只压三个预设政策对比,不是全表;输入置换是"同一次重训内的重复变换",绝不算独立生物学重复——这是防伪重复的关键防线。
b114Comparator Best@10 pHolmp_{\mathrm{Holm}} Frontier AUC pHolmp_{\mathrm{Holm}} BBBC036 AIDE 1.9×10−31.9\times 10^{-3} 1.231×10−11.231\times 10^{-1} CellForge 1×10−41\times 10^{-4} 2.9×10−32.9\times 10^{-3} HarmonyCell 8.06×10−28.06\times 10^{-2} 1.962×10−11.962\times 10^{-1} BBBC047 AIDE 4.9×10−34.9\times 10^{-3} 9.5×10−39.5\times 10^{-3} CellForge 1.81×10−21.81\times 10^{-2} 2.17×10−22.17\times 10^{-2} HarmonyCell 7.56×10−17.56\times 10^{-1} 1.293×10−11.293\times 10^{-1}
这段在干什么:这是附录里的一张"统计显著性结果表",报告 AIDE、CellForge、HarmonyCell 三个模型在 BBBC036、BBBC047 两个数据集上,Best@10 与 Frontier AUC 两项指标对比最优基线后的 Holm 校正 p 值。
需要解释的地方:
值得留意:表里数字越小代表差异越显著,但这段没说明以 0.05 为阈值、也没说方向(谁赢),需回正文对照。
b116Table 19 summarizes the competitive discovery setting, and Tables 20 and 21 give the complete trajectory and selected-predictor comparisons. Fixed-model references include the pre-tuned RealMLP (Holzmüller et al., 2024), TabM’s parameter-efficient ensemble (Gorishniy et al., 2025), and a joint-output adaptation of TabR (Gorishniy et al., 2024). Table 23 reports valid candidates, failures, repairs, success by budget, model-provider calls, tokens, GPU training time, and trajectory wall time; failed slots and provider attempts remain charged. The RBDE score provides a joint efficiency summary alongside its separate quality and cost columns.
1. 这段在干什么
这是附录里的表格导读段:说明候选表 19–21、23 各自装了什么,并交代对比基线从哪来。
2. 需要解释的地方
3. 值得留意
成本核算把失败也计入,所以效率对比不是"只算成功者"——这点容易被读漏。
b117Search Best@3 ↑\uparrow Best@5 ↑\uparrow Best@10 ↑\uparrow AUC ↑\uparrow Valid ↑\uparrow Calls Tokens (K) BBBC036 CellScientist 0.2833 ±\pm 0.0075 0.2856 ±\pm 0.0052 0.2924 ±\pm 0.0037 0.2845 ±\pm 0.0046 0.990 ±\pm 0.023 10.50 ±\pm 1.59 51.7 ±\pm 6.4 AIDE 0.2765 ±\pm 0.0036 0.2804 ±\pm 0.0033 0.2840 ±\pm 0.0030 0.2791 ±\pm 0.0023 1.000 ±\pm 0.000 11.10 ±\pm 0.98 41.8 ±\pm 3.8 CellForge 0.2708 ±\pm 0.0016 0.2757 ±\pm 0.0025 0.2772 ±\pm 0.0027 0.2739 ±\pm 0.0015 0.910 ±\pm 0.092 16.30 ±\pm 2.39 44.2 ±\pm 7.9 HarmonyCell 0.2803 ±\pm 0.0037 0.2845 ±\pm 0.0045 0.2877 ±\pm 0.0038 0.2820 ±\pm 0.0036 0.950 ±\pm 0.051 11.80 ±\pm 1.46 38.8 ±\pm 4.7 BBBC047 CellScientist 0.3052 ±\pm 0.0127 0.3084 ±\pm 0.0127 0.3097 ±\pm 0.0119 0.3058 ±\pm 0.0111 0.940 ±\pm 0.077 12.50 ±\pm 3.20 59.2 ±\pm 10.4 AIDE 0.2875 ±\pm 0.0015 0.2882 ±\pm 0.0014 0.2883 ±\pm 0.0012 0.2878 ±\pm 0.0013 0.990 ±\pm 0.023 12.30 ±\pm 1.17 45.2 ±\pm 4.4 CellForge 0.2891 ±\pm 0.0015 0.2901 ±\pm 0.0010 0.2917 ±\pm 0.0011 0.2899 ±\pm 0.0008 0.940 ±\pm 0.090 12.90 ±\pm 2.75 36.1 ±\pm 7.4 HarmonyCell 0.2918 ±\pm 0.0043 0.2965 ±\pm 0.0035 0.3070 ±\pm 0.0083 0.2964 ±\pm 0.0034 0.950 ±\pm 0.051 13.20 ±\pm 1.99 41.6 ±\pm 5.8
这段在干什么:这是一张结果表,报告两个数据集(BBBC036、BBBC047)上四个系统(CellScientist、AIDE、CellForge、HarmonyCell)的检索命中率与成本指标,用来支撑上一句提到的「质量与成本并列的效率总结」。
需要解释的地方:
值得留意:CellScientist 在两项 Best@k 和 AUC 上基本都最高,但 Calls 与 Tokens 并不总是最少——质量与成本存在取舍,这正是表格要并列展示的原因。Valid 一列差距不大,AIDE 甚至为 1.000。
b118Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.
这段在干什么:这是附录中表格的图例说明,告诉你正文字体样式代表什么排名。
需要解释的地方:
值得留意:上一段末尾那串带 ± 的数字就是被排名的候选值——± 后是误差范围,但按这条规则,误差本身不决定名次,只有均值参与比较。
b119Method Best@3 PCC ↑\uparrow Best@5 ↑\uparrow Best@10 ↑\uparrow Best MSE ↓\downarrow CP PCC ↑\uparrow L1000 PCC ↑\uparrow Frontier AUC ↑\uparrow Lift AUC ↑\uparrow BBBC036 CellScientist 0.2833 ±\pm 0.0075 0.2856 ±\pm 0.0052 0.2924 ±\pm 0.0037 0.0607 ±\pm 0.0001 0.2824 ±\pm 0.0015 0.2965 ±\pm 0.0052 0.2845 ±\pm 0.0046 0.0172 ±\pm 0.0046 AIDE 0.2765 ±\pm 0.0036 0.2804 ±\pm 0.0033 0.2840 ±\pm 0.0030 0.0612 ±\pm 0.0001 0.2762 ±\pm 0.0028 0.2875 ±\pm 0.0034 0.2791 ±\pm 0.0023 0.0118 ±\pm 0.0028 CellForge 0.2708 ±\pm 0.0016 0.2757 ±\pm 0.0025 0.2772 ±\pm 0.0027 0.0617 ±\pm 0.0003 0.2687 ±\pm 0.0038 0.2810 ±\pm 0.0032 0.2739 ±\pm 0.0015 0.0066 ±\pm 0.0020 HarmonyCell 0.2803 ±\pm 0.0037 0.2845 ±\pm 0.0045 0.2877 ±\pm 0.0038 0.0609 ±\pm 0.0002 0.2821 ±\pm 0.0034 0.2905 ±\pm 0.0045 0.2820 ±\pm 0.0036 0.0147 ±\pm 0.0038 BBBC047 CellScientist 0.3052 ±\pm 0.0127 0.3084 ±\pm 0.0127 0.3097 ±\pm 0.0119 0.0451 ±\pm 0.0004 0.3351 ±\pm 0.0048 0.2806 ±\pm 0.0212 0.3058 ±\pm 0.0111 0.0201 ±\pm 0.0108 AIDE 0.2875 ±\pm 0.0015 0.2882 ±\pm 0.0014 0.2883 ±\pm 0.0012 0.0458 ±\pm 0.0000 0.3269 ±\pm 0.0028 0.2433 ±\pm 0.0025 0.2878 ±\pm 0.0013 0.0021 ±\pm 0.0014 CellForge 0.2891 ±\pm 0.0015 0.2901 ±\pm 0.0010 0.2917 ±\pm 0.0011 0.0457 ±\pm 0.0001 0.3315 ±\pm 0.0040 0.2449 ±\pm 0.0049 0.2899 ±\pm 0.0008 0.0042 ±\pm 0.0016 HarmonyCell 0.2918 ±\pm 0.0043 0.2965 ±\pm 0.0035 0.3070 ±\pm 0.0083 0.0452 ±\pm 0.0003 0.3373 ±\pm 0.0041 0.2724 ±\pm 0.0151 0.2964 ±\pm 0.0034 0.0107 ±\pm 0.0039
这张表是消融/对比结果:在 BBBC036 和 BBBC047 两个任务上,横向比较四种方法、纵向看八项指标。
需要解释的地方
值得留意
b120Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.
这是表格的排版图例,告诉读者表里加粗、下划线、排名是怎么定的——不是新发现,是读表说明书。
排名是逐个任务或折内部比的,不是全表统一排;加粗=最优,下划线=次优。
b121Model Global PCC ↑\uparrow MSE ↓\downarrow CP PCC ↑\uparrow L1000 PCC ↑\uparrow Parameters Train (s) Selected epoch BBBC036 CellScientist 0.3101 ±\pm 0.0023 0.0586 ±\pm 0.0001 0.3107 ±\pm 0.0025 0.3098 ±\pm 0.0032 4.00M 7.2 ±\pm 1.2 21.8 ±\pm 17.0 AIDE 0.2984 ±\pm 0.0091 0.0593 ±\pm 0.0006 0.3085 ±\pm 0.0043 0.2950 ±\pm 0.0124 4.42M 6.3 ±\pm 1.2 4.8 ±\pm 1.0 CellForge 0.2942 ±\pm 0.0043 0.0593 ±\pm 0.0002 0.3052 ±\pm 0.0048 0.2902 ±\pm 0.0056 3.82M 6.2 ±\pm 1.3 2.0 ±\pm 0.0 HarmonyCell 0.3116 ±\pm 0.0013 0.0586 ±\pm 0.0001 0.3117 ±\pm 0.0024 0.3115 ±\pm 0.0020 2.06M 6.7 ±\pm 1.4 17.2 ±\pm 12.5 h0h_{0} 0.2894 ±\pm 0.0081 0.0597 ±\pm 0.0003 0.2984 ±\pm 0.0048 0.2860 ±\pm 0.0098 1.14M 5.9 ±\pm 1.2 3.2 ±\pm 0.6 Standard MLP 0.2908 ±\pm 0.0114 0.0594 ±\pm 0.0004 0.3058 ±\pm 0.0089 0.2845 ±\pm 0.0125 2.42M 5.8 ±\pm 1.2 1.8 ±\pm 0.6 RealMLP 0.2983 ±\pm 0.0202 0.0592 ±\pm 0.0007 0.3170 ±\pm 0.0129 0.2902 ±\pm 0.0260 3.47M 5.9 ±\pm 0.7 – TabM 0.3053 ±\pm 0.0039 0.0589 ±\pm 0.0001 0.3175 ±\pm 0.0063 0.3004 ±\pm 0.0044 2.37M 11.4 ±\pm 1.7 2.8 ±\pm 1.0 Joint TabR 0.2732 ±\pm 0.0033 0.0604 ±\pm 0.0001 0.2835 ±\pm 0.0036 0.2687 ±\pm 0.0044 2.17M 2.6 ±\pm 0.3 7.6 ±\pm 1.1 Ridge 0.0777 ±\pm 0.0000 0.2048 ±\pm 0.0000 0.0812 ±\pm 0.0000 0.0763 ±\pm 0.0000 4.14M 1.5 ±\pm 0.4 – BBBC047 CellScientist 0.3033 ±\pm 0.0011 0.0457 ±\pm 0.0000 0.3254 ±\pm 0.0010 0.2787 ±\pm 0.0015 1.78M 24.3 ±\pm 3.7 36.6 ±\pm 10.5 AIDE 0.2696 ±\pm 0.0047 0.0468 ±\pm 0.0002 0.3167 ±\pm 0.0062 0.2125 ±\pm 0.0101 1.20M 9.5 ±\pm 0.7 2.0 ±\pm 0.9 CellForge 0.2740 ±\pm 0.0035 0.0467 ±\pm 0.0001 0.3178 ±\pm 0.0061 0.2216 ±\pm 0.0060 2.53M 9.6 ±\pm 0.8 1.8 ±\pm 1.0 HarmonyCell 0.3021 ±\pm 0.0098 0.0458 ±\pm 0.0003 0.3265 ±\pm 0.0016 0.2743 ±\pm 0.0196 2.28M 26.0 ±\pm 9.2 41.8 ±\pm 27.1 h0h_{0} 0.2696 ±\pm 0.0047 0.0468 ±\pm 0.0002 0.3167 ±\pm 0.0062 0.2125 ±\pm 0.0101 1.20M 9.3 ±\pm 0.9 2.0 ±\pm 0.9 Standard MLP 0.2696 ±\pm 0.0059 0.0468 ±\pm 0.0002 0.3167 ±\pm 0.0062 0.2127 ±\pm 0.0067 2.53M 9.6 ±\pm 0.9 1.2 ±\pm 0.6 RealMLP 0.2831 ±\pm 0.0042 0.0463 ±\pm 0.0001 0.3259 ±\pm 0.0038 0.2326 ±\pm 0.0049 3.62M 13.7 ±\pm 0.8 – TabM 0.2852 ±\pm 0.0033 0.0463 ±\pm 0.0001 0.3282 ±\pm 0.0023 0.2350 ±\pm 0.0105 2.51M 16.7 ±\pm 2.1 2.6 ±\pm 0.7 Joint TabR 0.2533 ±\pm 0.0049 0.0475 ±\pm 0.0002 0.2951 ±\pm 0.0050 0.2027 ±\pm 0.0066 2.25M 10.8 ±\pm 1.1 4.6 ±\pm 0.7 Ridge 0.0492 ±\pm 0.0000 0.3880 ±\pm 0.0000 0.0547 ±\pm 0.0000 0.0438 ±\pm 0.0000 4.62M 1.3 ±\pm 0.2 –
这段在干什么:这是附录里的两张性能对比表(BBBC036 和 BBBC047),列出各模型在四个指标、参数量、训练时长上的均值±标准差,用来支撑正文中"搜索轨迹的可靠性与开销"这一部分。
需要解释的地方:
值得留意:Ridge 的 PCC 极低(约 0.05–0.08)但极稳定;HarmonyCell 参数最少(2.06M)却在 BBBC036 上综合最好。标准差本身差异很大,别只看均值排名。
b122Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.
这段在干什么:这是结果表格的图例,用来说明表中加粗、下划线等格式的含义。
需要解释的地方:加粗=该任务/折内表现最好的预测均值,下划线=第二好的,且二者是"不同的"模型;"平局共享排名"意味着并列同格式;区间上下界和诊断/资源列不参与排名。
值得留意:排名只在同一任务或折内比较,不跨任务;资源列(如上一行的 4.62M、1.3±0.2)不会因高低被加粗或划线。
b123Formulation Generated artifact Max output Slots Prediction-score policies Executable source 9000 10 Path-constrained Structured design card 2400 10 Falsification-guided Semantic design card 1800 10
这段是一张表格式汇总,列出四种轨迹生成策略(Formulation / Executable source / Path-constrained / Falsification-guided)各自产出的 artifact 类型、最大输出和 slots 数。
要解释的地方:这张表本身没有表头说明,各列对应关系是靠对齐推测的——比如「9000/2400/1800」「10/10/10」应是两列数值。四种策略与其 artifact 的配对也需靠上下文确认,原文没明说。值得留意的是,这段明显是表格被抽成了纯文本,列对齐已丢失,读时别把数字和策略错配。原文没交代这些数值的含义(是 token 上限?还是别的),这里没提。
b124Shared settings: model deepseek-v4-flash; temperature 0.2; reasoning effort not set; fallback disabled. Max output is the token limit per request.
这段是附录 C 里的共享实验参数声明,交代各条搜索轨迹共用的模型配置,属于复现所需的背景设定。
需要解释的地方:temperature 0.2 指采样随机性低,输出偏稳定;fallback disabled 指模型调用失败时不做备用切换;reasoning effort 是推理投入档位,此处「未设置」。这些是领域常识。
值得留意:上一段的表格在展示不同搜索策略的 token 与步数消耗,而这段把模型固定住了——说明各轨迹之间的差异来自策略本身,而非换了模型。
b125A. Execution quality and budget efficiency
1. 这段在干什么:这是附录 C 下的一个小标题(A 节),用来引出"执行质量与预算效率"这一子话题,起过渡/分节作用。
2. 需要解释的地方:
3. 值得留意:这只是一个标题行,正文尚未展开;它与前段(模型配置、token 上限)的衔接尚未在此段说明,别把它当成结论。
b126Method RBDE ↑\uparrow Valid rate ↑\uparrow Failed slots ↓\downarrow Repairs ↓\downarrow Success@3/5/10 ↑\uparrow BBBC036 CellScientist 0.0104 0.990 ±\pm 0.023 0.10 ±\pm 0.23 1.50 ±\pm 1.59 8/10/10 AIDE 0.0096 1.000 ±\pm 0.000 0.00 ±\pm 0.00 2.10 ±\pm 0.98 9/10/10 CellForge 0.0044 0.910 ±\pm 0.092 0.90 ±\pm 0.92 7.30 ±\pm 2.39 8/10/10 HarmonyCell 0.0102 0.950 ±\pm 0.051 0.50 ±\pm 0.51 2.80 ±\pm 1.46 9/10/10 BBBC047 CellScientist 0.0064 0.940 ±\pm 0.077 0.60 ±\pm 0.77 3.50 ±\pm 3.20 7/8/9 AIDE 0.0008 0.990 ±\pm 0.023 0.10 ±\pm 0.23 3.30 ±\pm 1.17 5/8/9 CellForge 0.0027 0.940 ±\pm 0.090 0.60 ±\pm 0.90 3.90 ±\pm 2.75 8/9/9 HarmonyCell 0.0057 0.950 ±\pm 0.051 0.50 ±\pm 0.51 4.20 ±\pm 1.99 8/10/10
这段在干什么:这是一张结果表,横向对比 CellScientist、AIDE、CellForge、HarmonyCell 四种方法在 BBBC036 和 BBBC047 两个数据集上的执行质量与预算效率。
需要解释的地方:
± 是标准差,表示多次运行的波动。值得留意:AIDE 在 BBBC036 上 Valid rate 为 1.000、Failed slots 为 0,但 RBDE 仅 0.0096,并非最高——说明"运行干净"和"效率高"是两回事。Success@k 跨越 5 到 10 时多有提升,暗示后段尝试仍能救回结果。
b128Method Calls Tokens (K) GPU training (s) Wall time (s) BBBC036 CellScientist 10.50 ±\pm 1.59 51.7 ±\pm 6.4 19.8 ±\pm 2.2 716.3 ±\pm 101.3 AIDE 11.10 ±\pm 0.98 41.8 ±\pm 3.8 16.6 ±\pm 2.5 559.0 ±\pm 212.1 CellForge 16.30 ±\pm 2.39 44.2 ±\pm 7.9 15.7 ±\pm 1.4 578.0 ±\pm 86.4 HarmonyCell 11.80 ±\pm 1.46 38.8 ±\pm 4.7 16.8 ±\pm 1.1 610.8 ±\pm 158.2 BBBC047 CellScientist 12.50 ±\pm 3.20 59.2 ±\pm 10.4 111.7 ±\pm 21.3 821.9 ±\pm 116.0 AIDE 12.30 ±\pm 1.17 45.2 ±\pm 4.4 60.3 ±\pm 3.9 533.0 ±\pm 152.3 CellForge 12.90 ±\pm 2.75 36.1 ±\pm 7.4 56.7 ±\pm 5.8 533.0 ±\pm 83.6 HarmonyCell 13.20 ±\pm 1.99 41.6 ±\pm 5.8 72.9 ±\pm 11.0 680.5 ±\pm 228.0
这份表格报告了四种方法在两个数据集上的搜索开销对比。
1. 这段在干什么:用表格列出各方法在两个数据集(BBBC036、BBBC047)上的运行成本,用于比较搜索效率与可靠性。
2. 需要解释的地方:
3. 值得留意:CellScientist 在 BBBC047 上 GPU 训练时间(111.7s)远高于其他方法,说明其搜索出的模型训练代价更大。
b129CP response block L1000 response block Model PCC @20 ↑\uparrow PCC @50 ↑\uparrow RMSE @20 ↓\downarrow RMSE @50 ↓\downarrow PCC @20 ↑\uparrow PCC @50 ↑\uparrow RMSE @20 ↓\downarrow RMSE @50 ↓\downarrow BBBC036 CellScientist 0.4620 ±\pm 0.0035 0.4454 ±\pm 0.0033 0.2411 ±\pm 0.0005 0.2385 ±\pm 0.0004 0.2828 ±\pm 0.0051 0.3167 ±\pm 0.0043 0.5135 ±\pm 0.0008 0.4497 ±\pm 0.0007 AIDE 0.4636 ±\pm 0.0024 0.4445 ±\pm 0.0025 0.2408 ±\pm 0.0006 0.2386 ±\pm 0.0004 0.2611 ±\pm 0.0200 0.3005 ±\pm 0.0160 0.5190 ±\pm 0.0054 0.4540 ±\pm 0.0048 CellForge 0.4607 ±\pm 0.0050 0.4430 ±\pm 0.0034 0.2412 ±\pm 0.0006 0.2388 ±\pm 0.0005 0.2598 ±\pm 0.0081 0.2966 ±\pm 0.0099 0.5176 ±\pm 0.0011 0.4531 ±\pm 0.0013 HarmonyCell 0.4633 ±\pm 0.0023 0.4464 ±\pm 0.0024 0.2409 ±\pm 0.0005 0.2383 ±\pm 0.0004 0.2823 ±\pm 0.0040 0.3170 ±\pm 0.0031 0.5137 ±\pm 0.0006 0.4497 ±\pm 0.0004 h0h_{0} 0.4526 ±\pm 0.0034 0.4357 ±\pm 0.0032 0.2426 ±\pm 0.0005 0.2399 ±\pm 0.0005 0.2555 ±\pm 0.0079 0.2983 ±\pm 0.0034 0.5210 ±\pm 0.0056 0.4545 ±\pm 0.0031 Standard MLP 0.4608 ±\pm 0.0077 0.4404 ±\pm 0.0070 0.2412 ±\pm 0.0011 0.2391 ±\pm 0.0009 0.2473 ±\pm 0.0048 0.2923 ±\pm 0.0111 0.5197 ±\pm 0.0006 0.4536 ±\pm 0.0017 RealMLP 0.4668 ±\pm 0.0081 0.4497 ±\pm 0.0080 0.2402 ±\pm 0.0013 0.2378 ±\pm 0.0011 0.2476 ±\pm 0.0332 0.2945 ±\pm 0.0320 0.5200 ±\pm 0.0053 0.4540 ±\pm 0.0047 TabM 0.4640 ±\pm 0.0023 0.4477 ±\pm 0.0021 0.2408 ±\pm 0.0005 0.2382 ±\pm 0.0003 0.2603 ±\pm 0.0079 0.3066 ±\pm 0.0033 0.5180 ±\pm 0.0031 0.4518 ±\pm 0.0016 Joint TabR 0.4433 ±\pm 0.0041 0.4235 ±\pm 0.0045 0.2437 ±\pm 0.0005 0.2414 ±\pm 0.0006 0.2389 ±\pm 0.0124 0.2812 ±\pm 0.0120 0.5231 ±\pm 0.0024 0.4567 ±\pm 0.0020 Ridge 0.1172 ±\pm 0.0000 0.1153 ±\pm 0.0000 0.4678 ±\pm 0.0000 0.4603 ±\pm 0.0000 0.0807 ±\pm 0.0000 0.0903 ±\pm 0.0000 0.9852 ±\pm 0.0000 0.8401 ±\pm 0.0000 BBBC047 CellScientist 0.4003 ±\pm 0.0028 0.4061 ±\pm 0.0018 0.2689 ±\pm 0.0002 0.2547 ±\pm 0.0002 0.3262 ±\pm 0.0025 0.3154 ±\pm 0.0019 0.3703 ±\pm 0.0004 0.3250 ±\pm 0.0003 AIDE 0.4127 ±\pm 0.0124 0.4078 ±\pm 0.0083 0.2676 ±\pm 0.0022 0.2549 ±\pm 0.0015 0.2577 ±\pm 0.0090 0.2486 ±\pm 0.0060 0.3789 ±\pm 0.0009 0.3323 ±\pm 0.0006 CellForge 0.4117 ±\pm 0.0095 0.4069 ±\pm 0.0032 0.2678 ±\pm 0.0010 0.2550 ±\pm 0.0007 0.2684 ±\pm 0.0051 0.2583 ±\pm 0.0041 0.3776 ±\pm 0.0009 0.3312 ±\pm 0.0007 HarmonyCell 0.4029 ±\pm 0.0049 0.4079 ±\pm 0.0026 0.2686 ±\pm 0.0009 0.2545 ±\pm 0.0004 0.3234 ±\pm 0.0156 0.3130 ±\pm 0.0148 0.3705 ±\pm 0.0021 0.3252 ±\pm 0.0016 h0h_{0} 0.4127 ±\pm 0.0124 0.4078 ±\pm 0.0083 0.2676 ±\pm 0.0022 0.2549 ±\pm 0.0015 0.2577 ±\pm 0.0090 0.2486 ±\pm 0.0060 0.3789 ±\pm 0.0009 0.3323 ±\pm 0.0006 Standard MLP 0.4125 ±\pm 0.0157 0.4083 ±\pm 0.0089 0.2678 ±\pm 0.0023 0.2548 ±\pm 0.0013 0.2626 ±\pm 0.0135 0.2531 ±\pm 0.0123 0.3782 ±\pm 0.0018 0.3316 ±\pm 0.0015 RealMLP 0.4142 ±\pm 0.0072 0.4137 ±\pm 0.0043 0.2667 ±\pm 0.0010 0.2536 ±\pm 0.0005 0.2883 ±\pm 0.0034 0.2798 ±\pm 0.0028 0.3750 ±\pm 0.0004 0.3287 ±\pm 0.0003 TabM 0.4152 ±\pm 0.0107 0.4142 ±\pm 0.0050 0.2666 ±\pm 0.0015 0.2536 ±\pm 0.0007 0.2784 ±\pm 0.0051 0.2685 ±\pm 0.0067 0.3763 ±\pm 0.0010 0.3300 ±\pm 0.0008 Joint TabR 0.3751 ±\pm 0.0078 0.3810 ±\pm 0.0061 0.2757 ±\pm 0.0007 0.2599 ±\pm 0.0006 0.2482 ±\pm 0.0101 0.2409 ±\pm 0.0084 0.3806 ±\pm 0.0012 0.3336 ±\pm 0.0009 Ridge 0.0833 ±\pm 0.0000 0.0808 ±\pm 0.0000 0.7616 ±\pm 0.0000 0.7251 ±\pm 0.0000 0.0583 ±\pm 0.0000 0.0583 ±\pm 0.0000 1.1990 ±\pm 0.0000 1.0275 ±\pm 0.0000
这张表是附录里的多数据集性能对比:BBBC036 与 BBBC047 两个数据集,各列 8 个指标——PCC@20/50、RMSE@20/50,前者在 CP 反应块、后者在 L1000 反应块下各算一遍。
关键概念:PCC 是预测与真实的相关性(越大越好,表头↑),RMSE 是误差(越小越好,表头↓)。"@20/@50"指只取前 20 或 50 个预测来算——这是领域常识里的 top-k 评估,不是本文定义。每个数字后的 ± 是多次运行的波动。
值得留意:Ridge 在 PCC 上远低于其余模型(如 0.1172),RMSE 却反常偏大,说明它基本没学到反应关系;而各 Cell 类模型之间差距很小。表中未见论文对此的具体解读。
b130Task Score PCC Path PCC Δ\Delta Tokens (K) Valid Repairs Linked BBBC036 0.2924 ±\pm 0.0037 0.2925 ±\pm 0.0019 +0.0001 21.9 ±\pm 1.4 100.0% 0.00 ±\pm 0.00 100/100 BBBC047 0.3097 ±\pm 0.0119 0.3008 ±\pm 0.0012 -0.0090 21.7 ±\pm 1.2 100.0% 0.00 ±\pm 0.00 100/100
这段是附录里的轨迹成本表,报告两个数据集在两种打分下的表现与代价。
1. 在干什么:用数据展示搜索轨迹的可靠性——任务分与路径分几乎一致,且100次全部有效、无需修复。
2. 解释:
3. 值得留意:BBBC047的Δ是负的(-0.0090),路径分反而略高于任务分,但作者未对此评论;两行均报100/100,说明全程零失败。
b132Boundary Predictor Global PCC ↑\uparrow Input Δ\DeltaPCC ↑\uparrow Support Fold 4 audit GEARS 0.6819 [0.6714, 0.6924] 0.3029 5/5 Initial model (h0h_{0}) 0.8960 [0.8909, 0.9011] 0.5425 5/5 CellScientist 0.9020 [0.8988, 0.9051] 0.5578 5/5 Fold 5 replication GEARS 0.6676 [0.6568, 0.6784] 0.3599 5/5 Initial model (h0h_{0}) 0.9051 [0.9023, 0.9079] 0.6193 5/5 CellScientist 0.9096 [0.9051, 0.9142] 0.6280 5/5
这张表是在做跨折叠复现检验:把 GEARS、初始模型 h₀、CellScientist 三者在 Fold 4 和 Fold 5 上的成绩并排摆出来。
需要解释的地方:
值得留意:两折中 h₀ 和 CellScientist 几乎持平(0.896 vs 0.902,0.905 vs 0.910),提升幅度很小——作者没明说,但这个"微弱优势"本身值得打个问号。
b133The transfer task is reconstructed from the count layer of the Norman K562 CRISPRa Perturb-seq experiment (GEO GSE133344) (Norman et al., 2019). Its source matrix contains 91,205 cells and 5,045 genes, including 7,353 control cells. Per-cell counts are library-size normalized to 10,000 and log-transformed, after which canonical condition means are centered by the control mean. The 284 source labels resolve to 237 canonical conditions: 105 observed singles, 131 doubles, and one control. The common discovery interface receives a pair-input bundle containing a 105-dimensional multi-hot perturbation identity and its matching mean single-perturbation anchor, and returns the full 5,045-gene response without external pathway features or cell-type covariates. GEARS retains its task-native graph inputs as described below.
1. 这段在干什么
交代迁移任务的来源与数据规模,并说明「通用发现接口」拿到什么输入、吐什么输出。
2. 需要解释的地方
3. 值得留意
b134The five-fold assignment is constructed without expression responses and balances pair count, source-cell count, and perturbation-gene incidence. All singles and controls remain fit-only with the Fold-1/2 doubles; Fold 3 supplies discovery feedback, and Folds 4 and 5 are successive held-out evaluation folds. Table 27 lists every held-out double-perturbation block. The comparison predicts unseen double combinations whose constituent single-gene responses are available during fitting.
1. 这段在干什么
交代数据划分方式:五折怎么分、哪几折用来训练(fit-only)、哪折给发现反馈、哪几折做留出评估——为下文"预测未见过的双扰动组合"设定前提。
2. 需要解释的地方
3. 值得留意
b135Each held-out boundary contains 26 unseen double-perturbation conditions. Before Fold 4 is opened, we construct 32 distinct complete derangements of the 26 row positions; each is a one-to-one permutation with no fixed point and is reused positionally on Fold 5. For each source condition, the registered intervention replaces both coordinates of the pair-input bundle (its 105-dimensional pair-identity vector and matching mean single-perturbation anchor) with the donor pair’s values while retaining the source response target. Replacing both coordinates prevents a residual predictor from retaining pair information through the anchor; consequently, the test measures dependence on the complete bundle rather than attributing the effect to the multi-hot coordinate alone. The maps use neither responses nor model outputs and do not require donor pairs to be gene-disjoint from source pairs.
这段在干什么
交代审计实验的操作细节:说清如何用错位置换构造“供体”干预,把双扰动测试从单扰动锚点里剥离出来。
需要解释的地方
值得留意
b136For refit jj and map rr, let LjL_{j} be joint-output MSE under the correct pair-input bundle and Lj,rshufL^{\mathrm{shuf}}_{j,r} the MSE after the bundle shuffle. We define gj,r=Lj,rshuf−Ljg_{j,r}=L^{\mathrm{shuf}}_{j,r}-L_{j} and g~j,r=gj,r/max(Lj,10−12)\tilde{g}_{j,r}=g_{j,r}/\max(L_{j},10^{-12}). A refit is qualified under the pair-input-use criterion when 32−1∑rg~j,r>0.00132^{-1}\sum_{r}\tilde{g}_{j,r}>0.001 and Q0.05({gj,r}r=132)>0Q_{0.05}(\{g_{j,r}\}_{r=1}^{32})>0. It is unsupported when the first quantity lies within [−0.001,0.001][-0.001,0.001] and inconclusive otherwise. The Global-PCC drop is reported as a continuous effect but does not enter this categorical rule. The 32 maps are repeated interventions within a refit; the five refits are the statistical units for model-performance intervals and support counts.
1. 这段在干什么
给出「配对输入使用」判据的量化定义:用打乱 bundle 后的 MSE 变化,判断模型是否真依赖该输入。
2. 需要解释的地方
3. 值得留意
32 个 map 只是 refit 内的重复,真正的统计单位是 5 个 refit——这影响后面区间的可信度,别把 32 当独立样本。
b137GEARS is evaluated through the shared task boundary while retaining its original multigene-perturbation formulation (Roohani et al., 2024): GO and coexpression graphs, native GNN/decoder, optimization objective, and validation-selected checkpoint. It shares count normalization, response target, fold roles, and held-out metrics with the other predictors. Table 28 records the architecture, graph scope, optimizer, and checkpoint rule used in the comparison.
交代 GEARS 作为对比模型的评测口径:保留其原有多基因扰动设定,只共享任务边界。
作者强调"保留原formulation"与"共享fold、指标",是想说明差异只来自方法本身。具体数值不在本段,见 Table 28。
b138Fold Role Double perturbations Source cells Component genes 1 Fit 26 7,301 40 2 Fit 26 7,849 41 3 Discovery/selection 27 6,817 42 4 Independent audit 26 6,541 40 5 Final replication 26 6,937 38
这段在干什么:用 Table 28 列出五折交叉验证中每一折的角色与数据规模,说明评估协议的划分。
需要解释的地方:
值得留意:审核与复现用的折完全不参与拟合和筛选,这是保证"独立"的关键;数据量在不同折间略有波动,并非均分。
b139Across ten trajectories, all 77 nonduplicate CellScientist candidates pass source, shape, and training checks; the 23 duplicate design cards still consume their candidate slots, and no Fold-4 or Fold-5 response enters proposal or selection. The trajectories account for 154 provider calls, 403,762 tokens, 17.8 H100 fit-minutes, and 50.5 minutes of end-to-end time. Held-out input-use testing is deterministic and uses zero model-provider calls.
承接上一段的逐轨迹统计表,汇总十条轨迹的算力/调用开销,并交代候选卡的去重与检查结果。
(约 140 字)
b140A. Search performance BB Best@3 Best@5 Best@10 Frontier AUC Δh0\Delta h_{0} [95% CI] 10 0.9192 0.9201 0.9210 0.9194 +0.0062 [+0.0036, +0.0087]
这段是附录 D 的结果表(表格被压成一行),汇报组合遗传扰动预测中「搜索性能」的指标。
1. 在干什么:给出不同预算下 Best@k 与 Frontier AUC 等数值,并报告 Δh₀ 及其 95% 置信区间。
2. 概念:Best@3/5/10 指前 3/5/10 名里命中真值的表现(领域常识,具体定义这段没提);AUC 是排序质量指标;Δh₀ 是某种效应量,正值表示提升,方括号是置信区间。
3. 值得留意:Δh₀ 的 CI 下限 +0.0036 > 0,即提升在统计上站得住,不是噪声;但各项 Best@k 数值几乎持平,别误读成有明显梯度。
b141B. Completion and resources Positive Executed Calls Tokens 10/10 77/100 154 403,762
这段在干什么:这是附录里的资源统计行,汇报跑完这套流程(Completion)的调用次数与消耗,紧接上一段模型指标(Best@k、AUC)之后。
需要解释的地方:Positive Executed Calls 10/10 指十次调用全部成功(领域常识:这类"成功数/总数"常用来表示执行可靠性);Tokens 77/100 指 token 用量七十七,上限或预算一百(具体含义这段没说明)。
值得留意:它只有指标名和数字,没有单位和说明,77/100 到底是用到上限还是达到某阈值,单看这行读不出来。
b142Predictor Global PCC ↑\uparrow MSE ↓\downarrow Condition PCC ↑\uparrow Rank corr. ↑\uparrow DEG-20 PCC ↑\uparrow DEG-50 PCC ↑\uparrow Full-state PCC ↑\uparrow Input-use support Fold 4 audit Matching mean 0.8810 0.0059 0.8966 0.4902 0.9702 0.9634 0.9915 supported Additive singles 0.8810 0.0046 0.8966 0.4902 0.9702 0.9634 0.9932 supported Ridge 0.8989 0.0033 0.8965 0.4584 0.9724 0.9665 0.9951 supported GEARS 0.6819 0.0122 0.6826 0.2512 0.9078 0.8915 0.9843 5/5 h0h_{0} 0.8960 0.0037 0.8999 0.4288 0.9664 0.9618 0.9945 5/5 CellScientist 0.9020 0.0039 0.9098 0.4345 0.9750 0.9697 0.9943 5/5 Fold 5 replication Matching mean 0.8982 0.0043 0.8799 0.4681 0.9614 0.9469 0.9938 supported Additive singles 0.8982 0.0033 0.8799 0.4681 0.9614 0.9469 0.9952 supported Ridge 0.8959 0.0029 0.8813 0.4375 0.9672 0.9546 0.9957 supported GEARS 0.6676 0.0115 0.6156 0.2267 0.8021 0.8096 0.9853 5/5 h0h_{0} 0.9051 0.0027 0.8844 0.4065 0.9542 0.9474 0.9961 5/5 CellScientist 0.9096 0.0028 0.8944 0.4071 0.9650 0.9554 0.9960 5/5
在附录里汇报「输入使用声明」的审计结果:把各预测器在 Fold 4 与 Fold 5 上的多项指标(PCC、MSE、DEG 等)列出,并给出支持与否的判定。
GEARS 两折都拿到「5/5」支持,但它的 Global PCC、MSE、Rank corr. 在各行里最差,说明支持认定并不等于整体预测最优。
b143The matched transfer analysis gives CellScientist, AIDE, CellForge, and HarmonyCell the same initial model, design language, Fold-3 feedback, physical budget, and model-selection rule; each policy generates its own candidates. CellScientist has the highest mean Best@10 and frontier AUC. Its paired Best@10 effect is +0.0029+0.0029 versus AIDE (95% CI [−0.0003,+0.0061][-0.0003,+0.0061]), +0.0058+0.0058 versus CellForge ([+0.0037,+0.0080][+0.0037,+0.0080]), and +0.0006+0.0006 versus HarmonyCell ([−0.0005,+0.0016][-0.0005,+0.0016]). After selection, all sources are refit under five paired seeds and audited without provider calls. Global-PCC ordering varies across folds, while all four selected models pass the pair-input-bundle test in 5/5 refits; CellScientist leads condition-, DEG-20-, and DEG-50-PCC at Fold 5.
这段接在上表之后,做配对效应量的文字总结。
在干什么:交代四个智能体用完全相同的初始模型、预算等条件,只放大差异在“各自生成候选”,然后比出 CellScientist 最强。
关键概念:Best@10 是取前 10 个候选里最好的那个分数;frontier AUC 衡量性能与预算权衡曲线下的面积;“配对”指同一初始条件下两两比较,比独立比更公平。95% CI 不含 0 才算显著优于对方。
值得留意:对 AIDE(区间含 −0.0003)和对 HarmonyCell(含 −0.0005)的区间都跨 0,即优势不稳健;只有对 CellForge 明显领先。后半段“无 provider 调用重训 5 次”的审计,四个模型全部 5/5 通过。
b144Policy Best@3 ↑\uparrow Best@5 ↑\uparrow Best@10 ↑\uparrow Frontier AUC ↑\uparrow Δh0\Delta h_{0} [95% CI] ↑\uparrow Positive ↑\uparrow Executed slots ↑\uparrow Calls ↓\downarrow Tokens ↓\downarrow CellScientist 0.9192 0.9201 0.9210 0.9194 +0.0062 [+0.0036,+0.0087] 10/10 77/100 154 403,762 AIDE 0.9164 0.9168 0.9181 0.9170 +0.0033 [+0.0011,+0.0054] 9/10 91/100 120 335,610 CellForge 0.9151 0.9152 0.9152 0.9151 +0.0003 [−-0.0004,+0.0010] 1/10 97/100 108 309,742 HarmonyCell 0.9191 0.9203 0.9204 0.9193 +0.0056 [+0.0028,+0.0084] 9/10 88/100 117 328,705
这是一张性能对比表,横向比四种方法(CellScientist、AIDE、CellForge、HarmonyCell)在组合基因扰动预测任务上的表现,是上一段文字结论的数据支撑。
作者用两个维度同时评判:不只是分数(Best@k、AUC),还有成本(Calls、Tokens)。CellScientist 分数最高但最贵(154 次调用、40 万 token);CellForge 最省、阳性却只有 1/10,Δh₀ 置信区间还跨 0——等于没提升。
b145Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.
1. 这段在干什么
这是表格的图例说明,交代表里加粗、下划线以及各列的排名规则——属于读数前的"规则声明"。
2. 需要解释的地方
3. 值得留意
上一段那串数字同时含预测值、区间、排名和资源开销,但套用本段规则时,只有预测均值参与加粗/下划线的名次,其余列一律不排。"distinct"一词说明重复值不计入前二。
b146A. Selected architecture Policy Architecture Width Depth Dropout PCC–MSE weight CellScientist Gated MLP 512 4 0.1 0.3 AIDE Gated MLP 512 4 0.2 0.2 CellForge Residual MLP 256 3 0.1 0.1 HarmonyCell Residual MLP 384 4 0.2 0.1
1. 这段在干什么
这是一张"选定架构"配置表,列出四个框架(CellScientist、AIDE、CellForge、HarmonyCell)各自用的网络结构和超参数,为后文比较"输入使用声明"提供统一的模型底座。
2. 需要解释的地方
3. 值得留意
四个框架架构并不统一(前两个 Gated、后两个 Residual),宽度深度也各异——即"选定的"基线本身就有差异,不是同一模型换名字。
b147B. Optimization and selection Policy Optimizer / learning rate Parameters Fold-3 PCC CellScientist AdamW / 0.001 5,850,037 0.9237 AIDE AdamW / 0.0005 5,850,037 0.9234 CellForge AdamW / 0.001 1,850,549 0.9195 HarmonyCell AdamW / 0.001 3,758,261 0.9247
1. 这段在干什么
紧接上一段的架构设置,这张表报告了四个模型(CellScientist、AIDE、CellForge、HarmonyCell)的优化与选择策略,即各自用什么优化器、学习率、参数量和 Fold-3 上的 PCC 表现。
2. 需要解释的地方
3. 值得留意
四个模型参数量不同(CellForge 最少,1,850,549),但 PCC 却非常接近(0.9195–0.9247),差距很小。表头写「B. Optimization and selection Policy」,但标题实际所在小节是 Appendix D,编号对不上——这段没解释原因。
b148Policy Global PCC ↑\uparrow [95% CI] MSE ↓\downarrow Condition PCC ↑\uparrow DEG-20 PCC ↑\uparrow DEG-50 PCC ↑\uparrow Bundle drop Bundle-use support Fold 4 audit CellScientist 0.9020 [0.8988,0.9051] 0.00388 0.9098 0.9750 0.9697 0.5578 5/5 AIDE 0.8981 [0.8920,0.9042] 0.00407 0.9038 0.9682 0.9637 0.5458 5/5 CellForge 0.8960 [0.8909,0.9011] 0.00372 0.8999 0.9664 0.9618 0.5425 5/5 HarmonyCell 0.8960 [0.8920,0.9000] 0.00369 0.9019 0.9693 0.9641 0.5506 5/5 Fold 5 replication CellScientist 0.9096 [0.9051,0.9142] 0.00277 0.8944 0.9650 0.9554 0.6280 5/5 AIDE 0.9125 [0.9088,0.9161] 0.00285 0.8912 0.9567 0.9488 0.6270 5/5 CellForge 0.9051 [0.9023,0.9079] 0.00267 0.8844 0.9542 0.9474 0.6193 5/5 HarmonyCell 0.9088 [0.9029,0.9148] 0.00255 0.8867 0.9597 0.9505 0.6294 5/5
这段在干什么:承接上表的超参配置,给出四个模型在 Fold 4 审计与 Fold 5 复现下的预测性能与输入使用情况。
需要解释的地方:PCC 是预测与真实的相关系数,越高越好;MSE 是均方误差,越低越好;方括号是 95% 置信区间。DEG-20/50 指只看表达变化最大的基因时的 PCC。Bundle drop 是去掉整个输入特征组后性能掉了多少,掉得越多说明用得越狠;Bundle-use support 是几次实验里该组被用上的次数(5/5 即全用上)。
值得留意:Fold 5 里 AIDE 的 PCC 最高,但 Bundle drop 反而略低;CellScientist 在 Fold 4 领先,到 Fold 5 就不是第一了。
b149Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.
这段是表格的图例说明,紧接上一条数据行,告诉你表里那些加粗/下划线是干嘛的。
这段在干什么:定义表格排版规则——每个 task 或 fold 内,预测均值最好/次好的分别加粗、加下划线;打平则并列同名次;区间和诊断、资源类列不参与排名。
需要解释的地方:
值得留意:排名只在组内有效,别跨任务横向比粗细;区间(方括号那项)和 5/5 这类列不排名,粗体与它们无关。
b150Comparator Δ\Delta Global PCC [95% CI] PCC wins Δ\Delta bundle drop [95% CI] Fold 4 audit AIDE +0.0039 [−-0.0044,+0.0122] 4/5 +0.0121 [−-0.0004,+0.0245] CellForge +0.0060 [−-0.0008,+0.0128] 4/5 +0.0154 [+0.0030,+0.0277] HarmonyCell +0.0060 [−-0.0009,+0.0129] 4/5 +0.0072 [−-0.0061,+0.0205] Fold 5 replication AIDE −-0.0028 [−-0.0103,+0.0046] 2/5 +0.0010 [−-0.0059,+0.0078] CellForge +0.0045 [−-0.0004,+0.0094] 5/5 +0.0087 [+0.0037,+0.0137] HarmonyCell +0.0008 [−-0.0056,+0.0072] 4/5 −-0.0015 [−-0.0131,+0.0101]
这段是全篇审计结果的总表:3 个基线模型(AIDE、CellForge、HarmonyCell)在 Fold 4 与 Fold 5 上的预测增益和「bundle drop」。
关键概念(领域常识):PCC 是预测值与真实值的相关性,越高越好;95% CI 是置信区间,跨 0 就说明该增益不显著;"wins" 指 5 个任务里赢了几次;"bundle drop" 指移除一组输入特征后性能掉多少,掉得越多说明这组特征越被真正依赖。
值得留意:Fold 4 上三家 ΔPCC 的区间都跨 0,即都不显著,但 bundle drop 里只有 CellForge 的区间不跨 0——增益不显著 ≠ 特征不被依赖,这正是"审计"的意义。Fold 5 上 AIDE 的 ΔPCC 变成了负的,5 次只赢 2 次。
b152After the final falsification-guided procedure is fixed, LINCS–Pilot1 provides an out-of-development-cohort test. It uses an observed control profile, compound identity, and dose to predict a joint 2,719-dimensional morphology–transcription response across 4,155 matched conditions and 1,108 Murcko groups. The ten trajectory-selected design instances are then fixed for transfer to cpg0004/LKCP/2017_12_05_Batch2, an independently acquired Cell Painting release (Weisbart et al., 2024). Its 134 files contain 11,904 rows; 8,958 treatment and 2,946 control rows yield 1,728 retained conditions and 64 compound-held-out groups. Response- and fold-blind interface quality control removes 110 non-finite or extreme release features from 1,781 common morphology columns, retaining a 1,671-dimensional same-plate control profile and response.
这段在干什么:介绍独立获取的外部验证集——把已固定的设计实例从 LINCS–Pilot1 迁移到另一套独立采集的 Cell Painting 数据上测试。
需要解释的地方:
值得留意:整段几乎全是数据规模与清洗细节(4,155、1,108、11,904、1,728、64 等),作者刻意把"外部独立性"铺垫得很重——这是判断小分子结论能否泛化的关键设计。
b153The scientific question and model designs precede LKCP transfer. The final dataset adapter is fixed after interface quality control and before LKCP training or held-out-fold access. The transferred designs, input-permutation construction, map seeds, threshold-construction rule, and decision rule are registered and hash-bound at that point. Each design is then refit under five paired seeds with no further source edit or model reselection. Compound permutations hold dose fixed; dose permutations use only observed doses of the same compound, require at least a 0.10.1 log-dose contrast, and retain the same eligibility rule. Across 22 cohorts ×\times 22 folds ×\times 1010 designs ×\times 55 seeds, the analysis contains 200 cohort–fold–design–seed refit boundaries. Each boundary reuses the 32 registered ordered compound–dose permutation pairs, yielding 6,400 matched within-refit permutation units; these maps are repeated measurements rather than independent statistical units. For U11U_{11} with both inputs correct, U01U_{01} with compound shuffled, U10U_{10} with dose shuffled, and U00U_{00} with both shuffled, the two-player allocation is ϕchem=12[(U11−U01)+(U10−U00)]\phi_{\mathrm{chem}}=\tfrac{1}{2}[(U_{11}-U_{01})+(U_{10}-U_{00})] and ϕdose=12[(U11−U10)+(U01−U00)]\phi_{\mathrm{dose}}=\tfrac{1}{2}[(U_{11}-U_{10})+(U_{01}-U_{00})]. These allocate the contrast U11−U00U_{11}-U_{00}; adding the residual U00−UanchorU_{00}-U_{\mathrm{anchor}} gives Full−-control-only to numerical precision. The component allocations thus refer to the fixed replacement baseline, while Full−-control-only measures the separately fitted predictors’ performance difference.
交代这套小分子泛化检验的分析协议与记账方式:先冻结设计、再重拟合,并用置换组合把预测贡献拆成「化合物」和「剂量」两块。
作者自己点明:这些置换图是重复测量,不是独立统计单元,别按 6,400 个样本算显著性。此外 φ 是相对固定的替换基线,而 Full−control-only 是另外单独拟合的,两者口径不同。
b154A. LINCS–Pilot1 Contrast Fold 4 [95% CI] Fold 5 [95% CI] Full−-control-only +0.0293 [+0.0281, +0.0305] +0.0233 [+0.0226, +0.0241] Compound identity +0.0175 [+0.0164, +0.0186] +0.0126 [+0.0117, +0.0134] Dose +0.0486 [+0.0473, +0.0498] +0.0467 [+0.0456, +0.0478] Both-shuffled residual -0.0368 [-0.0378, -0.0358] -0.0359 [-0.0370, -0.0349] Interaction +0.0050 [+0.0033, +0.0067] +0.0042 [+0.0027, +0.0058]
这段在报告 LINCS–Pilot1 对比实验在 Fold 4/5 两个交叉验证折上的结果。
术语:Fold 指交叉验证的折。「Full−control-only」是把完整模型减去只留对照的基线,衡量各预测器单独拟合后的性能差;「Compound identity」「Dose」分别代表化合物身份、剂量两类输入;「Both-shuffled residual」是两者都打乱后的残差;「Interaction」是二者的交互项。
留意:Dose 贡献最大(约 +0.049/+0.047),远高于 Compound identity;打乱后残差为负;交互项很小。
b155B. LKCP Batch 2 Contrast Fold 4 [95% CI] Fold 5 [95% CI] Full−-control-only +0.0104 [+0.0103, +0.0106] +0.0211 [+0.0205, +0.0217] Compound identity -0.00007000 [-0.00028998, +0.00014998] +0.0033 [+0.0027, +0.0039] Dose +0.0237 [+0.0228, +0.0247] +0.0341 [+0.0331, +0.0350] Both-shuffled residual -0.0132 [-0.0140, -0.0124] -0.0163 [-0.0170, -0.0157] Interaction +0.00004049 [-0.00001503, +0.00009601] +0.0002 [+0.0001, +0.0003]
接着上一段看,这张表在逐项拆 LKCP Batch 2 里各种“打乱对照”的结果。
1. 这段在干什么:补上 Batch 2 的 Fold 4、Fold 5 两列对照数据,量每种扰动对预测贡献的影响。
2. 需要解释的地方:各行是不同“打乱/对照”手法——只留对照、打乱化合物身份、打乱剂量、两者都打乱后看残差、以及交互项。方括号里是 95% 置信区间,即估计的不确定范围。数值是相对基准的“预测贡献变化”。
3. 值得留意:Compound identity 一行两列都极接近 0,且区间跨 0,说明打乱化合物身份几乎不改变贡献;Dose 一行数值最大(+0.0237、+0.0341),区间不跨 0。哪项才是主要驱动,从这两行就能读出取向。
b156Cohort Fold Mean effect Lower-effect quantile Threshold Reference margin Refits above Chemical identity LINCS–Pilot1 4 0.0200 0.0139 0.0112 +0.0028 35/50 LINCS–Pilot1 5 0.0147 0.0090 0.0100 -0.0010 10/50 LKCP Batch 2 4 0.00003872 -0.0016 0.0303 -0.0319 0/50 LKCP Batch 2 5 0.0032 0.0013 0.0307 -0.0294 0/50 Dose LINCS–Pilot1 4 0.0510 0.0467 0.0068 +0.0398 50/50 LINCS–Pilot1 5 0.0488 0.0452 0.0059 +0.0393 50/50 LKCP Batch 2 4 0.0238 0.0205 0.0060 +0.0145 50/50 LKCP Batch 2 5 0.0342 0.0267 0.0125 +0.0142 50/50
这张表在检验:小分子效应到底来自化学身份还是剂量。
值得留意:按化学身份分组,LKCP 两折 margin 为负、0/50 通过;按剂量分组,四个组合全部 50/50。
b158The full Fold-4 panel contains the Fold-3-selected source from each of the four agents together with three distinct diagnostic sources: control-only, compound-only, and the shared h0h_{0}/simple-concatenation source shown under both reference roles. Predictive estimates use a separate five-seed refit panel from Table 21; model source or design, data fingerprint, population, and every non-seed fitting setting are identical. Each within-policy source is fixed before Fold 4; the pre-specified across-policy score rule then uses mean Fold-4 Global PCC to choose the single prediction-score source carried to Fold 5, which includes its paired control-only comparator. Controlled fixtures check recognized input consumption, the one-key attention contradiction, missing declarations, abstention, and known input dependence. A fixed compound-strength sweep characterizes threshold crossing in those fixtures; Appendix L explains the criterion’s reference-relative interpretation.
1. 这段在干什么
交代 Fold-4 审计面板的组成(四个 agent 各自 Fold-3 入选源 + 三个诊断源)与选源规则,为进入 Fold 5 做过渡。
2. 需要解释的地方
3. 值得留意
预测估计另用五种子重拟合面板,且除随机种子外所有设置(源、数据指纹、群体、拟合参数)均保持一致——这是为了让差异只归因于种子。选源在 Fold 4 前就固定,避免事后挑选。具体数值这段没给。
b159A. Predictive performance Model Global PCC MSE CP PCC L1000 PCC BBBC036 CellScientist 0.3096 ±\pm 0.0029 0.0586 ±\pm 0.0001 0.3113 ±\pm 0.0014 0.3089 ±\pm 0.0041 AIDE 0.2970 ±\pm 0.0075 0.0593 ±\pm 0.0004 0.3052 ±\pm 0.0031 0.2944 ±\pm 0.0102 CellForge 0.2960 ±\pm 0.0036 0.0592 ±\pm 0.0002 0.3056 ±\pm 0.0043 0.2925 ±\pm 0.0037 HarmonyCell 0.3110 ±\pm 0.0016 0.0586 ±\pm 0.0001 0.3109 ±\pm 0.0020 0.3110 ±\pm 0.0017 h0h_{0} 0.2879 ±\pm 0.0040 0.0598 ±\pm 0.0004 0.2981 ±\pm 0.0016 0.2839 ±\pm 0.0056 Control-only 0.3098 ±\pm 0.0024 0.0586 ±\pm 0.0001 0.3094 ±\pm 0.0013 0.3099 ±\pm 0.0028 Chemical-only 0.2239 ±\pm 0.0002 0.0616 ±\pm 0.0000 0.2998 ±\pm 0.0003 0.1832 ±\pm 0.0002 Simple concat 0.2879 ±\pm 0.0040 0.0598 ±\pm 0.0004 0.2981 ±\pm 0.0016 0.2839 ±\pm 0.0056 BBBC047 CellScientist 0.3035 ±\pm 0.0008 0.0457 ±\pm 0.0000 0.3260 ±\pm 0.0016 0.2784 ±\pm 0.0017 AIDE 0.2697 ±\pm 0.0048 0.0469 ±\pm 0.0002 0.3148 ±\pm 0.0047 0.2152 ±\pm 0.0116 CellForge 0.2727 ±\pm 0.0019 0.0467 ±\pm 0.0001 0.3175 ±\pm 0.0035 0.2190 ±\pm 0.0059 HarmonyCell 0.2983 ±\pm 0.0110 0.0459 ±\pm 0.0003 0.3271 ±\pm 0.0030 0.2654 ±\pm 0.0231 h0h_{0} 0.2697 ±\pm 0.0048 0.0469 ±\pm 0.0002 0.3148 ±\pm 0.0047 0.2152 ±\pm 0.0116 Control-only 0.3019 ±\pm 0.0015 0.0458 ±\pm 0.0001 0.3223 ±\pm 0.0011 0.2789 ±\pm 0.0023 Chemical-only 0.2159 ±\pm 0.0012 0.0480 ±\pm 0.0000 0.2935 ±\pm 0.0022 0.0924 ±\pm 0.0028 Simple concat 0.2697 ±\pm 0.0048 0.0469 ±\pm 0.0002 0.3148 ±\pm 0.0047 0.2152 ±\pm 0.0116
这段在干什么:这是附录里的一张完整数值表,列出各模型在两个数据集(BBBC036、BBBC047)上的预测指标,供审计时逐项核对。
需要解释的地方:
值得留意:
b160B. Input-use evidence Model EchemE_{\rm chem} EcontrolE_{\rm control} Ichem,controlI_{\rm chem,control} Q/U/I Status BBBC036 CellScientist 0.0000 ±\pm 0.0000 0.1622 ±\pm 0.0095 0.0000 ±\pm 0.0000 0/5/0 Unsup. AIDE 0.0029 ±\pm 0.0017 0.1577 ±\pm 0.0118 -0.0021 ±\pm 0.0016 0/5/0 Unsup. CellForge 0.0089 ±\pm 0.0052 0.1464 ±\pm 0.0063 -0.0000 ±\pm 0.0007 0/5/0 Unsup. HarmonyCell 0.0000 ±\pm 0.0000 0.1693 ±\pm 0.0033 -0.0000 ±\pm 0.0000 0/5/0 Unsup. h0h_{0} 0.0193 ±\pm 0.0087 0.1359 ±\pm 0.0057 -0.0003 ±\pm 0.0007 1/2/2 Inc. Control-only 0.0000 ±\pm 0.0000 0.1719 ±\pm 0.0067 0.0000 ±\pm 0.0000 – N/C Chemical-only 0.0011 ±\pm 0.0002 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0/0/5 Inc. Simple concat 0.0193 ±\pm 0.0087 0.1359 ±\pm 0.0057 -0.0003 ±\pm 0.0007 1/2/2 Inc. BBBC047 CellScientist 0.0000 ±\pm 0.0000 0.1932 ±\pm 0.0044 0.0000 ±\pm 0.0000 0/5/0 Unsup. AIDE 0.0118 ±\pm 0.0052 0.1119 ±\pm 0.0121 0.0013 ±\pm 0.0009 0/1/4 Inc. CellForge 0.0100 ±\pm 0.0045 0.1114 ±\pm 0.0059 0.0012 ±\pm 0.0008 0/2/3 Inc. HarmonyCell 0.0006 ±\pm 0.0010 0.1866 ±\pm 0.0225 0.0009 ±\pm 0.0015 0/5/0 Unsup. h0h_{0} 0.0118 ±\pm 0.0052 0.1119 ±\pm 0.0121 0.0013 ±\pm 0.0009 0/1/4 Inc. Control-only 0.0000 ±\pm 0.0000 0.2030 ±\pm 0.0054 0.0000 ±\pm 0.0000 – N/C Chemical-only 0.0111 ±\pm 0.0010 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 5/0/0 Sup. Simple concat 0.0118 ±\pm 0.0052 0.1119 ±\pm 0.0121 0.0013 ±\pm 0.0009 0/1/4 Inc.
这张表给出两个数据集(BBBC036、BBBC047)上各模型的三类输入使用证据分数,用来核对模型是否真的在用化学信息。
需要解释的地方
值得留意
b161Aggregate status: Sup., claim-supported; Unsup., unsupported; Inc., inconclusive; N/C, not claimed.
这句是附录审计表格的表注,用来定义表格里状态列的缩写。
1. 这段在干什么:给读者一把"钥匙",说明状态列 Sup./Unsup./Inc./N/C 各代表什么意思,否则表格没法读。
2. 需要解释的地方:
(作者逐字给出了对应,非推测。)
3. 值得留意:它正好能解释上一段结尾的 5/0/0 Sup. 和 0/1/4 Inc.——中间那四个数字应分别对应这四类计数。但具体顺序,这段原文没有说,需回看表格表头确认。
b162A. Predictive performance Policy Model Global PCC ↑\uparrow MSE ↓\downarrow CP PCC ↑\uparrow L1000 PCC ↑\uparrow BBBC036 h0h_{0} reference h0h_{0} 0.2588 ±\pm 0.0059 0.0608 ±\pm 0.0004 0.2596 ±\pm 0.0100 0.2586 ±\pm 0.0094 Score-selected HarmonyCell 0.2932 ±\pm 0.0020 0.0590 ±\pm 0.0001 0.2909 ±\pm 0.0018 0.2942 ±\pm 0.0023 BBBC047 h0h_{0} reference h0h_{0} 0.2786 ±\pm 0.0031 0.0465 ±\pm 0.0003 0.3220 ±\pm 0.0091 0.2243 ±\pm 0.0105 Score-selected CellScientist 0.3153 ±\pm 0.0014 0.0452 ±\pm 0.0000 0.3393 ±\pm 0.0007 0.2872 ±\pm 0.0025
这段是审计面板的A 部分:预测性能,用数字支撑“选出的模型确实更好”。
关键概念:PCC 是皮尔逊相关系数,越高越好(↑);MSE 是均方误差,越低越好(↓)。h₀ 指零假设/基线,这里是“未经选择的参考模型”。两行数据分别是两个数据集(BBBC036、BBBC047),每行对比参考 h₀ 与“Score-selected”(评分选出的)模型。
值得留意:选出的模型(HarmonyCell、CellScientist)在 Global PCC、CP PCC、L1000 PCC 上都高于参考,MSE 更低,且标准差很小——说明提升稳定,不是偶然波动。但 BBBC047 参考的 L1000 PCC(0.2243)反而低于 BBBC036 的(0.2586),这处不对称表里没解释。
b163B. Input-use evidence and replication Model EchemE_{\rm chem} EcontrolE_{\rm control} Ichem,controlI_{\rm chem,control} Q/U/I Status Status replicated? BBBC036 h0h_{0} 0.0047 ±\pm 0.0031 0.1202 ±\pm 0.0085 0.0042 ±\pm 0.0011 0/5/0 Unsup. no HarmonyCell 0.0000 ±\pm 0.0000 0.1516 ±\pm 0.0035 -0.0000 ±\pm 0.0000 0/5/0 Unsup. yes BBBC047 h0h_{0} 0.0096 ±\pm 0.0037 0.1199 ±\pm 0.0109 0.0012 ±\pm 0.0007 0/1/4 Inc. yes CellScientist 0.0000 ±\pm 0.0000 0.2004 ±\pm 0.0061 0.0000 ±\pm 0.0000 0/5/0 Unsup. yes
这段是附录里的审计证据表:逐个模型列出「输入是否真被用到」的证据与复现结果。
关键列:\(E_{\rm chem}\) 是化学输入对预测的贡献(接近 0 表示没用上),\(E_{\rm control}\) 是对照量,\(I_{\rm chem,control}\) 是化学相对对照的影响;Q/U/I 是判为「存疑/未支持/不一致」的计数。Status 是审计结论,最后一列是有无复现。这是领域常识:贡献≈0 即模型没真正依赖该输入。
值得留意:HarmonyCell 与 CellScientist 的 \(E_{\rm chem}\) 全为 0.0000,却判 Unsup. 且复现 yes——「复现了」复现的是「未支持」这一结论,不是复现了模型有效。
b164Score-selected denotes prediction-score selection. Aggregate status: Unsup., unsupported; Inc., inconclusive.
这段是附录审计表格的表注/图例,用来定义前文表格里出现的缩写,属于阅读辅助而非论证内容。
需要解释的地方
值得留意
前一段末尾的 Inc.、Unsup. 正是靠这里才得到解释;yes 一列则未在表注中定义,含义这段没提到。
b165Task Score-selected model Selected-model PCC ↑\uparrow Control-only PCC ↑\uparrow Paired Δ\Delta [95% CI] BBBC036 HarmonyCell 0.2932 ±\pm 0.0020 0.2921 ±\pm 0.0017 +0.0011 [−-0.0008, +0.0030] BBBC047 CellScientist 0.3153 ±\pm 0.0014 0.3142 ±\pm 0.0012 +0.0011 [−-0.0005, +0.0027]
这张表是审计面板核心结果:对两个任务,比较按预测分数选出的模型与仅用对照的基线,看 PCC(预测相关性,领域常识:越大越好)差多少。
值得留意:两处 Δ 都很小(+0.0011),且 95% 置信区间都跨 0,即差距在统计上不显著——所以上一段标注为"Unsup./Inc."。这正是"发现—证伪—修正"里的证伪信号。表头未给这些符号的含义,这段没提。
b166A. BBBC036 Model Residual PCC ↑\uparrow Mean [95% CI] MSE ↓\downarrow Pert. effect ↑\uparrow h0h_{0} 0.0474 [0.0371, 0.0577] 0.0598 0.0381 CellScientist 0.0294 [0.0052, 0.0535] 0.0586 0.0000 AIDE 0.0301 [0.0236, 0.0366] 0.0593 0.0050 CellForge 0.0336 [0.0234, 0.0438] 0.0592 0.0223 HarmonyCell 0.0341 [0.0226, 0.0455] 0.0586 0.0000 Control-only Zero increment 0.0586 – Chemical-only 0.0052 [-0.0044, 0.0148] 0.0616 0.0012 Simple concat 0.0474 [0.0371, 0.0577] 0.0598 0.0381 B. BBBC047 Model Residual PCC ↑\uparrow Mean [95% CI] MSE ↓\downarrow Pert. effect ↑\uparrow h0h_{0} 0.0558 [0.0510, 0.0606] 0.0469 0.0205 CellScientist 0.0538 [0.0422, 0.0654] 0.0457 0.0000 AIDE 0.0558 [0.0510, 0.0606] 0.0469 0.0205 CellForge 0.0553 [0.0510, 0.0596] 0.0467 0.0174 HarmonyCell 0.0567 [0.0470, 0.0663] 0.0459 0.0020 Control-only Zero increment 0.0458 – Chemical-only 0.0457 [0.0382, 0.0532] 0.0480 0.0104 Simple concat 0.0558 [0.0510, 0.0606] 0.0469 0.0205
这一段是附录里的完整审计面板:把 BBBC036 和 BBBC047 两个数据集上,各模型(h₀、CellScientist、AIDE、CellForge、HarmonyCell)加几条对照基线,按三个指标逐行列出。
关键概念:Residual PCC 是残差相关性,越高越好(↑);MSE 是误差,越低越好(↓);"Pert. effect"是扰动效应,看模型是否真用上了扰动信息。Control-only、Chemical-only、Simple concat 是对照组。
值得留意:两表里 h₀ 与 Simple concat 的数值完全相同,且 Pert. effect 都偏高;而 CellScientist、HarmonyCell 的 Pert. effect 直接为 0.0000。这说明有些模型可能没真正利用扰动项,正是审计要查的疑点。
b168The fixed-context simulator uses f(z,c)=(az,c)f(z,c)=(az,c) and Y=(bz+e,c)Y=(bz+e,c). Matched replacement draws an independent normal z′z^{\prime} while fixing c,Yc,Y. The loss is the sum of squared errors over both channels, so the population paired loss increment is τ=2ab\tau=2ab; the mean squared Euclidean prediction distance is η=2a2\eta=2a^{2}, not RMS. Groups receive equal weight and rows equal weight within group. The experiment evaluates equal-group-weighted population contrasts; real-data row-weighted resampling is reported separately. Four classes are invariant (a,b)=(0,.2)(a,b)=(0,.2), sensitive without benefit (.2,0)(.2,0), beneficial (.2,.2)(.2,.2), and harmful (−.2,.2)(-.2,.2). Their loss truths are 0,0,.08,−.080,0,.08,-.08 and distance truths 0,.08,.08,.080,.08,.08,.08.
这段在干什么:用已知真值的模拟器校准审计指标,给出损失增量与预测距离两套"标准答案"。
需要解释的地方:
f(z,c)=(az,c)、Y=(bz+e,c):合成数据的生成式,z 是变量、c 是固定上下文,a 决定影响、b 决定结果侧。值得留意:距离真值只依赖 a,槽位组与行都等权,等组加权与真实数据的行加权重采样是分开报告的,别混。
b169The frozen 24=162^{4}=16 scenarios cross G∈{8,32}G\in\{8,32\} independent source groups, balanced size 8 or independently sampled sizes {2,4,8,32}\{2,4,8,32\} with equal probabilities, within-group correlation ρ∈{0,.6}\rho\in\{0,.6\}, and noise SD σ∈{.25,1}\sigma\in\{.25,1\}. Each scenario has 400 independent datasets (6400 total). Maps 32/128 are nested on the same datasets, not independent cohorts. Group-percentile intervals use 999 bootstrap draws; group-tt uses G−1G-1 degrees of freedom. Source/donor, bootstrap and context seeds are 2026091401, 2026091402 and 2026091403.
这段在干什么:交代「已知真值」实验的设计清单——场景怎么组合、每组算多大、随机种子取多少,为下一节的测量结果提供可复现的设置说明。
需要解释的地方:
值得留意:前段损失真值有负值,这段的 σ 与 ρ 只描述构造方式;「not independent cohorts」是作者主动提示的局限。
b170Map-only intervals target the donor-integrated contrast of one fixed dataset; group-aware intervals target the population expectation over independent groups. Their coverage is therefore reported against both truths. Positive evidence means a strictly positive lower bound of a two-sided 95% CI: the positive one-tail nominal Type-I error is .025.025, not .05.05. Power is positive-direction evidence for benefit and negative-direction evidence for harm. Prediction distance does not test target benefit. Invariant data and intervals are identically zero and are reported as a deterministic check, not “100% calibrated” stochastic coverage.
定义并区分两类区间(map-only / group-aware)各自对应的“真值”,并说明本节的覆盖率和证据判读规则。
b171Endpoint / class CI method Population coverage Conditional coverage Directed evidence rate 32 donor maps Null loss Map only [0.1500, 0.3000][0.1500,\,0.3000] [0.9275, 0.9675][0.9275,\,0.9675] [0.2850, 0.5000][0.2850,\,0.5000] Null loss Group tt [0.9200, 0.9775][0.9200,\,0.9775] [1.0000, 1.0000][1.0000,\,1.0000] [0.0150, 0.0725][0.0150,\,0.0725] Null loss Group bootstrap [0.8175, 0.9400][0.8175,\,0.9400] [1.0000, 1.0000][1.0000,\,1.0000] [0.0350, 0.1475][0.0350,\,0.1475] Beneficial loss Map only [0.1425, 0.3350][0.1425,\,0.3350] [0.9250, 0.9650][0.9250,\,0.9650] [0.7100, 1.0000][0.7100,\,1.0000] Beneficial loss Group tt [0.9100, 0.9650][0.9100,\,0.9650] [0.9975, 1.0000][0.9975,\,1.0000] [0.1200, 1.0000][0.1200,\,1.0000] Beneficial loss Group bootstrap [0.8350, 0.9425][0.8350,\,0.9425] [0.9975, 1.0000][0.9975,\,1.0000] [0.2650, 1.0000][0.2650,\,1.0000] Harmful loss Map only [0.1050, 0.2750][0.1050,\,0.2750] [0.9375, 0.9650][0.9375,\,0.9650] [0.6775, 1.0000][0.6775,\,1.0000] Harmful loss Group tt [0.8625, 0.9725][0.8625,\,0.9725] [1.0000, 1.0000][1.0000,\,1.0000] [0.0550, 1.0000][0.0550,\,1.0000] Harmful loss Group bootstrap [0.8300, 0.9400][0.8300,\,0.9400] [1.0000, 1.0000][1.0000,\,1.0000] [0.1900, 1.0000][0.1900,\,1.0000] Squared distance Map only [0.2175, 0.5250][0.2175,\,0.5250] [0.9400, 0.9700][0.9400,\,0.9700] [1.0000, 1.0000][1.0000,\,1.0000] Squared distance Group tt [0.8375, 0.9525][0.8375,\,0.9525] [1.0000, 1.0000][1.0000,\,1.0000] [0.9975, 1.0000][0.9975,\,1.0000] Squared distance Group bootstrap [0.7800, 0.9375][0.7800,\,0.9375] [0.9975, 1.0000][0.9975,\,1.0000] [1.0000, 1.0000][1.0000,\,1.0000] 128 donor maps Null loss Map only [0.0675, 0.1625][0.0675,\,0.1625] [0.9225, 0.9725][0.9225,\,0.9725] [0.3600, 0.5375][0.3600,\,0.5375] Null loss Group tt [0.9125, 0.9775][0.9125,\,0.9775] [1.0000, 1.0000][1.0000,\,1.0000] [0.0125, 0.0775][0.0125,\,0.0775] Null loss Group bootstrap [0.8200, 0.9350][0.8200,\,0.9350] [1.0000, 1.0000][1.0000,\,1.0000] [0.0350, 0.1450][0.0350,\,0.1450] Beneficial loss Map only [0.0750, 0.1750][0.0750,\,0.1750] [0.9175, 0.9650][0.9175,\,0.9650] [0.7675, 1.0000][0.7675,\,1.0000] Beneficial loss Group tt [0.9075, 0.9725][0.9075,\,0.9725] [1.0000, 1.0000][1.0000,\,1.0000] [0.1200, 1.0000][0.1200,\,1.0000] Beneficial loss Group bootstrap [0.8400, 0.9450][0.8400,\,0.9450] [1.0000, 1.0000][1.0000,\,1.0000] [0.2475, 1.0000][0.2475,\,1.0000] Harmful loss Map only [0.0400, 0.1450][0.0400,\,0.1450] [0.9125, 0.9600][0.9125,\,0.9600] [0.7025, 1.0000][0.7025,\,1.0000] Harmful loss Group tt [0.8600, 0.9700][0.8600,\,0.9700] [1.0000, 1.0000][1.0000,\,1.0000] [0.0500, 1.0000][0.0500,\,1.0000] Harmful loss Group bootstrap [0.8200, 0.9425][0.8200,\,0.9425] [1.0000, 1.0000][1.0000,\,1.0000] [0.1875, 1.0000][0.1875,\,1.0000] Squared distance Map only [0.0925, 0.2900][0.0925,\,0.2900] [0.9350, 0.9600][0.9350,\,0.9600] [1.0000, 1.0000][1.0000,\,1.0000] Squared distance Group tt [0.8425, 0.9500][0.8425,\,0.9500] [1.0000, 1.0000][1.0000,\,1.0000] [0.9975, 1.0000][0.9975,\,1.0000] Squared distance Group bootstrap [0.8025, 0.9350][0.8025,\,0.9350] [1.0000, 1.0000][1.0000,\,1.0000] [1.0000, 1.0000][1.0000,\,1.0000]
这一段的上一句刚说完「区间恒为零时只作确定性检查、不算 100% 校准」,这段紧跟着把已知真值场景下的三种 CI 方法摆开对比。
这段在干什么:汇报已知真值实验里各端点/类别在 32 与 128 donor maps 下的群体覆盖、条件覆盖和有向证据率,验证不确定性估计是否可靠。
需要解释的地方:
值得留意:Map only 的群体覆盖明显低于另两种(如 128 maps 下 Null loss 只有约 [0.0675, 0.1625]),但它的有向证据率反而更高;Group tt 和 bootstrap 的群体覆盖接近 1,方向率却常落到 0 附近。这个权衡是作者没说破的关键。
b172Each rate uses 400 independent datasets per scenario; ranges are minima/maxima over scenarios, not pooled coverage. The last column is positive Type I for null loss (nominal .025), correctly directed power for beneficial/harmful loss, and positive-distance evidence for distance. Conditional and population targets differ. Invariant cases are excluded from stochastic coverage.
交代上表结果的读法:每格数字怎么来的、代表什么,避免误读成汇总覆盖率。
不变量(invariant)案例不进入随机覆盖率统计——做跨场景比较时,样本范围其实被缩过。
b173GG/size/ρ\rho/σ\sigma Map only Group tt Group bootstrap tt Type-I MCSE tt Type-I Wilson 95% CI 8/B/0/0.25 0.1275 0.4600 0.9500 0.0400 0.8675 0.0950 0.00980.0098 [0.0248, 0.0640][0.0248,\,0.0640] 8/B/0/1 0.1175 0.4925 0.9550 0.0325 0.8950 0.0700 0.00890.0089 [0.0191, 0.0548][0.0191,\,0.0548] 8/B/0.6/0.25 0.0875 0.5075 0.9325 0.0600 0.8200 0.1450 0.01190.0119 [0.0406, 0.0877][0.0406,\,0.0877] 8/B/0.6/1 0.0675 0.4925 0.9650 0.0225 0.8525 0.0850 0.00740.0074 [0.0119, 0.0422][0.0119,\,0.0422] 8/I/0/0.25 0.1300 0.4675 0.9325 0.0650 0.8500 0.1275 0.01230.0123 [0.0447, 0.0935][0.0447,\,0.0935] 8/I/0/1 0.1550 0.3600 0.9775 0.0125 0.8825 0.0650 0.00560.0056 [0.0054, 0.0289][0.0054,\,0.0289] 8/I/0.6/0.25 0.0725 0.5375 0.9350 0.0650 0.8400 0.1375 0.01230.0123 [0.0447, 0.0935][0.0447,\,0.0935] 8/I/0.6/1 0.0825 0.4625 0.9600 0.0250 0.8300 0.1000 0.00780.0078 [0.0136, 0.0454][0.0136,\,0.0454] 32/B/0/0.25 0.1125 0.4275 0.9325 0.0300 0.9225 0.0375 0.00850.0085 [0.0172, 0.0517][0.0172,\,0.0517] 32/B/0/1 0.1350 0.4500 0.9625 0.0225 0.9350 0.0350 0.00740.0074 [0.0119, 0.0422][0.0119,\,0.0422] 32/B/0.6/0.25 0.0850 0.4650 0.9325 0.0575 0.9175 0.0675 0.01160.0116 [0.0386, 0.0848][0.0386,\,0.0848] 32/B/0.6/1 0.0875 0.4600 0.9600 0.0250 0.9350 0.0350 0.00780.0078 [0.0136, 0.0454][0.0136,\,0.0454] 32/I/0/0.25 0.1450 0.3975 0.9450 0.0425 0.9225 0.0500 0.01010.0101 [0.0267, 0.0670][0.0267,\,0.0670] 32/I/0/1 0.1625 0.4175 0.9450 0.0375 0.9225 0.0425 0.00950.0095 [0.0229, 0.0609][0.0229,\,0.0609] 32/I/0.6/0.25 0.0975 0.4900 0.9125 0.0775 0.8925 0.0850 0.01340.0134 [0.0551, 0.1079][0.0551,\,0.1079] 32/I/0.6/1 0.0750 0.4450 0.9375 0.0450 0.9025 0.0625 0.01040.0104 [0.0287, 0.0700][0.0287,\,0.0700]
这是一张已知真值的校准结果表,用来检查度量方法本身是否可靠。
1. 这段在干什么
在已知真值的合成场景下,报告各组(不同 GG/size/ρ/σ 配置)的统计量与检验结果,验证度量流程的可靠性。
2. 需要解释的地方
3. 值得留意
表头多列在原文里没有逐一定义;Type-I 大多在 0.02–0.08 间波动,判断是否合格需对照论文正文标准,这段没明说。
b174B/I: balanced/imbalanced. Type-I nominal .025. MCSE is p(1−p)/400\sqrt{p(1-p)/400} (worst possible .025); Wilson intervals are pointwise Monte Carlo intervals, not simultaneous bounds over 16 scenarios.
这段是 F.1 的图注/表注,交代上一段那串数字读法。
关键概念(领域常识):
值得留意:Wilson 区间是逐点的,不是覆盖 16 个场景的同时置信界——别当成多重比较校正后的区间读。
b175Endpoint / class Maps Truth Bias range Bias MCSE range Null loss 32 00 [−×10−3,×10−3][-8.41\!\times\!10^{-3},\,4.95\!\times\!10^{-3}] [×10−4,×10−3][3.84\!\times\!10^{-4},\,5.21\!\times\!10^{-3}] Null loss 128 00 [−×10−3,×10−3][-8.42\!\times\!10^{-3},\,5.48\!\times\!10^{-3}] [×10−4,×10−3][3.82\!\times\!10^{-4},\,5.14\!\times\!10^{-3}] Beneficial loss 32 0.080.08 [−×10−3,×10−3][-7.32\!\times\!10^{-3},\,4.58\!\times\!10^{-3}] [×10−4,×10−3][3.50\!\times\!10^{-4},\,5.14\!\times\!10^{-3}] Beneficial loss 128 0.080.08 [−×10−3,×10−3][-7.47\!\times\!10^{-3},\,5.06\!\times\!10^{-3}] [×10−4,×10−3][3.47\!\times\!10^{-4},\,5.06\!\times\!10^{-3}] Harmful loss 32 −0.08-0.08 [−×10−3,×10−3][-4.17\!\times\!10^{-3},\,6.55\!\times\!10^{-3}] [×10−4,×10−3][5.70\!\times\!10^{-4},\,5.48\!\times\!10^{-3}] Harmful loss 128 −0.08-0.08 [−×10−3,×10−3][-4.68\!\times\!10^{-3},\,6.51\!\times\!10^{-3}] [×10−4,×10−3][5.65\!\times\!10^{-4},\,5.40\!\times\!10^{-3}] Squared distance 32 0.080.08 [−×10−4,×10−4][-4.19\!\times\!10^{-4},\,7.09\!\times\!10^{-4}] [×10−4,×10−4][1.69\!\times\!10^{-4},\,7.60\!\times\!10^{-4}] Squared distance 128 0.080.08 [−×10−4,×10−4][-3.08\!\times\!10^{-4},\,5.67\!\times\!10^{-4}] [×10−4,×10−4][1.68\!\times\!10^{-4},\,7.49\!\times\!10^{-4}]
这段是 F.1 的已知真值检验表:对 8 种损失/样本量组合,报告偏差范围与 MCSE 范围,用来证明度量本身无系统性偏差。
关键概念(领域常识):
值得留意:Bias 范围远小于 MCSE 范围,说明误差主要来自随机性而非偏差;Squared distance 的量级整体小一个数量级——但这与损失尺度不同有关,原文未加说明。
b176All CI methods share the same point estimator. Bias MCSE is the empirical SD of estimation error divided by 400\sqrt{400}; full-precision per-scenario bias, RMSE, coverage and Wilson intervals are in the portable CSV.
1. 这段在干什么
补充说明不确定性检查的技术细节:交代置信区间(CI)方法的估计量来源和偏差误差的计算方式,并指引完整数据的位置。
2. 需要解释的地方
3. 值得留意
b177Class Amplitude Mean effect Mean QQ Support-rate range Invariant 00 00 00 [0, 0][0,\,0] Invariant 0.20.2 00 0.087920.08792 [0, 0][0,\,0] Invariant 2.02.0 00 8.792308.79230 [0, 0][0,\,0] Sensitive/no benefit 00 −0.00067-0.00067 0.077440.07744 [0, 0.007][0,\,0.007] Sensitive/no benefit 0.20.2 −0.00067-0.00067 0.120340.12034 [0, 0.003][0,\,0.003] Sensitive/no benefit 2.02.0 −0.00067-0.00067 8.792748.79274 [0, 0][0,\,0] Beneficial 00 0.079460.07946 0.082260.08226 [0, 1.000][0,\,1.000] Beneficial 0.20.2 0.079460.07946 0.122600.12260 [0, 0.110][0,\,0.110] Beneficial 2.02.0 0.079460.07946 8.792798.79279 [0, 0][0,\,0] Harmful 00 −0.07961-0.07961 0.082180.08218 [0, 0][0,\,0] Harmful 0.20.2 −0.07961-0.07961 0.122760.12276 [0, 0][0,\,0] Harmful 2.02.0 −0.07961-0.07961 8.792488.79248 [0, 0][0,\,0]
1. 这段在干什么
这是一张「已知真值」的对照表:人工构造的四种输入类别(Invariant / Sensitive-no benefit / Beneficial / Harmful),检查审计指标能不能把它们区分开——即方法的「体检」。
2. 需要解释的地方
3. 值得留意
Invariant 与 Sensitive 的 Mean effect 都是 0,靠生效与否区分;Amplitude 2.0 时 QQ 普遍飙到 8.79 而 support-rate 全塌成 [0,0]——高幅度下该指标基本失效,作者没点破。
b178Means average all 16 scenarios; support ranges remain scenario-specific. This replays the historical formula on negative squared-loss scores, not historical PCC. Changing orthogonal context changes its reference threshold while keeping the chemical contrast fixed. Non-support is not a false negative for the different proposition τ>0\tau>0; both map budgets and all original statuses remain in the portable CSV.
交代表格数字的统计口径与前提,防止读者把状态判定误读成 PCC 上的结论。
非 support 不等于命题 τ>0 的假阴性——两者不是一回事,别混读。
b179The group-aware methods still under-cover in some scenarios; the bootstrap is particularly weak with few groups, and positive null declarations can exceed the nominal one-tail rate. Independent normal donors, independent groups, and correct context matching are explicit simulation assumptions, not demonstrations of real biological conditional exchangeability or support. Finite reused donor pools and dependent biological groups require additional dependence handling. More maps reduce conditional Monte Carlo error but do not create more biological replicates.
1. 这段在干什么
这是对模拟实验局限性的自我批评:作者承认组感知方法在部分场景仍覆盖不足,并列出这些结果所依赖的假设与不能替代的东西。
2. 需要解释的地方
3. 值得留意
作者明确区分:加 maps 只减模拟误差,不增加生物学重复。这些独立正态、独立分组、上下文匹配都是模拟假设,不等于真实生物学的可交换性或支持度。
b181The complete intervention matrix reports search performance, held-out predictive performance, compound-identity effect, control-profile effect, target-loss gain, and input-use status in separate columns. Rows summarize the full pre-specified model sets, and −−-- marks contrasts that cannot be identified under the protocol. Response-block decompositions are diagnostic views of the same selected models and fitted checkpoints. Figure 6 follows the compound and control effects of the falsification-guided checkpoints across both held-out folds.
这段在干什么:交代完整干预矩阵里有哪些列、行汇总什么,并说明图表的分工——图6只看伪造引导检查点的两类效应。
需要解释的地方:
值得留意:矩阵列里包含 input-use status,这正是论文审计的对象;而分解只是"diagnostic views",别误当成独立证据。
b182A. Selection record Task Search setting Selected source Fold-3 PCC BBBC036 Prediction-score HarmonyCell 0.2985 BBBC036 Path-constrained 10-endpoint cohort 0.2925 ±\pm 0.0019 BBBC036 Falsification-guided 10-endpoint cohort 0.3006 ±\pm 0.0006 BBBC047 Prediction-score CellScientist 0.3231 BBBC047 Path-constrained 10-endpoint cohort 0.3008 ±\pm 0.0012 BBBC047 Falsification-guided 10-endpoint cohort 0.3238 ±\pm 0.0002
这段是附录里的一张选型结果表,记录每个数据集在三种任务搜索设置下选出的模型及其 Fold-3 PCC(皮尔逊相关系数,领域常识:衡量预测与真实值的相关程度,越高越好)。
关键概念:「Selection record」= 选型记录;三种设置——Prediction-score(只看预测分)、Path-constrained(路径约束,限定为 10-endpoint cohort)、Falsification-guided(用证伪引导筛选)。
值得留意:三个数据集/设置中,Falsification-guided 的 PCC 一致最高(BBBC036 为 0.3006,BBBC047 为 0.3238);且它带 ± 方差,说明跨多次运行平均,而 Prediction-score 只有单点值。表中数据集仅两个(BBBC036、BBBC047),统计口径不统一,勿过度外推。
b183B. Fold 4: held-out input-use evidence Search setting PCC EcmpdE_{\rm cmpd} EcmpdℒE^{\mathcal{L}}_{\rm cmpd} EctrlE_{\rm ctrl} Q/U/I BBBC036 Prediction-score 0.3110 ±\pm 0.0016 0.0000 ±\pm 0.0000 0.00000 ±\pm 0.00000 0.1693 ±\pm 0.0033 0/5/0 Path-constrained 0.3034 ±\pm 0.0029 0.0007 ±\pm 0.0006 0.00003 ±\pm 0.00003 0.1592 ±\pm 0.0050 0/10/0 Falsification-guided 0.3167 ±\pm 0.0013 0.0091 ±\pm 0.0010 0.00039 ±\pm 0.00004 0.1716 ±\pm 0.0042 0/10/0 BBBC047 Prediction-score 0.3035 ±\pm 0.0008 0.0000 ±\pm 0.0000 0.00000 ±\pm 0.00000 0.1932 ±\pm 0.0044 0/5/0 Path-constrained 0.2822 ±\pm 0.0017 0.0083 ±\pm 0.0016 0.00024 ±\pm 0.00005 0.1388 ±\pm 0.0052 0/6/4 Falsification-guided 0.3060 ±\pm 0.0004 0.0053 ±\pm 0.0002 0.00018 ±\pm 0.00001 0.1988 ±\pm 0.0027 0/10/0
这段是附录 G 的结果细表,把 Fold 4 在 BBBC036、BBBC047 两个数据集上的输入使用证据,按三种搜索策略逐行列出,供前面正文结论做支撑。
关键术语(领域常识):PCC 是预测与真值的相关性,越高越好;E_cmpd 等是一组“输入干预后模型输出变化”的指标,越接近 0 说明模型越不依赖该输入;Q/U/I 指被判定为合格/未用/需干预的输入个数。
值得留意:三种搜索里只有 Falsification-guided 在两个数据集上同时保住了较高 PCC 与较低 E_cmpd,且 Q/U/I 都是 0/10/0;Path-constrained 在 BBBC047 上 PCC 掉到 0.2822、I 出现 4,是三行里唯一有非零“需干预”的情况。
b184C. Fold 5: held-out input-use evidence Search setting PCC EcmpdE_{\rm cmpd} EcmpdℒE^{\mathcal{L}}_{\rm cmpd} EctrlE_{\rm ctrl} Q/U/I BBBC036 Prediction-score 0.2932 ±\pm 0.0020 0.0000 ±\pm 0.0000 0.00000 ±\pm 0.00000 0.1516 ±\pm 0.0035 0/5/0 Path-constrained 0.2886 ±\pm 0.0018 0.0039 ±\pm 0.0022 0.00016 ±\pm 0.00009 0.1423 ±\pm 0.0050 0/10/0 Falsification-guided 0.2897 ±\pm 0.0008 0.0022 ±\pm 0.0012 0.00009 ±\pm 0.00005 0.1515 ±\pm 0.0037 0/10/0 BBBC047 Prediction-score 0.3153 ±\pm 0.0014 0.0000 ±\pm 0.0000 0.00000 ±\pm 0.00000 0.2004 ±\pm 0.0061 0/5/0 Path-constrained 0.2922 ±\pm 0.0018 0.0074 ±\pm 0.0017 0.00022 ±\pm 0.00005 0.1477 ±\pm 0.0057 0/7/3 Falsification-guided 0.3173 ±\pm 0.0004 0.0048 ±\pm 0.0004 0.00017 ±\pm 0.00001 0.2051 ±\pm 0.0024 0/10/0
这一行是小节 Appendix G 表格中 Fold 5 的收尾数据。
1. 这段在干什么
汇报 5 折交叉验证里第 5 折(Fold 5)的留出输入使用证据,列出 BBBC036、BBBC047 两个数据集下三种搜索设置的表现,补齐前面各折的完整结果。
2. 需要解释的地方
3. 值得留意
BBBC047 行 Path-constrained 的 Q/U/I 是 0/7/3,而其他多为 0/10/0,是唯一非零 U/I 的行,别读漏。
b185A. Predictive performance Task Fold Control-only PCC Full PCC Δ\Delta Full−-Control BBBC036 Fold 4 0.3121 ±\pm 0.0011 0.3167 ±\pm 0.0013 0.0046 ±\pm 0.0006 BBBC036 Fold 5 0.2917 ±\pm 0.0006 0.2897 ±\pm 0.0008 -0.0020 ±\pm 0.0009 BBBC047 Fold 4 0.3023 ±\pm 0.0004 0.3060 ±\pm 0.0004 0.0037 ±\pm 0.0003 BBBC047 Fold 5 0.3145 ±\pm 0.0005 0.3173 ±\pm 0.0004 0.0028 ±\pm 0.0004
这张表在汇报干预实验的预测性能:对比只用对照组(Control-only)和用完整输入(Full)的 PCC,看差值 Δ。
术语:PCC 是皮尔逊相关系数,衡量预测与真实值的线性相关,越高越好(领域常识)。Fold 是交叉验证的折。±后面是标准差。
值得留意:Δ 列才是重点,但它普遍很小(0.0028~0.0046),Fold 5 甚至出现负值 −0.0020,说明"用全输入"并不总是更好——这正是审计要 Falsify 的地方。
b186B. Input-use evidence Task Fold EcmpdE_{\rm cmpd} EcmpdℒE^{\mathcal{L}}_{\rm cmpd} EctrlE_{\rm ctrl} Model Q/U/I BBBC036 Fold 4 0.0091 ±\pm 0.0010 0.00039 ±\pm 0.00004 0.1716 ±\pm 0.0042 0/10/0 BBBC036 Fold 5 0.0022 ±\pm 0.0012 0.00009 ±\pm 0.00005 0.1515 ±\pm 0.0037 0/10/0 BBBC047 Fold 4 0.0053 ±\pm 0.0002 0.00018 ±\pm 0.00001 0.1988 ±\pm 0.0027 0/10/0 BBBC047 Fold 5 0.0048 ±\pm 0.0004 0.00017 ±\pm 0.00001 0.2051 ±\pm 0.0024 0/10/0
这段是附录 G 的输入干预结果表,逐行列出各数据集划分在干预后的误差指标,用来支撑「输入是否被模型真正使用」的审计结论。
关键概念:领域常识——E_cmpd 是干预化合物后预测变化的大小,E_cmpd^L 是它的似然/对数形式,E_ctrl 是对照误差;Q/U/I 通常指 Quiet/Used/Ignored 一类判定计数(此处未展开说明)。
值得留意:四行 Q/U/I 全是 0/10/0,即全部判定为"使用",但各折误差数值差异不小,且 E_ctrl(0.15–0.21)远大于 E_cmpd(千分位级)。这说明什么,原文这段没说,需结合上下文判断。
b187Predictor Global PCC [95% CI] Chemical PCC drop [95% CI] Target-loss gain [95% CI] BBBC036 / Fold 4 Falsification-guided 0.3167 [0.3154, 0.3181] 0.0091 [0.0081, 0.0101] 0.00039 [0.00034, 0.00043] Residual MLP 0.3056 [0.3027, 0.3085] 0.0036 [0.0028, 0.0044] 0.00015 [0.00012, 0.00019] Joint Ridge 0.0892 [0.0867, 0.0917] 0.0153 [0.0147, 0.0159] 0.00299 [0.00290, 0.00308] BBBC036 / Fold 5 Falsification-guided 0.2897 [0.2890, 0.2905] 0.0022 [0.0010, 0.0034] 0.00009 [0.00004, 0.00014] Residual MLP 0.2861 [0.2829, 0.2892] 0.0007 [0.0001, 0.0013] 0.00003 [0.00000, 0.00005] Joint Ridge 0.1114 [0.1092, 0.1136] 0.0336 [0.0328, 0.0344] 0.00569 [0.00557, 0.00582] BBBC047 / Fold 4 Falsification-guided 0.3060 [0.3056, 0.3064] 0.0053 [0.0051, 0.0055] 0.00018 [0.00017, 0.00019] Residual MLP 0.3031 [0.3007, 0.3055] 0.0119 [0.0094, 0.0145] 0.00042 [0.00033, 0.00050] Joint Ridge 0.0774 [0.0765, 0.0783] 0.0082 [0.0079, 0.0085] 0.00420 [0.00416, 0.00423] BBBC047 / Fold 5 Falsification-guided 0.3173 [0.3169, 0.3177] 0.0048 [0.0044, 0.0052] 0.00017 [0.00015, 0.00018] Residual MLP 0.3154 [0.3135, 0.3174] 0.0113 [0.0095, 0.0132] 0.00040 [0.00033, 0.00047] Joint Ridge 0.0704 [0.0693, 0.0715] 0.0052 [0.0050, 0.0053] 0.00394 [0.00393, 0.00395]
这一节把附录前面汇总表里 BBBC047 Fold 5 那两行,按四个数据集/折展开成完整三列数值。
这段在干什么:把「完整输入干预结果」以表格形式铺开,列出各模型在四个数据集/折上的三项指标及 95% 置信区间。
需要解释的地方:
值得留意:Joint Ridge 的 PCC 明显低(0.07–0.11),但 Chemical PCC drop 反而最大——说明它整体预测差,却更依赖化学输入。Fold 5 的 drop 普遍比 Fold 4 小。
b188Fold Perturbation Context Interaction CP L1000 CP L1000 CP L1000 Path-constrained / BBBC036 F4 0.0012 ±\pm 0.0009 0.0004 ±\pm 0.0005 0.0250 ±\pm 0.0023 0.2164 ±\pm 0.0057 0.0003 ±\pm 0.0004 -0.0003 ±\pm 0.0006 F5 0.0008 ±\pm 0.0004 0.0053 ±\pm 0.0031 0.0290 ±\pm 0.0024 0.1935 ±\pm 0.0056 0.0005 ±\pm 0.0002 0.0013 ±\pm 0.0008 Path-constrained / BBBC047 F4 0.0130 ±\pm 0.0026 0.0020 ±\pm 0.0002 0.0911 ±\pm 0.0037 0.2178 ±\pm 0.0054 0.0026 ±\pm 0.0012 0.0007 ±\pm 0.0002 F5 0.0118 ±\pm 0.0026 0.0013 ±\pm 0.0005 0.1005 ±\pm 0.0054 0.2251 ±\pm 0.0054 0.0011 ±\pm 0.0012 -0.0001 ±\pm 0.0003 Falsification-guided / BBBC036 F4 0.0141 ±\pm 0.0023 0.0069 ±\pm 0.0007 0.0248 ±\pm 0.0015 0.2363 ±\pm 0.0041 0.0000 ±\pm 0.0000 0.0001 ±\pm 0.0000 F5 -0.0014 ±\pm 0.0016 0.0039 ±\pm 0.0013 0.0284 ±\pm 0.0016 0.2088 ±\pm 0.0035 0.0001 ±\pm 0.0000 0.0002 ±\pm 0.0001 Falsification-guided / BBBC047 F4 0.0088 ±\pm 0.0003 0.0014 ±\pm 0.0001 0.1338 ±\pm 0.0031 0.2759 ±\pm 0.0013 0.0002 ±\pm 0.0001 0.0001 ±\pm 0.0000 F5 0.0079 ±\pm 0.0007 0.0013 ±\pm 0.0002 0.1421 ±\pm 0.0025 0.2806 ±\pm 0.0014 -0.0000 ±\pm 0.0001 0.0001 ±\pm 0.0000
这段是附录 G 的完整输入干预结果表,铺列各模型/数据集组合下三种扰动(Fold、Context、Interaction)在 CP 与 L1000 两种读数上的效应值(均值±标准差),供正文结论查证。
关键概念:CP、L1000 是两种细胞表型读数(领域常识:L1000 是基因表达谱平台);Path-constrained、Falsification-guided 是两种模型发现路径。
留意:分为 BBBC036/BBBC047 两个数据集、F4/F5 两个折叠;Path-constrained 组的 Interaction 值近零甚至为负,而 Falsification-guided 组普遍更小。表格不带解读,具体论断需回正文。
b189Task Fold Δdose\Delta_{\rm dose} Compound Control Dose Compound×\times control Path-constrained BBBC036 4 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC036 5 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC047 4 – 0/6/4/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC047 5 – 0/7/3/0 10/0/0/0 0/0/0/10 0/10/0/0 Falsification-guided BBBC036 4 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC036 5 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC047 4 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC047 5 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0
这是附录 G 的干预结果表,逐行比较三种建模流程(Task、Fold、Path-constrained、Falsification-guided)下,各扰动类型干预实验的命中计数。
x/y/z/w,表示按干预类别分配的结果计数(顺序需回正文核对,这段未说明)。0/6/4/0、0/7/3/0,后者是 0/10/0/0;作者没明说这种差异的原因。b191Table 51 applies the deterministic checker to every retained prediction-score candidate and separates execution failure from checker abstention. Resolution coverage is the fraction of checked candidates whose cited source locations are all resolved. For candidates without a structured manifest declaring input use, the checker reports findings supported by pre-specified source locations. The one-key attention certificate is a code-level finding; complete-model replacement tests provide the corresponding behavioral evidence.
把确定性检查器跑遍所有保留下来的预测分数候选,区分「执行失败」和「检查器弃权」两类情况。
代码级发现和行为证据是两回事,作者明确把二者配对——别把前者当成行为已证的结论。
b192Candidate slots Code checker Coverage Policy Total Exec. Checked Resolved Singleton Abstain Resolved/ checked Checked/ slots BBBC036 CellScientist 100 99 99 60 15 39 60.6% 99.0% AIDE 100 100 100 100 0 0 100.0% 100.0% CellForge 100 91 91 91 0 0 100.0% 91.0% HarmonyCell 100 95 95 94 0 1 98.9% 95.0% BBBC047 CellScientist 100 94 94 48 17 46 51.1% 94.0% AIDE 100 99 99 99 0 0 100.0% 99.0% CellForge 100 94 94 94 0 0 100.0% 94.0% HarmonyCell 100 95 95 94 0 1 98.9% 95.0%
好,我们就盯着这张表看,逐列捋一遍。
这是附录里的数据表:列出两个数据集(BBBC036、BBBC047)下四个模型(CellScientist、AIDE、CellForge、HarmonyCell)的「全候选源码检查」结果,用覆盖率、检查数、解析数等指标评估各模型所提 input-use claim 的质量。
CellScientist 在 BBBC047 上 Resolved/checked 只有 51.1%,且 Singleton 高达 17——说明近半数 claim 无法比对验证,是被「单例」卡住的,而非检查失败。
正文这段没提到这些数的成因,别替它归因。
b193Search Slots Exec. Resolved Singleton Abstain Resolve % Exec. % BBBC036 Prediction-score 100 99 60 15 39 60.6% 99.0% Path-constrained 100 100 100 0 0 100.0% 100.0% BBBC047 Prediction-score 100 94 48 17 46 51.1% 94.0% Path-constrained 100 100 100 0 0 100.0% 100.0%
这是附录里的全候选源检查结果表,逐行给两个数据集(BBBC036、BBBC047)在两种搜索策略下的执行与解析统计,用来支撑“源检查”的审计结果。
Path-constrained 两数据集都是 100/100/100、0 弃权、100%,说明该策略几乎不缺执行;而 Prediction-score 在 BBBC047 上仅解析 48、弃权 46,解析率 51.1%,差异明显。
b194A deterministic hash selects model sources within each checker category for source-interface inspection. The inspection is restricted to pre-specified source locations and public fusion interfaces; fixed behavioral replacements remain applicable across all code-check categories.
这段在干什么:交代附录 H 的检查流程——用哈希挑出各类别里的模型源码,再送到指定位置和公开融合接口做检查。
需要解释的地方:
值得留意:检查范围被明确限定(restricted)在预先指定的位置和公开接口,即不覆盖全部代码;而行为替换对所有类别一视同仁,是流程里唯一"通吃"的部分。
b195A linked sample of 48 real candidate sources extends testing beyond the one-key attention case. Of these, 47 change their predictions after compound replacement at both boundaries; 20 have positive target-loss scaffold intervals at both. An implemented multi-tower source is exactly invariant. Source consumption, fitted dependence, and target benefit thus separate in diverse generated code (Appendix M).
用 48 个真实候选源码做更大范围的抽样检验,证明上一段“行为替换适用于所有代码检查类别”的说法不止于单键注意力这一个案例。
48 个里 47 个替换后预测会变、20 个两端都有正区间——作者用这些数字支撑“三者可分离”的结论;但只有 1 个实现的多塔源码完全不变,样本极小。数字差异的原因作者指向附录 M,本段未展开。
b197The prediction-space analysis measures normalized RMS change and cosine change between f(c,p,a)f(c,p,a) and predictions obtained after fixed input permutations. Table 54 reports prediction-change RMS divided by the target RMS after centering each output feature across evaluation rows, with a denominator floor of 10−1210^{-12}. Both RMS quantities average over rows and features within the reported response block. This complements target-based effects: near-zero distance indicates negligible prediction change on the tested replacements, whereas nonzero distance with near-zero target-loss gain indicates sensitivity with little measured benefit on the observed response. The calculation requires inference only over fixed checkpoints.
1. 这段在干什么
提出一种「预测空间」审计指标:看固定输入被替换后,模型输出本身变了多少,用来补足只看目标损失(target-loss)的视角。
2. 需要解释的地方
3. 值得留意
近零距离≈替换没影响;非零距离但目标增益也近零,说明模型敏感却没带来实际收益——这是作者区分"敏感性"与"有用性"的关键。
b198Model Compound replacement Control replacement Global CP L1000 Global BBBC036 CellScientist 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.3112 ±\pm 0.0246 Chemical-only 0.0560 ±\pm 0.0013 0.0580 ±\pm 0.0014 0.0552 ±\pm 0.0013 0.0000 ±\pm 0.0000 Control-only 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.3390 ±\pm 0.0199 BBBC047 CellScientist 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.3593 ±\pm 0.0207 Chemical-only 0.1253 ±\pm 0.0066 0.1467 ±\pm 0.0093 0.1025 ±\pm 0.0055 0.0000 ±\pm 0.0000 Control-only 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.4011 ±\pm 0.0199
这段是一张表格式的数据对比,用来审计模型预测到底依赖哪部分输入,是附录里支撑"预测空间依赖"的证据。
在干什么:用替换实验(换掉化合物 / 换掉对照)看各模型的预测变化,检验预测究竟依赖什么。
关键概念:CP 与 L1000/BBBC036 是任务或数据集名,数值是评分(带标准差)。"Compound replacement"是把化合物换掉,"Control replacement"是换掉对照组——领域常识:把真正有用的输入换掉,分数应明显下降。
容易读漏:CellScientist 在前两列(换化合物、换对照)全是 0.0000,却在 L1000 和 BBBC036 上有 0.31、0.36 这类非零值;Chemical-only 恰好相反——前几列有值,BBBC036 一列全是 0.0000。这说明两者依赖的来源不同,作者就此没在正文里多解释。
b200The main text follows one model from selection through source checking and held-out replication. This section complements that input-use analysis with two preselected trajectory-level run records: a representative successful search and a search containing code repair. Each record preserves the initial hypothesis, decisive score changes, component edits, failures, repair actions, prompt and source hashes, and final selection. The complete ten-slot traces show both successful and failed execution paths.
1. 这段在干什么
承上启下:正文只跟了一个模型的完整流程,附录这里补两个「轨迹级」记录——一个成功搜索、一个含代码修复,用来佐证输入使用分析。
2. 需要解释的地方
3. 值得留意
记录里同时保留了哈希值(prompt/source)和失败路径——作者强调失败样本也是证据,不只是成功案例。原文未说这两个记录具体结论。
b201The two cases are selected using Fold-3 run records only: a lower-median BBBC036 trajectory and the highest-scoring BBBC047 trajectory among those with at least one recorded candidate-code repair. Their records include the initial program, complete candidate sequence, parent hashes, structured decision context, candidate source, execution and repair outcomes, provider request/response hashes, and the versioned prompt constructor that defines the generation input.
交代附录 J 两个案例的挑选标准与记录内容:用 Fold-3 记录选出 BBBC036 的中位偏下轨迹和 BBBC047 的最高分轨迹。
选案例的前提是「至少有一次候选代码修复记录」,即两例都经历过失败—修复;BBBC047 是这类里的最高分。
b203The initial candidate is a reproducible concatenation–MLP joint predictor with joint input fusion, a shared response program, a joint readout, response loss, and AdamW optimization; it obtains Fold-3 Global PCC 0.26680.2668. The recorded decision context for the retained final revision is: “Make one final diagnosis-driven revision of the best incumbent; prioritize a complete, executable perturbation-response hypothesis over a hyperparameter-only change.” The selected multi-head cross-modal-attention candidate uses separate control, compound, and dose encoders, cross-modal attention fusion, a shared response program, joint readout, block-weighted MSE, AdamW, and cosine scheduling, reaching 0.29280.2928. The trace contains 11 documented provider calls (53,096 reported tokens), each with a request hash, response hash, status, and token record.
这段在干什么:展示一次多模态发现轨迹的起点与终点——从初始的拼接-MLP 联合预测器(Fold-3 Global PCC 0.2668)到最终选定的多头跨模态注意力候选(0.2928),并交代支撑这次修订的 provider 调用记录。
需要解释的地方:
值得留意:作者保留的最终修订决策上下文强调的是“做一次完整的、可执行的扰动-响应假设”,而非只调超参——这暗示轨迹的筛选标准偏向机制性探索,而非单纯刷分。
b205The second trace begins from the same initial program at Fold-3 Global PCC 0.28810.2881. Its structured agenda specifies separate control, compound, and dose encoders; compound-query/control-key-value interaction with a dose residual; a shared response program; joint response prediction; CP/L1000 block-weighted MSE (0.4/0.6); AdamW; and cosine scheduling. Its selected first revision implements this design, resolves an output-shape mismatch by returning a [batch,target][\mathrm{batch},\mathrm{target}] tensor, and reaches 0.32300.3230. The complete trace contains 13 provider calls (63,572 reported tokens); a later slot exhausts all three repair attempts and remains recorded as a failed slot.
这段在干什么:展示第二条发现轨迹——从同一初始程序出发,按结构化议程做出第一次修订,性能从 0.2881 升到 0.3230,并交代整条轨迹的开销与一处失败。
需要解释的地方:
值得留意:性能提升只来自"第一次修订",而整条轨迹有 13 次调用、6.3 万 token,且存在一个彻底失败的槽——作者是在用这条"含修复"的轨迹说明流程并非总能成功。
b207Cell Painting, L1000, Perturb-seq, and sci-Plex measure morphological and transcriptional responses across chemical and genetic interventions (Bray et al., 2016; Caicedo et al., 2017; Subramanian et al., 2017; Dixit et al., 2016; Srivatsan et al., 2020). Matched resources and harmonized collections broaden this coverage (Haghighi et al., 2022; Chandrasekaran et al., 2024; Peidli et al., 2024). Predictors use latent-state transfer, compositional models, neural transport, genetic graphs, and foundation models (Lotfollahi et al., 2019; Lotfollahi et al., 2023; Bunne et al., 2023; Roohani et al., 2024; Cui et al., 2024); cycleCDR adds cycle-consistency constraints to learn transferable perturbation representations (Huang and Liu, 2024). CIPHER combines unperturbed-cell covariance with the specified perturbation to predict responses, illustrating the predictive value of baseline cellular structure (Kuznets-Speck et al., 2025). CellAudit asks whether a selected predictor uses the particular perturbation input invoked by its design.
1. 这段在干什么
这是扩展相关工作,梳理细胞扰动数据集、现有扰动预测方法,最后落到本文的 CellAudit:追问预测器是否真用了它声称的扰动输入。
2. 需要解释的地方
3. 值得留意
作者先承认现有方法(含 CIPHER)确用扰动信息,再让 CellAudit 做"使用审计"——重点是验证而非再提一个预测器。
b208PerturbBench and PertEval-scFM standardize predictive comparisons across splits and baselines (Wu et al., 2025; Wenteler et al., 2025). Deep predictors need not outperform linear references, and common metrics can reward systematic variation shared across perturbations (Ahlmann-Eltze et al., 2025; Viñas Torné et al., 2026). Other studies show that well-calibrated predictive metrics and informative foundation-model representations reveal gains over simple baselines (Miller et al., 2025; Hasanaj et al., 2025; Cole et al., 2026). In-the-wild evaluation further examines context and perturbation shifts (Mao et al., 2026). Together, these studies establish the importance of metrics, representations, and transfer conditions. CellAudit tests a complementary property: whether registered input-use claims are supported from cited source computation through fitted dependence to target-relevant predictive contribution.
做文献定位:先承认已有基准/指标/表征研究已确立重要性,再声明 CellAudit 检验的是一个互补性质——输入使用声明能否从源计算一路支持到预测贡献。
作者用"Together…"先收束前人三条线(指标、表征、迁移条件),再一句"complementary property"把 CellAudit 摘出去——别把前人的结论误当成本文结果。本段未给任何数字或结论。
b209Scientific agents span laboratory workflows, program search, and iterative machine-learning experimentation (Boiko et al., 2023; Romera-Paredes et al., 2024; Huang et al., 2024; Li et al., 2024; Lu et al., 2026; Jiang et al., 2025). CellScientist, CellForge, and HarmonyCell apply related search procedures to cellular prediction (Li et al., 2026; Tang et al., 2025; Huang et al., 2026). The existing CellScientist policy supplies candidate models in these experiments; CellAudit links their registered input-use claims to deterministic tests of source consumption, fitted dependence, and target-relevant predictive contribution.
这是在铺相关工作时给本文定位:先列科学智能体的既有工作,再把它自己的 CellAudit 与 CellScientist 区分开。
CellScientist 不是被批判的对象,而是提供候选模型的现有策略,CellAudit 是在它上面加审计环节——承接上一段那三种支撑层级。
b210Permutation reliance measures performance changes after input disruption; conditional permutations account for dependence among inputs (Fisher et al., 2019; Chamma et al., 2023). Shortcut and underspecification studies show that benchmark success can coexist with unintended or unstable decision rules (D’Amour et al., 2022; Geirhos et al., 2020; Lapuschkin et al., 2019; DeGrave et al., 2021). Attribution sanity checks, removal-based evaluation, and behavioral suites test explanations and capabilities under targeted interventions (Adebayo et al., 2018; Hooker et al., 2019; Ribeiro et al., 2020). ConceptSMILE audits concept-explanation reliability through input perturbations and local surrogate modeling (Mollapour et al., 2026); POPPER tests free-form hypotheses through agentic sequential falsification with statistical error control (Huang et al., 2025). CellAudit follows one selected cellular-response model and its registered input-use claim from the cited code to complete fitted predictions and later-data evaluation. This links source implementation, fitted dependence, and target-loss changes under registered replacements, distinguishing a blocked pathway, sensitivity without predictive benefit, and a contribution that recurs on new data.
这是扩展相关工作,把 CellAudit 放进已有方法谱系里做定位:前面综述扰动/归因等已有做法,最后一句才点出本文的不同。
末尾三分类——"被阻断的通路 / 无预测收益的敏感性 / 新数据上复现的贡献"——是本文区别于上述所有方法的关键,容易读漏。
b212An explicit input route establishes a possible computation; its fitted contribution is tested with the checkpoint and observed targets held fixed. We adapt permutation reliance and behavioral testing to the registered biological inputs (Fisher et al., 2019; Ribeiro et al., 2020). For a chemical-response model, the 2×22\times 2 test combines the correct or shuffled perturbation with the correct or an alternative control profile. Let mpcm_{pc} denote the response score with both correct inputs, mp~cm_{\tilde{p}c} the score after a matched perturbation shuffle, mpc~m_{p\tilde{c}} the score after a control-profile swap, and mp~c~m_{\tilde{p}\tilde{c}} the score after both. For a higher-is-better metric mm,
提出一个 2×2 因子置换检验,用扰动与控制两种输入的正确/打乱组合,测量每个输入对模型输出的实际贡献。
原文停在「For a higher-is-better metric $m$」——公式被截断,这段没给出具体对比式。此处是在检验「输入是否真被用到」,而非只验证输入路径存在。
b213These are factorial contrasts in predictive score under the registered replacement distribution: EpertE_{\mathrm{pert}} averages perturbation-replacement effects across the two tested contexts, while Ipert,contextI_{\mathrm{pert,context}} measures their nonadditivity in score. For compound inputs we write EcmpdE_{\rm cmpd} and EctrlE_{\rm ctrl} in tables. The loss contrast EcmpdℒE^{\mathcal{L}}_{\rm cmpd} similarly averages the increase after compound replacement over both context states. Each checkpoint uses 32 fixed permutations constructed without response values or model scores. Continuous effects and their intervals are the primary quantitative evidence. For a per-model stability summary, let τ\tau be the 95th percentile of absolute pairwise differences among the corresponding reference scores. A contribution is qualified if its effect’s empirical 2.5th percentile exceeds τ\tau, unsupported if its 97.5th percentile does not exceed τ\tau, and inconclusive otherwise. These labels summarize effect stability relative to map-induced score variation; an unsupported effect may still be positive. The empirical quantiles describe replacement variability within a checkpoint; Appendix B specifies the reference scores and separate across-model uncertainty.
1. 这段在干什么
定义附录里用的因子对比记号(E、I),并给出一套把「效应大小」与「稳定性」对照的判断规则,用于给每个化合物输入打稳定性标签。
2. 需要解释的地方
3. 值得留意
b214Reference-score variation can include variation induced by the other input. Holding the compound effect fixed while increasing this variation can therefore change the status. An existing synthetic check separates invariant, sensitive-without-benefit, target-relevant, and harmful predictors. All 96 target-relevant instances have positive mean effects, while 32 of 96 exceed the reference-relative criterion, identically with 32, 128, or 512 maps. The continuous measurements recover predictive benefit, whereas the categorical rule asks whether that benefit exceeds the specified reference variation. Its threshold is an operational effect scale. The additional 6,400-dataset study in Appendix F.1 tests estimator bias, interval coverage, and sensitivity to the reference scale; it distinguishes conditional replacement uncertainty from population uncertainty and retains the observed finite-group undercoverage.
这一段先给个定位:它是在说明评估标准本身的脆弱性——参考分数的波动会「污染」判定结果。
需要解释的地方:所谓「reference-score variation」指参照基准本身也在变动,而基准一动,同一个 predictor 可能就从「target-relevant」掉到别的类别里。作者举了个领域常识性的例子:把 compound effect 固定住、只抬高波动,状态就会翻。那个「96 个实例、32 个超标」的合成检验,是在验证分类规则和连续测量结果是否一致。
值得留意:两个数字要分清——「96 个全有正效应」但只有「32 个超过标准」,且 32、128、512 三种 map 数下都一模一样。作者没明说的是:连续测量能捞回真实收益,而类别规则只是在问「收益够不够超过那个操作性的阈值」。
b217The frozen replay changes one encoded coordinate while holding the source row’s other inputs fixed. Eligible donors satisfy the recorded metadata constraints and produce an actual input change. Table 57 reports the supported populations separately from the original context-averaged factorial summaries. Score-selected BBBC047 remains exactly invariant; the path-constrained and guided panels retain positive compound effects on both boundaries. On LKCP, dose intervals are positive and compound intervals include zero. For the replay and real-source audit below, 95% intervals resample source groups while conditioning on the selected checkpoint set, donor pools, and 128 replacement maps; BBBC and LINCS use Murcko source scaffolds, and LKCP uses registered BRD identifier groups. The matched-reference and control-bank intervals instead describe paired training-seed variation at the fixed task split. Individual intervals and positive-interval counts are descriptive and unadjusted.
1. 这段在干什么
交代单坐标重放实验的设定(只改一个编码坐标、其余输入冻结),并报告各面板的区间结果,同时说明不同区间分别刻画什么来源的不确定性。
2. 需要解释的地方
3. 值得留意
b218Original-entry-weighted means: score-selected BBBC047 has 5 entries per fold; path-constrained and guided BBBC047 have 10 each; LINCS and LKCP have 50 each. LKCP’s 50 entries contain 30 unique weight hashes. Path-constrained and guided checkpoints are fixed across the two boundary evaluations. BBBC047 has no registered observed-support dose contrast.
这段在干什么:交代单坐标回放实验用的数据规模与权重构成,为后续评估设前提。
需要解释的地方:
值得留意:样本量差很大(5 vs 10 vs 50),且 LKCP 的 50 条去重后只剩 30 个唯一权重——实际独立信息比表面少。BBBC047 那句是在说明它无法做该类对比。
b220The four input subsets use the same masked-MLP family, loss, fit budget, and five paired training seeds. Each selected checkpoint is evaluated on both held-out folds. Here cc denotes context, pp compound representation, and aa dose. Dose improves the LINCS reference; its increment is small on BBBC and LKCP. On LKCP Fold 5, the full reference has a positive PCC increment over g(c,a)g(c,a), while its MSE interval includes zero. These trained-subset comparisons describe the fixed reference family; within-checkpoint replacement measures use in the separately selected agent models.
1. 这段在干什么
交代"匹配输入-参考族"实验的公平性设置,并汇报在固定参考族下四个输入子集的评估结果,属于方法说明兼结果过渡。
2. 需要解释的地方
3. 值得留意
b221Paired increments and their 95% intervals appear in Table 59; the portable source preserves all per-reference and paired intervals at full precision.
这段是交代数据出处:配对增量及其 95% 区间放在表 59,并声明可移植源文件保留了全部逐参考与配对的完整精度区间。
概念:
留意:原文只说「完整精度存在源文件里」,正文表格很可能是四舍五入过的;要拿精确数值得去看那个源文件,而不是表 59。
b223A fixed, dataset–policy-balanced and source-structure-stratified sample links 48 source records to 96 same-checkpoint boundary evaluations. Implemented source consumption coexists with distinct fitted outcomes: an implemented MultiTowerFusion source is exactly invariant on both boundaries, whereas 47 sources change predictions and 20 have positive target-loss intervals on both boundaries. Unresolved source locations also admit complete-model behavioral tests.
1. 这段在干什么
交代实证样本的构造方式,并给出「实现层面被使用」与「预测层面是否有贡献」的对照结果——即后文审计统计的取样基础。
2. 需要解释的地方
3. 值得留意
「implemented」(代码里真被调用)不等于「有贡献」:文中那个 MultiTowerFusion 实现后被用,却在两个边界上都完全不变。另有 47 个改变预测、20 个在两个边界上都有正损失区间。无定位的来源也能做整模型行为测试。
b224The middle behavioral category has positive prediction distance but lacks a positive target-loss interval on at least one fold. Source status describes the cited computation. Counts describe this structurally sampled set and its scripted source classifications. Claim origin remains task_contract where no candidate-authored manifest was emitted.
1. 这段在干什么
给中间那类候选模型下定义:有正预测距离,但至少在某一折上没有正的目标损失区间。后面几句交代来源状态、计数和申报出处。
2. 需要解释的地方
3. 值得留意
它是在补齐前段"两类边界"之外的中间情形,作者没明说这类算通过还是失败,但"lacks"一词暗示它是被单独标记的一档,不是简单的达标或不达标。
b226The existing control-bank comparison holds the response target Y−RY-R fixed and changes whether the context input reuses reference bank RR or uses the other bank. Reference A/B assignment is fixed before fitting and counterbalanced within cell-line, replicate, and control-layout strata (24 plate groups per orientation). Three input sets, two bank assignments, and five paired seeds give 30 fitted checkpoints. All six shared-minus-disjoint PCC intervals include zero. This comparison measures control-reference reuse on sci-Plex.
1. 这段在干什么
报告一个对照实验(control-bank comparison):固定响应目标 Y−R,只切换上下文输入用的是不是参考库 R,以测量 sci-Plex 上的"control-reference reuse"。
2. 需要解释的地方
3. 值得留意
六个区间全含零,说明这个对照里参考库复用没有产生显著差异——但作者只说"测量了",没直接下结论。
b228Predictor F4 PCC ↑\uparrow F4 MSE ↓\downarrow F5 PCC ↑\uparrow F5 MSE ↓\downarrow Anchor g(c,a)g(c,a) 0.9664 ±\pm 0.0001 0.0308 ±\pm 0.0001 0.9728 ±\pm 0.0002 0.0251 ±\pm 0.0001 Score 0.9756 ±\pm 0.0027 0.0225 ±\pm 0.0025 0.9764 ±\pm 0.0015 0.0219 ±\pm 0.0013 Audit 0.9779 ±\pm 0.0004 0.0204 ±\pm 0.0004 0.9778 ±\pm 0.0002 0.0205 ±\pm 0.0002
这段在干什么:这是一张结果表,横向比三种模型(Anchor / Score / Audit)在 F4、F5 两个预测器上的表现,用 PCC(越高越好)和 MSE(越低越好)两组指标呈现,配合上段"共享减不共享"的对比,展示 Audit 的预测效果。
需要解释的地方:
值得留意:Audit 在两个预测器、两个指标上都优于另两者;而标准差普遍很小,说明差距不是偶然波动。
b229The anchor is shared within each seed. F4 and F5 are held-out combination boundaries of the sci-Plex development source.
上一段刚列完几组指标数字,这段是给这些数字补实验设定的。
1. 在干什么:交代这些预测结果所依据的数据划分口径,是表注式的说明。
2. 解释:anchor(锚点)——领域常识里通常指作为基准的那份参照数据或配置;这里说它「在每个 seed 内共享」,即同一随机种子的各次实验都对比同一个锚点,避免锚点差异混进结果。held-out combination boundaries(留出的组合边界)——指刻意留出、不参与训练的那部分组合条件,用于检验外推。
3. 值得留意:F4、F5 是留出的,不是训练用的;而 anchor 反而是共享的。两者性质不同,别把 F4/F5 当成普通数据点读。
b230Set Feedback Chemical PCC drop Chemical loss gain Dose PCC drop Dose loss gain F4 Score 0.0203 ±\pm 0.0046 0.0186 ±\pm 0.0042 0.0068 ±\pm 0.0026 0.0062 ±\pm 0.0024 F4 Audit 0.0240 ±\pm 0.0016 0.0220 ±\pm 0.0015 0.0090 ±\pm 0.0006 0.0082 ±\pm 0.0006 F5 Score 0.0070 ±\pm 0.0016 0.0064 ±\pm 0.0015 0.0039 ±\pm 0.0017 0.0035 ±\pm 0.0015 F5 Audit 0.0086 ±\pm 0.0005 0.0079 ±\pm 0.0005 0.0049 ±\pm 0.0004 0.0045 ±\pm 0.0004
用一组对照数字,说明在 F4、F5 两个留出边界上,审计(Audit)比原始评分(Score)表现更强。这是正文结论的附表证据。
四列里每一列,Audit 的值都高于对应的 Score,且±波动普遍更小——即更稳且更高,作者没在原文里点破。
b231Positive values favor the correct input over its legal replacement. Each displayed coordinate metric is positive in 5/5 endpoints in both arms at both boundaries.
1. 这段在干什么
给上一段那串数字(各端点、各臂的指标)下判读规则:往哪个方向读才算"对",结论一句话说的是 5/5 全为正。
2. 需要解释的地方
3. 值得留意
b232Set Metric Mean Δ\Delta Paired tt 95% CI Joint bootstrap 95% CI Wins F4 PCC +2.364 [-1.277, +6.005] [+0.253, +5.619] 4/5 F4 MSE -2.159 [-5.482, +1.164] [-5.136, -0.234] 4/5 F5 PCC +1.460 [-0.496, +3.416] [+0.202, +3.021] 3/5 F5 MSE -1.333 [-3.116, +0.450] [-2.748, -0.187] 3/5
这段在干什么:展示 F4、F5 两个模型组在 PCC 和 MSE 两个指标上的匹配反馈结果,用效应量、置信区间和胜出次数汇报预测层面的变化。
需要解释的地方:
值得留意:Paired t 区间都跨 0(如 F4 PCC 的 [-1.277, +6.005]),而 bootstrap 区间不跨 0,两者结论不一致——作者没说哪个更可信,读时别只看"Wins 4/5"。
b233Descriptive paired comparison. Paired tt: five trajectories, df=4. Joint bootstrap: 4,000 shared trajectory–compound draws. Wins favor higher PCC or lower MSE. Both uncertainty summaries retain all five trajectory pairs.
这段用三句话交代配对比较的统计口径:先给 t 检验(5 条轨迹,df=4),再用联合 bootstrap 做 4,000 次共享轨迹–化合物抽样;并说明胜负判据是 PCC 更高或 MSE 更低,两种不确定性汇总都保留全部 5 对轨迹。
概念上,df=4 是配对 t 检验的自由度(5 对减 1,领域常识);bootstrap 是重抽样估不确定性的通用方法;PCC 是皮尔逊相关系数。
留意:判据是"越高/越低越好",不是数值大小;4,000 次是抽样次数,非样本量。
b234Feedback Selected PCC Valid Failed Repairs LLM calls Reported tokens Training wall s Score 0.9757 25/25 0 11 36 ≥\geq286,995 194.93 Audit 0.9781 25/25 0 13 38 592,347 217.41
这段是附录里的对比表,比较两轮实验(Feedback 与 Audit)的预测与审计结果。
b235Selected PCC averages the five discovery-selection scores. Calls include repairs; the Score token total is a lower bound with one missing usage receipt. Five shared anchors contribute a further 32.61 training seconds, counted once.
1. 这段在干什么
给上一段表格里的数字做口头备注:说明这些分数、调用次数、token 和训练时间是怎么算出来的,属于对结果的口径澄清。
2. 需要解释的地方
3. 值得留意
Score 的 token 总数是下界,因为缺了一条用量凭证——真实值应更高。32.61 秒是共享锚点额外加的训练时间,与上一段 194.93/217.41 的口径衔接。
b237We analyze Score and Audit feedback using all five paired discovery trajectories and their shared context-plus-dose anchors. The descriptive comparison holds model space, trainer, selection rule, and candidate budget fixed, and reports every selected endpoint, paired effect, and uncertainty interval. Within the sci-Plex development source, compound–cell-line roles were reassigned and frozen before response preprocessing. F4 and F5 evaluate the same selected checkpoints on held-out combinations.
1. 这段在干什么
交代 Score/Audit 反馈对比的实验设置:用五条配对轨迹,控制变量,报告全部端点与不确定性,并说明 sci-Plex 数据里角色在预处理前已冻结。
2. 需要解释的地方
3. 值得留意
控制项列得很全(模型空间、训练器、选择规则、候选预算),但共用的只有锚点——两套反馈真正的差别只剩反馈本身,这是对比能成立的关键。
b238The fit, selection, F4, and F5 packages contain 2,666, 891, 438, and 442 treatment wells. Complete compound ×\times cell-line groups, including their doses and replicate wells, remain within one role; evaluated compounds and cell lines each appear among the fit marginals. The 2,000 target genes are selected by variance among fit treatment wells, with fit-only feature scaling. Control bank A supplies inputs and bank B the reference; they contain distinct wells matched on plate, cell line, and replicate, with control pools reused across these development roles. The target is the complete gene-expression profile.
1. 这段在干什么
交代数据划分的细节:各角色(fit/selection/F4/F5)的样本量、化合物×细胞系组合归入单一角色、目标基因与对照库的选取方式。是方法可复现性的补充说明。
2. 需要解释的地方
3. 值得留意
"remain within one role"意味着同一组合不跨角色——这是防止数据泄漏的关键设计;作者未明说,但读者应意识到。
b239Score supplies selection metrics, sampled learning curves, model size, timing, and past designs. Audit adds grouped errors, component-source clues, and chemical/dose replacement effects. This comparison estimates the complete enhanced-feedback package. Both arms share g(c,a)g(c,a) within seed, five new candidate slots, and three repair opportunities per slot. The language model may revise encoders, fusion, readout, capacity, regularization, and residual scale; a direct predictor or frozen-anchor residual is equally available. Each new component is trained from scratch. Completed selection diagnostics inform subsequent proposals; final checkpoints are frozen before F4/F5 evaluation.
1. 这段在干什么
交代两组对比(Score 与 Audit)各自提供什么信息,并列出二者共享的实验设置,为评估"完整增强反馈包"做铺垫。
2. 需要解释的地方
3. 值得留意
两臂除反馈内容外其余条件(同种子、五个候选槽、每槽三次修复机会、从头训练、最终 checkpoint 冻结后才评估)全部对齐——即差异被归因于反馈本身。
b240AdamW uses learning rate 10−310^{-3}, weight decay 10−510^{-5}, batch size 128, gradient clipping at 5, a 100-epoch maximum, and early-stopping patience 15. The shared loss is fit-standardized target MSE plus 0.02(1−PCC)0.02(1-\mathrm{PCC}), using flattened Pearson correlation. The common limit is five million trainable parameters. Epoch and endpoint selection prioritize selection Global PCC, then lower MSE; endpoint ties additionally prefer fewer parameters and earlier slots. The anchor stays frozen in residual models.
1. 这段在干什么
交代实验的统一训练配置——优化器、损失、规模上限、选点规则,让后续所有模型在公平条件下比较。
2. 需要解释的地方
1−PCC 越小越好。3. 值得留意
选点顺序是先看 Global PCC、再看 MSE,平局才比参数量和更早的 slot——这是明确但容易读漏的优先级。
b241SD uses denominator n−1n-1 across five trajectories; paired two-sided Student-tt intervals use df=4. Compound-only bootstrap fixes the fitted trajectories, trajectory-only bootstrap fixes the evaluated compounds, and joint bootstrap resamples both; each uses 4,000 shared-index draws. Appendix N.2 compares these intervals with complete leave-one-pair-out and exact sign-flip analyses. Maps are averaged within endpoints before trajectory inference. RMS measures sensitivity; PCC drop is correct-input minus replacement PCC, and loss gain is replacement minus correct-input MSE. Positive continuous effects are distinct from individual qualification or biological-mechanism evidence.
这段在干什么:这是附录的“记录说明”段,交代本研究所用统计口径(区间与重采样方案),把方法细节存档备查。
需要解释的地方:
值得留意:三种 bootstrap 区分“固定谁、重抽谁”,决定了不确定性归因于哪一方;末句提醒正向连续效应≠个体资格或机制证据。
b242Recorded model identifier: deepseek-v41. All compared calls use this serving identity. Trajectory indices 1–5 correspond, in order, to seeds 2026091201, 2026091202, 2026091203, 2026091204, 2026091205. Full-precision endpoint, cost, and paired-effect tables accompany the renderer under results/feedback_comparison/.
这段在干什么:这是一条记录性说明,交代复核时用的模型版本、轨迹编号与随机种子的对应关系,以及配套数据放在哪。它不承担论证,属于可复现性交代。
需要解释的地方:
值得留意:五个种子与轨迹 1–5 是按顺序一一对应的,别读乱;完整表格没写进正文,只在 results/feedback_comparison/ 目录里,正文只给指针。
(正文没提这些数字的含义或结论。)
b243Relative MSE reduction against the anchor is computed within seed as 1−MSEendpoint,s/MSEanchor,s1-\mathrm{MSE}_{\mathrm{endpoint},s}/\mathrm{MSE}_{\mathrm{anchor},s} and then averaged. A ratio of arm mean MSEs is reported separately when used. Costs retain failed attempts and recorded interruptions. Shared anchor training totals 32.611546 seconds and is counted once, separately from each condition’s candidate-training costs. Trainer and worker wall times overlap; missing token usage produces a lower bound.
1. 这段在干什么
交代评测指标的算法与成本口径——怎么算相对 MSE 降幅、成本怎么计、锚点训练时间怎么摊。
2. 需要解释的地方
1 − 端点MSE / 锚点MSE,再跨 seed 平均;即「比基线好了百分之多少」。3. 值得留意
b244Tr. Model Slot PCC MSE C-R C-P C-L D-R D-P D-L Fold 4 1 Anchor 0 0.96652 30.768 0.000 0.000 0.000 28.653 1.152 1.042 1 Score 5 0.97557 22.551 141.301 20.580 18.805 84.343 6.798 6.212 1 Audit 5 0.97806 20.278 142.171 23.162 21.173 94.176 8.838 8.081 2 Anchor 0 0.96646 30.836 0.000 0.000 0.000 25.007 0.985 0.896 2 Score 3 0.97740 20.889 136.127 21.774 19.882 84.750 7.580 6.923 2 Audit 4 0.97809 20.252 155.045 25.294 23.142 101.848 9.453 8.650 3 Anchor 0 0.96647 30.813 0.000 0.000 0.000 31.861 1.311 1.189 3 Score 3 0.97570 22.432 138.732 20.644 18.823 96.003 7.982 7.278 3 Audit 3 0.97817 20.181 138.099 22.698 20.763 91.844 8.613 7.880 4 Anchor 0 0.96646 30.820 0.000 0.000 0.000 26.818 1.066 0.960 4 Score 5 0.97807 20.274 157.142 25.704 23.535 96.377 9.070 8.307 4 Audit 5 0.97725 21.023 145.497 22.784 20.858 90.603 8.193 7.502 5 Anchor 0 0.96629 30.977 0.000 0.000 0.000 17.193 0.640 0.580 5 Score 4 0.97115 26.587 117.806 13.009 11.887 48.269 2.392 2.186 5 Audit 2 0.97815 20.203 159.745 26.108 23.941 102.759 9.697 8.893 Fold 5 1 Anchor 0 0.97265 25.253 0.000 0.000 0.000 27.811 0.712 0.647 1 Score 5 0.97586 22.332 100.896 7.672 7.048 65.922 3.595 3.306 1 Audit 5 0.97785 20.495 90.349 8.720 7.978 71.763 5.036 4.612 2 Anchor 0 0.97297 24.955 0.000 0.000 0.000 24.562 0.606 0.553 2 Score 3 0.97777 20.567 80.270 7.558 6.905 71.301 4.866 4.452 2 Audit 4 0.97772 20.619 90.797 8.556 7.835 74.695 4.693 4.303 3 Anchor 0 0.97297 24.957 0.000 0.000 0.000 31.737 0.839 0.764 3 Score 3 0.97626 21.947 92.154 7.588 6.919 75.540 4.604 4.202 3 Audit 3 0.97815 20.230 86.705 8.379 7.691 71.783 4.883 4.487 4 Anchor 0 0.97270 25.230 0.000 0.000 0.000 26.502 0.688 0.621 4 Score 5 0.97771 20.623 89.824 8.217 7.526 77.179 5.209 4.776 4 Audit 5 0.97758 20.744 88.060 7.995 7.320 70.291 4.508 4.131 5 Anchor 0 0.97292 24.999 0.000 0.000 0.000 17.108 0.399 0.363 5 Score 4 0.97425 23.789 75.832 4.122 3.769 40.870 1.079 0.988 5 Audit 2 0.97786 20.504 99.006 9.405 8.635 83.693 5.505 5.060
这段在干什么:这是附录里的完整数据表,逐折、逐随机种子列出 Anchor/Score/Audit 三种模型在 PCC、MSE 和六项成本指标上的原始记录,供复核前面正文的对比结论。
需要解释的地方:PCC 是预测值和真值的相关系数(越大越准,领域常识);MSE 是均方误差(越小越好,领域常识);C-R/C-P/C-L 和 D-R/D-P/D-L 是各类训练或贡献成本指标,具体定义这段没给,需回正文查。
值得留意:Anchor 行的成本列全为 0,说明它是零成本基线,不是真实测量值;Audit 的 PCC 和 MSE 通常略优于 Score,但成本也更高,得分与代价之间是有取舍的。
b245PCC is unscaled; all other numeric outcomes are in 10−310^{-3} units. C/D: chemical/dose; R: prediction RMS; P: PCC drop; L: target-loss gain. Slot 0 is the shared anchor. RMS has no preferred direction.
上一段末尾刚给完一条完整记录(含 RMS、PCC 等数字),这段是它的读表说明——交代表里各列的单位、缩写和方向,属方法性脚注。
需要解释的地方
值得留意
除 PCC 外,所有数值都乘了 10⁻³,比较数字时别忽略这个量纲。
b246Tr. Arm Slot Sel. PCC Repair Calls HTTP Tokens Train s Work s Diag. s 1 Score 5 0.97491 1 6 7 ≥\geq50,685 48.09 67.26 7.10 1 Audit 5 0.97828 2 7 7 106,239 41.24 60.03 12.10 2 Score 3 0.97757 3 8 8 62,703 41.96 60.07 7.50 2 Audit 4 0.97783 3 8 8 125,047 42.32 61.01 12.94 3 Score 3 0.97628 4 9 9 70,709 40.04 58.06 7.36 3 Audit 3 0.97825 3 8 8 136,386 44.25 61.58 12.15 4 Score 5 0.97852 3 8 8 68,757 35.40 52.75 8.73 4 Audit 5 0.97798 3 8 8 129,227 39.60 55.93 12.24 5 Score 4 0.97145 0 5 5 34,141 29.44 46.61 6.75 5 Audit 2 0.97830 2 7 7 95,448 50.00 67.66 11.96
这段是附录里的原始记录表:把 5 组模型各自的 Score 与 Audit 两次运行逐行列出,供读者核对上一段摘要背后的完整数据。
需要解释的地方
值得留意
Audit 行普遍 PCC 更高,但 Calls、Tokens、耗时也大幅上升(如第 5 组 Token 从 34,141 涨到 95,448)——精度提升伴随成本膨胀。表中只给数据,未作评价。
b247Every trajectory completed 5/5 new candidates and had zero fully failed slots. Train is known completed trainer wall time; Work is candidate worker wall time; Diag. is postfit diagnostic wall time. Logical calls and physical HTTP attempts are distinct. Complete token breakdowns and missing-cost flags are in the CSV.
1. 这段在干什么
它是 N.1 小节的数据说明,交代这张表里各列的定义,以及样本完整性。
2. 需要解释的地方
3. 值得留意
作者先声明"零完全失败",但下一句就点出逻辑调用与 HTTP 尝试不同——即可能有重试或部分失败被藏在这两个计数差里。完整 token 拆分与 missing-cost 标记不在正文,在 CSV。
b248Set Metric Mean Δ\Delta Paired 95% CI +/–/0 F4 PCC +2.364 [-1.277, +6.005] 4/1/0 F4 MSE -2.159 [-5.482, +1.164] 1/4/0 F4 Chem. RMS +9.890 [-16.189, +35.969] 3/2/0 F4 Chem. PCC drop +3.667 [-3.577, +10.912] 4/1/0 F4 Chem. loss gain +3.389 [-3.269, +10.048] 4/1/0 F4 Dose RMS +14.298 [-16.033, +44.628] 3/2/0 F4 Dose PCC drop +2.194 [-1.639, +6.028] 4/1/0 F4 Dose loss gain +2.020 [-1.496, +5.536] 4/1/0 F5 PCC +1.460 [-0.496, +3.416] 3/2/0 F5 MSE -1.333 [-3.116, +0.450] 2/3/0 F5 Chem. RMS +3.189 [-13.712, +20.089] 2/3/0 F5 Chem. PCC drop +1.580 [-1.069, +4.228] 4/1/0 F5 Chem. loss gain +1.458 [-0.979, +3.896] 4/1/0 F5 Dose RMS +8.283 [-16.535, +33.100] 3/2/0 F5 Dose PCC drop +1.054 [-1.483, +3.592] 3/2/0 F5 Dose loss gain +0.974 [-1.355, +3.303] 3/2/0
这段在干什么:这是 N.1 节的完整数据表,逐行记录两组模型(F4、F5)在 8 个指标上「score–audit」反馈对比的均值差、配对 95% 置信区间和 +/-/0 计数。
需要解释的地方:Δ 是两次评分之差;Paired 95% CI 是配对样本的置信区间,跨 0 说明差异不稳健;「+/–/0」是各次配对里改善/变差/持平的出现次数,三者之和即样本量(F4 多为 5,F5 各行为 5)。
值得留意:多数 CI 都跨 0,看着均值大(如 F4 Dose RMS +14.298)其实不稳定;F4 PCC、F5 各 PCC 系列几乎全为 4/1/0,方向一致但样本极小,别把「4:1」当成显著。
b249Descriptive effect estimates over all five paired trajectories. Signs count positive, negative, and exactly zero differences, not wins. Lower MSE is favorable; RMS has no favorable direction.
这段是 N.1 小节的总说明,交代下面五条配对轨迹的数字该怎么读。
关键概念:描述性效应估计——只汇总差值、不做显著性检验;3/2/0 是正/负/恰好为零的计数,不是胜负场次。MSE 越小越好,所以方向可判;RMS 没有"好"的方向(领域常识:它是误差量纲的量,无所谓正负优劣)。
值得留意:上段那些 +1.054、+0.974 都是"越大越好"的正向指标,与 MSE 相反,别混着看。
b250Set Metric Paired tt CI Compound-only CI Joint CI F4 PCC [-1.277, +6.005] [+1.300, +3.704] [+0.253, +5.619] F4 MSE [-5.482, +1.164] [-3.389, -1.197] [-5.136, -0.234] F5 PCC [-0.496, +3.416] [+1.006, +1.975] [+0.202, +3.021] F5 MSE [-3.116, +0.450] [-1.800, -0.924] [-2.748, -0.187]
你正在看的这段,其实是上一句的「数据正文」。
这段在干什么
它把 F4、F5 两个模型在 PCC、MSE 两个指标下的配对 t 检验置信区间列全,供读者与正文图表对照。
需要解释的地方
值得留意
区间跨没跨 0 才是重点:跨 0 说明效应不稳健,别只盯着区间宽窄。具体哪些跨 0 需要你逐行核对,这段没替你说。
b251Bootstrap intervals use 4,000 draws. No coordinate-effect bootstrap interval is substituted from predictive statistics, and no new confirmatory tests are introduced.
紧接上面那串区间,这段是给整套审计数字定规矩的,不是新结果。
1. 这段在干什么:交代区间怎么算、边界在哪——只报 bootstrap 区间,不拿预测统计量去顶替坐标效应区间,也不加新的确证检验。
2. 需要解释的地方:
3. 值得留意:「is substituted(不被替换)」是刻意划界——作者在声明坐标效应的区间一律来自坐标效应本身,绝不拿预测侧的数字来凑。这句是防读者误读上面那三列区间怎么来的。
b252Set Coordinate Eligible wells Maps Unique First 32 unique F4 Chemical 438/438 128 128 32 F4 Dose 428/438 128 1 1 F5 Chemical 434/442 128 128 32 F5 Dose 436/442 128 1 1
1. 这段在干什么
这是一张审计汇总表,逐条列出「集合—坐标—合格孔数—映射数—唯一数」,用来完整记录分数审计的反馈对比结果。
2. 需要解释的地方
这是领域常识:表中「合格孔数」写作 438/438 这类分数,前者是实际合格数,后者是总数;128 疑似每组映射的固定规模;F4/F5 是两个实验集合,Chemical 与 Dose 是同一集合下的两种坐标口径。
3. 值得留意
Chemical 行映射为 128,Dose 行却只有 1,两者差异极大——原文未解释原因。另外 Dose 的 432 与 1 两个数字含义也不明,作者没有交代。
b253Replacements match plate, cell line, replicate, and treatment time. Chemical replacement holds dose fixed; dose replacement holds compound fixed and requires at least 0.1 log-dose difference. No cross-group fallback is used. Discovery used 32 maps; final evaluation used 128, with only one distinct dose map. Unsupported wells are excluded, not assigned zero effect.
1. 这段在干什么
交代「替换」的配对规则和实验设置,是上段那张统计表的配套说明。
2. 需要解释的地方
3. 值得留意
最终评估 128 个 map,却只有 1 个剂量 map(对应上表 Dose 行);不支持的孔是剔除而非记 0,避免引入假信号。
b255This retrospective sensitivity analysis retains seeds 2026091201–2026091205 and all eight metrics at both boundaries. Audit minus Score is computed within each selected-trajectory pair; F4 and F5 reuse the same selected checkpoints and do not make N=10N=10. Means and descriptive two-sided 95% t4t_{4} intervals use all five pairs. All 32 sign assignments are enumerated for the absolute unstudentized mean under a pair-exchangeability/sign-symmetry null, not a randomized-assignment guarantee. The smallest attainable two-sided tail probability is 2/32=.06252/32=.0625; no multiplicity-adjusted confirmatory claim is made.
这一段在做稳健性/敏感性分析:用固定5个种子、8个指标,检验“Audit−Score”结果是否只是特定选点造成的。
关键概念:
值得留意:最小双侧尾部概率只有 .0625,即5对样本下无法做到 p<.05;作者明确不做多重校正后的确证声明,结论只是探索性。
b256Fold Metric Mean Paired-tt 95% CI Median +/−/0+/-/0 Tail F4 Global PCC 2.3642.364 [−1.277, 6.005][-1.277,\,6.005] 2.4652.465 4/1/0 6/32 F4 MSE −2.159-2.159 [−5.482, 1.164][-5.482,\,1.164] −2.251-2.251 1/4/0 6/32 F4 Chemical RMS 9.8909.890 [−16.189, 35.969][-16.189,\,35.969] 0.8700.870 3/2/0 12/32 F4 Chemical loss gain 3.3893.389 [−3.269, 10.048][-3.269,\,10.048] 2.3682.368 4/1/0 8/32 F4 Chemical PCC drop 3.6673.667 [−3.577, 10.912][-3.577,\,10.912] 2.5822.582 4/1/0 8/32 F4 Dose RMS 14.29814.298 [−16.033, 44.628][-16.033,\,44.628] 9.8339.833 3/2/0 10/32 F4 Dose loss gain 2.0202.020 [−1.496, 5.536][-1.496,\,5.536] 1.7271.727 4/1/0 6/32 F4 Dose PCC drop 2.1942.194 [−1.639, 6.028][-1.639,\,6.028] 1.8731.873 4/1/0 6/32 F5 Global PCC 1.4601.460 [−0.496, 3.416][-0.496,\,3.416] 1.8881.888 3/2/0 8/32 F5 MSE −1.333-1.333 [−3.116, 0.450][-3.116,\,0.450] −1.716-1.716 2/3/0 8/32 F5 Chemical RMS 3.1893.189 [−13.712, 20.089][-13.712,\,20.089] −1.763-1.763 2/3/0 24/32 F5 Chemical loss gain 1.4581.458 [−0.979, 3.896][-0.979,\,3.896] 0.9300.930 4/1/0 4/32 F5 Chemical PCC drop 1.5801.580 [−1.069, 4.228][-1.069,\,4.228] 0.9990.999 4/1/0 4/32 F5 Dose RMS 8.2838.283 [−16.535, 33.100][-16.535,\,33.100] 3.3943.394 3/2/0 20/32 F5 Dose loss gain 0.9740.974 [−1.355, 3.303][-1.355,\,3.303] 0.2850.285 3/2/0 12/32 F5 Dose PCC drop 1.0541.054 [−1.483, 3.592][-1.483,\,3.592] 0.2790.279 3/2/0 12/32
这段是逐层核对稳健性:把 F4、F5 两对轨迹在各指标上的配对差异摆出来,看结论是否普遍成立。
关键术语:*Mean* 是配对差值均值;*Paired-t* 是配对 t 统计量;*95% CI* 是置信区间;*Median* 是中位差;+/−/0 是正向/负向/无变化计数;*Tail* 是列尾概率(分母 32)。
值得留意:多个指标的 95% CI 跨零(如 F5 Chemical RMS 的 [−13.7, 20.1]),说明方向不稳;Tail 列分母多为 32,与上段 2/32 的离散下限一致——效应再大也难达显著,故不作确认性声明。
b257Tail is the exact sign-flip count out of 32. All 16 paired-tt intervals contain zero and all exact tails are at least .125. RMS measures sensitivity; target-loss gain and PCC drop measure changes in loss and correlation.
这段是五组轨迹对稳健性检验的收尾说明,给上表数据做注解。
关键概念:
值得留意:作者其实在说没有稳健的差异——16 个区间全含 0、所有 tail 至少 .125,即翻不出显著结果。但这段只提"度量什么",没解释表里 3/2/0 的含义。
b258Fold Metric Compound only Trajectory only Joint F4 Global PCC [1.300, 3.704][1.300,\,3.704] [0.144, 4.828][0.144,\,4.828] [0.253, 5.619][0.253,\,5.619] F4 MSE [−3.389,−1.197][-3.389,\,-1.197] [−4.408,−0.132][-4.408,\,-0.132] [−5.136,−0.234][-5.136,\,-0.234] F4 Chemical RMS [6.425, 13.160][6.425,\,13.160] [−4.738, 28.126][-4.738,\,28.126] [−5.652, 27.694][-5.652,\,27.694] F4 Chemical loss gain [2.032, 5.060][2.032,\,5.060] [−0.482, 8.094][-0.482,\,8.094] [−0.456, 8.402][-0.456,\,8.402] F4 Chemical PCC drop [2.201, 5.452][2.201,\,5.452] [−0.534, 8.787][-0.534,\,8.787] [−0.508, 9.067][-0.508,\,9.067] F4 Dose RMS [4.676, 22.279][4.676,\,22.279] [−2.007, 34.959][-2.007,\,34.959] [−3.228, 38.642][-3.228,\,38.642] F4 Dose loss gain [0.689, 3.753][0.689,\,3.753] [0.039, 4.490][0.039,\,4.490] [0.066, 5.631][0.066,\,5.631] F4 Dose PCC drop [0.750, 4.083][0.750,\,4.083] [0.028, 4.884][0.028,\,4.884] [0.068, 6.187][0.068,\,6.187] F5 Global PCC [1.006, 1.975][1.006,\,1.975] [0.302, 2.618][0.302,\,2.618] [0.202, 3.021][0.202,\,3.021] F5 MSE [−1.800,−0.924][-1.800,\,-0.924] [−2.392,−0.274][-2.392,\,-0.274] [−2.748,−0.187][-2.748,\,-0.187] F5 Chemical RMS [−1.073, 7.471][-1.073,\,7.471] [−6.751, 14.920][-6.751,\,14.920] [−7.875, 15.416][-7.875,\,15.416] F5 Chemical loss gain [0.824, 2.180][0.824,\,2.180] [0.248, 3.260][0.248,\,3.260] [0.143, 3.537][0.143,\,3.537] F5 Chemical PCC drop [0.892, 2.364][0.892,\,2.364] [0.266, 3.528][0.266,\,3.528] [0.146, 3.843][0.146,\,3.843] F5 Dose RMS [1.196, 15.103][1.196,\,15.103] [−3.716, 25.621][-3.716,\,25.621] [−5.151, 28.995][-5.151,\,28.995] F5 Dose loss gain [0.324, 1.866][0.324,\,1.866] [−0.261, 2.557][-0.261,\,2.557] [−0.253, 3.122][-0.253,\,3.122] F5 Dose PCC drop [0.349, 2.021][0.349,\,2.021] [−0.294, 2.768][-0.294,\,2.768] [−0.284, 3.397][-0.284,\,3.397]
这段在干什么:把稳健性检验从之前的两条轨迹扩到全部五对轨迹,用区间汇总 F4、F5 六项指标的波动范围。
需要解释的地方:
值得留意:三列区间明显逐列变宽,Joint 最宽;Chemical/Dose RMS 在 Trajectory only 和 Joint 下下界转负,而 Compound only 保持正。
b2594000 draws, seed 69313. Compound weights are shared across all paired arms and trajectories; trajectory draws resample the five paired indices. Joint draws resample both units. The 128 donor maps are fixed. Sufficient statistics are pooled before PCC; compound PCCs are not averaged. Compound-only intervals condition on five selected trajectories; trajectory-only intervals condition on the observed compound set, while joint intervals vary both empirical units.
这段在交代重采样方案与区间口径,确保五对轨迹下的稳健性结果可比。
关键概念:
值得留意:三类区间的条件对象不同——compound-only 固定五条轨迹,trajectory-only 固定观测到的化合物集,joint 两者都变,所以数值不可直接横向比较。128 张 donor map 是固定的,不参与抽样。
b260Trajectory-only percentile bootstrap and paired-tt inference target the same conditional mean paired contrast; their different conclusions reflect the five-point empirical bootstrap distribution versus Student-tt sampling assumptions, not different estimands. Compound-only and joint intervals additionally have different resampling scopes. All predictive bootstrap intervals exclude zero, whereas the paired-tt intervals and exact sign-flip tails do not provide the same evidence. Increasing bootstrap draws does not add independent trajectories or establish broad generalization.
这段在干什么:对比几种统计推断方法为何结论不一致,指出差异来自假设与重采样范围,而非估计目标不同。
需要解释的地方:
值得留意:作者明说「增加自助抽样次数并不能带来独立轨迹或广泛泛化」——别把区间变窄误当成证据变强。
b261Fold Metric All-five mean Fifth-pair d5/5d_{5}/5 Omit-fifth mean All five LOO means F4 Global PCC 2.3642.364 1.3991.399 1.2061.206 [1.206, 3.160][1.206,\,3.160] F4 MSE −2.159-2.159 −1.277-1.277 −1.103-1.103 [−2.886,−1.103][-2.886,\,-1.103] F4 Chemical RMS 9.8909.890 8.3888.388 1.8771.877 [1.877, 15.273][1.877,\,15.273] F4 Chemical loss gain 3.3893.389 2.4112.411 1.2231.223 [1.223, 4.906][1.223,\,4.906] F4 Chemical PCC drop 3.6673.667 2.6202.620 1.3091.309 [1.309, 5.314][1.309,\,5.314] F4 Dose RMS 14.29814.298 10.89810.898 4.2494.249 [4.249, 19.315][4.249,\,19.315] F4 Dose loss gain 2.0202.020 1.3411.341 0.8480.848 [0.848, 2.726][0.848,\,2.726] F4 Dose PCC drop 2.1942.194 1.4611.461 0.9160.916 [0.916, 2.962][0.916,\,2.962] F5 Global PCC 1.4601.460 0.7220.722 0.9220.922 [0.922, 1.859][0.922,\,1.859] F5 MSE −1.333-1.333 −0.657-0.657 −0.845-0.845 [−1.697,−0.845][-1.697,\,-0.845] F5 Chemical RMS 3.1893.189 4.6354.635 −1.808-1.808 [−1.808, 6.622][-1.808,\,6.622] F5 Chemical loss gain 1.4581.458 0.9730.973 0.6060.606 [0.606, 1.875][0.606,\,1.875] F5 Chemical PCC drop 1.5801.580 1.0571.057 0.6540.654 [0.654, 2.030][0.654,\,2.030] F5 Dose RMS 8.2838.283 8.5658.565 −0.352-0.352 [−0.352, 12.075][-0.352,\,12.075] F5 Dose loss gain 0.9740.974 0.8150.815 0.1990.199 [0.199, 1.379][0.199,\,1.379] F5 Dose PCC drop 1.0541.054 0.8850.885 0.2110.211 [0.211, 1.493][0.211,\,1.493]
这段是五组轨迹对的稳健性汇总表:每行一个「模型×指标」,看五对全纳入和只留第五对时结论是否一致。
需要解释的
d5/5:第五对单独贡献占五对均值多少,>1 说明它比平均更大。LOO means:留一法,逐对剔除后重算,看区间是否跨 0。值得留意
d5/5 为负(−1.808、−0.352),且 LOO 区间跨 0——这两项在第五对上方向相反、不稳定。d5/5 为正且 LOO 区间不含 0,说明主要结论不靠单对支撑。b262The fifth-pair column is its contribution to the five-pair mean, not its whole paired difference. Every leave-one-out estimate and t3t_{3} interval is retained in the portable CSV. F5 chemical and dose RMS have positive all-five means but negative omit-fifth means; the F5 chemical-RMS median is negative, whereas the dose-RMS median is positive. All five pairs remain in the primary estimate.
这段在干什么:解释第五对轨迹在五对均值里的权重口径,并说明留一法结果的符号分歧。
需要解释的地方:
值得留意:F5 化学与剂量的 RMS 出现"全五对均值为正、去掉第五对为负"的翻转,且两者中位数符号相反——说明结论对是否包含 F5 敏感。但作者仍把五对全留在主估计里。
b263Retained full predictions independently verify global PCC/MSE and eligible-source correct statistics. Full counterfactual prediction arrays were not retained for this legacy feedback package, so coordinate reconstruction verifies stored sufficient statistics but is not a new independent full-array counterfactual check. These additional descriptions do not reselect checkpoints, retrain models, or revise historical qualification labels.
这段在给结论加限定:说明保留的是完整预测,但反事实数组没留。
需要解释的地方:
值得留意:作者主动承认证据层级不同——全局统计可信,但没做新的完整反事实检验。末句强调这些描述不改变既有结论,防止误读为重新筛选。
b265The two cases are the first two registered trajectory seeds (2026091201 and 2026091202), selected by registration order. Each retains the common initial predictor and all five new audit-feedback proposals. Parent links denote design ancestry: every new candidate is trained from scratch, while the shared context-plus-dose predictor g(c,a)g(c,a) is frozen. This anchor concatenates 2,000 context features with dose and uses a 2001–256–256–2000 MLP with 1,092,816 trainable parameters. Local metrics describe the development-selection partition. The public case package contains the hash-verified anchor source, actual system/user prompts, proposals, candidate source, diffs, and received feedback.
这段在干什么
登记两个轨迹种子案例(2026091201/02),交代它们的设计来源、冻结锚点与公开包内容。
需要解释的地方
值得留意
承上句强调"不改历史标签",这里进一步点明锚点冻结、候选从头训练——即审计改动只发生在候选侧,参照系始终不变。
b266In the first trajectory, slot 2 is the incumbent after its child slot 3 fails to improve selection PCC. Before slot 4, the actual prompt contains slot 2’s best epoch (5), sampled learning curve, group-error diagnostics, source-check findings, and positive compound and dose replacement effects. The model hypothesizes overfitting and proposes reducing encoder capacity; compound memorization is a proposed explanation, not a separately measured biological or learning mechanism. Inspection of the emitted source verifies that each chemical/context encoder changes from two Linear layers to one, with dropout 0.1 added in the encoders and fusion. Dose-conditioned FiLM and the rank-128 bilinear interaction remain. Trainable parameters decrease from 4,056,400 to 2,754,896; total parameters including the fixed anchor are 5,149,216 and 3,847,712. Slot 2 versus slot 4 selection PCC is 0.97808647 versus 0.97818047; MSE is 0.02028296 versus 0.02019429; compound target-loss gain is 0.01846193 versus 0.01949683. The static checker reports syntactic return-dependency hints while leaving functional usefulness unresolved; fixed-model input replacement supplies the separate behavioral evidence.
这段在干什么:它是第一条轨迹的实验记录——报告 slot 3 输给 slot 2、slot 4 提出“过拟合、缩减编码器容量”的假设,然后核对改后源码、参数量和指标。
需要解释的地方:
值得留意:slot 4 的 PCC 和 MSE 都略好于 slot 2,但幅度很小;参数量从 4,056,400 降到 2,754,896(含固定锚点则是 5,149,216→3,847,712)。
b267Both first-case sources zero-initialize their output readout. Under the fixed residual wrapper this initializes the newly fitted predictor at g(c,a)g(c,a). Slot 5 adds a dose-conditioned scalar interaction gate and an extra projected chemical feature before fusion, then wins by selection PCC. The emitted chem_direct feature enters the nonlinear fusion, whereas its proposal calls it a direct additive target-space path. Slot 5 has higher selection PCC but lower compound target-loss gain than slot 4; selection and input contribution therefore remain distinct. The same frozen slot-5 checkpoint is subsequently evaluated on the two held-out sci-Plex partitions.
回答上一段遗留的问题:固定模型下换掉输入后,slot 5 的行为证据是什么——即它赢在选中指标,而非输入贡献本身。
slot 5 的提案称 chem_direct 是「直接加性路径」,实际它进入非线性融合——提案与实现的描述对不上,这正是「审计」的落点。另外,选中指标高≠输入贡献大,两者须分开看。
b268For seed 2026091202, slot 2 decreases selection PCC from its slot-1 parent (0.977817 to 0.977703). Slot 3 improves on slot 2 but remains below slot 1. Slot 4 explicitly returns to slot 1 as its design parent and is selected at PCC 0.977826; its MSE (0.02053288) is slightly worse than slot 1’s (0.02052894), consistent with PCC-first selection. Slot 5 adds explicit pairwise projections but falls to PCC 0.977506 and is not selected. Its rationale claims that zero-initialized new terms preserve the parent, but the shared from-scratch training contract does not inherit fitted parent weights: this claim is preserved as an agent statement in the public proposal, not endorsed as a measured guarantee.
这段在干什么:记录 seed 2026091202 下各 slot 的选择结果,展示选代轨迹,并揭穿 slot 5 的一个说法。
需要解释的地方:
值得留意:slot 1→2 下降、3 回升但仍不如 1、4 直接回到 slot 1;选 4 时 PCC 更好但 MSE 更差,作者点明这是「PCC-first」取向。slot 5 的理由说零初始化能保留父本,作者反驳:共享训练契约并不继承已拟合权重——这只是 agent 声明,不是被验证的保证。
b269The first case uses one JSON-contract repair at each of slots 4 and 5; the second uses one at slot 4 and two at slot 5. These logged failures are completion-token-limit/JSON-contract failures. Lower-scoring, successfully trained candidates are retained without score-driven repair. The case package preserves these outcomes and the actual branch structure.
这段在对比两个案例的修复开销,并交代这些失败被如实记录、未做分数驱动的修补。
关键概念
值得留意