Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

全文来源:arxiv  · 全文共 222 段,已全部带读  · 单段均价 ¥0.002504

b001AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction–claim gap is especially consequential in agentic model discovery, where language-model agents generate and revise predictors using predominantly score-based feedback. We introduce CellAudit, which audits registered input-use claims through three distinct questions: whether the input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed responses. On a paired morphology–transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor achieves a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 yet remains exactly invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key–value attention; the invariance persists after refitting with physically disjoint control wells for inputs and target references. In a stratified audit of 48 generated candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining predictive gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback shows higher mean held-out predictive performance and larger mean compound and dose contributions across five paired trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort further shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not persist. CellAudit therefore adds a falsification layer to agentic model discovery, moving from generate–score–revise toward discover–falsify–revise. Project page: https://limengran98.github.io/CellAudit/.

讲解

1. 这段在干什么

这是全文开篇:提出"预测好 ≠ 真用了输入"这一核心 gap,并引出 CellAudit 的三问审计框架和主要实验结果。

2. 需要解释的地方

  • input-use claim(输入使用声明):模型声称"我用了化合物信息",本文要查这声明是否属实。
  • held-out PCC:留出集上的预测相关性,越高越好(领域常识)。
  • exactly invariant to compound replacement:换掉化合物,预测结果一点不变——说明它根本没用到化合物。
  • falsification(证伪):主动找反例来推翻声明,对应标题的 discover–falsify–revise。

3. 值得留意

  • 0.3153 vs 0.3142:两者预测分数几乎一样,但一个"用"了化合物、一个毫无反应——差距极小却结论相反,这正是"预测–声明 gap"最扎眼的地方。
  • 只知分数反馈(score-based feedback)是 agent 被误导的根源,这是本文动机所在。

1 Introduction

b003AI virtual cells aim to predict cellular responses to chemical and genetic interventions and ultimately support intervention design (Bunne et al., 2024). Yet a model can predict well by exploiting cellular context, control profiles, or other systematic structure while making little or no use of the supplied perturbation information (Ahlmann-Eltze et al., 2025; Viñas Torné et al., 2026). Predictive performance alone therefore does not establish perturbation use (D’Amour et al., 2022; Geirhos et al., 2020; Lapuschkin et al., 2019; DeGrave et al., 2021). We call a mismatch between a registered claim about input use and the model’s implementation or fitted behavior the prediction–claim gap.

这段是 Introduction 的问题提出:先给 AI 虚拟细胞的目标(预测干预反应、辅助设计),再指出一个隐患,最后命名全文核心概念。

关键概念

  • "预测好 ≠ 用到了扰动信息":模型可能靠细胞背景、对照组等系统性结构蒙对,真正输入的扰动却没被用上。领域常识:这叫"捷径学习"。
  • prediction–claim gap:作者自造的术语,指"声称用了某输入"与"代码实现或拟合行为"对不上。

值得留意

性能指标无法证明用了扰动,这是后文要用代码审计 + 贡献归因去"证伪"的动机;注意它是"对不上"而非"造假"。

b004This gap poses an additional challenge for agentic model discovery (Huang et al., 2024; Li et al., 2024; Jiang et al., 2025; Lu et al., 2026). Systems such as CellScientist (Li et al., 2026), CellForge (Tang et al., 2025), and HarmonyCell (Huang et al., 2026) generate executable predictors, train them, observe validation feedback, and iteratively revise or select candidates. When feedback primarily rewards predictive performance, search can favor high-scoring predictors without testing whether they use the inputs specified by their designs. As generated models become more numerous and diverse, manual verification becomes increasingly difficult. Agentic discovery therefore needs scalable tests of the claims attached to its selected models, not only mechanisms for proposing better predictors.

这段在干什么

承接上文提出的「预测–声明缺口」,指出它在 *agentic model discovery* 场景下更难查,从而论证为什么需要可扩展的检验机制,而不只是更好的预测器。

需要解释的地方

  • agentic model discovery:让 AI 智能体自动生成、训练、迭代模型(领域常识)。
  • executable predictors:能实际跑起来做预测的模型代码。
  • validation feedback:训练中用验证集给的回馈信号。
  • predictive performance:预测准不准。
  • 本段逻辑:反馈只看准不准,就会奖励「高分但没用对输入」的模型。

值得留意

  • 作者的落点不是「怎么造更好的模型」,而是怎么审计已选模型的声明——这是全文主旨。
  • 模型变多、变杂 → 人工核查难,这句是「需要自动化检验」的理由。

b005We introduce CellAudit, a framework that links registered input-use claims to evidence from source code and fitted models. It asks three non-equivalent questions: can the input enter the cited computation, do fitted predictions depend on it, and does that dependence improve prediction of the observed response? These correspond to source consumption, fitted dependence, and target-relevant predictive contribution. A cited pathway can be mathematically inactive; an executable pathway may have little effect after fitting; and input dependence need not provide predictive benefit. To distinguish these cases, CellAudit combines source inspection with matched input replacement (Fisher et al., 2019; Chamma et al., 2023) on frozen checkpoints, testing both whole-model dependence and target-relevant contribution. During development, the same diagnostics can guide revision; final claims are evaluated only after the selected source, checkpoint, and audit rules are frozen.

这段提出 CellAudit 框架:把"注册的输入用途声明"和源代码、已拟合模型里的证据连起来,作为对上一段"需要可扩展检验"的回应。

关键概念:它问三个不等价的问题——输入能否进入所指计算(源代码层面)、拟合后的预测是否依赖它(依赖层面)、这种依赖是否改善对观测响应的预测(目标贡献层面)。领域常识:三者可以脱节,比如某通路数学上不激活、可执行但拟合后影响很小、依赖却无预测收益。

值得留意:方法是在冻结的 checkpoint 上做 matched input replacement;诊断在开发期可指导修改,但最终声明只在源码、checkpoint、审计规则都冻结后才评估。

b006CellAudit reveals these distinctions in agent-generated cellular-response predictors. On a paired morphology–transcriptomics perturbation task built from BBBC047 (Bray et al., 2017; Haghighi et al., 2022), an agent-selected model reaches a mean held-out Global PCC of 0.3153, compared with 0.3142 for a control-profile-only predictor, while remaining exactly invariant to compound replacement. Source inspection identifies a singleton key–value attention pathway whose query cannot transmit compound information. The model remains invariant after refitting with physically disjoint control wells for inputs and target references. A broader audit separates fitted dependence from evidence of predictive contribution. In a stratified sample of 48 generated candidates across the two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 have registered target-loss scaffold intervals that lie entirely above zero on both folds.

这段在干什么:用具体例子展示 CellAudit 如何区分「模型拟合出的依赖」与「真正有预测贡献的证据」——一个反例加一次更广的审计。

关键概念:

  • Global PCC:全局皮尔逊相关系数,衡量预测与真实值的整体相关(领域常识)。
  • compound replacement:替换化合物输入,看预测是否变化,用来检验模型是否真用了化合物信息。
  • scaffold intervals / 靶点损失区间:这里指 bootstrap/重采样得到的置信区间。

值得留意:关键区分是「换了输入预测就变」≠「该输入有真实预测贡献」。48 个候选里 47 个对化合物替换敏感,但只有 20 个的目标损失区间在两折上都整体大于零——即多数只是拟合了关联,而非稳健贡献。

b007These diagnostics also distinguish what a revision needs to address. Making a perturbation pathway executable does not by itself ensure a target-relevant contribution. On BBBC047, falsification-guided revisions learn a residual over a control-profile-only predictor, recovering positive compound contributions together with predictive gains. In matched sci-Plex searches (Srivatsan et al., 2020), audit-enriched feedback produces higher mean held-out predictive performance and larger mean compound and dose contributions than score-only feedback. However, paired confidence intervals for these differences across five trajectories include zero, so the evidence for improved discovery remains suggestive rather than conclusive.

讲解

1. 这段在干什么

承接上段的“诊断结果”,说明诊断也能指出修订该改哪儿:能跑通≠有贡献,并对比两种反馈策略的效果。

2. 需要解释的地方

  • falsification-guided / audit-enriched feedback:用“证伪”“审计”信号来指导修订,而非只看分数(score-only)。
  • residual over a control-profile-only predictor:在只含对照的预测器上加一个残差修正项。
  • paired confidence interval 含零:差异的置信区间跨过 0,统计学上不能判定显著。

3. 值得留意

作者自己留了退路:BBBC047 上看似有增益,但 sci-Plex 上跨五条轨迹的配对置信区间含零,结论只是“suggestive rather than conclusive”——这是领域常识层面的谨慎,不是论文额外下的定论。

b008Finally, we distinguish predictive generalization from claim generalization. The latter asks whether a registered input-use claim remains supported when a fixed model design is refit on an independently acquired cohort. When LINCS-selected designs are fixed and refit on LKCP (Keenan et al., 2018; Subramanian et al., 2017; Weisbart et al., 2024), predictive gains and dose contribution persist, whereas compound-identity contribution changes substantially. A separate CRISPRa evaluation applies the same auditing principle to unseen genetic combinations (Norman et al., 2019). Together, these studies trace input-use claims from source code through fitted dependence to predictive contribution, use the resulting evidence to guide revision, and re-evaluate those claims on new data.

讲解

1. 这段在干什么

作为引言的收尾,把讨论从「预测泛化」推进到更难的「论断泛化」:换一套独立数据重训后,原来的输入使用论断还站得住吗?并总结全文的研究路线。

2. 需要解释的地方

  • 预测泛化 vs. 论断泛化(领域常识式区分):前者问模型在新数据上准不准;后者问「模型用了某类输入」这个结论本身还成不成立——设计固定、换队列重训后再看。
  • LINCS / LKCP / CRISPRa(均为领域常用数据集/扰动平台名):文中只说在 LKCP 上重训、CRISPRa 用于未见过的基因组合,未展开细节。
  • 贡献:指某类输入对预测结果的实际影响大小。

3. 值得留意

关键不对称——预测增益和剂量贡献能保持,但化合物身份贡献变化很大。这说明「有用」和「靠什么有用」是两回事,正是审计要抓的点。

2 Related work

b010Cellular perturbation models use latent transfer, compositional representations, neural transport, genetic graphs, and foundation models to predict chemical or genetic responses (Lotfollahi et al., 2019; Lotfollahi et al., 2023; Bunne et al., 2023; Roohani et al., 2024; Cui et al., 2024). Benchmarks standardize evaluation across splits and baselines (Wu et al., 2025; Wenteler et al., 2025), while recent studies examine linear references, shared response variation, calibrated metrics, and distribution shifts (Ahlmann-Eltze et al., 2025; Viñas Torné et al., 2026; Miller et al., 2025; Mao et al., 2026). These works ask how well perturbation responses can be predicted and evaluated; CellAudit asks whether registered input-use claims are supported by a model’s source code and fitted behavior.

讲解

1. 这段在干什么

交代相关工作的版图,最后一句划清界限:别人问「预测得准不准」,CellAudit 问「模型声称用了的输入,代码和拟合行为真的支持吗」。

2. 需要解释的地方

  • perturbation(扰动):对细胞施加化学或基因干预,看它怎么反应(领域常识)。
  • Benchmarks:统一划分数据、统一基线的评测标准。
  • registered input-use claims:模型预先声明的「我用了哪些输入」。

3. 值得留意

前半句全在罗列方法和基准,读起来像文献堆砌,但它是为最后一句的反差做铺垫——前人的关注点是「预测性能」,本文的关注点是「输入使用声明是否属实」,这个对照才是这段的真正落点。另外,代码与拟合行为是两个独立证据源,缺一不可。

b011Scientific agents generate and revise executable hypotheses, programs, and models (Boiko et al., 2023; Romera-Paredes et al., 2024; Jiang et al., 2025). Permutation reliance, conditional replacement, and behavioral testing probe fitted-model behavior under targeted input interventions (Fisher et al., 2019; Chamma et al., 2023; Ribeiro et al., 2020), while ConceptSMILE studies explanation reliability and POPPER automates hypothesis falsification (Mollapour et al., 2026; Huang et al., 2025). Unlike reliance tests that begin from prediction behavior, CellAudit starts with a registered input-use claim and its cited computation, then follows it through source consumption, fitted dependence, and target-relevant predictive contribution. This evidence can also guide model revision. Appendix K provides additional context.

这段在干什么:摆在相关工作里做定位对比——先列出科学智能体和依赖测试两条线,再一句"Unlike…"标出 CellAudit 的差异:从注册的输入使用声明出发,而非从预测行为出发。

需要解释的地方:

  • Permutation reliance / conditional replacement:领域常识,指打乱或替换某输入、看模型输出变多少,来判断模型是否真用了它。
  • fitted dependence:模型拟合完成后表现出的依赖关系。
  • target-relevant predictive contribution:该输入对目标预测实际的贡献。

值得留意:CellAudit 走的是"声明→引用代码→源码消费→拟合依赖→预测贡献"这条链,且证据还能反过来指导模型修改;末句把细节推给附录 K。

3 Problem setup: predictive performance and registered input-use claims

b013Each example contains biological context cc, perturbation representation pp, optional attributes aa such as dose, and response y=[y(1),…,y(M)]y=[y^{(1)},\ldots,y^{(M)}]. Source ss and fitted parameters θ\theta define fs​(c,p,a,θ)f_{s}(c,p,a;\theta); outputs may include imaging, molecular, or other readouts. Before search, the task registers input coordinates and admissible interventions. Each audit records whether an input-use claim comes from a candidate or the task contract and, after selection, fixes its cited computation and test mapping. A prediction–claim gap occurs when the cited source contradicts a registered claim or fitted predictions are invariant to input replacement. Predictive performance affects the importance of the case, not the definition; target-relevant contribution is separate.

这段在干什么

定义后续审计的形式化框架:每个样本的输入/输出是什么、模型由谁定义、claim 从哪来、什么算“预测—claim 差距”。

需要解释的地方

  • cc/p/a/y:生物背景、扰动、剂量等属性、响应向量;ss 与拟合参数 θθ 共同定义模型 fsf_s。
  • input-use claim:模型“用了哪些输入”的主张,可能来自候选模型,也可能来自任务约定(task contract)。
  • prediction–claim gap:引用的来源与已登记的 claim 矛盾,或替换输入后预测不变。

值得留意

作者划了两条线:性能只影响案例权重,不影响 gap 的定义;“目标相关贡献”是另一回事,别混为一谈。

b014Data processing, budgets, evaluation, and selection rules are fixed before search. Folds 1–2 fit candidates; Fold 3 provides feedback and selects checkpoints. Selected checkpoints are evaluated on Folds 4 and 5, except prediction-score discovery also selects a source across policies on Fold 4 and refits it for Fold 5. Norman instead selects sources before Fold 4 and refits the same sources on Fold 5 without reselection. Appendices B and N.1 specify study-specific refitting and reuse rules.

这段交代评估协议:数据处理、预算、评估与选择规则在搜索前就固定,属于防数据泄漏和防调参作弊的常规设计(领域常识)。

  • Folds 1–2:拟合候选;Fold 3:反馈并挑 checkpoint。
  • Fold 4/5:选中后评估。区别在:prediction-score 那条线在 Fold 4 上跨策略选源、再在 Fold 5 重拟合;Norman 则在 Fold 4 前选源、Fold 5 直接重拟合、不再选。
  • 细节见附录 B 与 N.1。

留意:这段是"规则先定死",与上段"性能不影响定义"呼应——性能只动案例重要性,不动定义。

4.1 Agentic model search with fixed data and evaluation

b017The existing CellScientist policy (Li et al., 2026) starts from a common predictor h0h_{0}. It receives the current model, development diagnostics, execution status, and compact history, and proposes a hypothesis and executable revision within the registered permissions. Shared preflight and training code provide up to three repairs per failed proposal; persistent failures consume a slot. Each executable candidate trains once, and a fixed Fold-3 rule selects the endpoint. Records preserve parent links, emitted hypotheses and declarations, source, diagnostics, repairs, checkpoint, and costs. The language model proposes revisions; input-use tests are deterministic.

讲解

1. 这段在干什么

描述 CellScientist 这个 agent 搜索策略的具体运行机制:给定固定数据和评估,它如何从初始模型出发提出并执行修改。

2. 需要解释的地方

  • predictor h₀:起始模型/预测器,搜索的出发点。
  • preflight:领域常识,指运行前的检查代码。
  • permissions:agent 被允许做的修改范围(原文只说"注册的权限",具体内容这段没说)。
  • repairs:提案失败后的修复机会,最多三次;连续失败会占掉一个"槽位"。
  • Fold-3 rule:用第 3 折数据选定最终模型。
  • input-use tests are deterministic:检验"用了哪些输入"的测试是确定性的,不靠语言模型。

3. 值得留意

分工很关键——语言模型只负责"提修改",确定性测试负责"验输入使用"。这一句点出了全文核心:把"发现、证伪、修正"的验证环节与生成环节分开,避免全靠 LLM 判断。

4.2 Checking registered input-use claims in source code

b019Given registered inputs and cited source locations, CellAudit checks the computation attached to each input-use statement. Its abstract-syntax-tree checker returns implemented, implementation-contradicted, or unresolved. Coverage is the fraction of checked candidates whose cited locations are all resolved; unresolved code remains eligible for behavioral testing. With one key–value pair, attention has constant normalized weight and cannot transmit compound information through its query (Vaswani et al., 2017). This localizes a recognized implementation defect. Whole-model replacement then establishes whether fitted predictions depend on the input, including through other routes; target-loss contrasts test whether that dependence helps prediction.

这段在干什么

讲 CellAudit 怎么先静态检查源码里的输入使用声明,再对没查清的做行为测试。

需要解释的地方

  • AST 检查器:领域常识,指直接分析代码语法结构,而非运行程序。
  • 三态结果:implemented(确实实现了)、implementation-contradicted(实现与声明矛盾)、unresolved(查不清)。
  • coverage:这里指被检查候选中,引用位置全部解析成功的比例;没解析的代码还能继续做行为测试。

值得留意

  • 静态查不清 ≠ 有问题,unresolved 只是转交给行为测试。
  • attention 那个例子是"已知缺陷"的定位,属方法演示,不是说模型本身出错。
  • 最后两句是整体替换和 target-loss 对比,用来验证预测是否真依赖该输入。

4.3 Testing fitted dependence and predictive contribution

b021With checkpoint and targets fixed, CellAudit replaces one registered input xqx_{q} while retaining the others. Prediction distance measures target-blind dependence; zero identifies invariance on the tested replacements. To test predictive benefit, it computes the decrease in a higher-is-better score mm and the increase in loss:

这段在交代评测协议:固定其余输入,只替换一个已注册输入,看预测怎么变。是方法过渡到结果的一步。

解释

  • target-blind dependence:不看标签,仅看换输入后预测是否改变。“distance=0”即没变,说明模型对该替换不变。
  • higher-is-better score $m$:领域常识,指越大越好的评分指标。

留意

  • 两种测法方向相反:评分看下降,损失看上升,都是“变差”。
  • 原文未给出具体分数、损失名或结果数值。

b022Positive effects indicate benefit from the observed input relative to its registered replacements. Replacement distributions and their observed-support coverage are specified by task. We distinguish single-coordinate contrasts Δ​mq\Delta m_{q} from factorial effects EqE_{q}, which average over the other input’s two states in the 2×22\times 2 compound–context audit. Loss effects EqℒE^{\mathcal{L}}_{q} use the same averaging. The transfer analysis separates compound and dose allocations (Lundberg and Lee, 2017). Appendix L gives the full contrasts.

这段在干什么:承接上文「计算分数下降与损失上升」,定义怎么比对:把观测输入和它的「注册替代品」比,正效应即观测输入更有益。

需要解释的地方:

  • 单坐标对比 Δm_q:只动一个输入、看分数变化;
  • 因子效应 E_q:2×2(化合物×情境)审计里,把另一个输入的两个状态平均掉——领域常识:这是析因设计里求「主效应」的常规做法;
  • 损失效应 E_q^ℒ:同样做平均,只是换成损失指标。

值得留意:两种对比(单坐标 vs 因子平均)不是一回事,别混用;具体替换分布和覆盖范围「由任务指定」,本文没给,细节都在附录 L。

b023Maps are repeated measurements within a checkpoint, averaged before across-model inference. Continuous contributions have paired training-seed or trajectory intervals; fixed-model source-group resampling tests variation across biological groups. Known-truth checks distinguish fixed-dataset and population uncertainty (Appendix F.1). BBBC, LINCS, and LKCP separately report whether the effect quantile interval lies above, below, or across a reference threshold. A positive mean contribution can occur in any category. Appendix B defines the threshold, quantiles, and original category labels; Norman uses a separate loss-based criterion (Appendix D).

好,我们看这段。它紧接上一段“分离化合物与剂量分配”,开始讲怎么做统计检验。

1. 这段在干什么:交代本节的检验方法——先平均重复测量,再做跨模型推断,并逐项说明各数据集的报告方式。

2. 需要解释的地方:

  • 重复测量先平均:同一 checkpoint 内多次测量的值先取均值,再拿去跨模型比较,避免重复数据被当独立样本(这是统计常识)。
  • 区间:连续贡献看配对训练种子或轨迹的区间;固定模型则对生物组做重采样,看组间变异。
  • BBBC/LINCS/LKCP:分别报告效应分位区间是高于、低于还是跨越参考阈值。

3. 值得留意:作者特意强调“正的平均贡献可以出现在任何一类”——别把正负号直接当成“高于阈值”。阈值、分位数、类别标签的定义在附录 B,Norman 用的是另一套基于损失的判据(附录 D),这段只点名字、没展开。

4.4 Using falsification evidence to guide the next search

b025The audit distinguishes two reasons to revise a model. A contradicted source path motivates an executable route; weak target contribution motivates learning an increment over context. Path-constrained discovery addresses the first by compiling structured design cards into models with an explicit compound route. BBBC falsification-guided discovery addresses the second through residual modeling and input-use feedback:

讲解

1. 这段在干什么

承上启下:把"修订模型"分成两类原因,并预告两种对应方法。

2. 需要解释的地方

  • *contradicted source path*:源路径被证伪,即数据来源的那条逻辑被推翻。
  • *weak target contribution*:某个输入对预测目标贡献太弱。
  • *executable route*:可执行的(代码)路径;*compound route*:复合路径。
  • *residual modeling*:残差建模,即先有基础预测,再学"补差"。

3. 值得留意

两种"修订理由"与两种方法是一一对应的(first/second),别读成两件独立的事。

b026where g⁡(c)g(c) is fitted on Folds 1–2 and fixed. The compiler enforces centered perturbation/dose branches, context gates, and bias-free readouts, giving zero residual at their joint reference input. Training and Fold-3 feedback use prediction error and replacement loss changes. BBBC047 selects by predictive score among qualifying candidates; BBBC036 prioritizes compound loss gain among predictively noninferior candidates. Each trajectory evaluates the anchor plus nine proposals; the harness selects scale α\alpha and checkpoint on Fold 3 and freezes that model for Folds 4 and 5. These formulations test the joint design changes; Appendix B specifies their rules.

这段在干什么

交代 BBBC 两条线(047/036)的具体选择规则与训练/冻结流程,说明假证据如何落到下一轮搜索。

需要解释的地方

  • g(c):在 Fold 1–2 上拟合后固定,作为后续残差建模的基准,不再更新。
  • compiler 约束(领域常识式解释):居中扰动/剂量分支、上下文门、无偏读出,使参照输入处残差为零。
  • qualifying / predictively noninferior:先过资格线,再在合格者中比分数或损失增益。
  • α 与 checkpoint:在 Fold 3 上选,随后冻结用于 Fold 4–5。

值得留意

两条线排序准则不同:047 看预测分,036 在"预测不劣"前提下看复合损失增益——作者未细说阈值,规则在 Appendix B。

b027The sci-Plex comparison instead shares a context-plus-dose anchor g⁡(c,a)g(c,a) and permissions for direct predictors or anchor residuals. Score feedback supplies metrics and learning curves; Audit adds grouped errors, source checks, and development-set compound/dose replacements. After each fit, CellAudit computes these diagnostics and inserts them into the agent’s next prompt, linking input-path findings and measured contributions to model revision. Loss, optimizer, training limits, and PCC-first selection stay fixed. Paired outcomes evaluate the complete audit-enriched feedback package; selected sources and checkpoints are frozen before both held-out evaluations.

讲解

这段在干什么:说明 sci-Plex 对比实验里用什么反馈驱动“下一轮模型修正”——反馈分 Score 和 Audit 两层,并交代评估怎么保持公平。

需要解释的地方:

  • context-plus-dose anchor:以“上下文+剂量”作为基准锚点(领域常识:anchor 指参照基线)。
  • anchor residuals:相对锚点的残差,即模型可预测“偏离基线多少”。
  • audit-enriched feedback package:把审计诊断结果打包塞进下一次 prompt。

值得留意:Loss、优化器、训练上限、PCC-first 选择全部固定,来源与 checkpoint 在两次留出评估前冻结——作者刻意只改反馈,排除其他变量干扰。

5.1 Development on linked multimodal perturbation tasks

b030Development uses linked tasks from one paired cellular-profile release: BBBC036, the known-bioactive subset, and its parent BBBC047 (Haghighi et al., 2022; Bray et al., 2017; Subramanian et al., 2017). Each predicts joint morphology–transcription responses from a same-plate control profile, compound representation, and dose. Response-blind Bemis–Murcko scaffold folds (Bemis and Murcko, 1996) hold out compound families. Global PCC on concatenated responses selects models; MSE and modality-specific metrics diagnose gains. Appendix A gives provenance, transformations, fold inventories, and a complementary plate-family split.

1. 这段在干什么

交代实验的数据与评测设计:用同一批配对细胞数据里的两个任务,规定输入、划分方式、选模指标,为下一段结果做铺垫。

2. 需要解释的地方

  • BBBC036 / BBBC047:同一细胞成像数据集(Broad Bioimage Benchmark Collection)的一对,前者是已确认有生物活性的子集,后者是它的"母集"。
  • 配对/联合响应:同一化合物既要看形态、又要看转录,两者一起预测。
  • Bemis–Murcko 骨架划分:领域常识,按分子骨架把化合物家族整体留出,而不是随机拆散——防止"见过近亲"导致虚高。
  • PCC / MSE:皮尔逊相关(选模主指标)、均方误差(诊断)。

3. 值得留意

任务标注为"response-blind"——划分时不看响应值;控制来自同板;细节都推给附录 A。

b031Prediction-score discovery compares CellScientist, AIDE, CellForge, and HarmonyCell policies (Li et al., 2026; Jiang et al., 2025; Tang et al., 2025; Huang et al., 2026) under a matched initial model, executor, feedback, and ten-candidate budget, with ten trajectories per task and five paired refits per selected model. Both follow-up formulations retain the folds and budget and carry all ten selected designs through Folds 4 and 5; failures, costs, and repeated designs are retained (Appendix C).

讲解

1. 这段在干什么

说明"预测分数发现"实验的对照设置:四个方法在相同初始模型、执行器、反馈和十候选预算下比较,并交代重复次数与后续两种变体的延续方式。

2. 需要解释的地方

  • 配对 refit(paired refit):同一选中模型用同一批数据折叠再拟合一次,便于成对比较差异。
  • Budgets/folds:预算指候选数上限;folds 是数据划分,这里沿用同一套划分,后续变体跑到 Fold 4、5。
  • Failures、costs、repeated designs 被保留:指不丢弃失败或重复结果,全部计入(细节在附录 C)。

3. 值得留意

"匹配"(matched)的是初始模型、执行器、反馈和预算——公平比较的前提;但十轨迹、五配对 refit 属于领域常见设定,具体理由这段未提。

b032For BBBC047, each plate’s control wells are split into equal-size banks A/B. With the target reference fixed, shared and disjoint arms take the input profile from the same or the other bank, so their contrast isolates physical control-well overlap. Both orientations refit the score-selected source, a control-only source, and one fixed guided design under five seeds; all 60 checkpoints are fixed before evaluation on Folds 4 and 5 (Appendix A.1).

这段在干什么:交代 BBBC047 的对照设计——用同/异控制孔库(shared/disjoint arms)来隔离控制孔重叠带来的混杂,并说明评估前的固定流程。

需要解释的地方:

  • bank A/B:每块板的控制孔均分成两份,做对照用。
  • shared/disjoint arms:两组输入分别来自同一库或另一库,二者对比就只差"物理控制孔是否重叠"这一件事——这是领域常识里的"消融/隔离变量"思路。
  • refit / checkpoints:重新训练模型,保存权重快照;这些快照在评估前就固定好,避免边看结果边调。

值得留意:作者特意强调"所有 60 个 checkpoint 在评估前就冻结",这是在防"事后调整"的嫌疑。

b033We predict 2,000-gene well-level responses in sci-Plex3 (Srivatsan et al., 2020), holding out compound–cell-line combinations, with whole treatment wells in one partition, while both marginals occur in training; gene selection and scaling use training wells only. Both conditions use the same DeepSeek serving model, five paired trajectory seeds, five new candidate slots per trajectory, and a shared anchor. All ten selected checkpoints are evaluated on two held-out combination partitions of this source. Compound replacements match dose and experimental context; dose replacements retain compound identity (Appendix N.1).

这段在干什么

交代所选 checkpoint 的评测数据集与评测设置:在 sci-Plex3 的留出组合划分上做评测。

需要解释的地方

  • held-out 组合:训练里各单因素都见过,但“化合物×细胞系”这个搭配没见过(领域常识:留出组合测试)。基因选择和缩放只用训练孔,避免泄漏。
  • marginals:指两个单因素各自的边缘分布,训练里都有。
  • anchor:共享锚点,原文未展开其含义。

值得留意

所有 10 个 checkpoint 在评测前已固定;化合物替换与剂量替换的对照条件不同。

5.2 Testing input-use claims on new cohorts and perturbation types

b035The fixed falsification-guided protocol first runs on LINCS–Pilot1 (Keenan et al., 2018; Subramanian et al., 2017). Its ten selected designs and sources are then frozen and refit on independently acquired LKCP Batch 2 (Weisbart et al., 2024), with response- and fold-blind interface adaptation fixed before training. Compound permutations retain dose; dose permutations stay within compound (Appendix E). The Norman CRISPRa task predicts 5,045-gene responses for unseen double-perturbation combinations (Norman et al., 2019). Its registered input bundles multi-hot pair identity with the matching mean single-perturbation anchor. GEARS supplies a task-native reference (Roohani et al., 2024).

讲解

1. 这段在干什么

交代实验怎么落地:先在 LINCS–Pilot1 上跑固定协议、冻结设计,再原样搬到 LKCP Batch 2 上重训验证;同时说明两个扰动任务(Norman、GEARS)的设置和参照。

2. 需要解释的地方

  • falsification-guided protocol:上一段说的那套"先假定输入被用到,再设计能证伪的实验"的流程。
  • frozen and refit:设计选好就锁死,换到新数据集只重新训练,不改设计——防止事后调参。
  • response- and fold-blind:接口适配对响应值和数据划分都不知情(领域常识:避免信息泄漏)。
  • multi-hot pair identity:用 0/1 向量表示一对扰动里各是谁。

3. 值得留意

  • "Compound permutations retain dose; dose permutations stay within compound"——置换不是随便打乱,两类置换各守住一个不变量(Appendix E/N.1 有细节)。
  • GEARS 只是"task-native reference",即对照组,不是说它用同一套协议。

6.1 CellAudit detects compound invariance in a high-scoring predictor

b038CellAudit’s automated audit identifies the BBBC047 source selected on Fold 4 as compound-invariant. Its source-selection score is 0.30350.3035; mean Global PCC reaches 0.31530.3153 on held-out replication Fold 5. Compound shuffling changes PCC by 0.00000.0000 in every refit, whereas control replacement lowers PCC by 0.19320.1932 on Fold 4 and 0.20040.2004 on Fold 5 (Table 1). Control-only reaches 0.31420.3142 on Fold 5; Full−-control-only is +0.0011+0.0011 (95% CI [−0.0005,+0.0027][-0.0005,+0.0027]). Prediction-space distance is exactly zero under compound replacement and large under control replacement (Appendix I).

这段在干什么:用一组具体数字,给出 CellAudit 对高分预测器的一条审计结论——该模型是「化合物不变的」。

需要解释的地方:

  • *compound-invariant*:把化合物标签打乱,预测结果完全不变,说明模型根本没用到化合物信息。
  • *PCC*:预测与真值的相关性,越高越好。
  • *shuffling vs. replacement*:打乱化合物 vs. 换成对照,前者无影响、后者大幅掉分。

值得留意:打乱化合物 PCC 变化恰为 0、预测空间距离恰为 0,等于模型完全没依赖化合物;真正起作用的是对照。

b039The source checker identifies a concrete computation to revise: the cited compound-conditioned cross-attention has one key and one value, making its normalized weight constant. The behavioral test establishes complete-model invariance; the source result localizes an inactive cited pathway. A compound-only positive control produces a shuffle effect of 0.0111±0.00100.0111\pm 0.0010, with all five refits above the reference threshold.

这段在干什么:承上给出定位结果——源检查指出该被引通路存在一处具体可修正的计算,行为测试则确认整个模型对该化合物替换完全不变。

需要解释的地方:cross-attention 的 key/value(领域常识:注意力里用来算权重的两路向量);只有一个 key 和一个 value,归一化权重就成了常数——即这一路对化合物条件实际不生效;阳性对照是验证实验本身有效的参照;shuffle effect 指打乱后造成的效应量。

值得留意:源检查与行为测试结论一致,一个定位、一个确证;0.0111 的效应量连同五次重拟合全部过阈值,说明阳性对照确实生效,反衬出化合物条件那一路是真的没起作用。

b040A source-stratified sample of 48 real candidates separates these outcomes further (Figure 2a,b). An implemented multi-tower source is exactly invariant; 47 candidates show fitted dependence on both folds, but only 20 have positive target-loss scaffold intervals on both. Source consumption and fitted dependence thus leave predictive contribution unresolved for many candidates (Appendix M.3). CellScientist supplies a competitive discovery setting: mean Best@10 and frontier AUC lead both linked tasks, with intervals versus HarmonyCell crossing zero (Table 19).

这段在干什么:用一批真实候选样本,进一步验证“源码里用了”和“拟合上依赖”都不能直接等同于“真有预测贡献”。

需要解释的地方:

  • source-stratified sample:按源码结构分层抽的48个真实候选,用来把结论分得更细。
  • exactly invariant / fitted dependence:前者指模型结构上完全不依赖;后者指训练后数值上表现出依赖。
  • positive target-loss scaffold intervals on both:两折都显示对目标损失有正向贡献,才算真“有用”。

值得留意:48个里只有20个两折都正向——说明“被用上”“拟合依赖”与“实际贡献”之间存在大缺口。

b041Control: control-only predictor. Reference: effect quantile interval above/below/across the reference threshold (Appendix B); distinct from a positive mean contribution. aOne-key attention cannot transmit compound identity. bCompiler-enforced compound path. cExplicit compound-residual path.

这段在干什么

这是表格的脚注说明:定义「Control」「Reference」两个对照基线,并解释表内 a/b/c 三个角标各自的含义。

需要解释的地方

  • control-only predictor:只用对照条件训练出的预测器,作为比较基准。
  • effect quantile interval:效应量的分位区间;用来判断某效应是高于、低于还是跨越参考阈值。
  • a/b/c 角标:分别标注三种情形——单键注意力无法传递化合物身份、编译器强制走化合物路径、显式走化合物-残差路径。

值得留意

「distinct from a positive mean contribution」是在提醒:落在区间外 ≠ 有正向平均贡献,两者别混为一谈。这段不是正文论述,读表时容易直接跳过。

6.2 The mismatch persists with disjoint control references

b043We test whether the mismatch persists when control inputs and CP target references use disjoint physical wells. Newly reconstructed shared and disjoint arms have identical targets within each bank orientation (Figure 2c,d). In the disjoint arm, the score-selected source reaches Global PCC 0.30850.3085 and 0.31930.3193, while control-only reaches 0.30820.3082 and 0.31850.3185 on Folds 4 and 5. The score-selected source remains exactly compound-invariant in every refit and evaluation. All six shared–disjoint PCC comparisons have five-seed intervals crossing zero. The detected mismatch therefore persists without physical control-well reuse at this reference-construction layer.

这段在干什么:它检验前面的错配是不是由「对照组与目标组共用同一批物理孔」造成的——换成互不重叠的孔后,错配依然存在。

需要解释的地方:

  • disjoint physical wells:对照组输入和目标参考取自完全不同的孔(不共用),排除"共用孔导致假错配"这一替代解释。
  • Global PCC:全局皮尔逊相关系数,领域常识,衡量预测与真实值的线性吻合度,越高越好。
  • compound-invariant:对化合物身份不敏感,即换化合物预测不变。

值得留意:两组数值极接近(0.3085 vs 0.3082、0.3193 vs 0.3185),且「五种子区间均跨零」说明差异不显著——作者据此强调错配不是孔复用造成的假象。

b044The fixed guided design retains both predictive and compound-use increments under the same separation (Table 6). Relative to disjoint control-only, its PCC gain is +0.0028+0.0028 (95%95\% paired-seed CI [+0.0017,+0.0038][+0.0017,+0.0038]) on Fold 4 and +0.0025+0.0025 ([+0.0012,+0.0037][+0.0012,+0.0037]) on Fold 5. Compound replacement lowers PCC by 0.00430.0043 and 0.00330.0033 and increases target loss by 1.5106×10−41.5106\times 10^{-4} and 1.1939×10−41.1939\times 10^{-4}; each mean effect has a positive five-seed interval. Effects are slightly smaller than in the shared arm. Same-plate context and the original scaffold folds are retained (Appendix A.1).

这段在干什么:验证固定引导设计在分离对照下依然保持预测增益与复合使用增益,回应上段"错配持续存在"的结论。

需要解释的地方:

  • PCC gain:预测相关系数的提升,正数代表预测更准。
  • paired-seed CI:同一随机种子配对比较后的置信区间,全为正说明增益稳定。
  • Compound replacement:把复合(化合物)信息替换后,PCC 下降、target loss 上升——反向证明复合信息有用。
  • shared arm / disjoint control:两种对照构造方式,这里用的是后者。

值得留意:所有区间都为正,效果虽比 shared 组小,但方向一致;同板上下文与原始 scaffold 折被保留,说明结论不是在无控制条件下取得的。

6.3 An explicit compound route and its fitted contribution

b046Path-constrained discovery addresses the source defect by compiling an explicit compound route: all 200 candidate evaluations pass source and interface checks without candidate-code repair. On BBBC047, mean factorial compound effects are 0.00830.0083 and 0.00740.0074, with positive mean loss effects on Folds 4 and 5. Relative to the reference threshold, Fold 4 has six below-threshold and four overlapping cases; Fold 5 has seven and three. Control-profile effects exceed the threshold throughout (Table 1). An executable route therefore admits compound dependence without ensuring a consistently above-reference contribution in each model.

这段在干什么

它检验「显式复合路径」这个修复方案:200 个候选全部通过源码与接口检查,说明源码缺陷被绕过了,但拟合出的复合效应在 Fold 4、5 上只是「部分」超过参考阈值。

需要解释的地方

  • factorial compound effects(因子复合效应):领域常识,指用因子设计估计「换用复合体系」这一因素带来的影响大小。
  • 参考阈值 / below-threshold:判定某模型是否算「真正依赖复合物」的对照线;低于它即不显著。
  • control-profile effects:对照组(非复合)层面的效应,用作对照。

值得留意

  • 「可执行 ≠ 有效应」:通过检查只保证能跑,不保证每个模型都稳定超过阈值——这正是最后一句的落点。
  • 效应数值很小(0.0083、0.0074),Fold 4 有 6/10 低于阈值,说明阳性并非一致。

6.4 Guided residual discovery combines predictive gains with target-relevant compound contribution

b048Falsification-guided discovery next asks the perturbation residual to improve prediction beyond a fixed control-profile-only model. On BBBC047, the full predictor exceeds that baseline by +0.0037+0.0037 on Fold 4 and +0.0028+0.0028 on Fold 5. Mean factorial compound effects are 0.00530.0053 and 0.00480.0048, with positive loss effects (Table 1); joint model–scaffold intervals for these increments are above zero. The explicit-route models have larger compound effects but lower predictive scores. Guided discovery combines higher scores with positive target-relevant effects, although none of its ten models exceeds the reference-relative criterion.

这段在干什么

论证「假说引导的残差发现」在 BBBC047 上既提升了预测分数(超过仅用对照的基线),又带来正向的目标相关化合物效应——但十模型都未过参考相对标准。

需要解释的地方

  • perturbation residual(扰动残差):领域常识,指模型在扰动条件下仍没解释掉的剩余部分,这里让残差模型去补足预测。
  • control-profile-only model:只用对照样本特征做的固定基线模型。
  • factorial compound effects:此处指化合物对损失的效应量(原文说"positive loss effects")。

值得留意

分数增益(+0.0037/+0.0028)和效应量(0.0053/0.0048)是不同量纲的两类指标,别当成一回事。另外"joint intervals above zero"只说明这两个增量不为零,并不等于超过参考标准——末句正是这个转折。

b049For the original guided checkpoints, observed-support replacements retain positive contributions on BBBC047 while holding context and dose fixed. They cover 3341/43803341/4380 and 2927/38782927/3878 rows on Folds 4 and 5: compound PCC drops are 0.005340.00534 and 0.004040.00404, and loss gains are 1.84×10−41.84\times 10^{-4} and 1.44×10−41.44\times 10^{-4}, all with positive source-scaffold intervals. The score-selected model remains exactly invariant on these same eligible rows (Table 57).

这段在干什么:承接上段,报告引导式检查点在观测支持替换下的稳健性——贡献为正、性能只小幅下降。

需要解释的地方:

  • 观测支持替换:把输入换成数据里真实出现过的取值,看结论是否还成立(领域常识式说法)。
  • PCC 降幅、loss gain:衡量替换后预测变差多少,数值小即影响小。
  • source-scaffold 区间为正:来源-骨架层面的置信区间不含零。
  • score-selected model 不变:这是对照组,说明另一种选择方式对这些行不敏感。

值得留意:合格行占比极小(3341/4380、2927/3878),贡献结论只覆盖少数行,作者没明说其代表性有限。

b050Fixed residual learners clarify the trade-off: Ridge has strong compound sensitivity but low joint-response scores, and on BBBC047 the residual MLP has larger compound effects but lower scores than the guided models, which lead both references in mean Global PCC at all four BBBC evaluations. On BBBC036 Fold 5, their gain over the context anchor is −0.0020-0.0020, with model–scaffold intervals for predictive and compound increments crossing zero (Table 16).

这段拿固定残差学习器做对照。作用:论证 guided 模型更好——它在四个 BBBC 评估上平均 Global PCC 都领先两个参照物。

关键概念:残差学习器=在已有预测基础上只学"补差"的模型(领域常识);Global PCC 是整体预测相关性指标;模型–scaffold 区间是评估波动范围,跨过零意味着效应不稳健。

留意:BBBC036 Fold 5 上它们只比上下文锚点提升 −0.0020,且两类增量区间都跨零——作者用这个"没提升"的例子反衬 guided 的增益。

b051Known-function experiments separate invariance, sensitivity without benefit, benefit, and harm across 6,400 independent simulated datasets. On the sensitive-without-benefit null, map-only intervals attain near-nominal coverage for the fixed-dataset conditional effect but severely undercover the population effect. Source-group intervals improve population coverage, with finite-group undercoverage remaining. Holding compound benefit fixed while changing reference-context variation can also change qualification. These checks support reporting continuous contribution, across-model uncertainty, and reference-relative status as distinct results (Appendix F.1).

讲解

这段在干什么:总结"已知函数"模拟实验的四种判别情形,并据此提出报告建议——贡献度、跨模型不确定性和参照相对状态应作为三种独立结果分别汇报。

需要解释的地方:

  • *map-only 区间*:只用最大后验点估计推的区间(领域常识)。
  • *nominal coverage*:名义覆盖率,即区间本该达到的覆盖概率。
  • *fixed-dataset 条件效应 vs. population 效应*:前者只看当前这份数据,后者要看整个总体——同一区间对前者够用、对后者严重不足。
  • *reference-context*:参照上下文,指比较基准的设定。

值得留意:benefit 固定不变时,仅改动参照上下文的变化就能改变"是否合格"的判定,说明结论对基准选择高度敏感;6,400 这个规模也提示结论来自大量独立模拟,而非单次实验。

6.5 Audit-enriched feedback yields higher mean held-out prediction

b053Do audit measurements help when returned to an agent? Figure 3a retains all five paired sci-Plex search frontiers. The context-plus-dose anchor already predicts much of the whole-expression profile, making incremental error reduction informative. Audit feedback yields higher mean PCC and lower mean MSE on both later partitions (Figure 3b): MSE is 9.58%9.58\% lower than Score on Fold 4 and 6.10%6.10\% lower on Fold 5. Both improve on the anchor.

1. 这段在干什么

论证把审计测量结果反馈给 agent 后,预测效果确实变好了——用 Figure 3 的 PCC、MSE 数字支撑这一点。

2. 需要解释的地方

  • PCC(皮尔逊相关系数):衡量预测曲线与真实曲线形状是否贴合,越高越好;MSE 是均方误差,越低越好(领域常识)。
  • anchor:即"context-plus-dose anchor",是作为基线/对照的预测。原文说它已能预测大部分整体表达谱,所以后续误差下降才有意义——"incremental error reduction informative"就是说在强基线上再降才有信息量。
  • Score:这里的对照方法(用分数反馈),审计反馈是与它比。
  • Fold 4 / Fold 5:数据的不同划分。

3. 值得留意

  • 只报了 MSE 的降幅(9.58%、6.10%),PCC 只说"higher",没给具体数。
  • 两个分区上审计反馈都优于 Score,且都优于 anchor——是双重比较。

b054The higher mean predictive performance of Audit is accompanied by higher mean compound and dose contributions (Figure 3c). Compound target-loss gain rises from 0.01860.0186 to 0.02200.0220 on Fold 4 and from 0.00640.0064 to 0.00790.0079 on Fold 5. Paired tt intervals for these Audit–Score contribution contrasts include zero at both folds (Table 68). Correct inputs help prediction in all five endpoints of both conditions. Audit–Score PCC is +0.0024+0.0024 (paired 95%95\% CI [−0.0013,+0.0060][-0.0013,+0.0060]) on Fold 4 and +0.0015+0.0015 ([−0.0005,+0.0034][-0.0005,+0.0034]) on Fold 5. Both means, and the favorable MSE and compound loss-gain directions, persist after removing any one trajectory pair (Figure 3d), whereas prediction-RMS differences can reverse after one omission. Percentile-bootstrap predictive intervals exclude zero even when resampling only trajectories, so uncertainty depends on interval construction at five pairs (Appendix N.2). Both conditions complete all 25 new candidates; Audit uses 38 versus 36 logical model calls and 592,347 versus at least 286,995 reported tokens (Table 65).

6.5 段讲解

这段在干什么:在上一段报告主预测误差(MSE)更低的基线上,这段补上贡献度与相关性证据,说明 Audit 的优势不只在单一指标,而是多角度一致。

需要解释的地方:

  • compound/dose contribution(gain):化合物与剂量对预测的贡献增益,越大说明模型越依赖正确的输入。领域常识。
  • Paired t 区间 / PCC:前者是配对均值差的置信区间,含 0 即差异不显著;PCC 是预测值与真实值的相关。
  • 去掉任一对轨迹:逐对剔除的稳健性检验。

值得留意:贡献增益方向一致,但 t 区间跨零、PCC 区间也跨零——作者没宣称显著。真正排除零的是百分位自助区间,而这是否成立取决于区间构造方式(仅 5 对)。

b055In the first registered Audit trajectory, CellAudit returns source checks, compound/dose replacement effects, grouped errors, and learning curves for the bilinear incumbent and a non-improving challenger. The agent returns to that incumbent and proposes simpler encoders with dropout, retaining dose-conditioned FiLM. Parameters decrease while selection PCC and compound loss gain increase (Figure 1D); search continues to a distinct final endpoint (Appendix O).

这段在干什么:描述第一次登记审计的实际过程——审计工具给出诊断,agent 据此回到原模型并改出更简单的编码器,得到更好的结果。

需要解释的地方:

  • incumbent / challenger:领域常识,指现任模型(当前最优)和挑战者(试图取代它的新模型)。
  • source checks、compound/dose replacement effects:审计工具返回的四类诊断信息,原文只说"返回了什么",未展开。
  • FiLM、dropout、PCC:都是模型结构或指标术语,领域常识;这里只需知道改后的模型参数更少、选择 PCC 和 compound loss 的表现变好。

值得留意:challenger 是"non-improving"——它没赢,agent 因此退回原模型继续改,而不是接受挑战者。这是"审计→修正"闭环的关键一步,但原文没点明。

6.6 Predictive gains and input-use support generalize differently under independent acquisition

b057To test claim generalization, the frozen LINCS-selected designs are refit on LKCP without source revision, model reselection, or audit-rule changes. They improve over the control-profile-only baseline by +0.0104+0.0104 (95% CI [+0.0103,+0.0106][+0.0103,+0.0106]) on Fold 4 and +0.0211+0.0211 (95% CI [+0.0205,+0.0217][+0.0205,+0.0217]) on Fold 5. Figure 4a,b separates compound identity from dose. On Folds 4 and 5, the dose-replacement Shapley allocation is +0.0486+0.0486 and +0.0467+0.0467 Global PCC on LINCS and +0.0237+0.0237 and +0.0341+0.0341 on LKCP, whereas the compound-identity allocation is +0.0175+0.0175 and +0.0126+0.0126 on LINCS but −0.000070-0.000070 and +0.0033+0.0033 on LKCP. All 50 trajectory-endpoint refits pass the dose-use criterion at every cohort–fold boundary, but none passes the compound-identity criterion at either LKCP boundary. LINCS shows why average effects and per-model decisions must remain separate: its compound-identity criterion is passed by 35/50 refits on Fold 4 but 10/50 on Fold 5 (mean lower effect quantile 0.01390.0139 to 0.00900.0090), although average compound contribution remains positive.

这段在干什么

测试前文冻结的设计在独立数据集 LKCP 上能否复现:预测增益能泛化,但“输入用途”是否成立要分开看。

需要解释的地方

  • refit:固定设计后重训模型,不改架构、不调规则。
  • Shapley allocation:领域常识,一种把总体贡献分摊给各输入因素的方法,这里用来分“剂量”与“化合物身份”。
  • Global PCC / 分位数:整体预测相关性指标;两者一个是均值层面、一个是逐模型层面。

值得留意

平均效应为正,不代表每个模型都通过判据:LINCS 上通过化合物身份判据的重训模型从 35/50 掉到 10/50,作者明确提醒二者要分开看。

6.7 The audit detects dependence on the perturbation-pair input bundle

b059On Norman, all five CellScientist refits pass the registered pair-input-bundle criterion on both held-out folds; shuffling the multi-hot pair identity together with its matching single-perturbation anchor reduces PCC by 0.55780.5578 and 0.62800.6280. Mean PCC exceeds h0h_{0} by +0.0060+0.0060 and +0.0045+0.0045 and the observed-single additive baseline by +0.0209+0.0209 and +0.0114+0.0114 (Table 26; Appendix D).

这段在给「pair bundle 审计」下结论:在 Norman 数据上,五个 CellScientist 重拟合都通过了成对输入捆绑的检验。

关键概念:multi-hot pair identity 指把「哪一对扰动」编码成向量;shuffle 是把它和对应的单扰动锚点一起打乱,看性能掉多少——掉得越多,说明模型越依赖这对输入。PCC 是预测与真实的相关系数。h₀ 是零假设基线,这里指随机打乱后的水平。

值得留意:打乱后 PCC 掉 0.56 和 0.63,是很大的跌幅,说明模型确实靠这对身份信息;而它相对加性基线只高 0.006 和 0.0045,增益其实很小。这两组数字一对比,才是这节真正的张力。

7 Discussion and conclusion

b061CellAudit shows that predictive success and claimed input use should be evaluated separately. Source consumption, fitted dependence, and target-relevant contribution can diverge: on BBBC047, a high-scoring source remains compound-invariant after disjoint control-reference refitting, while falsification-guided revisions recover positive compound contributions with predictive gains. The broader candidate audit shows this is not specific to singleton attention. In sci-Plex, audit-enriched feedback improves mean prediction and input contributions, but paired intervals over five trajectories span zero. Independent-acquisition refits further show that predictive generalization need not imply claim generalization.

讲解

这段在干什么:讨论段收尾——总结本文核心论点:预测好≠用了所声称的输入,并用多个数据集佐证这不是个例。

需要解释的地方:

  • source consumption / fitted dependence / target-relevant contribution:三个层次——模型读没读某输入、拟合中依赖它没有、它对预测目标是否真有贡献。领域常识:三者可脱节。
  • falsification-guided revisions:用证伪手段找出并修正错误的输入使用主张。
  • disjoint control-reference refitting:换一批独立的对照参考数据重训。

值得留意:作者主动承认 sci-Plex 的配对区间跨零(即提升不显著),这是诚实而非示弱;末句「预测泛化不蕴含主张泛化」是本段最重的结论。

b062These conclusions are conditional on the fitted model, evaluation setting, and registered replacement distribution. Positive contribution is distinct from passing the reference-relative criterion, and support for one fit should not be inherited by later refits. Observed-support replacements avoid unobserved combinations but do not establish biological exchangeability, mechanism, or causal effects. The sci-Plex study evaluates the complete audit-feedback package rather than individual diagnostics. CellAudit therefore adds a falsification layer to agentic model discovery, tracing input-use claims from cited computation to fitted behavior and predictive contribution.

1. 这段在干什么

收束全文:把前面结论限定条件摆明,定位 CellAudit 的贡献是给智能体模型发现加了一层"证伪"。

2. 需要解释的地方

  • conditional on:这些结论只在"该拟合模型、该评测设定、该注册替代分布"下成立,不能外推。
  • positive contribution ≠ 通过参考相对判据:领域常识里,单个输入的"预测贡献"和"是否通过某个相对标准"是两回事。
  • observed-support replacements:用已观测到的组合做替换,避开没观测过的组合。
  • sci-Plex:这是一个数据集/研究(原文只提名字,不展开)。

3. 值得留意

  • "support for one fit should not be inherited by later refits"——一次拟合成立,不自动传给后续重拟合,很容易读漏。
  • 作者自己划清界限:观测支撑的替换不等于生物可交换性、机制或因果。

Appendix A Data provenance, transformations, and task statistics

b065BBBC036 is the known-bioactive subset of its parent BBBC047 screen; the two tasks are therefore linked cohorts from one paired cellular-profile release rather than independent acquisitions. Their source collections in cpg0003-rosetta are CDRPBIO-BBBC036-Bray and CDRP-BBBC047-Bray, and the abbreviated task names retain the formal BBBC accession mapping (Ljosa et al., 2012). Inputs come from released replicate-level Cell Painting and L1000 profile tables (Weisbart et al., 2024; Haghighi et al., 2022). The morphology screen originates from the CDRP/BBBC047 Cell Painting resource (Bray et al., 2017), and the transcriptional profiles use the L1000 platform (Subramanian et al., 2017). Both assays apply matched compound–dose conditions on parallel plates and are joined at the condition level.

讲解

1. 这段在干什么

交代两个数据集 BBBC036 和 BBBC047 的来龙去脉,说明它们是同源配对,不是各自独立采集的。

2. 需要解释的地方

  • BBBC036 / BBBC047:两个公开筛选数据集,前者是后者中"已知有生物活性"的子集(领域常识:BBBC 是 Broad Bioimage Benchmark Collection 的编号)。
  • Cell Painting / L1000:前者是形态学成像筛选,后者是转录组表达谱平台;两者在同一批化合物–剂量下平行做。
  • condition level:按"同一化合物+同一剂量"这一条件把两份数据对齐拼接。

3. 值得留意

两个任务共用同一次细胞画像发布,属于"linking cohorts"。这一点很关键——后续审计输入使用声明时,"任务相关"可能来自同源而非独立,容易被忽略。

b066Table 2 reports both the released profile rows and the matched samples used for modeling. BBBC036 begins with 21,122 CP and 6,929 L1000 replicate profiles and retains 1,916 matched compound–dose conditions; BBBC047 begins with 153,386 CP and 68,120 L1000 profiles and retains 20,081 matched conditions. The matched subsets span 47 CP/22 L1000 treatment plates for BBBC036 and 273 CP/360 L1000 treatment plates for BBBC047. Before aggregation, 143 BBBC047 L1000 treatment profiles without a same-plate control are removed.

这段交代两个数据集从原始 profile 到建模样本的筛选过程。

关键概念:CP 和 L1000 是两种不同的检测平台,同一 compound–dose 条件在两平台上各测一遍,所以能按 condition 配对;treatment plate 是实验批次单位。

值得留意:CP 与 L1000 的样本量差距悬殊(如 BBBC047 15 万 vs 6.8 万),最终保留的 matched conditions 远少于任一平台原始数;末尾一句说明 L1000 里缺同板对照的 profile 会被剔除,但只提了 BBBC047。

b067For each condition, cc is the featurewise median of same-plate control-only CP profiles. Compound identity pp is represented by a 2,048-bit radius-2 Morgan fingerprint (Rogers and Hahn, 2010), while the separate attribute aa is the log-transformed observed dose. Before treatment filtering and condition aggregation, numeric profile columns are converted to float32 and any column containing a non-finite value in the released profile table is removed. CP treatment responses undergo the empirical same-plate-control mid-CDF transform, are centered by subtracting 0.50.5, and are aggregated by condition median; L1000 responses are centered by the same-plate control median and likewise aggregated by condition median. The resulting joint targets contain 591 CP and 977 L1000 outputs for BBBC036, and 701 CP and 977 L1000 outputs for BBBC047, yielding 1,568 and 1,678 dimensions. For training, each target coordinate is centered and scaled using its Folds 1–2 mean and population standard deviation; scales at or below 10−810^{-8} are replaced by one. Reported metrics invert this standardization, retaining the control-transformed response scale. Bemis–Murcko groups (Bemis and Murcko, 1996) are assigned without response values and remain disjoint across all five folds. Table 3 lists the complete fold inventory and experimental roles.

讲解

这段在干什么:交代数据预处理流水线——控制基线怎么定、化合物怎么表示、响应怎么变换聚合,最后给出两套数据集的目标维度。

需要解释的地方:

  • CP:领域常识指 Cell Painting 形态学画像;L1000 是转录组表达谱。
  • mid-CDF 变换:把处理组响应按同板对照分布做秩映射,抵消板间批次差。
  • Morgan 指纹:把分子结构编码成 2048 位向量。
  • Bemis–Murcko:按骨架分组的分子切分法。

值得留意:标准化参数只用 Folds 1–2 估计,且报告指标时被反变换回控制变换后的尺度;骨架分组"无响应值"且跨五折不相交,是为防泄漏。

b068For the complementary plate-family/context transfer, connected components are formed from released CP treatment-plate identifiers, their same-plate control contexts, linked L1000 treatment plates, and matched-condition counts, without reading responses. The construction yields six components for BBBC036 and 64 for BBBC047, assigning every CP/context and linked L1000 plate to exactly one role. Fit/selection/audit/replication condition counts are 640/320/637/319 for BBBC036 and 8,264/3,868/3,826/4,123 for BBBC047. BBBC036 therefore provides a six-family transfer case, while BBBC047 supplies broader 64-family coverage.

这段在干什么

交代“互补的板系/情境迁移”数据是怎么组装的,并给出两个数据集划分出的组件数与各角色条件数。

需要解释的地方

  • connected components(连通分量):领域常识,指按共享标识把板子连成一组组,组内相连、组间无关。
  • 不读 responses:只用标识和匹配计数搭结构,不看实验响应值——即避免用结果反推分组。
  • fit/selection/audit/replication:四种实验角色(拟合/选择/审计/复现)各自的条件数。

值得留意

作者特别强调“without reading responses”,这是防止信息泄漏的关键措辞,容易读漏。

b069Cohort CP source profiles L1000 source profiles A. Released source profiles BBBC036 21,122 (17,594/3,528) 6,929 (3,451/3,478) BBBC047 153,386 (126,814/26,572) 68,120 (64,642/3,478)

这段在干什么

这是附录里的一张数据表片段,列出 CP 与 L1000 两种来源画像下、BBBC036 与 BBBC047 两个队列的样本量。

需要解释的地方

  • Cohort CP / L1000 source profiles:领域常识上,CP(Cell Painting)和 L1000 是两种细胞成像/表达谱数据来源;这里指各队列对应的画像文件。
  • 括号里的 (x/y):原表只给数字,未说明分子分母含义,这段没明确解释。

值得留意

BBBC047 的 L1000 分母两个数都是 3,478,与上一段提到的家族覆盖数(六族 vs 64 族)呼应,但具体对应关系原文未点明。

b070Parentheses: treatment/control profile rows.

这段在干什么:这是表格的脚注说明——括号里的数字表示 treatment(处理组)/ control(对照组)两类的样本行数。

需要解释的地方:「profile」在这里指表达谱样本;「treatment/control」是实验分组,即受药物等干预的组 vs 未干预的对照组。这是生物实验的领域常识,具体定义这段没展开。

值得留意:上一段那些数字(如 17,594/3,528)加起来正好等于括号外的总数,读表时可自行核对;两个数据集 BBBC036 和 BBBC047 的 control 数竟然都是 3,478,是巧合还是共用对照,这段没说。

b071Cohort Matched conditions Unique SMILES Murcko groups CP/context dim. L1000 dim. Joint target dim. B. Matched samples and model dimensions BBBC036 1,916 1,916 1,196 591 977 1,568 BBBC047 20,081 20,068 4,540 701 977 1,678

1. 这段在干什么

这是附录里的数据统计表:列出两个数据集(BBBC036、BBBC047)的样本配对情况、分子多样性指标和各模块的维度数。起补充数据支撑作用。

2. 需要解释的地方

  • Cohort:实验批次/队列,这里指两个数据集。
  • SMILES:领域常识,一种用文本表示分子结构的编码。
  • Murcko groups:领域常识,把分子按核心骨架归类后的组数,用来衡量化学多样性。
  • CP/context、L1000、Joint target dim.:各类输入/输出特征的维度。
  • Matched conditions / samples:处理组与对照组配成对后的数量。

3. 值得留意

BBBC047 中 Matched conditions(20,081)比 Unique SMILES(20,068)略多——配对后可复用片段,这与上段"括号内为处理/对照行"的说明呼应。原文只给数字,未解释差异。

b072Statistic Fold 1 Fold 2 Fold 3 Fold 4 Fold 5 BBBC036 Matched conditions 285 509 356 411 355 Murcko groups 206 244 258 265 223 CP replicate profiles 2,222 3,995 2,779 3,213 2,803 L1000 replicate profiles 519 919 626 739 648 BBBC047 Matched conditions 3,788 3,921 4,114 4,380 3,878 Murcko groups 853 912 941 913 921 CP replicate profiles 15,891 17,060 17,263 18,390 16,429 L1000 replicate profiles 10,722 10,692 11,531 12,252 10,815

这段在干什么:用表格列出两个数据集(BBBC036、BBBC047)在5个交叉验证折上的四种样本计数。

需要解释的地方:

  • *折(Fold)*:交叉验证的5份划分,每折轮流当测试集,这里是每折的样本量。
  • *Matched conditions*:配对实验条件数。
  • *Murcko groups*:按 Murcko 骨架(领域常识:药物化学里表示分子核心结构的算法)聚类的分子组数。
  • *CP / L1000 replicate profiles*:两种表达谱平台的重复样本数。
  • 四个指标是同一批数据的不同"粒度"。

值得留意:BBBC047 各指标远大于 BBBC036,说明两数据集规模差异大;每折数值不等,说明划分并非等分。

b073Frozen source Audit Global PCC Replication Global PCC Audit CP PCC Audit L1000 PCC Parameters BBBC036 h0h_{0} -0.0030 [-0.0400, 0.0339] 0.0484 [0.0223, 0.0746] 0.0255 -0.0130 1,144,864 CellScientist 0.1921 [0.1903, 0.1939] 0.1567 [0.1517, 0.1617] 0.2758 0.1568 3,999,776 AIDE 0.1746 [0.1495, 0.1998] 0.1497 [0.1193, 0.1800] 0.2520 0.1400 4,419,872 CellForge -0.0091 [-0.0397, 0.0215] 0.0640 [0.0187, 0.1093] 0.0821 -0.0435 3,824,672 HarmonyCell 0.0657 [0.0378, 0.0936] 0.0102 [-0.0426, 0.0629] 0.0486 0.0719 2,061,343 BBBC047 h0h_{0} 0.1805 [0.1769, 0.1840] 0.2014 [0.1941, 0.2087] 0.2688 0.0494 1,201,294 CellScientist 0.1781 [0.1716, 0.1846] 0.2036 [0.1940, 0.2132] 0.2615 0.0469 1,777,486 AIDE 0.1805 [0.1769, 0.1840] 0.2014 [0.1941, 0.2087] 0.2688 0.0494 1,201,294 CellForge 0.1771 [0.1716, 0.1826] 0.2011 [0.1963, 0.2059] 0.2652 0.0444 2,531,982 HarmonyCell 0.1808 [0.1699, 0.1916] 0.1963 [0.1921, 0.2006] 0.2689 0.0484 2,284,875

讲解

这段在干什么:这是附录里的一张结果表,逐行列出五种模型(h₀ 到 HarmonyCell)在两个数据集(BBBC036、BBBC047)上的四项 PCC 指标和参数量。

需要解释的地方:PCC 是皮尔逊相关系数(领域常识,衡量预测与真实的线性相关)。列名里 "Frozen source" 是冻结源模型,其余几列是不同层面的相关性检查;方括号是置信区间。

值得留意:表中 h₀ 在 BBBC047 一行的数值与 AIDE 完全相同,参数量也一样,值得核对是否为复制粘贴错误。另外 CellForge 在 BBBC036 出现负值,与其余模型方向不同。

b074Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.

这段在干什么:这是表格的图注,规定表里加粗和下划线分别代表同一任务/折中排名第一、第二的预测均值。

需要解释的地方:加粗=该任务最好的;下划线=第二好的,同列中"distinct"即去掉并列后分别取两档。平局并列共享同一排名;区间上下界和诊断/资源类列不参与排名,故不加标记。

值得留意:排名只在"同一任务或同一折内"比较,跨任务不横比;平局共享名次意味着可能出现没有下划线的列。

b075The plate-family audit crosses the observed or permuted compound representation with the observed or an alternative control profile from the same plate family. Each selected model is refit under five fixed seeds, and each refit averages 32 permutations constructed without response values. Compounds are permuted within the observed control-profile group whenever possible, covering every held-out BBBC036 row and 98.8%98.8\% of BBBC047 rows; each remaining row receives a fixed donor with a different compound from the same held-out partition. Table 5 separates compound, control-profile, and interaction effects on Folds 4 and 5.

讲解

1. 这段在干什么

交代「plate-family audit」的实验设计:怎么组合表征与控制 profile、怎么重复采样,并说明 Table 5 用这些数据拆解三类效应。

2. 需要解释的地方

  • plate-family:同一实验板来源的样本,属领域常识。
  • permutation(置换):打乱标签/表征当"无信号"对照,看模型是否仍表现好,属领域常识。
  • held-out / donor:留出未参与训练的数据;这里的 donor 指替补的来源行。

3. 值得留意

  • 置换是「without response values」构造的——即不看真实响应,这点容易漏。
  • BBBC047 有约 1.2% 的行没被置换覆盖,改用同分区不同化合物的固定 donor,这是作者未强调的例外。

b076Model Global PCC Compound effect Control effect Interaction BBBC036 / Audit h0h_{0} -0.0030 [-0.0400, 0.0339] 0.0025 [0.0020, 0.0031] 0.0012 [-0.0126, 0.0150] -0.0005 [-0.0008, -0.0002] CellScientist 0.1921 [0.1903, 0.1939] 0.0000 [0.0000, 0.0000] 0.0146 [0.0114, 0.0178] 0.0000 [0.0000, 0.0000] AIDE 0.1746 [0.1495, 0.1998] 0.0024 [0.0019, 0.0028] 0.0260 [0.0076, 0.0445] -0.0008 [-0.0012, -0.0003] CellForge -0.0091 [-0.0397, 0.0215] 0.0018 [0.0011, 0.0025] -0.0086 [-0.0169, -0.0004] -0.0005 [-0.0010, 0.0001] HarmonyCell 0.0657 [0.0378, 0.0936] 0.0019 [0.0002, 0.0037] 0.0768 [0.0632, 0.0905] -0.0028 [-0.0044, -0.0013] BBBC047 / Audit h0h_{0} 0.1805 [0.1769, 0.1840] 0.0046 [0.0039, 0.0053] 0.0249 [0.0228, 0.0271] 0.0008 [0.0005, 0.0012] CellScientist 0.1781 [0.1716, 0.1846] 0.0000 [0.0000, 0.0000] 0.0295 [0.0243, 0.0346] 0.0000 [0.0000, 0.0000] AIDE 0.1805 [0.1769, 0.1840] 0.0046 [0.0039, 0.0053] 0.0249 [0.0228, 0.0271] 0.0008 [0.0005, 0.0012] CellForge 0.1771 [0.1716, 0.1826] 0.0050 [0.0039, 0.0061] 0.0208 [0.0163, 0.0252] 0.0014 [0.0004, 0.0023] HarmonyCell 0.1808 [0.1699, 0.1916] ×10−96.892\!\times\!10^{-9} [−1.390,2.768]×10−8[-1.390,2.768]\!\times\!10^{-8} 0.0365 [0.0245, 0.0484] ×10−81.211\!\times\!10^{-8} [−2.216,4.637]×10−8[-2.216,4.637]\!\times\!10^{-8} BBBC036 / Held-out replication h0h_{0} 0.0484 [0.0223, 0.0746] 0.0026 [×10−6,0.0053][9.873\!\times\!10^{-6},0.0053] 0.0043 [-0.0117, 0.0203] 0.0002 [-0.0006, 0.0009] CellScientist 0.1567 [0.1517, 0.1617] 0.0000 [0.0000, 0.0000] -0.0070 [-0.0088, -0.0053] 0.0000 [0.0000, 0.0000] AIDE 0.1497 [0.1193, 0.1800] 0.0013 [-0.0004, 0.0031] -0.0016 [-0.0168, 0.0136] -0.0001 [-0.0004, 0.0002] CellForge 0.0640 [0.0187, 0.1093] 0.0034 [0.0026, 0.0043] 0.0018 [-0.0087, 0.0123] 0.0001 [-0.0006, 0.0009] HarmonyCell 0.0102 [-0.0426, 0.0629] -0.0063 [-0.0100, -0.0025] -0.0038 [-0.0350, 0.0274] -0.0037 [-0.0080, 0.0005] BBBC047 / Held-out replication h0h_{0} 0.2014 [0.1941, 0.2087] 0.0030 [0.0023, 0.0038] 0.0136 [0.0077, 0.0194] 0.0002 [-0.0010, 0.0013] CellScientist 0.2036 [0.1940, 0.2132] 0.0000 [0.0000, 0.0000] 0.0231 [0.0168, 0.0293] 0.0000 [0.0000, 0.0000] AIDE 0.2014 [0.1941, 0.2087] 0.0030 [0.0023, 0.0038] 0.0136 [0.0077, 0.0194] 0.0002 [-0.0010, 0.0013] CellForge 0.2011 [0.1963, 0.2059] 0.0037 [0.0030, 0.0044] 0.0088 [0.0029, 0.0148] -0.0004 [-0.0010, 0.0002] HarmonyCell 0.1963 [0.1921, 0.2006] ×10−93.073\!\times\!10^{-9} [−5.460,11.61]×10−9[-5.460,11.61]\!\times\!10^{-9} 0.0205 [0.0142, 0.0268] ×10−92.608\!\times\!10^{-9} [−4.632,9.848]×10−9[-4.632,9.848]\!\times\!10^{-9}

这段是附录里的一张完整数值表,紧接上一段"把效应拆成三块"的说法,把每个模型在四个数据集上的拆分结果全摆出来。

这段在干什么:用 Table 5 给出各模型($h_0$、CellScientist、AIDE、CellForge、HarmonyCell)的 Global PCC,及其分解出的 compound、control、interaction 三项效应和区间。

需要解释的地方:Global PCC 是预测与真值的整体相关性(领域常识);方括号是置信区间;中间三列把整体表现拆成"化合物本身""对照谱""两者交互"三部分贡献。

值得留意:CellScientist 的 control 效应和交互几乎恒为 0.0000,AIDE 在 BBBC047 两行与 $h_0$ 数值完全相同——这两点容易一扫而过,但可能正对应正文的审计结论。

b077Fixed residual baselines isolate a perturbation-specific function class while keeping the same task. They retain the selected control-profile predictor g⁡(c)g(c) and learn one joint residual h⁡(p,a)h(p,a) using either multi-output Ridge or a shallow multi-output MLP. Both baselines use Fold 3 for selection and are refit under the same five seeds. Their Fold-4 and Fold-5 predictive scores, compound effects, and target-loss gains are reported together in Table 48.

这段在干什么:介绍两个“固定残差基线”,用来对照扰动本身的效应:保留同一个预测器 g(c),只学一个联合残差 h(p,a)。

需要解释的地方:

  • g(c):原有对照预测器,吃对照特征 c,这段里它被固定不动。领域常识:特征 c/p/a 的含义需看前文。
  • h(p,a):额外加上的残差项,输入是 p 和 a(前文定义)。
  • 多输出 Ridge / 浅层多输出 MLP:两种实现 h 的模型,输出多个目标。
  • Fold 3 选择、五种子重拟合:用第 3 折调选择,换 5 个随机种子重训。
  • Table 48:结果汇总表。

值得留意:两个基线只在 h 的模型类上不同(Ridge vs. MLP),g 和任务保持一致——这正是“隔离出扰动专属函数类”的意思;Fold 4/5 才报告分数,说明 3 折用于选择、后两折用于评估。

b078Two additional small-molecule cohorts test the final procedure beyond BBBC development. We first apply it to paired morphology–transcription LINCS–Pilot1, then transfer its ten trajectory-selected design instances to independently acquired Cell Painting LKCP Batch 2 under the same compound, dose, and control-profile inputs. Appendix E records data provenance, the order in which models and rules were fixed, the input permutations, and the compound–dose decomposition.

这章在交代数据来源:用两个额外的小分子队列,检验前面在 BBBC 上定下来的流程能不能推广出去。

几个词:小分子队列指一批用化合物处理的实验样本。LINCS–Pilot1 是配对好的「形态+转录」数据(领域常识:形态指细胞成像,转录指基因表达)。LKCP Batch 2 是另一批独立采集的 Cell Painting 图像数据。

值得留意:「ten trajectory-selected design instances」是从最开始那份数据里挑出来、再原样搬到新数据上的十个设计实例,化合物、剂量、对照输入都保持不变——这正是「迁移检验」的关键,读的时候别把它当成又跑了一遍新实验。

A.1 Physical separation of control inputs and target references

b080Model Fold Shared PCC Disjoint PCC Shared PCC drop Disjoint PCC drop Shared Loss gain Disjoint Loss gain Score F4 0.30810.3081 0.30850.3085 00 00 00 00 Control only F4 0.30840.3084 0.30820.3082 00 00 00 00 Guided F4 0.31110.3111 0.31100.3110 0.00470.0047 0.00430.0043 ×10−41.65\!\times\!10^{-4} ×10−41.51\!\times\!10^{-4} Score F5 0.31910.3191 0.31930.3193 00 00 00 00 Control only F5 0.31870.3187 0.31850.3185 00 00 00 00 Guided F5 0.32140.3214 0.32100.3210 0.00360.0036 0.00330.0033 ×10−41.31\!\times\!10^{-4} ×10−41.19\!\times\!10^{-4}

这段在干什么

这是附录里的一张对照表,报告各模型在 Shared / Disjoint 两种划分下的 PCC、PCC 降幅和 Loss 增益。

需要解释的地方

  • PCC:预测值与真值的相关性指标,领域常识。
  • Shared / Disjoint:控制输入的共享与不相交两种划分方式。
  • Loss gain:损失上的增益,表中约为 $10^{-4}$ 量级。

值得留意

Shared 与 Disjoint 两列数值几乎相同,说明划分方式几乎不影响结果。三组 Score/Control only/Guided 中,只有 Guided 的 drop 与 gain 非零,Score 与 Control only 全为 0。

b081Means of five seeds, averaging directions A/B within seed; compound replacement uses 32 fixed maps. PCC drop and loss gain are correct-minus-replacement PCC and replacement-minus-correct MSE. Guided minus control-only disjoint-arm PCC: F4 0.00280.0028 [0.0017, 0.0038][0.0017,\,0.0038], F5 0.00250.0025 [0.0012, 0.0037][0.0012,\,0.0037] (paired-seed 95% tt CIs).

这段在干什么:报告引导组与控制组之间 PCC 差值的汇总统计,用五个种子的均值和配对置信区间来支撑“控制输入与目标引用可分离”这一主张。

需要解释的地方:

  • PCC drop / loss gain:这里定义得很明确——“正确减替换的 PCC”和“替换减正确的 MSE”,不是泛泛的相关性。
  • paired-seed 95% t CI:领域常识,指按种子配对后算出的 95% 置信区间;两组的区间都跨在 0 以上(下界 0.0017、0.0012),不包含 0。

值得留意:表格里 F5 出现两次(0.3214 / 0.3210),差值 0.0036 恰好是上一段末尾的数;而这里的 0.0025 是“引导减仅控制”的分离臂口径,两者不是同一个量,别混。

b082The shared and disjoint arms are newly constructed BBBC047 comparisons; the historical shared-control result is not their absolute baseline. Physical control wells are hash-sorted within each plate and well-row stratum and allocated to balanced, disjoint A/B banks (seed 2026091401). Within direction A, the target uses bank A as reference: shared context uses A, disjoint context uses B; direction B reverses these roles. CP targets are treated-well empirical mid-CDF values relative to the direction’s reference bank, centered by 0.50.5 and aggregated by condition medians. The 701 CP features are fixed from the historical schema; no new feature selection is performed. Shared/disjoint arms have byte-identical targets and row identities within direction, and L1000 targets are unchanged. A/B are averaged within each seed, not counted as independent replicas. This intervention separates physical wells at the reconstructed CP reference layer; it does not establish that all upstream preprocessing or biological dependence has been removed. No historical qualification label is changed.

讲解

1. 这段在干什么

交代 A/B 对照臂的构造细节,说明这只是一次「物理隔离」干预,不改变历史标签。

2. 需要解释的地方

  • A/B 双银行:把同一板内对照孔按哈希排序后,均分成两组不重叠的孔,互为参照,避免"自己跟自己比"。
  • mid-CDF 中心化到 0.5:把处理孔在参照组里的分位排名当作目标值,中位对齐。
  • CP / L1000:领域常识——两种细胞表型/转录组测量面板,非本文提出。

3. 值得留意

作者主动划界:只说隔离了物理孔,不声称清除了上游预处理或生物学依赖;A/B 取平均后不算独立重复。

b083Quantity Value Interpretation Predictor refits / anchors 60 / 20 80 training stages; 120 fold evaluations Seeds / directions 5 / 2 A/B are paired inside each seed F4 retained support 3341 / 4380 76.28% of evaluation rows F5 retained support 2927 / 3878 75.48% of evaluation rows F4 / F5 source scaffolds 913 / 921 Bootstrap over source-scaffold labels Maps / scaffold draws 32 / 2000 All refits and donor maps held fixed Donor kernel total variation 0 (both folds) Every eligible compound contributes one row Guided coordinate signs 40 / 40 positive RMS, loss gain and PCC drop, each record Score/control coordinates 80 / 80 exact zero All three compound coordinates Physical controls, A / B 9144 / 9144 Unique plate/well identities; banks disjoint CP plates / retained conditions 273 / 20081 0 conditions excluded

讲解

1. 这段在干什么

这是一张审计配置清单:用表格列出评估中固定了哪些量、保留了多少样本,作用是证明控制输入与目标引用在物理上确实被分开了。

2. 需要解释的地方

  • Refits / anchors:模型重新拟合的次数与锚定训练阶段数。
  • Retained support:保留下来的评估行数占比,F4、F5 各一。
  • Donor kernel total variation:一个衡量映射平滑度的量,值为 0 表示没有额外扰动。
  • Guided / control coordinates:受引导的坐标(为正值)、对照坐标(恰为零),两者对比说明信号只从该来的地方进。

3. 值得留意

表里「Maps / scaffold draws = 32 / 2000,所有 refit 与 donor map 保持固定」意味着打乱只发生在标签层面,配置本身没变——这正是"物理分离"主张的关键。

b084Only row-uniform maps were evaluated. An independent metadata calculation proves equality of row-uniform and compound-balanced donor probabilities on all 6268 retained source pools; pre-generated compound-balanced maps are not an additional evaluated sensitivity analysis. Kernel equality is not biological exchangeability.

讲解

1. 这段在干什么

这是一段方法上的“澄清与边界声明”,回应读者可能提的质疑:为什么只评估 row-uniform 地图、没有评估 compound-balanced 地图。

2. 需要解释的地方

  • row-uniform vs. compound-balanced:两种给供体分配概率的方案(领域常识:具体定义此处没展开)。
  • metadata calculation:作者用另一套独立计算,证明两种方案在全部 6268 个保留的 source pool 上概率相等。
  • sensitivity analysis:敏感性分析,即换一种设定看结论是否稳健。

3. 值得留意

两层关键声明:一、compound-balanced 地图并非“漏做的分析”,因为概率已被证明相等,预先生成的版本不算额外的敏感性分析;二、作者特意划清界限——核(kernel)相等不等于生物学可交换性,即数学上的等价不能直接推到生物意义上的可互换。这句话很短,但很可能是作者在防过度解读。

b085Separately from the direction-specific bank targets, a full-bank CP mid-CDF reconstruction was compared with historical normalized CP ranks. It recorded a maximum absolute discrepancy of 0.191406250.19140625 and mean absolute discrepancy ×10−62.76\!\times\!10^{-6}, with 99.9799%99.9799\% agreement at 10−610^{-6}. This discrepancy was retained, not repaired by choosing a favorable target transform. The portable metadata retains the complete check.

讲解

1. 这段在干什么

报告一次"全库重建 vs 历史排名"的对齐检查结果:两者差异极小(最大绝对差 0.191,平均约 2.76×10⁻⁶,99.97% 一致),并声明没为了好看去挑变换把差异抹平。

2. 需要解释的地方

  • CP mid-CDF reconstruction:用重建出的"中位累积分布"位置,去对应历史排名——本质是两套排序口径之间的核对(领域常识:CDF 即累积分布)。
  • 10⁻⁶ 一致:不是误差为零,而是落在给定精度阈值内。

3. 值得留意

作者强调"差异被保留、未被修复",这是在主动承担不完美,以反驳"调参数凑结果"的质疑;末尾"metadata 保留完整检查"意味着可复现。这是承上一段"kernel 相等≠生物可换"的同一防守姿态——但具体机制这段没展开。

b086Model Fold Dir. Arm Global PCC MSE CP PCC CP MSE L1000 PCC L1000 MSE Fold 4 Score F4 A Shared 0.30820.3082 0.04600.0460 0.33310.3331 0.05320.0532 0.28020.2802 0.04080.0408 Score F4 A Disjoint 0.30870.3087 0.04600.0460 0.33340.3334 0.05320.0532 0.28070.2807 0.04080.0408 Score F4 B Shared 0.30790.3079 0.04600.0460 0.33170.3317 0.05320.0532 0.28100.2810 0.04080.0408 Score F4 B Disjoint 0.30840.3084 0.04600.0460 0.33170.3317 0.05320.0532 0.28180.2818 0.04080.0408 Control only F4 A Shared 0.30930.3093 0.04600.0460 0.33270.3327 0.05320.0532 0.28280.2828 0.04080.0408 Control only F4 A Disjoint 0.30860.3086 0.04600.0460 0.33220.3322 0.05330.0533 0.28180.2818 0.04080.0408 Control only F4 B Shared 0.30740.3074 0.04600.0460 0.33030.3303 0.05330.0533 0.28130.2813 0.04080.0408 Control only F4 B Disjoint 0.30770.3077 0.04600.0460 0.33040.3304 0.05330.0533 0.28180.2818 0.04080.0408 Guided F4 A Shared 0.31220.3122 0.04590.0459 0.33840.3384 0.05300.0530 0.28230.2823 0.04080.0408 Guided F4 A Disjoint 0.31150.3115 0.04590.0459 0.33730.3373 0.05310.0531 0.28210.2821 0.04080.0408 Guided F4 B Shared 0.31000.3100 0.04590.0459 0.33500.3350 0.05310.0531 0.28140.2814 0.04080.0408 Guided F4 B Disjoint 0.31050.3105 0.04590.0459 0.33560.3356 0.05310.0531 0.28170.2817 0.04080.0408 Fold 5 Score F5 A Shared 0.32190.3219 0.04540.0454 0.34920.3492 0.05360.0536 0.28950.2895 0.03960.0396 Score F5 A Disjoint 0.32220.3222 0.04540.0454 0.35000.3500 0.05350.0535 0.28890.2889 0.03960.0396 Score F5 B Shared 0.31630.3163 0.04550.0455 0.33890.3389 0.05370.0537 0.28940.2894 0.03960.0396 Score F5 B Disjoint 0.31650.3165 0.04550.0455 0.33860.3386 0.05370.0537 0.29030.2903 0.03960.0396 Control only F5 A Shared 0.32170.3217 0.04540.0454 0.34800.3480 0.05360.0536 0.29040.2904 0.03960.0396 Control only F5 A Disjoint 0.32180.3218 0.04550.0455 0.34930.3493 0.05360.0536 0.28890.2889 0.03960.0396 Control only F5 B Shared 0.31560.3156 0.04550.0455 0.33730.3373 0.05380.0538 0.28970.2897 0.03960.0396 Control only F5 B Disjoint 0.31530.3153 0.04560.0456 0.33640.3364 0.05390.0539 0.29010.2901 0.03960.0396 Guided F5 A Shared 0.32490.3249 0.04540.0454 0.35360.3536 0.05340.0534 0.29050.2905 0.03960.0396 Guided F5 A Disjoint 0.32360.3236 0.04540.0454 0.35250.3525 0.05350.0535 0.28920.2892 0.03960.0396 Guided F5 B Shared 0.31780.3178 0.04550.0455 0.34160.3416 0.05370.0537 0.28950.2895 0.03960.0396 Guided F5 B Disjoint 0.31840.3184 0.04550.0455 0.34200.3420 0.05360.0536 0.29030.2903 0.03960.0396

这段在干什么:用一张大表汇报 Shared 与 Disjoint 两种输入模式下、各模型架构在两个 Fold 上的六项指标,为「物理分离控制输入与目标引用」提供证据。

需要解释的地方:

  • Fold:交叉验证的折,这里是 Fold 4 和 Fold 5。
  • Arm:实验分支,分 Score / Control only / Guided。
  • Shared vs Disjoint:两者是否共用同一段输入,这正是本小节要检验的变量。
  • PCC / MSE:相关系数与均方误差,一高一低为好。

值得留意:Shared 与 Disjoint 的数值几乎一致,差异都在小数点后第三四位;L1000 MSE 列基本恒为 0.0408/0.0396,几乎不区分任何条件——这两点是全表的重点,但作者在本段没有点明。

b087Model Fold Dir. Arm Compound RMS Compound loss gain Compound PCC drop Fold 4 Score F4 A Shared 00 00 00 Score F4 A Disjoint 00 00 00 Score F4 B Shared 00 00 00 Score F4 B Disjoint 00 00 00 Control only F4 A Shared 00 00 00 Control only F4 A Disjoint 00 00 00 Control only F4 B Shared 00 00 00 Control only F4 B Disjoint 00 00 00 Guided F4 A Shared 0.01260.0126 ×10−41.75\!\times\!10^{-4} 0.00500.0050 Guided F4 A Disjoint 0.01180.0118 ×10−41.48\!\times\!10^{-4} 0.00410.0041 Guided F4 B Shared 0.01180.0118 ×10−41.55\!\times\!10^{-4} 0.00450.0045 Guided F4 B Disjoint 0.01180.0118 ×10−41.54\!\times\!10^{-4} 0.00440.0044 Fold 5 Score F5 A Shared 00 00 00 Score F5 A Disjoint 00 00 00 Score F5 B Shared 00 00 00 Score F5 B Disjoint 00 00 00 Control only F5 A Shared 00 00 00 Control only F5 A Disjoint 00 00 00 Control only F5 B Shared 00 00 00 Control only F5 B Disjoint 00 00 00 Guided F5 A Shared 0.01250.0125 ×10−41.43\!\times\!10^{-4} 0.00390.0039 Guided F5 A Disjoint 0.01160.0116 ×10−41.17\!\times\!10^{-4} 0.00320.0032 Guided F5 B Shared 0.01170.0117 ×10−41.19\!\times\!10^{-4} 0.00330.0033 Guided F5 B Disjoint 0.01170.0117 ×10−41.21\!\times\!10^{-4} 0.00340.0034

这段是一张消融结果表,用来对比不同条件下模型的表现。

这段在干什么:按 Fold 4/5 分组,列出 Score、Control only、Guided 三类模型在 Shared/Disjoint 两种设置下的三项指标,用数字说明哪些输入真正有用。

需要解释的地方:

  • Dir. 列(Shared/Disjoint):指控制输入与目标参考是否物理分离,领域常识里 Disjoint 是分开存放,避免模型偷看目标。
  • Score / Control only / Guided:三种实验条件,Guided 是加了引导的版本。
  • Compound RMS / loss gain / Compound PCC drop:三个评估指标,数值越大代表贡献越明显。

值得留意:Score 和 Control only 两组的数值全是 0,只有 Guided 行有非零值(如 0.0126、×10⁻⁴、0.0050),说明前两类模型在这段里没有可测得的贡献。

b088Contrast Fold Mean Seed 95% CI Scaffold 95% CI Global PCC (×103\times 10^{3}) Score: D−-S F4 0.4270.427 [−0.130, 0.984][-0.130,\,0.984] [−0.070, 0.945][-0.070,\,0.945] Score: D−-S F5 0.2430.243 [−0.583, 1.069][-0.583,\,1.069] [−0.158, 0.638][-0.158,\,0.638] Control only: D−-S F4 −0.186-0.186 [−0.769, 0.398][-0.769,\,0.398] [−0.794, 0.338][-0.794,\,0.338] Control only: D−-S F5 −0.136-0.136 [−1.679, 1.407][-1.679,\,1.407] [−0.687, 0.393][-0.687,\,0.393] Guided: D−-S F4 −0.099-0.099 [−0.972, 0.773][-0.972,\,0.773] [−0.442, 0.225][-0.442,\,0.225] Guided: D−-S F5 −0.327-0.327 [−0.963, 0.309][-0.963,\,0.309] [−0.640, 0.004][-0.640,\,0.004] Score−-control, S F4 −0.295-0.295 [−1.223, 0.633][-1.223,\,0.633] [−1.724, 1.066][-1.724,\,1.066] Score−-control, S F5 0.4280.428 [−0.662, 1.517][-0.662,\,1.517] [−0.938, 1.539][-0.938,\,1.539] Guided−-control, S F4 2.6882.688 [2.059, 3.317][2.059,\,3.317] [1.012, 4.541][1.012,\,4.541] Guided−-control, S F5 2.6742.674 [1.609, 3.739][1.609,\,3.739] [0.272, 5.761][0.272,\,5.761] Score−-control, D F4 0.3180.318 [−0.514, 1.149][-0.514,\,1.149] [−0.853, 1.491][-0.853,\,1.491] Score−-control, D F5 0.8070.807 [−0.037, 1.650][-0.037,\,1.650] [−0.554, 1.931][-0.554,\,1.931] Guided−-control, D F4 2.7752.775 [1.731, 3.818][1.731,\,3.818] [1.264, 4.468][1.264,\,4.468] Guided−-control, D F5 2.4832.483 [1.227, 3.739][1.227,\,3.739] [0.187, 5.548][0.187,\,5.548] MSE (×105\times 10^{5}) Score: D−-S F4 −1.079-1.079 [−2.884, 0.726][-2.884,\,0.726] [−2.847, 0.536][-2.847,\,0.536] Score: D−-S F5 −0.564-0.564 [−3.258, 2.131][-3.258,\,2.131] [−1.846, 0.775][-1.846,\,0.775] Control only: D−-S F4 1.9231.923 [−1.011, 4.858][-1.011,\,4.858] [0.232, 3.870][0.232,\,3.870] Control only: D−-S F5 1.4901.490 [−4.983, 7.964][-4.983,\,7.964] [−0.383, 3.538][-0.383,\,3.538] Guided: D−-S F4 1.0471.047 [−2.859, 4.954][-2.859,\,4.954] [−0.192, 2.309][-0.192,\,2.309] Guided: D−-S F5 1.7801.780 [−0.662, 4.222][-0.662,\,4.222] [0.555, 2.980][0.555,\,2.980] Score−-control, S F4 −0.598-0.598 [−3.662, 2.465][-3.662,\,2.465] [−4.858, 4.209][-4.858,\,4.209] Score−-control, S F5 −2.943-2.943 [−7.456, 1.569][-7.456,\,1.569] [−6.772, 1.784][-6.772,\,1.784] Guided−-control, S F4 −7.236-7.236 [−11.010,−3.463][-11.010,\,-3.463] [−13.811,−1.332][-13.811,\,-1.332] Guided−-control, S F5 −7.329-7.329 [−14.377,−0.282][-14.377,\,-0.282] [−18.153, 0.990][-18.153,\,0.990] Score−-control, D F4 −3.601-3.601 [−6.091,−1.110][-6.091,\,-1.110] [−7.476, 0.520][-7.476,\,0.520] Score−-control, D F5 −4.998-4.998 [−8.488,−1.507][-8.488,\,-1.507] [−8.982,−0.105][-8.982,\,-0.105] Guided−-control, D F4 −8.112-8.112 [−12.543,−3.681][-12.543,\,-3.681] [−14.199,−2.832][-14.199,\,-2.832] Guided−-control, D F5 −7.040-7.040 [−12.443,−1.636][-12.443,\,-1.636] [−17.585, 0.849][-17.585,\,0.849]

这段是一张数值表,列出各组对比的 PCC、MSE 及两种置信区间。

在干什么:用统一格式汇总消融实验的量化结果,支撑正文关于「控制输入/目标参考分离」的论证。

需解释:Fold 指数据划分折,Seed 指随机种子;Scaffold 是另一种划分方式(领域常识:按分子骨架划分以测泛化)。PCC 为相关,MSE 为误差。表头 D−S 是「去脚手架 vs 留脚手架」差异;Score−control、Guided−control 等为消融组相减。

值得留意:Guided−control 两行数值大(约 2.7)且置信区间不跨 0,而 Score/Control only 的区间多跨 0——但作者没在原文点明。

b089D−-S: disjoint minus shared wells. Within-arm contrasts subtract control only. Seed CIs use five A/B-averaged paired values (t4t_{4}). Scaffold CIs share each draw across all arms/models/directions and fix all five refits, 32 maps and the donor pool; 2000 draws, seed 2026091433+fold2026091433+\mathrm{fold}. PCC is recomputed after pooling sufficient statistics, never averaged across scaffold PCCs. These are source-scaffold, not compound-cluster, intervals; scaffold labels do not ensure fully independent biological groups.

讲解

这段在干什么:交代区间估计的构建方式——用 D−S(去共享井)做对照、用 bootstrap 重采样(2000 次)算置信区间,并说明为什么这些区间是 source-scaffold 级别而非真正独立的生物学分组。

需要解释的地方:

  • D−S(disjoint minus shared):只保留不共享的井,臂内对比时只减对照。
  • Seed CI / Scaffold CI:两种置信区间的构造方式,后者把每次抽样固定共享到所有臂/模型/方向。
  • PCC:皮尔逊相关系数(领域常识),先合并充分统计量再算,不逐 scaffold 平均。
  • scaffold:指细胞模型里的"脚手架"分组。

值得留意:作者主动声明这些区间不保证生物学独立性——这是重要的自我限制,别当成严格的统计推断。

b090Contrast Fold Mean Seed 95% CI Scaffold 95% CI Compound RMS (×103\times 10^{3}) Score: D−-S F4 00 [0, 0][0,\,0] [0, 0][0,\,0] Score: D−-S F5 00 [0, 0][0,\,0] [0, 0][0,\,0] Control only: D−-S F4 00 [0, 0][0,\,0] [0, 0][0,\,0] Control only: D−-S F5 00 [0, 0][0,\,0] [0, 0][0,\,0] Guided: D−-S F4 −0.414-0.414 [−1.567, 0.738][-1.567,\,0.738] [−0.495,−0.333][-0.495,\,-0.333] Guided: D−-S F5 −0.406-0.406 [−1.563, 0.750][-1.563,\,0.750] [−0.490,−0.327][-0.490,\,-0.327] Score−-control, S F4 00 [0, 0][0,\,0] [0, 0][0,\,0] Score−-control, S F5 00 [0, 0][0,\,0] [0, 0][0,\,0] Guided−-control, S F4 12.22012.220 [11.209, 13.231][11.209,\,13.231] [11.249, 13.186][11.249,\,13.186] Guided−-control, S F5 12.08512.085 [11.011, 13.158][11.011,\,13.158] [11.456, 12.679][11.456,\,12.679] Score−-control, D F4 00 [0, 0][0,\,0] [0, 0][0,\,0] Score−-control, D F5 00 [0, 0][0,\,0] [0, 0][0,\,0] Guided−-control, D F4 11.80511.805 [9.826, 13.784][9.826,\,13.784] [10.884, 12.725][10.884,\,12.725] Guided−-control, D F5 11.67811.678 [9.740, 13.616][9.740,\,13.616] [11.106, 12.221][11.106,\,12.221] Compound loss gain (×105\times 10^{5}) Score: D−-S F4 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.36\!\times\!10^{-12},\,2.64\!\times\!10^{-12}] Score: D−-S F5 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.71\!\times\!10^{-12},\,1.87\!\times\!10^{-12}] Control only: D−-S F4 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.91\!\times\!10^{-12},\,1.53\!\times\!10^{-12}] Control only: D−-S F5 00 [0, 0][0,\,0] [−×10−12,×10−12][-1.46\!\times\!10^{-12},\,3.05\!\times\!10^{-12}] Guided: D−-S F4 −1.403-1.403 [−4.179, 1.374][-4.179,\,1.374] [−2.134,−0.809][-2.134,\,-0.809] Guided: D−-S F5 −1.163-1.163 [−3.432, 1.107][-3.432,\,1.107] [−1.938,−0.376][-1.938,\,-0.376] Score−-control, S F4 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.43\!\times\!10^{-12},\,2.50\!\times\!10^{-12}] Score−-control, S F5 00 [0, 0][0,\,0] [−×10−12,×10−12][-1.39\!\times\!10^{-12},\,3.26\!\times\!10^{-12}] Guided−-control, S F4 16.50916.509 [13.377, 19.640][13.377,\,19.640] [10.136, 24.392][10.136,\,24.392] Guided−-control, S F5 13.10213.102 [10.729, 15.475][10.729,\,15.475] [5.985, 20.741][5.985,\,20.741] Score−-control, D F4 00 [0, 0][0,\,0] [−×10−12,×10−12][-1.32\!\times\!10^{-12},\,3.12\!\times\!10^{-12}] Score−-control, D F5 00 [0, 0][0,\,0] [−×10−12,×10−12][-2.57\!\times\!10^{-12},\,2.01\!\times\!10^{-12}] Guided−-control, D F4 15.10615.106 [9.906, 20.305][9.906,\,20.305] [9.267, 22.336][9.267,\,22.336] Guided−-control, D F5 11.93911.939 [7.423, 16.455][7.423,\,16.455] [5.390, 18.963][5.390,\,18.963]

这段是两张消融表,直接支撑上段“scaffold 标签不保证独立性”的论点。

在干什么:把各类方法(Score、Control、Guided)在 Compound RMS 和 loss gain 两个指标上的 D−S 差值列出来,做证据。

需要解释:

  • D−S:Compound 划分减 Scaffold 划分的差,衡量两种数据划分下表现差多少。
  • Fold F4/F5:不同数据折。
  • Guided−control:Guided 相对对照组的净提升,是这里最该看的行。

值得留意:纯 Score 和 Control 行的 D−S 全是 0,置信区间覆盖零;只有 Guided 行才出现非零值(如 12.220、11.805)。说明差异只在方法引入引导后才显现。两列 CI(Seed/Scaffold)在 Guided 行常不重叠,值得对比。

b091D−-S: disjoint minus shared wells. Within-arm contrasts subtract control only. Seed CIs use five A/B-averaged paired values (t4t_{4}). Scaffold CIs share each draw across all arms/models/directions and fix all five refits, 32 maps and the donor pool; 2000 draws, seed 2026091433+fold2026091433+\mathrm{fold}. PCC is recomputed after pooling sufficient statistics, never averaged across scaffold PCCs. These are source-scaffold, not compound-cluster, intervals; scaffold labels do not ensure fully independent biological groups.

这段在干什么

说明这套方法里对照组怎么切分(D−-S)以及置信区间怎么算——即结果可信度的计算口径。

需要解释的地方

  • D−-S(disjoint minus shared):把「不重叠」与「共享」的井分开的对照方式。
  • within-arm contrast:同一臂内部做对比,只减去对照。
  • seed CI / scaffold CI:两种不同重采样得到的置信区间。种子层面用 5 个 A/B 平均的配对值;脚手架层面一次抽样同时牵动所有臂、模型、方向,并固定全部五次重拟、32 张图谱和供体池,抽 2000 次。
  • PCC(领域常识):皮尔逊相关系数;这里是先合并充分统计量再重算,而不是把各 scaffold 的 PCC 平均。

值得留意

  • scaffold 区间不等于化合物聚类区间,标签并不能保证生物学上完全独立——这是作者主动声明的局限。

b092Contrast Fold Mean Seed 95% CI Scaffold 95% CI Score: D−-S F4 00 [0, 0][0,\,0] [−×10−13,×10−13][-1.50\!\times\!10^{-13},\,1.83\!\times\!10^{-13}] Score: D−-S F5 00 [0, 0][0,\,0] [−×10−13,×10−13][-1.89\!\times\!10^{-13},\,1.78\!\times\!10^{-13}] Control only: D−-S F4 00 [0, 0][0,\,0] [−×10−13,×10−13][-1.61\!\times\!10^{-13},\,1.83\!\times\!10^{-13}] Control only: D−-S F5 00 [0, 0][0,\,0] [−×10−13,×10−13][-2.33\!\times\!10^{-13},\,1.44\!\times\!10^{-13}] Guided: D−-S F4 −0.424-0.424 [−1.191, 0.342][-1.191,\,0.342] [−0.627,−0.248][-0.627,\,-0.248] Guided: D−-S F5 −0.337-0.337 [−0.944, 0.270][-0.944,\,0.270] [−0.559,−0.115][-0.559,\,-0.115] Score−-control, S F4 00 [0, 0][0,\,0] [−×10−13,×10−13][-2.05\!\times\!10^{-13},\,1.11\!\times\!10^{-13}] Score−-control, S F5 00 [0, 0][0,\,0] [−×10−13,×10−13][-2.28\!\times\!10^{-13},\,1.33\!\times\!10^{-13}] Guided−-control, S F4 4.7084.708 [3.832, 5.583][3.832,\,5.583] [2.949, 6.840][2.949,\,6.840] Guided−-control, S F5 3.6253.625 [2.955, 4.296][2.955,\,4.296] [1.650, 5.741][1.650,\,5.741] Score−-control, D F4 00 [0, 0][0,\,0] [−×10−13,×10−13][-2.22\!\times\!10^{-13},\,1.33\!\times\!10^{-13}] Score−-control, D F5 00 [0, 0][0,\,0] [−×10−13,×10−13][-1.89\!\times\!10^{-13},\,1.67\!\times\!10^{-13}] Guided−-control, D F4 4.2834.283 [2.823, 5.744][2.823,\,5.744] [2.673, 6.238][2.673,\,6.238] Guided−-control, D F5 3.2883.288 [2.042, 4.535][2.042,\,4.535] [1.471, 5.215][1.471,\,5.215]

讲解

1. 这段在干什么

用一张对照表检验:把输入和目标做物理隔离后,各类预测分数是否还站得住。

2. 需要解释的地方

  • Contrast:两组的差值;Fold:数据切分折,F4/F5 是其中两折。
  • Mean:均值;Seed / Scaffold 95% CI:重随机种子 / 按骨架划分下的 95% 置信区间(领域常识)。
  • D−S:折叠间差异。

3. 值得留意

所有含 Score 的对比 Mean 都是 0,CI 里是 10⁻¹³ 级,即隔离后分数信号基本归零;而 Guided 相关对比仍显著偏离 0(如 4.708),差异不在一个量级。

b093D−-S: disjoint minus shared wells. Within-arm contrasts subtract control only. Seed CIs use five A/B-averaged paired values (t4t_{4}). Scaffold CIs share each draw across all arms/models/directions and fix all five refits, 32 maps and the donor pool; 2000 draws, seed 2026091433+fold2026091433+\mathrm{fold}. PCC is recomputed after pooling sufficient statistics, never averaged across scaffold PCCs. These are source-scaffold, not compound-cluster, intervals; scaffold labels do not ensure fully independent biological groups.

讲解

这段在干什么:交代统计推断的口径——(D−S) 是「减去共享孔」的差集,并按「臂内」和「scaffold 重采样」两套规则分别构造置信区间。

需要解释的地方:

  • *Within-arm contrasts subtract control only*:同一条臂内做对比时,只减对照,不跨臂相减。
  • *Seed CI / Scaffold CI*:前者按 5 个 A/B 平均后的配对值(t₄ 分布)算;后者每次重采样在所有臂/模型/方向间共享同一抽法,并固定那 5 次重拟合、32 张图与供体池;共 2000 次抽样,种子含 fold。
  • *PCC 先池化充分统计量再算,绝不把各 scaffold 的 PCC 取平均*——这是领域常识里避免平均相关系数的做法。

值得留意:作者自认这是 source-scaffold 区间,不是 compound-cluster 区间;scaffold 标签不等于完全独立的生物分组。

b094The guided model has positive compound RMS, target-loss gain and PCC drop in all 40 evaluations, whereas Score and control only are exactly invariant. The guided disjoint-arm PCC advantage over control only remains positive under both reported interval methods. However, all guided disjoint-minus-shared coordinate differences include zero under the five-seed interval while their fixed-refit source-scaffold intervals are negative; the two uncertainty scopes do not establish an unchanged effect. The F5 guided-minus-control MSE contrast likewise crosses zero under the source-scaffold interval despite a negative seed interval. A positive prediction RMS indicates input use, not target benefit. RMS is the mean over maps of the square root of mean squared prediction distance, not the square root after averaging maps.

这段在干什么

汇报引导模型在全部 40 次评估中都有正向效应,但两种区间方法结论不一致,故不能断言效应不变。

需要解释的地方

  • RMS / PCC / MSE:均为预测指标。RMS 这里是「先对每张图算均方距离、开方,再对图取平均」——不是先平均再开方。
  • 区间方法:用不同重采样范围(五种子 vs. source-scaffold)估不确定度,是领域常识术语。

值得留意

作者自陈:正向预测 RMS 只说明「用了输入」,不等于「对目标有好处」——这是防止误读的关键限定。

Appendix B Exact protocol, decision rules, and statistical units

b096The fixed executor owns joint inputs (c,p,a)(c,p,a), the response target, data transformations, folds, evaluator, checkpoint rule, and training configuration; candidate code cannot alter data loading, targets, folds, or metrics. Discovery accesses Folds 1–3 only. Global PCC is one Pearson correlation after flattening all post-search rows and all dimensions of the inverse-standardized joint output. CP and L1000 PCC use the same operation within each response block. MSE is averaged over every element of the inverse-standardized joint output. At search step tt, the policy maps the current model, Fold-3 diagnostics, and compact search history to a proposed revision,

这段在干什么:交代实验的固定约束与评价口径——执行器锁死输入、目标、折与指标,发现阶段只能用 Fold 1–3,并逐条定义 PCC/MSE 怎么算。

需要解释的地方:

  • 固定执行器:数据和训练配置由它统一掌管,候选代码只能改模型本身,不能动数据或指标(防止偷偷改考卷)。
  • Global PCC:先把搜索后的所有行、输出所有维度摊平,再算一次 Pearson 相关。
  • 逆标准化:把标准化过的输出还原回原尺度再算误差(领域常识)。
  • search step t:策略按当前模型 + Fold-3 诊断 + 精简历史,提出下一版修改。

值得留意:发现阶段只碰 Fold 1–3,说明 Fold 4 之后是留给验证的——原文这里没明说,但"only"是刻意的。

b097where edits are limited to architecture, fusion, response readout, loss, optimizer, and scheduler. The experiment log stores the proposal’s parent, hypothesis, source or structured design, declared input use, execution and repair results, checkpoint, metrics, provider usage, and compute. Within each ten-slot prediction-score trajectory, the incumbent changes only under a strict improvement in Fold-3 Global PCC. After all ten trajectories finish, the first, within-policy selection chooses one source per search policy from that policy’s complete executable candidate pool by maximum Fold-3 Global PCC, with lower Fold-3 MSE, fewer parameters, and stable candidate identity as tie-breakers. Neither Fold 4 nor Fold 5 enters this within-policy selection. The four policy-selected agent sources and three distinct diagnostic sources are then refit under five paired seeds and evaluated on Fold 4; the h0h_{0} source is also displayed under its simple-concatenation role. A second, pre-specified across-policy score rule chooses the agent source with the highest mean Fold-4 PCC for transfer to Fold 5. Within source selection, Fold 4 enters only this disclosed across-policy choice, while Fold 5 never enters. The path-constrained and falsification-guided studies instead carry all ten Fold-3 winners through Folds 4 and 5 without Fold-4 reselection. Before Fold 4, each study freezes the candidate set and registers the source-check rules, input-intervention construction, map seeds, threshold-construction rule, status rule, and aggregation. The realized model-specific thresholds are computed from the 32 reference scores on the fold being audited. For the prediction-score panel, Fold 5 independently refits the frozen source on Folds 1–2, selects its checkpoint on Fold 3 under the same paired-seed schedule, and evaluates Fold 5 once. Thus Fold 5 does not reuse Fold-4 weights and never changes the selected source. LINCS and LKCP subsequently test the final procedure outside BBBC development.

这段在干什么

紧接着上一段的搜索步骤,交代实验协议:编辑范围、日志记录、候选如何按 Fold-3 择优、如何在 Fold-4/5 上检验。

需要解释的地方

  • Fold-3/4/5 Global PCC:交叉验证的折与全局皮尔逊相关,领域常识,即衡量预测与真值的吻合度。
  • incumbent:当前占位的候选,只有 Fold-3 PCC 严格提升才换。
  • within-policy / across-policy:先在每种策略内部选一个源,再跨策略选一个。
  • refit:固定来源后重新训练。

值得留意

  • Fold 4、5 的分工被作者刻意隔开:within-policy 阶段两者都不参与,across-policy 阶段只进 Fold 4,Fold 5 始终不进。这是防泄漏设计,别读漏。
  • 阈值由 32 个参考分数算出,这是作者定的具体数,不是通用规则。
  • 末句 LINCS/LKCP 是 BBBC 之外的最终检验。

b098The compound intervention first seeks a different-fingerprint derangement within rows sharing the same observed control-profile vector. If that class derangement is impossible, a deterministic partition-wide donor with a different fingerprint is used and recorded. Only the compound representation changes: the source control profile and dose remain fixed. The resulting effect is measured under this control-stratified, dose-preserving replacement distribution. The control intervention maps each row to a different observed control-vector class; donor reuse is allowed to ensure that every control vector changes. Dose maps replace only the scalar dose, with observed-dose support checked against the source compound’s released BRD identity and available cell/time annotations. None of these maps uses responses, model predictions, or evaluation scores to choose donors.

讲解

1. 这段在干什么

这段在定义三种"干预"的精确操作规则——分别改化合物、改对照、改剂量,并强调选供体时不看响应或模型分数。

2. 需要解释的地方

"不同指纹的错排(derangement)"是领域常识:指每个元素都不留在原位的重新排列,这里要求供体行的"指纹"与原行不同,即不能照抄自己。"对照分层"指只在与源行共享同一对照向量的行里找供体;找不到才退而用全分区供体。"保存剂量的替换分布"指换对照时剂量不动,换剂量时只动标量剂量。

3. 值得留意

最后一句"不使用响应、预测或评分选供体"是作者在防一种质疑:如果选供体时偷看了结果,后面的归因就有循环论证之嫌。另外剂量替换要对照源化合物已发布的 BRD 身份和细胞/时间注释来核验——这是可复现性上的约束。

b099For the primary scaffold-fold compound maps, within-control-group derangements cover 411/411411/411 and 355/355355/355 rows on BBBC036, 4330/43804330/4380 and 3831/38783831/3878 on BBBC047, 847/847847/847 and 868/868868/868 on LINCS, and 378/405378/405 and 378/432378/432 on LKCP, respectively on Folds 4 and 5. The remaining rows use partition-wide donors. These counts specify control matching; the response contrasts retain the source dose, including when the donor compound was not observed at that dose. The plate-family analysis uses its separately specified maps in Appendix A.

这段在干什么:交代主分析中化合物映射的匹配完成度,说明哪些行用了组内错配、哪些改用分组宽供体,并声明对照方式。

需要解释的地方:

  • derangement(错配):把供体化合物打乱重配,避免用自身当对照;这是领域常识。
  • 411/411 这类分数:分子是成功错配的行数,分母是总行数。
  • partition-wide donors:剩余行改用更大范围(整个分组)的供体,而非组内。

值得留意:作者强调对照只换供体、响应仍保留原始剂量(哪怕该剂量下没观测到供体化合物);具体映射规则在附录 A,不在本段。

b100For the 2×22\times 2 audit, the model-specific reference values are derived from the intervened predictions rather than from responses used to choose a permutation. For compound use they are bj,r(chem)=12​(mp~​c,r+mp~​c~,r)b^{(\mathrm{chem})}_{j,r}=\tfrac{1}{2}(m_{\tilde{p}c,r}+m_{\tilde{p}\tilde{c},r}); for control-profile use, bj,r(context)=12​(mp​c~,r+mp~​c~,r)b^{(\mathrm{context})}_{j,r}=\tfrac{1}{2}(m_{p\tilde{c},r}+m_{\tilde{p}\tilde{c},r}); and for their interaction, bj,r(int)=mp~​c,r−mp~​c~,rb^{(\mathrm{int})}_{j,r}=m_{\tilde{p}c,r}-m_{\tilde{p}\tilde{c},r}. Dose uses the prediction score under a within-compound dose permutation. In every case, τj,q\tau_{j,q} is the 95th percentile of absolute pairwise differences among the 32 fixed reference values. The lower and upper bounds used for the status are the empirical 2.5th and 97.5th percentiles of the 32 effect values; the interaction uses absolute effects. A model is qualified when the lower bound exceeds τj,q\tau_{j,q}, unsupported when the upper bound does not exceed τj,q\tau_{j,q}, and inconclusive otherwise. Inputs without enough legal interventions are reported as not identifiable. For the five-refit prediction-score panel, each refit receives its own status and an aggregate status requires agreement in at least four of five refits; if no category reaches four of five, the aggregate is inconclusive. Path-constrained and falsification-guided results report status counts across ten selected models without a vote. The 50-refit coordinate analysis is descriptive at the refit level and uses the ten trajectory-selected endpoint instances as the policy-distribution units. These instances contain six unique configurations; repeated configurations are retained because they are independent trajectory selections. For the BBBC, LINCS, and LKCP panels, the 32 permutations are repeated measurements within a trained model, not independent sample units. The sci-Plex feedback study uses 32 development maps and 128 final held-out maps, which are likewise repeated measurements; Norman uses a separate set of 32 complete derangements (Appendix D).

这段在干什么

这是一段操作手册式的协议说明:交代 2×2 审计中各类"参考值"怎么算、状态怎么判定、以及各面板的统计单元是什么。

需要解释的地方

  • 参考值 b:由"干预后的预测"推出来,而非用来选置换的响应。chem/context/int 三种分别对应化合物、对照、交互。
  • τ:32 个固定参考值两两绝对差的第 95 百分位,相当于噪声门槛。
  • 状态判定:下界>τ→qualified;上界≤τ→unsupported;否则 inconclusive。交互用绝对效应。
  • 统计单元:置换和 refit 是同一模型内的重复测量,不是独立样本(领域常识:重复测量不能当独立样本算)。

值得留意

"not identifiable"(干预不足)和"inconclusive"是两种不同结果,别混。聚合状态需五次 refit 中至少四次一致。

b101The falsification-guided trainer evaluates α∈{0.05,0.10,0.20,0.35,0.50,0.75,1.00}\alpha\in\{0.05,0.10,0.20,0.35,0.50,0.75,1.00\} on Fold 3. Predictive noninferiority allows at most a 0.0030.003 decrease from the control-only predictor in Global PCC, a 0.0050.005 decrease in either response block, and a 0.00050.0005 increase in MSE. Search qualification uses eight fixed Fold-3 replacement maps. The empirical 2.5th percentile of their compound-related Δ​ℒ\Delta\mathcal{L} values (increases in inverse-standardized joint-output MSE) must exceed both the 95th percentile of absolute pairwise differences among the replaced-input losses and 0.0010.001 times the reference loss. Here reference loss is the current candidate’s Fold-3 MSE with correct inputs. Dose uses the same rule when eligible, with reference loss restricted to eligible rows. These eight development maps are separate from the 32 held-out audit maps. The discovery code groups equal Morgan fingerprints for its replacement maps. Its dose gate requires at least 32 eligible rows and eight fingerprint groups in both the fit partition and Fold 3, with at least 0.10.1 log-dose contrast. For each candidate, epoch/scale selection maximizes Fold-3 Global PCC among qualifying states, breaking ties by lower MSE; if none qualifies, the best predictive state is retained with its qualification status recorded.

逐段带读

1. 这段在干什么

这是附录方法论部分:把「证伪引导训练器」的筛选规则(α 网格、非劣性容差、置换图显著性检验、剂量门槛、epoch/scale 选择)逐条写死,让协议可复现。

2. 需要解释的地方

  • α 网格:一组可调的显著性水平候选值,在 Fold 3 上试。
  • 预测非劣性:新模型允许比「仅对照」略差,但有上限——Global PCC 至多降 0.003、单响应块至多降 0.005、MSE 至多升 0.0005。
  • 资格检验:用 8 张固定置换图算 Δℒ,其 2.5 百分位要同时超过两个阈值(成对差分的 95 百分位、参考损失的 0.001 倍)。
  • 参考损失:当前候选在正确输入下的 Fold-3 MSE。
  • Morgan 指纹:领域常识,一种分子结构编码;相同指纹会被分到同一组。
  • epoch/scale 选择:在合格状态里挑 Global PCC 最高,平手取更低 MSE。

3. 值得留意

  • 这 8 张开发图和 32 张审计图是分开的——防止调参泄漏到最终审计。
  • 资格不过时并非丢弃,而是「保留最佳预测状态并记录未合格」,这点容易被读漏。

b102Trajectory selection is distinct from this within-candidate checkpoint rule. BBBC047 selects the highest-PCC qualifying candidate, with lower MSE as tie-breaker, or reports no qualified endpoint. BBBC036 prefers predictively noninferior candidates and ranks them by compound-related loss increase, then Global PCC, lower MSE, and stable candidate identity. If none is noninferior, that ordering retains a predictive-fallback endpoint from the executable candidates. Strict qualification is reported separately for BBBC036. These cohort-specific rules select the ten sources before either post-search boundary; held-out results never reselect them.

讲解

1. 这段在干什么

承接上句的"候选内检查点规则",说明轨迹选择是另一套规则:不同队列(BBBC036/047)用不同排序挑出十个源。

2. 需要解释的地方

  • PCC:预测值与真值的相关系数,越高越准(领域常识)。
  • tie-breaker:并列时的次级排序依据。
  • noninferior:非劣,即不比别的候选差。

3. 值得留意

排序规则各队列不同,且都发生在"两个边界"之前——即held-out 结果不参与重选,这是防止数据泄漏的关键。

b103Held-out dose testing requires at least 64 eligible rows and coverage of at least 10% of the partition. The prediction-score and path-constrained tests accept any observed within-compound dose change; the falsification-guided test additionally requires eight compounds and an absolute log-dose difference of 0.10.1. Compound identity here is the BRD compound identifier, with sample/batch suffixes removed, rather than a fingerprint-equivalence class. On BBBC047, the resulting coverage is 0/43800/4380 rows on Fold 4 and 14/387814/3878 rows from seven compounds on Fold 5 (0.3610%0.3610\%); neither boundary meets the minimum, and neither contains a legal 0.10.1-log-dose contrast. Dose is therefore not identifiable in the BBBC047 panels (Table 50). Equal-fingerprint grouping alone would merge distinct released identities. On LINCS and LKCP, all replacement dose values in the 32 frozen maps are observed for the source BRD compound and its available cell/time annotations; the original input tensors and effect estimates are retained. The held-out PCC threshold is computed from fold-specific reference-score variation and is separate from the Fold-3 loss gate.

这段在干什么:交代 held-out dose 测试的准入门槛(≥64 行、覆盖 ≥10%),并报告 BBBC047 上 coverage 太低、剂量不可识别,因此不做该检验。

需要解释的地方:

  • *coverage*:满足剂量对比条件的行占该 partition 的比例。
  • *BRD compound identifier*:领域常识——Broad 研究所的化合物编号,去掉样本/批次后缀。
  • *log-dose 0.1 contrast*:对数剂量差至少 0.1 的组内比较。

值得留意:Fold 4 是 0 行,Fold 5 只有七个化合物、0.361%,两边都没过线;作者特意说用指纹分组会合并不同 release 身份。

b104The prediction-score discovery study contains 80 ten-slot trajectories: two tasks, four search methods, and ten independent trajectory seeds. Each trajectory’s budget includes the common initial predictor, and discovery intervals treat paired trajectory identifiers as the statistical unit because all policies share the same candidate-training seed schedule within an identifier. Selected-model comparisons use five paired fitting seeds on the relevant held-out fold. Path-constrained and falsification-guided studies retain all ten trajectory winners and report their distributions. All summaries are computed from fixed experiment logs that retain failed slots, repairs, and provider attempts.

这段交代实验的统计口径,是协议的操作性收尾。

术语:

  • ten-slot 轨迹:每条搜索路径有 10 个候选槽位,预算共享同一个初始预测器。
  • 统计单元:不按单个模型算,而按「配对的轨迹标识」算——因为同一标识下各策略共用同一套候选训练种子(领域常识:配对设计可消掉种子噪声)。
  • 选择模型比较:在对应留出折上用 5 个配对拟合种子比对。
  • 路径约束/证伪引导研究:保留全部 10 个轨迹胜者,报告分布而非单点。

值得留意:所有汇总都基于保留失败槽、修复和 provider 尝试的固定日志——即不清理失败样本,这是审计可复现的关键。

b105The accompanying package separates three reproducibility targets. Numerical source records and deterministic renderers regenerate reported tables. Frozen-source refitting trains new parameters under the recorded data and selection rules. Checkpoint replay takes the selected source, matching fitted weights, prepared data, and replacement specification to recompute predictions and input-use effects by inference. Its standalone feedback entry point verifies checkpoint and data hashes and writes fresh outputs. Biological matrices and fitted weights are supplied as external inputs; tabulated records include checkpoint identities and counterfactual sufficient statistics. The package’s docs/REPRODUCIBILITY.md gives the corresponding commands and required assets.

讲解

这段在干什么:交代配套代码包如何实现三种可复现目标,并说明数据、权重等资产从哪来。

需要解释的地方:

  • 三种复现目标:① 数值源记录 + 确定性渲染器 → 重出表格;② 冻结源重拟合 → 用记录的数据和选择规则重新训练参数;③ 检查点回放 → 用选定的源、匹配的权重和数据重算预测与输入使用效应。
  • 反事实充分统计量:领域常识,指做反事实推断足够用的汇总量。

值得留意:生物矩阵和拟合权重是外部输入,并不打包在代码里;要完全复现得自己备齐这些资产。

b106Comparator Best@10 Δ\Delta [95% CI]; pp Frontier AUC Δ\Delta [95% CI]; pp BBBC036 AIDE +0.0084 [+0.0044, +0.0124] p=9.638×10−4p=9.638\times 10^{-4} +0.0054 [-0.0003, +0.0111] p=6.157×10−2p=6.157\times 10^{-2} CellForge +0.0151 [+0.0107, +0.0196] p=3.24×10−5p=3.24\times 10^{-5} +0.0106 [+0.0056, +0.0155] p=9.601×10−4p=9.601\times 10^{-4} HarmonyCell +0.0046 [-0.0007, +0.0100] p=8.064×10−2p=8.064\times 10^{-2} +0.0025 [-0.0016, +0.0066] p=1.962×10−1p=1.962\times 10^{-1} BBBC047 AIDE +0.0214 [+0.0105, +0.0324] p=1.636×10−3p=1.636\times 10^{-3} +0.0180 [+0.0078, +0.0282] p=3.18×10−3p=3.18\times 10^{-3} CellForge +0.0181 [+0.0057, +0.0304] p=9.051×10−3p=9.051\times 10^{-3} +0.0159 [+0.0046, +0.0271] p=1.086×10−2p=1.086\times 10^{-2} HarmonyCell +0.0028 [-0.0168, +0.0223] p=7.56×10−1p=7.56\times 10^{-1} +0.0094 [-0.0033, +0.0220] p=1.293×10−1p=1.293\times 10^{-1}

讲解

1. 这段在干什么

这是附录 B 的对比结果表:把 AIDE、CellForge、HarmonyCell 三个方法在 BBBC036 / BBBC047 两个数据集上的 Best@10 和 Frontier AUC 提升幅度、置信区间和 p 值列出来,用来支撑正文里"方法确实带来增益"的统计证据。

2. 需要解释的地方

  • Δ [95% CI]; p:相对比较器的提升量(Δ)、95% 置信区间,以及显著性检验的 p 值。区间跨 0 或 p 偏大,说明该提升不稳健。
  • Best@10 / Frontier AUC:两个评价指标,属于领域常识——前者看最优前 10 个候选的表现,后者看整体前沿曲线下的面积。

3. 值得留意

HarmonyCell 的多行 Δ 区间都跨 0、p 值偏大(如 0.756、0.129),说明它的增益在这两个数据集上并不显著;CellForge 则普遍区间不含 0。这张表是"用统计说话",别只看 Δ 的正负号。

b107Comparator Δ\Delta Global PCC 95% CI pp nn BBBC036 HarmonyCell -0.0015 [-0.0037, +0.0007] 1.228×10−11.228\times 10^{-1} 5 AIDE +0.0117 [+0.0030, +0.0205] 2.027×10−22.027\times 10^{-2} 5 CellForge +0.0158 [+0.0095, +0.0222] 2.307×10−32.307\times 10^{-3} 5 TabM +0.0048 [+0.0016, +0.0081] 1.445×10−21.445\times 10^{-2} 5 RealMLP +0.0118 [-0.0103, +0.0338] 2.13×10−12.13\times 10^{-1} 5 Standard MLP +0.0193 [+0.0072, +0.0313] 1.128×10−21.128\times 10^{-2} 5 h0h_{0} +0.0207 [+0.0130, +0.0283] 1.733×10−31.733\times 10^{-3} 5 Joint TabR +0.0369 [+0.0343, +0.0395] 2.576×10−62.576\times 10^{-6} 5 Ridge +0.2324 [+0.2301, +0.2347] 1.033×10−91.033\times 10^{-9} 5 BBBC047 HarmonyCell +0.0012 [-0.0083, +0.0108] 7.394×10−17.394\times 10^{-1} 5 AIDE +0.0337 [+0.0283, +0.0390] 6.271×10−56.271\times 10^{-5} 5 CellForge +0.0293 [+0.0256, +0.0330] 2.57×10−52.57\times 10^{-5} 5 TabM +0.0181 [+0.0150, +0.0211] 7.867×10−57.867\times 10^{-5} 5 RealMLP +0.0202 [+0.0154, +0.0250] 3.062×10−43.062\times 10^{-4} 5 Standard MLP +0.0337 [+0.0275, +0.0398] 1.103×10−41.103\times 10^{-4} 5 h0h_{0} +0.0337 [+0.0283, +0.0390] 6.271×10−56.271\times 10^{-5} 5 Joint TabR +0.0500 [+0.0451, +0.0549] 9.339×10−69.339\times 10^{-6} 5 Ridge +0.2541 [+0.2530, +0.2552] 3.077×10−113.077\times 10^{-11} 5

这是两套数据集(BBBC036、BBBC047)上各模型的 ΔGlobal PCC 与显著性结果汇总表,接在上一段显著性数据之后。

在干什么:用表格列出各对比模型相对基线的整体相关变化、置信区间与 p 值,支撑“预测贡献”审计结论。

解释:ΔGlobal PCC 指相对基线整体皮尔逊相关系数的增量,正数代表更好;95% CI 是不确定性区间;p 值越小越可能非偶然;n=5 是重复次数。

留意:Ridge 的 Δ 最大(0.23、0.25)却并非主结论重点;HarmonyCell 在 BBBC047 上几乎为零;RealMLP 的 CI 跨零,接近不显著——别只看点估计。

b108For chemical-generalization contrasts, complete held-out Murcko-scaffold clusters are jointly resampled across all five paired refits, and Global PCC is recomputed from cluster sufficient statistics. This preserves within-scaffold response dependence and quantifies uncertainty over held-out chemical families. The Holm step-down adjustment applies to the three registered discovery-policy contrasts reported in Table 18; the broader frozen-predictor comparison is descriptive and reports its unadjusted paired intervals explicitly (Holm, 1979).

这一段在交代统计推断的决策规则:重采样怎么做、Holm 校正用在哪些对比上、哪些只是描述性报告。

需要解释的地方

  • Murcko-scaffold 簇:按分子骨架(领域常识:Murcko 骨架是把侧链去掉后的核心结构)分组,整簇一起重采样,不拆散。
  • 充分统计量(sufficient statistics):不重算每个分子,只用簇级汇总值重算 Global PCC,等价但更快。
  • Holm step-down:多重比较校正法(领域常识),控制假阳性。
  • 描述性对比报未校正区间:即不声称显著性。

值得留意

  • 整簇重采样是为了保住"同一骨架内响应相关"这一结构,不是随机的分子级 bootstrap。
  • Holm 只覆盖 Table 18 的三个对比,更广的 frozen-predictor 比较不算在内,二者证据强度不同。

b109Contrast Evaluation Δ\Delta Global PCC [95% CI] BBBC036 CellScientist −- control-only Audit -0.0002 [-0.0020, +0.0015] CellScientist −- HarmonyCell Audit -0.0014 [-0.0032, +0.0003] score-selected −- h0h_{0} Replication +0.0344 [+0.0195, +0.0514] BBBC047 CellScientist −- control-only Audit +0.0016 [+0.0001, +0.0032] CellScientist −- HarmonyCell Audit +0.0052 [+0.0034, +0.0069] score-selected −- h0h_{0} Replication +0.0366 [+0.0308, +0.0424]

逐段带读

1. 这段在干什么

这是一张对比评估表,报告两个数据集(BBBC036/BBBC047)下三类对比的全局 PCC 及 95% 置信区间,是上一段"报告未校正配对区间"的具体数值落实。

2. 关键概念

  • PCC:皮尔逊相关系数(领域常识),衡量预测与真值的线性相关。
  • 95% CI:区间不含 0 即视为该对比稳健。
  • Audit / Replication:前者是审计性对比(对照基线),后者是复现性对比(score-selected − h₀)。

3. 值得留意

Replication 行在两个数据集区间均明显为正(+0.0344、+0.0366)且不跨 0;而 Audit 各行区间几乎都贴着 0 甚至跨 0,说明前者才是稳健信号,后者的差异不宜过度解读——这正是上句"descriptive、不做多重校正"的含义。

b110For the central falsification-guided study, effects are first computed from all 32 fixed within-model input permutations. The conditional interval resamples held-out Murcko scaffolds while holding the ten selected models fixed; the hierarchical sensitivity interval jointly resamples selected models and scaffolds. Models and scaffolds are the resampling units, whereas the 32 permutations remain fixed repeated measurements within a model.

这段交代中心证伪研究的统计口径。

1. 这段在干什么:说明效应值与两类区间的计算方式,即“用什么单位重采样”。

2. 需要解释的地方:

  • 固定置换:32 种输入排列是先定死的重复测量,不参与重采样。
  • 条件区间:只重采样留出的 Murcko 骨架,十个选定模型不动。
  • 层次敏感性区间:模型和骨架一起重采样。
  • 重采样单位:模型、骨架是“抽签单位”,置换不是。(统计学术语属领域常识)

3. 值得留意:两类区间差别就在“模型动不动”,这点最易读漏。

b111Quantity Estimate Scaffold 95% CI Model×\timesscaffold 95% CI BBBC036 / Fold 4 Full −- anchor PCC +0.0046 [-0.0042, +0.0165] [-0.0043, +0.0164] Compound PCC drop +0.0091 [+0.0001, +0.0211] [+1.78×10−5+1.78\times 10^{-5}, +0.0213] Compound target-loss gain +0.00039 [+0.00002, +0.00089] [+0.00002, +0.00090] BBBC036 / Fold 5 Full −- anchor PCC -0.0020 [-0.0104, +0.0052] [-0.0102, +0.0053] Compound PCC drop +0.0022 [-0.0050, +0.0090] [-0.0051, +0.0092] Compound target-loss gain +0.00009 [-0.00022, +0.00038] [-0.00022, +0.00039] BBBC047 / Fold 4 Full −- anchor PCC +0.0037 [+0.0018, +0.0060] [+0.0018, +0.0060] Compound PCC drop +0.0053 [+0.0033, +0.0078] [+0.0033, +0.0078] Compound target-loss gain +0.00018 [+0.00011, +0.00026] [+0.00011, +0.00026] BBBC047 / Fold 5 Full −- anchor PCC +0.0028 [+0.0003, +0.0061] [+0.0002, +0.0061] Compound PCC drop +0.0048 [+0.0024, +0.0079] [+0.0023, +0.0079] Compound target-loss gain +0.00017 [+0.00009, +0.00026] [+0.00009, +0.00026]

这段在干什么:这是一张数值明细表,逐条列出两个数据集(BBBC036/BBBC047)在两个 fold 下三种指标的点估计与两种 95% CI。

需要解释的地方:CI 是置信区间,即估计的不确定范围。表里给了两种:Scaffold 与 Model×scaffold——领域常识是,换一种重采样单位,区间会略变,用来检验结论对"统计单位"选择是否敏感。

值得留意:同一指标两列 CI 几乎相同,说明换单位影响很小;但 BBBC036/Fold 5 的 Compound PCC drop 等区间跨越 0,即该 fold 下效应不显著,与 BBBC047 四条几乎都远离 0 形成对比。

b112The larger BBBC047 cohort retains positive hierarchical intervals for all three quantities on both held-out folds: Full−-control-only PCC, compound-shuffle PCC drop, and compound target-loss gain. On BBBC036, Fold-4 compound and target-loss intervals remain positive, whereas Fold-5 intervals include zero. Scaffold resampling therefore shows which compound-family effects persist without selecting models or permutations again.

讲解

1. 这段在干什么

报告稳健性检验的结果:把化合物按骨架重新抽样(scaffold resampling)后,哪些效应仍站得住。

2. 需要解释的地方

  • 层级区间(hierarchical intervals):领域常识,指考虑了数据分组结构的置信/可信区间,判断"是否为正"就是看区间是否跨过零。这里三个量的区间全为正,说明效应稳定。
  • 三个量:Full−control-only PCC(完整模型比只用对照的模型好多少)、compound-shuffle PCC drop(打乱化合物标签后性能掉多少)、compound target-loss gain(化合物对靶点损失的正贡献)。
  • 保留/跨零:区间全正=效应在;跨零=不排除无效。

3. 值得留意

BBBC047(大队列)三个量在两个 held-out fold 上全正;BBBC036 只有 Fold-4 正,Fold-5 跨零——同一数据不同折结论不一致,作者没直说,但这正是稳健性检验要暴露的边界。

b113Trajectory-level Best@10 Global PCC is the primary discovery outcome, with frontier AUC as the co-primary search-efficiency summary. Each contrast pairs the shared trajectory seed and reports its effect estimate, two-sided paired interval, and complete ten-trajectory distribution. Holm adjustment is applied to the three pre-specified policy contrasts within each task and outcome (Table 18). Response blocks, response-sensitive metrics, cost records, code-check outcomes, and input-shuffle effects retain their own estimands and statistical units rather than being pooled into one omnibus claim. Input permutations are repeated transformations within a trained refit and are never treated as independent biological replicates.

这段在干什么:在协议收尾处规定统计口径——主结局、配对方式、多重比较校正,并划清各类指标的统计单元边界。

需要解释的地方:

  • Best@10 / 前沿 AUC:前者是发现质量主指标,后者衡量搜寻效率。
  • 配对区间:同一轨迹种子下做对比,再给双侧区间。
  • Holm 校正:领域常识,多重检验时控制假阳性、比 Bonferroni 稍宽松。
  • 统计单元:每类指标各自独立计数,不混成一个总论断。

值得留意:Holm 只压三个预设政策对比,不是全表;输入置换是"同一次重训内的重复变换",绝不算独立生物学重复——这是防伪重复的关键防线。

b114Comparator Best@10 pHolmp_{\mathrm{Holm}} Frontier AUC pHolmp_{\mathrm{Holm}} BBBC036 AIDE 1.9×10−31.9\times 10^{-3} 1.231×10−11.231\times 10^{-1} CellForge 1×10−41\times 10^{-4} 2.9×10−32.9\times 10^{-3} HarmonyCell 8.06×10−28.06\times 10^{-2} 1.962×10−11.962\times 10^{-1} BBBC047 AIDE 4.9×10−34.9\times 10^{-3} 9.5×10−39.5\times 10^{-3} CellForge 1.81×10−21.81\times 10^{-2} 2.17×10−22.17\times 10^{-2} HarmonyCell 7.56×10−17.56\times 10^{-1} 1.293×10−11.293\times 10^{-1}

这段在干什么:这是附录里的一张"统计显著性结果表",报告 AIDE、CellForge、HarmonyCell 三个模型在 BBBC036、BBBC047 两个数据集上,Best@10 与 Frontier AUC 两项指标对比最优基线后的 Holm 校正 p 值。

需要解释的地方:

  • "Best@10 / Frontier AUC"是两种排序/性能评价指标。
  • "$p_{\mathrm{Holm}}$"指用 Holm 方法做多重比较校正后的 p 值——这是统计领域常识:同时检验多个假设时,为避免假阳性膨胀而对 p 值阈值收紧。

值得留意:表里数字越小代表差异越显著,但这段没说明以 0.05 为阈值、也没说方向(谁赢),需回正文对照。

Appendix C Search trajectories, reliability, and cost

b116Table 19 summarizes the competitive discovery setting, and Tables 20 and 21 give the complete trajectory and selected-predictor comparisons. Fixed-model references include the pre-tuned RealMLP (Holzmüller et al., 2024), TabM’s parameter-efficient ensemble (Gorishniy et al., 2025), and a joint-output adaptation of TabR (Gorishniy et al., 2024). Table 23 reports valid candidates, failures, repairs, success by budget, model-provider calls, tokens, GPU training time, and trajectory wall time; failed slots and provider attempts remain charged. The RBDE score provides a joint efficiency summary alongside its separate quality and cost columns.

讲解

1. 这段在干什么

这是附录里的表格导读段:说明候选表 19–21、23 各自装了什么,并交代对比基线从哪来。

2. 需要解释的地方

  • Fixed-model references:不做搜索、直接拿现成模型当参照的基线(RealMLP、TabM、TabR 都是已发表的表格/预测模型,属领域常识)。
  • RBDE score:作者自评指标,同时给质量、成本两列和一个综合效率分。
  • failed slots 仍计费:失败的候选和调模型的尝试照样算进成本。

3. 值得留意

成本核算把失败也计入,所以效率对比不是"只算成功者"——这点容易被读漏。

b117Search Best@3 ↑\uparrow Best@5 ↑\uparrow Best@10 ↑\uparrow AUC ↑\uparrow Valid ↑\uparrow Calls Tokens (K) BBBC036 CellScientist 0.2833 ±\pm 0.0075 0.2856 ±\pm 0.0052 0.2924 ±\pm 0.0037 0.2845 ±\pm 0.0046 0.990 ±\pm 0.023 10.50 ±\pm 1.59 51.7 ±\pm 6.4 AIDE 0.2765 ±\pm 0.0036 0.2804 ±\pm 0.0033 0.2840 ±\pm 0.0030 0.2791 ±\pm 0.0023 1.000 ±\pm 0.000 11.10 ±\pm 0.98 41.8 ±\pm 3.8 CellForge 0.2708 ±\pm 0.0016 0.2757 ±\pm 0.0025 0.2772 ±\pm 0.0027 0.2739 ±\pm 0.0015 0.910 ±\pm 0.092 16.30 ±\pm 2.39 44.2 ±\pm 7.9 HarmonyCell 0.2803 ±\pm 0.0037 0.2845 ±\pm 0.0045 0.2877 ±\pm 0.0038 0.2820 ±\pm 0.0036 0.950 ±\pm 0.051 11.80 ±\pm 1.46 38.8 ±\pm 4.7 BBBC047 CellScientist 0.3052 ±\pm 0.0127 0.3084 ±\pm 0.0127 0.3097 ±\pm 0.0119 0.3058 ±\pm 0.0111 0.940 ±\pm 0.077 12.50 ±\pm 3.20 59.2 ±\pm 10.4 AIDE 0.2875 ±\pm 0.0015 0.2882 ±\pm 0.0014 0.2883 ±\pm 0.0012 0.2878 ±\pm 0.0013 0.990 ±\pm 0.023 12.30 ±\pm 1.17 45.2 ±\pm 4.4 CellForge 0.2891 ±\pm 0.0015 0.2901 ±\pm 0.0010 0.2917 ±\pm 0.0011 0.2899 ±\pm 0.0008 0.940 ±\pm 0.090 12.90 ±\pm 2.75 36.1 ±\pm 7.4 HarmonyCell 0.2918 ±\pm 0.0043 0.2965 ±\pm 0.0035 0.3070 ±\pm 0.0083 0.2964 ±\pm 0.0034 0.950 ±\pm 0.051 13.20 ±\pm 1.99 41.6 ±\pm 5.8

这段在干什么:这是一张结果表,报告两个数据集(BBBC036、BBBC047)上四个系统(CellScientist、AIDE、CellForge、HarmonyCell)的检索命中率与成本指标,用来支撑上一句提到的「质量与成本并列的效率总结」。

需要解释的地方:

  • Best@3/5/10:领域常识,指前 3/5/10 个候选里是否命中正确答案的比例,越高越好。
  • AUC / Valid:AUC 是排序质量指标;Valid 是产出结果合法的比例。
  • Calls / Tokens:成本列,分别指模型调用次数和消耗的 token 数(K 表示千)。
  • ± 数字:多次重复实验的标准差,衡量稳定性。

值得留意:CellScientist 在两项 Best@k 和 AUC 上基本都最高,但 Calls 与 Tokens 并不总是最少——质量与成本存在取舍,这正是表格要并列展示的原因。Valid 一列差距不大,AIDE 甚至为 1.000。

b118Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.

这段在干什么:这是附录中表格的图例说明,告诉你正文字体样式代表什么排名。

需要解释的地方:

  • Bold/underline:粗体=每个任务/折中最好的预测均值,下划线=第二好的。这是论文里表格排版的常见约定(领域常识)。
  • Distinct:不同方法若结果相同,算同一个名次,不重复占位。
  • Ties share a rank:并列就共享名次。
  • Unranked:区间上下界、诊断/资源类列不参与排名,不加粗下划线。

值得留意:上一段末尾那串带 ± 的数字就是被排名的候选值——± 后是误差范围,但按这条规则,误差本身不决定名次,只有均值参与比较。

b119Method Best@3 PCC ↑\uparrow Best@5 ↑\uparrow Best@10 ↑\uparrow Best MSE ↓\downarrow CP PCC ↑\uparrow L1000 PCC ↑\uparrow Frontier AUC ↑\uparrow Lift AUC ↑\uparrow BBBC036 CellScientist 0.2833 ±\pm 0.0075 0.2856 ±\pm 0.0052 0.2924 ±\pm 0.0037 0.0607 ±\pm 0.0001 0.2824 ±\pm 0.0015 0.2965 ±\pm 0.0052 0.2845 ±\pm 0.0046 0.0172 ±\pm 0.0046 AIDE 0.2765 ±\pm 0.0036 0.2804 ±\pm 0.0033 0.2840 ±\pm 0.0030 0.0612 ±\pm 0.0001 0.2762 ±\pm 0.0028 0.2875 ±\pm 0.0034 0.2791 ±\pm 0.0023 0.0118 ±\pm 0.0028 CellForge 0.2708 ±\pm 0.0016 0.2757 ±\pm 0.0025 0.2772 ±\pm 0.0027 0.0617 ±\pm 0.0003 0.2687 ±\pm 0.0038 0.2810 ±\pm 0.0032 0.2739 ±\pm 0.0015 0.0066 ±\pm 0.0020 HarmonyCell 0.2803 ±\pm 0.0037 0.2845 ±\pm 0.0045 0.2877 ±\pm 0.0038 0.0609 ±\pm 0.0002 0.2821 ±\pm 0.0034 0.2905 ±\pm 0.0045 0.2820 ±\pm 0.0036 0.0147 ±\pm 0.0038 BBBC047 CellScientist 0.3052 ±\pm 0.0127 0.3084 ±\pm 0.0127 0.3097 ±\pm 0.0119 0.0451 ±\pm 0.0004 0.3351 ±\pm 0.0048 0.2806 ±\pm 0.0212 0.3058 ±\pm 0.0111 0.0201 ±\pm 0.0108 AIDE 0.2875 ±\pm 0.0015 0.2882 ±\pm 0.0014 0.2883 ±\pm 0.0012 0.0458 ±\pm 0.0000 0.3269 ±\pm 0.0028 0.2433 ±\pm 0.0025 0.2878 ±\pm 0.0013 0.0021 ±\pm 0.0014 CellForge 0.2891 ±\pm 0.0015 0.2901 ±\pm 0.0010 0.2917 ±\pm 0.0011 0.0457 ±\pm 0.0001 0.3315 ±\pm 0.0040 0.2449 ±\pm 0.0049 0.2899 ±\pm 0.0008 0.0042 ±\pm 0.0016 HarmonyCell 0.2918 ±\pm 0.0043 0.2965 ±\pm 0.0035 0.3070 ±\pm 0.0083 0.0452 ±\pm 0.0003 0.3373 ±\pm 0.0041 0.2724 ±\pm 0.0151 0.2964 ±\pm 0.0034 0.0107 ±\pm 0.0039

这张表是消融/对比结果:在 BBBC036 和 BBBC047 两个任务上,横向比较四种方法、纵向看八项指标。

需要解释的地方

  • Best@k:从 k 个候选里取最好那个的分数(领域常识,表里没定义),k 越大通常越高,所以列间不可直接横比。
  • PCC:皮尔逊相关系数,越高越好;MSE 越低越好(↑↓ 已标出)。
  • ± 后面是标准差,反映多次运行的波动。

值得留意

  • 全部数字都只在这两个 cell 任务上,别推广到其他数据集。
  • CellScientist 在多数列占优,但 BBBC047 的 CP PCC 一列 HarmonyCell(0.3373)反而更高,别只看前排。
  • 标准差有时接近甚至超过方法间差距(如 BBBC047 Best@3),差异是否显著这段没说。

b120Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.

这段在干什么

这是表格的排版图例,告诉读者表里加粗、下划线、排名是怎么定的——不是新发现,是读表说明书。

需要解释的地方

  • predictive mean:预测均值,每个任务/折的平均预测结果。
  • distinct:先去掉重复值再排,避免并列把名次挤掉。
  • Ties share a rank:数值相同就并列同名次。
  • interval bounds / diagnostic / resource columns:置信区间上下界、诊断类、资源消耗类列,不参与排名。

值得留意

排名是逐个任务或折内部比的,不是全表统一排;加粗=最优,下划线=次优。

b121Model Global PCC ↑\uparrow MSE ↓\downarrow CP PCC ↑\uparrow L1000 PCC ↑\uparrow Parameters Train (s) Selected epoch BBBC036 CellScientist 0.3101 ±\pm 0.0023 0.0586 ±\pm 0.0001 0.3107 ±\pm 0.0025 0.3098 ±\pm 0.0032 4.00M 7.2 ±\pm 1.2 21.8 ±\pm 17.0 AIDE 0.2984 ±\pm 0.0091 0.0593 ±\pm 0.0006 0.3085 ±\pm 0.0043 0.2950 ±\pm 0.0124 4.42M 6.3 ±\pm 1.2 4.8 ±\pm 1.0 CellForge 0.2942 ±\pm 0.0043 0.0593 ±\pm 0.0002 0.3052 ±\pm 0.0048 0.2902 ±\pm 0.0056 3.82M 6.2 ±\pm 1.3 2.0 ±\pm 0.0 HarmonyCell 0.3116 ±\pm 0.0013 0.0586 ±\pm 0.0001 0.3117 ±\pm 0.0024 0.3115 ±\pm 0.0020 2.06M 6.7 ±\pm 1.4 17.2 ±\pm 12.5 h0h_{0} 0.2894 ±\pm 0.0081 0.0597 ±\pm 0.0003 0.2984 ±\pm 0.0048 0.2860 ±\pm 0.0098 1.14M 5.9 ±\pm 1.2 3.2 ±\pm 0.6 Standard MLP 0.2908 ±\pm 0.0114 0.0594 ±\pm 0.0004 0.3058 ±\pm 0.0089 0.2845 ±\pm 0.0125 2.42M 5.8 ±\pm 1.2 1.8 ±\pm 0.6 RealMLP 0.2983 ±\pm 0.0202 0.0592 ±\pm 0.0007 0.3170 ±\pm 0.0129 0.2902 ±\pm 0.0260 3.47M 5.9 ±\pm 0.7 – TabM 0.3053 ±\pm 0.0039 0.0589 ±\pm 0.0001 0.3175 ±\pm 0.0063 0.3004 ±\pm 0.0044 2.37M 11.4 ±\pm 1.7 2.8 ±\pm 1.0 Joint TabR 0.2732 ±\pm 0.0033 0.0604 ±\pm 0.0001 0.2835 ±\pm 0.0036 0.2687 ±\pm 0.0044 2.17M 2.6 ±\pm 0.3 7.6 ±\pm 1.1 Ridge 0.0777 ±\pm 0.0000 0.2048 ±\pm 0.0000 0.0812 ±\pm 0.0000 0.0763 ±\pm 0.0000 4.14M 1.5 ±\pm 0.4 – BBBC047 CellScientist 0.3033 ±\pm 0.0011 0.0457 ±\pm 0.0000 0.3254 ±\pm 0.0010 0.2787 ±\pm 0.0015 1.78M 24.3 ±\pm 3.7 36.6 ±\pm 10.5 AIDE 0.2696 ±\pm 0.0047 0.0468 ±\pm 0.0002 0.3167 ±\pm 0.0062 0.2125 ±\pm 0.0101 1.20M 9.5 ±\pm 0.7 2.0 ±\pm 0.9 CellForge 0.2740 ±\pm 0.0035 0.0467 ±\pm 0.0001 0.3178 ±\pm 0.0061 0.2216 ±\pm 0.0060 2.53M 9.6 ±\pm 0.8 1.8 ±\pm 1.0 HarmonyCell 0.3021 ±\pm 0.0098 0.0458 ±\pm 0.0003 0.3265 ±\pm 0.0016 0.2743 ±\pm 0.0196 2.28M 26.0 ±\pm 9.2 41.8 ±\pm 27.1 h0h_{0} 0.2696 ±\pm 0.0047 0.0468 ±\pm 0.0002 0.3167 ±\pm 0.0062 0.2125 ±\pm 0.0101 1.20M 9.3 ±\pm 0.9 2.0 ±\pm 0.9 Standard MLP 0.2696 ±\pm 0.0059 0.0468 ±\pm 0.0002 0.3167 ±\pm 0.0062 0.2127 ±\pm 0.0067 2.53M 9.6 ±\pm 0.9 1.2 ±\pm 0.6 RealMLP 0.2831 ±\pm 0.0042 0.0463 ±\pm 0.0001 0.3259 ±\pm 0.0038 0.2326 ±\pm 0.0049 3.62M 13.7 ±\pm 0.8 – TabM 0.2852 ±\pm 0.0033 0.0463 ±\pm 0.0001 0.3282 ±\pm 0.0023 0.2350 ±\pm 0.0105 2.51M 16.7 ±\pm 2.1 2.6 ±\pm 0.7 Joint TabR 0.2533 ±\pm 0.0049 0.0475 ±\pm 0.0002 0.2951 ±\pm 0.0050 0.2027 ±\pm 0.0066 2.25M 10.8 ±\pm 1.1 4.6 ±\pm 0.7 Ridge 0.0492 ±\pm 0.0000 0.3880 ±\pm 0.0000 0.0547 ±\pm 0.0000 0.0438 ±\pm 0.0000 4.62M 1.3 ±\pm 0.2 –

这段在干什么:这是附录里的两张性能对比表(BBBC036 和 BBBC047),列出各模型在四个指标、参数量、训练时长上的均值±标准差,用来支撑正文中"搜索轨迹的可靠性与开销"这一部分。

需要解释的地方:

  • PCC:皮尔逊相关系数,衡量预测与真实值的线性相关,越高越好;MSE:均方误差,越低越好。
  • ± 后面是多次运行的标准差,反映稳定性。
  • h₀:通常指无学习的基线/初始模型(领域常见做法,此处论文未解释)。
  • Selected epoch:被选中的训练轮数;"–"表示该项缺失。

值得留意:Ridge 的 PCC 极低(约 0.05–0.08)但极稳定;HarmonyCell 参数最少(2.06M)却在 BBBC036 上综合最好。标准差本身差异很大,别只看均值排名。

b122Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.

这段在干什么:这是结果表格的图例,用来说明表中加粗、下划线等格式的含义。

需要解释的地方:加粗=该任务/折内表现最好的预测均值,下划线=第二好的,且二者是"不同的"模型;"平局共享排名"意味着并列同格式;区间上下界和诊断/资源列不参与排名。

值得留意:排名只在同一任务或折内比较,不跨任务;资源列(如上一行的 4.62M、1.3±0.2)不会因高低被加粗或划线。

b123Formulation Generated artifact Max output Slots Prediction-score policies Executable source 9000 10 Path-constrained Structured design card 2400 10 Falsification-guided Semantic design card 1800 10

这段是一张表格式汇总,列出四种轨迹生成策略(Formulation / Executable source / Path-constrained / Falsification-guided)各自产出的 artifact 类型、最大输出和 slots 数。

要解释的地方:这张表本身没有表头说明,各列对应关系是靠对齐推测的——比如「9000/2400/1800」「10/10/10」应是两列数值。四种策略与其 artifact 的配对也需靠上下文确认,原文没明说。值得留意的是,这段明显是表格被抽成了纯文本,列对齐已丢失,读时别把数字和策略错配。原文没交代这些数值的含义(是 token 上限?还是别的),这里没提。

b124Shared settings: model deepseek-v4-flash; temperature 0.2; reasoning effort not set; fallback disabled. Max output is the token limit per request.

这段是附录 C 里的共享实验参数声明,交代各条搜索轨迹共用的模型配置,属于复现所需的背景设定。

需要解释的地方:temperature 0.2 指采样随机性低,输出偏稳定;fallback disabled 指模型调用失败时不做备用切换;reasoning effort 是推理投入档位,此处「未设置」。这些是领域常识。

值得留意:上一段的表格在展示不同搜索策略的 token 与步数消耗,而这段把模型固定住了——说明各轨迹之间的差异来自策略本身,而非换了模型。

b125A. Execution quality and budget efficiency

讲解

1. 这段在干什么:这是附录 C 下的一个小标题(A 节),用来引出"执行质量与预算效率"这一子话题,起过渡/分节作用。

2. 需要解释的地方:

  • Execution quality:指 agent 跑搜索任务时结果的可靠程度(领域常识)。
  • Budget efficiency:指在给定成本/调用预算下能跑多少、效果如何(领域常识)。
  • 注意原文只给了标题,具体内容这段没提。

3. 值得留意:这只是一个标题行,正文尚未展开;它与前段(模型配置、token 上限)的衔接尚未在此段说明,别把它当成结论。

b126Method RBDE ↑\uparrow Valid rate ↑\uparrow Failed slots ↓\downarrow Repairs ↓\downarrow Success@3/5/10 ↑\uparrow BBBC036 CellScientist 0.0104 0.990 ±\pm 0.023 0.10 ±\pm 0.23 1.50 ±\pm 1.59 8/10/10 AIDE 0.0096 1.000 ±\pm 0.000 0.00 ±\pm 0.00 2.10 ±\pm 0.98 9/10/10 CellForge 0.0044 0.910 ±\pm 0.092 0.90 ±\pm 0.92 7.30 ±\pm 2.39 8/10/10 HarmonyCell 0.0102 0.950 ±\pm 0.051 0.50 ±\pm 0.51 2.80 ±\pm 1.46 9/10/10 BBBC047 CellScientist 0.0064 0.940 ±\pm 0.077 0.60 ±\pm 0.77 3.50 ±\pm 3.20 7/8/9 AIDE 0.0008 0.990 ±\pm 0.023 0.10 ±\pm 0.23 3.30 ±\pm 1.17 5/8/9 CellForge 0.0027 0.940 ±\pm 0.090 0.60 ±\pm 0.90 3.90 ±\pm 2.75 8/9/9 HarmonyCell 0.0057 0.950 ±\pm 0.051 0.50 ±\pm 0.51 4.20 ±\pm 1.99 8/10/10

这段在干什么:这是一张结果表,横向对比 CellScientist、AIDE、CellForge、HarmonyCell 四种方法在 BBBC036 和 BBBC047 两个数据集上的执行质量与预算效率。

需要解释的地方:

  • RBDE:核心指标,越高越好(↑),衡量修复预算的利用效率。
  • Valid rate:有效槽位比例;Failed slots / Repairs:失败数与修复次数,越低越好(↓)。
  • Success@3/5/10:前 3/5/10 次尝试内成功的次数比。
  • 数字后 ± 是标准差,表示多次运行的波动。

值得留意:AIDE 在 BBBC036 上 Valid rate 为 1.000、Failed slots 为 0,但 RBDE 仅 0.0096,并非最高——说明"运行干净"和"效率高"是两回事。Success@k 跨越 5 到 10 时多有提升,暗示后段尝试仍能救回结果。

b128Method Calls Tokens (K) GPU training (s) Wall time (s) BBBC036 CellScientist 10.50 ±\pm 1.59 51.7 ±\pm 6.4 19.8 ±\pm 2.2 716.3 ±\pm 101.3 AIDE 11.10 ±\pm 0.98 41.8 ±\pm 3.8 16.6 ±\pm 2.5 559.0 ±\pm 212.1 CellForge 16.30 ±\pm 2.39 44.2 ±\pm 7.9 15.7 ±\pm 1.4 578.0 ±\pm 86.4 HarmonyCell 11.80 ±\pm 1.46 38.8 ±\pm 4.7 16.8 ±\pm 1.1 610.8 ±\pm 158.2 BBBC047 CellScientist 12.50 ±\pm 3.20 59.2 ±\pm 10.4 111.7 ±\pm 21.3 821.9 ±\pm 116.0 AIDE 12.30 ±\pm 1.17 45.2 ±\pm 4.4 60.3 ±\pm 3.9 533.0 ±\pm 152.3 CellForge 12.90 ±\pm 2.75 36.1 ±\pm 7.4 56.7 ±\pm 5.8 533.0 ±\pm 83.6 HarmonyCell 13.20 ±\pm 1.99 41.6 ±\pm 5.8 72.9 ±\pm 11.0 680.5 ±\pm 228.0

这份表格报告了四种方法在两个数据集上的搜索开销对比。

1. 这段在干什么:用表格列出各方法在两个数据集(BBBC036、BBBC047)上的运行成本,用于比较搜索效率与可靠性。

2. 需要解释的地方:

  • Method Calls:方法调用次数,即搜索过程共执行了多少步。
  • Tokens (K):消耗的 token 数(单位千)。
  • GPU training / Wall time:模型训练耗时 / 端到端总耗时(秒)。
  • ± 后是标准差,表示多次运行的波动。

3. 值得留意:CellScientist 在 BBBC047 上 GPU 训练时间(111.7s)远高于其他方法,说明其搜索出的模型训练代价更大。

b129CP response block L1000 response block Model PCC @20 ↑\uparrow PCC @50 ↑\uparrow RMSE @20 ↓\downarrow RMSE @50 ↓\downarrow PCC @20 ↑\uparrow PCC @50 ↑\uparrow RMSE @20 ↓\downarrow RMSE @50 ↓\downarrow BBBC036 CellScientist 0.4620 ±\pm 0.0035 0.4454 ±\pm 0.0033 0.2411 ±\pm 0.0005 0.2385 ±\pm 0.0004 0.2828 ±\pm 0.0051 0.3167 ±\pm 0.0043 0.5135 ±\pm 0.0008 0.4497 ±\pm 0.0007 AIDE 0.4636 ±\pm 0.0024 0.4445 ±\pm 0.0025 0.2408 ±\pm 0.0006 0.2386 ±\pm 0.0004 0.2611 ±\pm 0.0200 0.3005 ±\pm 0.0160 0.5190 ±\pm 0.0054 0.4540 ±\pm 0.0048 CellForge 0.4607 ±\pm 0.0050 0.4430 ±\pm 0.0034 0.2412 ±\pm 0.0006 0.2388 ±\pm 0.0005 0.2598 ±\pm 0.0081 0.2966 ±\pm 0.0099 0.5176 ±\pm 0.0011 0.4531 ±\pm 0.0013 HarmonyCell 0.4633 ±\pm 0.0023 0.4464 ±\pm 0.0024 0.2409 ±\pm 0.0005 0.2383 ±\pm 0.0004 0.2823 ±\pm 0.0040 0.3170 ±\pm 0.0031 0.5137 ±\pm 0.0006 0.4497 ±\pm 0.0004 h0h_{0} 0.4526 ±\pm 0.0034 0.4357 ±\pm 0.0032 0.2426 ±\pm 0.0005 0.2399 ±\pm 0.0005 0.2555 ±\pm 0.0079 0.2983 ±\pm 0.0034 0.5210 ±\pm 0.0056 0.4545 ±\pm 0.0031 Standard MLP 0.4608 ±\pm 0.0077 0.4404 ±\pm 0.0070 0.2412 ±\pm 0.0011 0.2391 ±\pm 0.0009 0.2473 ±\pm 0.0048 0.2923 ±\pm 0.0111 0.5197 ±\pm 0.0006 0.4536 ±\pm 0.0017 RealMLP 0.4668 ±\pm 0.0081 0.4497 ±\pm 0.0080 0.2402 ±\pm 0.0013 0.2378 ±\pm 0.0011 0.2476 ±\pm 0.0332 0.2945 ±\pm 0.0320 0.5200 ±\pm 0.0053 0.4540 ±\pm 0.0047 TabM 0.4640 ±\pm 0.0023 0.4477 ±\pm 0.0021 0.2408 ±\pm 0.0005 0.2382 ±\pm 0.0003 0.2603 ±\pm 0.0079 0.3066 ±\pm 0.0033 0.5180 ±\pm 0.0031 0.4518 ±\pm 0.0016 Joint TabR 0.4433 ±\pm 0.0041 0.4235 ±\pm 0.0045 0.2437 ±\pm 0.0005 0.2414 ±\pm 0.0006 0.2389 ±\pm 0.0124 0.2812 ±\pm 0.0120 0.5231 ±\pm 0.0024 0.4567 ±\pm 0.0020 Ridge 0.1172 ±\pm 0.0000 0.1153 ±\pm 0.0000 0.4678 ±\pm 0.0000 0.4603 ±\pm 0.0000 0.0807 ±\pm 0.0000 0.0903 ±\pm 0.0000 0.9852 ±\pm 0.0000 0.8401 ±\pm 0.0000 BBBC047 CellScientist 0.4003 ±\pm 0.0028 0.4061 ±\pm 0.0018 0.2689 ±\pm 0.0002 0.2547 ±\pm 0.0002 0.3262 ±\pm 0.0025 0.3154 ±\pm 0.0019 0.3703 ±\pm 0.0004 0.3250 ±\pm 0.0003 AIDE 0.4127 ±\pm 0.0124 0.4078 ±\pm 0.0083 0.2676 ±\pm 0.0022 0.2549 ±\pm 0.0015 0.2577 ±\pm 0.0090 0.2486 ±\pm 0.0060 0.3789 ±\pm 0.0009 0.3323 ±\pm 0.0006 CellForge 0.4117 ±\pm 0.0095 0.4069 ±\pm 0.0032 0.2678 ±\pm 0.0010 0.2550 ±\pm 0.0007 0.2684 ±\pm 0.0051 0.2583 ±\pm 0.0041 0.3776 ±\pm 0.0009 0.3312 ±\pm 0.0007 HarmonyCell 0.4029 ±\pm 0.0049 0.4079 ±\pm 0.0026 0.2686 ±\pm 0.0009 0.2545 ±\pm 0.0004 0.3234 ±\pm 0.0156 0.3130 ±\pm 0.0148 0.3705 ±\pm 0.0021 0.3252 ±\pm 0.0016 h0h_{0} 0.4127 ±\pm 0.0124 0.4078 ±\pm 0.0083 0.2676 ±\pm 0.0022 0.2549 ±\pm 0.0015 0.2577 ±\pm 0.0090 0.2486 ±\pm 0.0060 0.3789 ±\pm 0.0009 0.3323 ±\pm 0.0006 Standard MLP 0.4125 ±\pm 0.0157 0.4083 ±\pm 0.0089 0.2678 ±\pm 0.0023 0.2548 ±\pm 0.0013 0.2626 ±\pm 0.0135 0.2531 ±\pm 0.0123 0.3782 ±\pm 0.0018 0.3316 ±\pm 0.0015 RealMLP 0.4142 ±\pm 0.0072 0.4137 ±\pm 0.0043 0.2667 ±\pm 0.0010 0.2536 ±\pm 0.0005 0.2883 ±\pm 0.0034 0.2798 ±\pm 0.0028 0.3750 ±\pm 0.0004 0.3287 ±\pm 0.0003 TabM 0.4152 ±\pm 0.0107 0.4142 ±\pm 0.0050 0.2666 ±\pm 0.0015 0.2536 ±\pm 0.0007 0.2784 ±\pm 0.0051 0.2685 ±\pm 0.0067 0.3763 ±\pm 0.0010 0.3300 ±\pm 0.0008 Joint TabR 0.3751 ±\pm 0.0078 0.3810 ±\pm 0.0061 0.2757 ±\pm 0.0007 0.2599 ±\pm 0.0006 0.2482 ±\pm 0.0101 0.2409 ±\pm 0.0084 0.3806 ±\pm 0.0012 0.3336 ±\pm 0.0009 Ridge 0.0833 ±\pm 0.0000 0.0808 ±\pm 0.0000 0.7616 ±\pm 0.0000 0.7251 ±\pm 0.0000 0.0583 ±\pm 0.0000 0.0583 ±\pm 0.0000 1.1990 ±\pm 0.0000 1.0275 ±\pm 0.0000

这张表是附录里的多数据集性能对比:BBBC036 与 BBBC047 两个数据集,各列 8 个指标——PCC@20/50、RMSE@20/50,前者在 CP 反应块、后者在 L1000 反应块下各算一遍。

关键概念:PCC 是预测与真实的相关性(越大越好,表头↑),RMSE 是误差(越小越好,表头↓)。"@20/@50"指只取前 20 或 50 个预测来算——这是领域常识里的 top-k 评估,不是本文定义。每个数字后的 ± 是多次运行的波动。

值得留意:Ridge 在 PCC 上远低于其余模型(如 0.1172),RMSE 却反常偏大,说明它基本没学到反应关系;而各 Cell 类模型之间差距很小。表中未见论文对此的具体解读。

b130Task Score PCC Path PCC Δ\Delta Tokens (K) Valid Repairs Linked BBBC036 0.2924 ±\pm 0.0037 0.2925 ±\pm 0.0019 +0.0001 21.9 ±\pm 1.4 100.0% 0.00 ±\pm 0.00 100/100 BBBC047 0.3097 ±\pm 0.0119 0.3008 ±\pm 0.0012 -0.0090 21.7 ±\pm 1.2 100.0% 0.00 ±\pm 0.00 100/100

这段是附录里的轨迹成本表,报告两个数据集在两种打分下的表现与代价。

1. 在干什么:用数据展示搜索轨迹的可靠性——任务分与路径分几乎一致,且100次全部有效、无需修复。

2. 解释:

  • Task/Path PCC:Pearson相关系数,衡量"结局"与"路径"打分的一致性,领域常识。
  • Δ:两者之差,越接近0越好。
  • Tokens(K):搜索消耗的token数(千计),衡量成本。
  • Valid/Repairs/Linked:有效比例、修复次数、成功链接的模型数。

3. 值得留意:BBBC047的Δ是负的(-0.0090),路径分反而略高于任务分,但作者未对此评论;两行均报100/100,说明全程零失败。

Appendix D Input-use claims in combinatorial genetic-perturbation prediction

b132Boundary Predictor Global PCC ↑\uparrow Input Δ\DeltaPCC ↑\uparrow Support Fold 4 audit GEARS 0.6819 [0.6714, 0.6924] 0.3029 5/5 Initial model (h0h_{0}) 0.8960 [0.8909, 0.9011] 0.5425 5/5 CellScientist 0.9020 [0.8988, 0.9051] 0.5578 5/5 Fold 5 replication GEARS 0.6676 [0.6568, 0.6784] 0.3599 5/5 Initial model (h0h_{0}) 0.9051 [0.9023, 0.9079] 0.6193 5/5 CellScientist 0.9096 [0.9051, 0.9142] 0.6280 5/5

这张表是在做跨折叠复现检验:把 GEARS、初始模型 h₀、CellScientist 三者在 Fold 4 和 Fold 5 上的成绩并排摆出来。

需要解释的地方:

  • PCC:预测与真实值的相关性,越高越好(这里是领域常识)。
  • ΔPCC:相比基线提升多少,看模型有没有真的加价值。
  • 方括号:置信区间。
  • Support 5/5:五次审计全部通过,是稳健性背书。

值得留意:两折中 h₀ 和 CellScientist 几乎持平(0.896 vs 0.902,0.905 vs 0.910),提升幅度很小——作者没明说,但这个"微弱优势"本身值得打个问号。

b133The transfer task is reconstructed from the count layer of the Norman K562 CRISPRa Perturb-seq experiment (GEO GSE133344) (Norman et al., 2019). Its source matrix contains 91,205 cells and 5,045 genes, including 7,353 control cells. Per-cell counts are library-size normalized to 10,000 and log-transformed, after which canonical condition means are centered by the control mean. The 284 source labels resolve to 237 canonical conditions: 105 observed singles, 131 doubles, and one control. The common discovery interface receives a pair-input bundle containing a 105-dimensional multi-hot perturbation identity and its matching mean single-perturbation anchor, and returns the full 5,045-gene response without external pathway features or cell-type covariates. GEARS retains its task-native graph inputs as described below.

讲解

1. 这段在干什么

交代迁移任务的来源与数据规模,并说明「通用发现接口」拿到什么输入、吐什么输出。

2. 需要解释的地方

  • Perturb-seq(领域常识):CRISPR 扰动 + 单细胞测序,读出每个细胞被扰动后的基因表达。
  • library-size 归一化 / log 变换(领域常识):消除测序深度差异、压缩数值范围的标准预处理。
  • 多热(multi-hot):105 维向量里,被扰动的基因位置标 1,其余为 0。
  • anchor(锚):配套提供的单扰动平均响应,给模型当参照。

3. 值得留意

  • 接口不含通路特征和细胞类型协变量,只有扰动身份和锚——是刻意的约束。
  • 284 个原始标签压成 237 个条件;GEARS 例外,仍用自带图输入。
  • 上一段那串数字与这段无关,这段没提到模型性能。

b134The five-fold assignment is constructed without expression responses and balances pair count, source-cell count, and perturbation-gene incidence. All singles and controls remain fit-only with the Fold-1/2 doubles; Fold 3 supplies discovery feedback, and Folds 4 and 5 are successive held-out evaluation folds. Table 27 lists every held-out double-perturbation block. The comparison predicts unseen double combinations whose constituent single-gene responses are available during fitting.

讲解

1. 这段在干什么

交代数据划分方式:五折怎么分、哪几折用来训练(fit-only)、哪折给发现反馈、哪几折做留出评估——为下文"预测未见过的双扰动组合"设定前提。

2. 需要解释的地方

  • 五折划分:常见做法(领域常识),把数据切成5份轮流当测试,但这里的折数被赋予了不同角色,不只是评测。
  • fit-only:只参与拟合,不参与评估。
  • held-out:留出、不参与训练,用于检验。
  • discovery feedback:用于发现(如筛选假设)的反馈信号,而非最终打分。
  • double-perturbation:同时扰动两个基因的组合。

3. 值得留意

  • 划分不依赖表达响应,且刻意平衡了配对数量、来源细胞数、扰动基因出现频率——是为公平比较做的设计。
  • 关键前提:待预测的双组合,其单基因响应在拟合时已知——这是任务难度设定,别读漏。
  • 表27列了全部留出双扰动块,原文未给具体内容。

b135Each held-out boundary contains 26 unseen double-perturbation conditions. Before Fold 4 is opened, we construct 32 distinct complete derangements of the 26 row positions; each is a one-to-one permutation with no fixed point and is reused positionally on Fold 5. For each source condition, the registered intervention replaces both coordinates of the pair-input bundle (its 105-dimensional pair-identity vector and matching mean single-perturbation anchor) with the donor pair’s values while retaining the source response target. Replacing both coordinates prevents a residual predictor from retaining pair information through the anchor; consequently, the test measures dependence on the complete bundle rather than attributing the effect to the multi-hot coordinate alone. The maps use neither responses nor model outputs and do not require donor pairs to be gene-disjoint from source pairs.

这段在干什么

交代审计实验的操作细节:说清如何用错位置换构造“供体”干预,把双扰动测试从单扰动锚点里剥离出来。

需要解释的地方

  • *derangement(错排)*:排列的一种领域常识,指每个元素都不在自己原位、且一一对应的置换;这里“无固定点”即为此。
  • *两个坐标*:配对输入包含两样东西——105维的配对身份向量,和与之匹配的“平均单扰动锚点”。
  • *donor pair(供体对)*:被借来提供数值的那一对条件。
  • 做法是:把这两个坐标都换成供体的值,但保留原条件的响应目标。

值得留意

  • 把两个坐标一起换,是为防止预测器从锚点里偷偷留住配对信息;所以测的是对“整个打包输入”的依赖,而不只是多热那一位。
  • 映射不靠响应、不靠模型输出,也不要求供体与来源基因不重叠——都是这段明说的。
  • 32个错排先在Fold 4用,再原样按位置复用到Fold 5——这段没解释为何复用。

b136For refit jj and map rr, let LjL_{j} be joint-output MSE under the correct pair-input bundle and Lj,rshufL^{\mathrm{shuf}}_{j,r} the MSE after the bundle shuffle. We define gj,r=Lj,rshuf−Ljg_{j,r}=L^{\mathrm{shuf}}_{j,r}-L_{j} and g~j,r=gj,r/max⁡(Lj,10−12)\tilde{g}_{j,r}=g_{j,r}/\max(L_{j},10^{-12}). A refit is qualified under the pair-input-use criterion when 32−1​∑rg~j,r>0.00132^{-1}\sum_{r}\tilde{g}_{j,r}>0.001 and Q0.05​({gj,r}r=132)>0Q_{0.05}(\{g_{j,r}\}_{r=1}^{32})>0. It is unsupported when the first quantity lies within [−0.001,0.001][-0.001,0.001] and inconclusive otherwise. The Global-PCC drop is reported as a continuous effect but does not enter this categorical rule. The 32 maps are repeated interventions within a refit; the five refits are the statistical units for model-performance intervals and support counts.

讲解

1. 这段在干什么

给出「配对输入使用」判据的量化定义:用打乱 bundle 后的 MSE 变化,判断模型是否真依赖该输入。

2. 需要解释的地方

  • $g_{j,r}$:打乱前后 MSE 之差,即"打乱让性能掉了多少"。
  • $\max(L_j,10^{-12})$:归一化分母,$10^{-12}$ 是防除零的小常数。
  • $Q_{0.05}$:32 个 $g$ 的 5% 分位数(领域常识:分位数是排序后的位置值)。
  • 三条结论:均值比 >0.001 且有分位支撑 → 合格;落在 ±0.001 → 不支持;否则不确定。

3. 值得留意

32 个 map 只是 refit 内的重复,真正的统计单位是 5 个 refit——这影响后面区间的可信度,别把 32 当独立样本。

b137GEARS is evaluated through the shared task boundary while retaining its original multigene-perturbation formulation (Roohani et al., 2024): GO and coexpression graphs, native GNN/decoder, optimization objective, and validation-selected checkpoint. It shares count normalization, response target, fold roles, and held-out metrics with the other predictors. Table 28 records the architecture, graph scope, optimizer, and checkpoint rule used in the comparison.

这段在干什么

交代 GEARS 作为对比模型的评测口径:保留其原有多基因扰动设定,只共享任务边界。

需要解释的地方

  • 共享任务边界:即各家模型输入输出、划分和指标对齐,便于公平比较(领域常识)。
  • native GNN/decoder、优化目标、checkpoint:都沿用 GEARS 原版,未做改动。
  • held-out metrics:在留出集上算的指标。

值得留意

作者强调"保留原formulation"与"共享fold、指标",是想说明差异只来自方法本身。具体数值不在本段,见 Table 28。

b138Fold Role Double perturbations Source cells Component genes 1 Fit 26 7,301 40 2 Fit 26 7,849 41 3 Discovery/selection 27 6,817 42 4 Independent audit 26 6,541 40 5 Final replication 26 6,937 38

讲解

这段在干什么:用 Table 28 列出五折交叉验证中每一折的角色与数据规模,说明评估协议的划分。

需要解释的地方:

  • Fold(折):把数据切成几份轮流做验证,领域常识。
  • Role(角色):各折承担的任务不同——1、2 折用于拟合(Fit),3 折做发现/筛选,4 折做独立审核,5 折做最终复现。这解释了为何同一份数据要分角色,而非简单对半分。
  • Double perturbations / Source cells / Component genes:三列数据规模,分别指双基因扰动数、来源细胞数、组件基因数。

值得留意:审核与复现用的折完全不参与拟合和筛选,这是保证"独立"的关键;数据量在不同折间略有波动,并非均分。

b139Across ten trajectories, all 77 nonduplicate CellScientist candidates pass source, shape, and training checks; the 23 duplicate design cards still consume their candidate slots, and no Fold-4 or Fold-5 response enters proposal or selection. The trajectories account for 154 provider calls, 403,762 tokens, 17.8 H100 fit-minutes, and 50.5 minutes of end-to-end time. Held-out input-use testing is deterministic and uses zero model-provider calls.

这段在干什么

承接上一段的逐轨迹统计表,汇总十条轨迹的算力/调用开销,并交代候选卡的去重与检查结果。

需要解释的地方

  • source / shape / training checks:领域常识——指代码来源、数据形状、训练能否跑通三类自动检查。
  • duplicate design cards:重复设计卡;虽重复,仍占用候选名额,不算浪费。
  • Fold-4/5 response:交叉验证的折;这里强调这两折数据从未进入提案或筛选,即无泄漏。

值得留意

  • "零 provider 调用"完成 held-out 测试,说明审计环节不依赖外部模型,可复现。
  • 重复卡"仍消耗名额"是主动披露的损耗,不是忽略。

(约 140 字)

b140A. Search performance BB Best@3 Best@5 Best@10 Frontier AUC Δ​h0\Delta h_{0} [95% CI] 10 0.9192 0.9201 0.9210 0.9194 +0.0062 [+0.0036, +0.0087]

这段是附录 D 的结果表(表格被压成一行),汇报组合遗传扰动预测中「搜索性能」的指标。

1. 在干什么:给出不同预算下 Best@k 与 Frontier AUC 等数值,并报告 Δh₀ 及其 95% 置信区间。

2. 概念:Best@3/5/10 指前 3/5/10 名里命中真值的表现(领域常识,具体定义这段没提);AUC 是排序质量指标;Δh₀ 是某种效应量,正值表示提升,方括号是置信区间。

3. 值得留意:Δh₀ 的 CI 下限 +0.0036 > 0,即提升在统计上站得住,不是噪声;但各项 Best@k 数值几乎持平,别误读成有明显梯度。

b141B. Completion and resources Positive Executed Calls Tokens 10/10 77/100 154 403,762

这段在干什么:这是附录里的资源统计行,汇报跑完这套流程(Completion)的调用次数与消耗,紧接上一段模型指标(Best@k、AUC)之后。

需要解释的地方:Positive Executed Calls 10/10 指十次调用全部成功(领域常识:这类"成功数/总数"常用来表示执行可靠性);Tokens 77/100 指 token 用量七十七,上限或预算一百(具体含义这段没说明)。

值得留意:它只有指标名和数字,没有单位和说明,77/100 到底是用到上限还是达到某阈值,单看这行读不出来。

b142Predictor Global PCC ↑\uparrow MSE ↓\downarrow Condition PCC ↑\uparrow Rank corr. ↑\uparrow DEG-20 PCC ↑\uparrow DEG-50 PCC ↑\uparrow Full-state PCC ↑\uparrow Input-use support Fold 4 audit Matching mean 0.8810 0.0059 0.8966 0.4902 0.9702 0.9634 0.9915 supported Additive singles 0.8810 0.0046 0.8966 0.4902 0.9702 0.9634 0.9932 supported Ridge 0.8989 0.0033 0.8965 0.4584 0.9724 0.9665 0.9951 supported GEARS 0.6819 0.0122 0.6826 0.2512 0.9078 0.8915 0.9843 5/5 h0h_{0} 0.8960 0.0037 0.8999 0.4288 0.9664 0.9618 0.9945 5/5 CellScientist 0.9020 0.0039 0.9098 0.4345 0.9750 0.9697 0.9943 5/5 Fold 5 replication Matching mean 0.8982 0.0043 0.8799 0.4681 0.9614 0.9469 0.9938 supported Additive singles 0.8982 0.0033 0.8799 0.4681 0.9614 0.9469 0.9952 supported Ridge 0.8959 0.0029 0.8813 0.4375 0.9672 0.9546 0.9957 supported GEARS 0.6676 0.0115 0.6156 0.2267 0.8021 0.8096 0.9853 5/5 h0h_{0} 0.9051 0.0027 0.8844 0.4065 0.9542 0.9474 0.9961 5/5 CellScientist 0.9096 0.0028 0.8944 0.4071 0.9650 0.9554 0.9960 5/5

这段在干什么

在附录里汇报「输入使用声明」的审计结果:把各预测器在 Fold 4 与 Fold 5 上的多项指标(PCC、MSE、DEG 等)列出,并给出支持与否的判定。

需要解释的地方

  • PCC:预测值与真实值的相关性,越高越好(↑);MSE:均方误差,越低越好(↓)。
  • DEG-20/50、Full-state PCC:指在不同基因子集上算的 PCC(领域常识:DEG 即差异表达基因)。
  • Input-use support:审计结论,判断模型是否真用到了应有输入。

值得留意

GEARS 两折都拿到「5/5」支持,但它的 Global PCC、MSE、Rank corr. 在各行里最差,说明支持认定并不等于整体预测最优。

b143The matched transfer analysis gives CellScientist, AIDE, CellForge, and HarmonyCell the same initial model, design language, Fold-3 feedback, physical budget, and model-selection rule; each policy generates its own candidates. CellScientist has the highest mean Best@10 and frontier AUC. Its paired Best@10 effect is +0.0029+0.0029 versus AIDE (95% CI [−0.0003,+0.0061][-0.0003,+0.0061]), +0.0058+0.0058 versus CellForge ([+0.0037,+0.0080][+0.0037,+0.0080]), and +0.0006+0.0006 versus HarmonyCell ([−0.0005,+0.0016][-0.0005,+0.0016]). After selection, all sources are refit under five paired seeds and audited without provider calls. Global-PCC ordering varies across folds, while all four selected models pass the pair-input-bundle test in 5/5 refits; CellScientist leads condition-, DEG-20-, and DEG-50-PCC at Fold 5.

这段接在上表之后,做配对效应量的文字总结。

在干什么:交代四个智能体用完全相同的初始模型、预算等条件,只放大差异在“各自生成候选”,然后比出 CellScientist 最强。

关键概念:Best@10 是取前 10 个候选里最好的那个分数;frontier AUC 衡量性能与预算权衡曲线下的面积;“配对”指同一初始条件下两两比较,比独立比更公平。95% CI 不含 0 才算显著优于对方。

值得留意:对 AIDE(区间含 −0.0003)和对 HarmonyCell(含 −0.0005)的区间都跨 0,即优势不稳健;只有对 CellForge 明显领先。后半段“无 provider 调用重训 5 次”的审计,四个模型全部 5/5 通过。

b144Policy Best@3 ↑\uparrow Best@5 ↑\uparrow Best@10 ↑\uparrow Frontier AUC ↑\uparrow Δ​h0\Delta h_{0} [95% CI] ↑\uparrow Positive ↑\uparrow Executed slots ↑\uparrow Calls ↓\downarrow Tokens ↓\downarrow CellScientist 0.9192 0.9201 0.9210 0.9194 +0.0062 [+0.0036,+0.0087] 10/10 77/100 154 403,762 AIDE 0.9164 0.9168 0.9181 0.9170 +0.0033 [+0.0011,+0.0054] 9/10 91/100 120 335,610 CellForge 0.9151 0.9152 0.9152 0.9151 +0.0003 [−-0.0004,+0.0010] 1/10 97/100 108 309,742 HarmonyCell 0.9191 0.9203 0.9204 0.9193 +0.0056 [+0.0028,+0.0084] 9/10 88/100 117 328,705

这段在干什么

这是一张性能对比表,横向比四种方法(CellScientist、AIDE、CellForge、HarmonyCell)在组合基因扰动预测任务上的表现,是上一段文字结论的数据支撑。

需要解释的地方

  • Best@k:预测里前 k 个最优候选的命中率,越大越好(↑)。
  • Frontier AUC:整体排序质量的指标(领域常识:AUC 越接近 1 越好)。
  • Δh₀ [95% CI]:加了输入后指标相对基线的提升量及置信区间;区间不含 0 才算稳。
  • Positive:阳性命中数,格式是「命中/总数」。
  • Calls / Tokens:调模型次数与消耗的 token,↓表示越省越好。

值得留意

作者用两个维度同时评判:不只是分数(Best@k、AUC),还有成本(Calls、Tokens)。CellScientist 分数最高但最贵(154 次调用、40 万 token);CellForge 最省、阳性却只有 1/10,Δh₀ 置信区间还跨 0——等于没提升。

b145Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.

讲解

1. 这段在干什么

这是表格的图例说明,交代表里加粗、下划线以及各列的排名规则——属于读数前的"规则声明"。

2. 需要解释的地方

  • bold/underline:每个任务或折内,加粗=预测均值第一名,下划线=第二名(只数"不同值")。
  • ties share a rank:并列就同名次,不跳号。
  • interval bounds:区间上下界那一列。
  • diagnostic/resource columns:诊断类、资源类列(如你上一段看到的 1/10、88/100、117、328,705 这些)——不参与排名。

3. 值得留意

上一段那串数字同时含预测值、区间、排名和资源开销,但套用本段规则时,只有预测均值参与加粗/下划线的名次,其余列一律不排。"distinct"一词说明重复值不计入前二。

b146A. Selected architecture Policy Architecture Width Depth Dropout PCC–MSE weight CellScientist Gated MLP 512 4 0.1 0.3 AIDE Gated MLP 512 4 0.2 0.2 CellForge Residual MLP 256 3 0.1 0.1 HarmonyCell Residual MLP 384 4 0.2 0.1

讲解

1. 这段在干什么

这是一张"选定架构"配置表,列出四个框架(CellScientist、AIDE、CellForge、HarmonyCell)各自用的网络结构和超参数,为后文比较"输入使用声明"提供统一的模型底座。

2. 需要解释的地方

  • Gated MLP / Residual MLP:两种多层感知机变体。"Gated"门控、"Residual"残差连接,都是领域常识,指不同的层间信息传递方式。
  • Width / Depth / Dropout:宽度(每层神经元数)、深度(层数)、丢弃率,训练超参数。
  • PCC–MSE weight:把相关系数(PCC)和均方误差(MSE)两项损失加权组合时的权重。
  • 表头似乎因换行而串行了,读时把每行的名字、架构、宽度、深度、Dropout、权重依次对应。

3. 值得留意

四个框架架构并不统一(前两个 Gated、后两个 Residual),宽度深度也各异——即"选定的"基线本身就有差异,不是同一模型换名字。

b147B. Optimization and selection Policy Optimizer / learning rate Parameters Fold-3 PCC CellScientist AdamW / 0.001 5,850,037 0.9237 AIDE AdamW / 0.0005 5,850,037 0.9234 CellForge AdamW / 0.001 1,850,549 0.9195 HarmonyCell AdamW / 0.001 3,758,261 0.9247

讲解

1. 这段在干什么

紧接上一段的架构设置,这张表报告了四个模型(CellScientist、AIDE、CellForge、HarmonyCell)的优化与选择策略,即各自用什么优化器、学习率、参数量和 Fold-3 上的 PCC 表现。

2. 需要解释的地方

  • 优化器 / 学习率:训练时更新权重的算法(这里都是 AdamW)和每次更新的步长(领域常识)。
  • PCC:皮尔逊相关系数,衡量预测值与真实值的线性相关程度,越接近 1 越好(领域常识)。
  • Fold-3:交叉验证中的第 3 折。

3. 值得留意

四个模型参数量不同(CellForge 最少,1,850,549),但 PCC 却非常接近(0.9195–0.9247),差距很小。表头写「B. Optimization and selection Policy」,但标题实际所在小节是 Appendix D,编号对不上——这段没解释原因。

b148Policy Global PCC ↑\uparrow [95% CI] MSE ↓\downarrow Condition PCC ↑\uparrow DEG-20 PCC ↑\uparrow DEG-50 PCC ↑\uparrow Bundle drop Bundle-use support Fold 4 audit CellScientist 0.9020 [0.8988,0.9051] 0.00388 0.9098 0.9750 0.9697 0.5578 5/5 AIDE 0.8981 [0.8920,0.9042] 0.00407 0.9038 0.9682 0.9637 0.5458 5/5 CellForge 0.8960 [0.8909,0.9011] 0.00372 0.8999 0.9664 0.9618 0.5425 5/5 HarmonyCell 0.8960 [0.8920,0.9000] 0.00369 0.9019 0.9693 0.9641 0.5506 5/5 Fold 5 replication CellScientist 0.9096 [0.9051,0.9142] 0.00277 0.8944 0.9650 0.9554 0.6280 5/5 AIDE 0.9125 [0.9088,0.9161] 0.00285 0.8912 0.9567 0.9488 0.6270 5/5 CellForge 0.9051 [0.9023,0.9079] 0.00267 0.8844 0.9542 0.9474 0.6193 5/5 HarmonyCell 0.9088 [0.9029,0.9148] 0.00255 0.8867 0.9597 0.9505 0.6294 5/5

这段在干什么:承接上表的超参配置,给出四个模型在 Fold 4 审计与 Fold 5 复现下的预测性能与输入使用情况。

需要解释的地方:PCC 是预测与真实的相关系数,越高越好;MSE 是均方误差,越低越好;方括号是 95% 置信区间。DEG-20/50 指只看表达变化最大的基因时的 PCC。Bundle drop 是去掉整个输入特征组后性能掉了多少,掉得越多说明用得越狠;Bundle-use support 是几次实验里该组被用上的次数(5/5 即全用上)。

值得留意:Fold 5 里 AIDE 的 PCC 最高,但 Bundle drop 反而略低;CellScientist 在 Fold 4 领先,到 Fold 5 就不是第一了。

b149Bold/underline: best/second distinct displayed predictive mean within each task or fold. Ties share a rank; interval bounds and diagnostic/resource columns are unranked.

这段是表格的图例说明,紧接上一条数据行,告诉你表里那些加粗/下划线是干嘛的。

这段在干什么:定义表格排版规则——每个 task 或 fold 内,预测均值最好/次好的分别加粗、加下划线;打平则并列同名次;区间和诊断、资源类列不参与排名。

需要解释的地方:

  • predictive mean:模型预测值的平均,是这里排序的依据(领域常识:表格常在多指标中挑一个主指标排名)。
  • task / fold:任务或交叉验证折,是"分组比较"的单位,即排名在组内比,不跨组。
  • Ties share a rank:数值相同则名次一样。

值得留意:排名只在组内有效,别跨任务横向比粗细;区间(方括号那项)和 5/5 这类列不排名,粗体与它们无关。

b150Comparator Δ\Delta Global PCC [95% CI] PCC wins Δ\Delta bundle drop [95% CI] Fold 4 audit AIDE +0.0039 [−-0.0044,+0.0122] 4/5 +0.0121 [−-0.0004,+0.0245] CellForge +0.0060 [−-0.0008,+0.0128] 4/5 +0.0154 [+0.0030,+0.0277] HarmonyCell +0.0060 [−-0.0009,+0.0129] 4/5 +0.0072 [−-0.0061,+0.0205] Fold 5 replication AIDE −-0.0028 [−-0.0103,+0.0046] 2/5 +0.0010 [−-0.0059,+0.0078] CellForge +0.0045 [−-0.0004,+0.0094] 5/5 +0.0087 [+0.0037,+0.0137] HarmonyCell +0.0008 [−-0.0056,+0.0072] 4/5 −-0.0015 [−-0.0131,+0.0101]

这段是全篇审计结果的总表:3 个基线模型(AIDE、CellForge、HarmonyCell)在 Fold 4 与 Fold 5 上的预测增益和「bundle drop」。

关键概念(领域常识):PCC 是预测值与真实值的相关性,越高越好;95% CI 是置信区间,跨 0 就说明该增益不显著;"wins" 指 5 个任务里赢了几次;"bundle drop" 指移除一组输入特征后性能掉多少,掉得越多说明这组特征越被真正依赖。

值得留意:Fold 4 上三家 ΔPCC 的区间都跨 0,即都不显著,但 bundle drop 里只有 CellForge 的区间不跨 0——增益不显著 ≠ 特征不被依赖,这正是"审计"的意义。Fold 5 上 AIDE 的 ΔPCC 变成了负的,5 次只赢 2 次。

Appendix E Independent-acquisition test of small-molecule claim generalization

b152After the final falsification-guided procedure is fixed, LINCS–Pilot1 provides an out-of-development-cohort test. It uses an observed control profile, compound identity, and dose to predict a joint 2,719-dimensional morphology–transcription response across 4,155 matched conditions and 1,108 Murcko groups. The ten trajectory-selected design instances are then fixed for transfer to cpg0004/LKCP/2017_12_05_Batch2, an independently acquired Cell Painting release (Weisbart et al., 2024). Its 134 files contain 11,904 rows; 8,958 treatment and 2,946 control rows yield 1,728 retained conditions and 64 compound-held-out groups. Response- and fold-blind interface quality control removes 110 non-finite or extreme release features from 1,781 common morphology columns, retaining a 1,671-dimensional same-plate control profile and response.

这段在干什么:介绍独立获取的外部验证集——把已固定的设计实例从 LINCS–Pilot1 迁移到另一套独立采集的 Cell Painting 数据上测试。

需要解释的地方:

  • out-of-development-cohort:不在建模队列里的样本,即真正的"外部测试集"。
  • Murcko group:领域常识,指按分子骨架聚类分组,用于化合物层面的留出。
  • response- and fold-blind QC:不看预测响应和倍数变化做质控,避免信息泄漏。

值得留意:整段几乎全是数据规模与清洗细节(4,155、1,108、11,904、1,728、64 等),作者刻意把"外部独立性"铺垫得很重——这是判断小分子结论能否泛化的关键设计。

b153The scientific question and model designs precede LKCP transfer. The final dataset adapter is fixed after interface quality control and before LKCP training or held-out-fold access. The transferred designs, input-permutation construction, map seeds, threshold-construction rule, and decision rule are registered and hash-bound at that point. Each design is then refit under five paired seeds with no further source edit or model reselection. Compound permutations hold dose fixed; dose permutations use only observed doses of the same compound, require at least a 0.10.1 log-dose contrast, and retain the same eligibility rule. Across 22 cohorts ×\times 22 folds ×\times 1010 designs ×\times 55 seeds, the analysis contains 200 cohort–fold–design–seed refit boundaries. Each boundary reuses the 32 registered ordered compound–dose permutation pairs, yielding 6,400 matched within-refit permutation units; these maps are repeated measurements rather than independent statistical units. For U11U_{11} with both inputs correct, U01U_{01} with compound shuffled, U10U_{10} with dose shuffled, and U00U_{00} with both shuffled, the two-player allocation is ϕchem=12​[(U11−U01)+(U10−U00)]\phi_{\mathrm{chem}}=\tfrac{1}{2}[(U_{11}-U_{01})+(U_{10}-U_{00})] and ϕdose=12​[(U11−U10)+(U01−U00)]\phi_{\mathrm{dose}}=\tfrac{1}{2}[(U_{11}-U_{10})+(U_{01}-U_{00})]. These allocate the contrast U11−U00U_{11}-U_{00}; adding the residual U00−UanchorU_{00}-U_{\mathrm{anchor}} gives Full−-control-only to numerical precision. The component allocations thus refer to the fixed replacement baseline, while Full−-control-only measures the separately fitted predictors’ performance difference.

这段在干什么

交代这套小分子泛化检验的分析协议与记账方式:先冻结设计、再重拟合,并用置换组合把预测贡献拆成「化合物」和「剂量」两块。

需要解释的地方

  • LKCP 迁移:此前搭好的模型框架被搬到这里复用,故需先定协议。
  • U₁₁/U₀₁/U₁₀/U₀₀:两个输入(化合物、剂量)分别正确或打乱的四种组合。
  • φ_chem、φ_dose:把 U₁₁−U₀₀ 这个总对比,按「各打乱一个输入」的差值平分成两份贡献。

值得留意

作者自己点明:这些置换图是重复测量,不是独立统计单元,别按 6,400 个样本算显著性。此外 φ 是相对固定的替换基线,而 Full−control-only 是另外单独拟合的,两者口径不同。

b154A. LINCS–Pilot1 Contrast Fold 4 [95% CI] Fold 5 [95% CI] Full−-control-only +0.0293 [+0.0281, +0.0305] +0.0233 [+0.0226, +0.0241] Compound identity +0.0175 [+0.0164, +0.0186] +0.0126 [+0.0117, +0.0134] Dose +0.0486 [+0.0473, +0.0498] +0.0467 [+0.0456, +0.0478] Both-shuffled residual -0.0368 [-0.0378, -0.0358] -0.0359 [-0.0370, -0.0349] Interaction +0.0050 [+0.0033, +0.0067] +0.0042 [+0.0027, +0.0058]

这段在报告 LINCS–Pilot1 对比实验在 Fold 4/5 两个交叉验证折上的结果。

术语:Fold 指交叉验证的折。「Full−control-only」是把完整模型减去只留对照的基线,衡量各预测器单独拟合后的性能差;「Compound identity」「Dose」分别代表化合物身份、剂量两类输入;「Both-shuffled residual」是两者都打乱后的残差;「Interaction」是二者的交互项。

留意:Dose 贡献最大(约 +0.049/+0.047),远高于 Compound identity;打乱后残差为负;交互项很小。

b155B. LKCP Batch 2 Contrast Fold 4 [95% CI] Fold 5 [95% CI] Full−-control-only +0.0104 [+0.0103, +0.0106] +0.0211 [+0.0205, +0.0217] Compound identity -0.00007000 [-0.00028998, +0.00014998] +0.0033 [+0.0027, +0.0039] Dose +0.0237 [+0.0228, +0.0247] +0.0341 [+0.0331, +0.0350] Both-shuffled residual -0.0132 [-0.0140, -0.0124] -0.0163 [-0.0170, -0.0157] Interaction +0.00004049 [-0.00001503, +0.00009601] +0.0002 [+0.0001, +0.0003]

接着上一段看,这张表在逐项拆 LKCP Batch 2 里各种“打乱对照”的结果。

1. 这段在干什么:补上 Batch 2 的 Fold 4、Fold 5 两列对照数据,量每种扰动对预测贡献的影响。

2. 需要解释的地方:各行是不同“打乱/对照”手法——只留对照、打乱化合物身份、打乱剂量、两者都打乱后看残差、以及交互项。方括号里是 95% 置信区间,即估计的不确定范围。数值是相对基准的“预测贡献变化”。

3. 值得留意:Compound identity 一行两列都极接近 0,且区间跨 0,说明打乱化合物身份几乎不改变贡献;Dose 一行数值最大(+0.0237、+0.0341),区间不跨 0。哪项才是主要驱动,从这两行就能读出取向。

b156Cohort Fold Mean effect Lower-effect quantile Threshold Reference margin Refits above Chemical identity LINCS–Pilot1 4 0.0200 0.0139 0.0112 +0.0028 35/50 LINCS–Pilot1 5 0.0147 0.0090 0.0100 -0.0010 10/50 LKCP Batch 2 4 0.00003872 -0.0016 0.0303 -0.0319 0/50 LKCP Batch 2 5 0.0032 0.0013 0.0307 -0.0294 0/50 Dose LINCS–Pilot1 4 0.0510 0.0467 0.0068 +0.0398 50/50 LINCS–Pilot1 5 0.0488 0.0452 0.0059 +0.0393 50/50 LKCP Batch 2 4 0.0238 0.0205 0.0060 +0.0145 50/50 LKCP Batch 2 5 0.0342 0.0267 0.0125 +0.0142 50/50

这张表在检验:小分子效应到底来自化学身份还是剂量。

  • Cohort/Fold:数据集的两个划分。
  • Mean effect:平均效应;Lower-effect quantile:效应下分位。
  • Threshold:判定门槛;Reference margin=Mean effect−Threshold。
  • Refits above:超过门槛的重拟合数,如 35/50。

值得留意:按化学身份分组,LKCP 两折 margin 为负、0/50 通过;按剂量分组,四个组合全部 50/50。

Appendix F Full selected-model audit panel and calibration

b158The full Fold-4 panel contains the Fold-3-selected source from each of the four agents together with three distinct diagnostic sources: control-only, compound-only, and the shared h0h_{0}/simple-concatenation source shown under both reference roles. Predictive estimates use a separate five-seed refit panel from Table 21; model source or design, data fingerprint, population, and every non-seed fitting setting are identical. Each within-policy source is fixed before Fold 4; the pre-specified across-policy score rule then uses mean Fold-4 Global PCC to choose the single prediction-score source carried to Fold 5, which includes its paired control-only comparator. Controlled fixtures check recognized input consumption, the one-key attention contradiction, missing declarations, abstention, and known input dependence. A fixed compound-strength sweep characterizes threshold crossing in those fixtures; Appendix L explains the criterion’s reference-relative interpretation.

讲解

1. 这段在干什么

交代 Fold-4 审计面板的组成(四个 agent 各自 Fold-3 入选源 + 三个诊断源)与选源规则,为进入 Fold 5 做过渡。

2. 需要解释的地方

  • control-only / compound-only:分别只含对照、只含化合物的诊断源,用来做对比。
  • PCC:皮尔逊相关系数,领域常识,衡量预测与真实的线性相关。
  • fixtures:受控的测试样例,用来检验模型行为是否如声明。

3. 值得留意

预测估计另用五种子重拟合面板,且除随机种子外所有设置(源、数据指纹、群体、拟合参数)均保持一致——这是为了让差异只归因于种子。选源在 Fold 4 前就固定,避免事后挑选。具体数值这段没给。

b159A. Predictive performance Model Global PCC MSE CP PCC L1000 PCC BBBC036 CellScientist 0.3096 ±\pm 0.0029 0.0586 ±\pm 0.0001 0.3113 ±\pm 0.0014 0.3089 ±\pm 0.0041 AIDE 0.2970 ±\pm 0.0075 0.0593 ±\pm 0.0004 0.3052 ±\pm 0.0031 0.2944 ±\pm 0.0102 CellForge 0.2960 ±\pm 0.0036 0.0592 ±\pm 0.0002 0.3056 ±\pm 0.0043 0.2925 ±\pm 0.0037 HarmonyCell 0.3110 ±\pm 0.0016 0.0586 ±\pm 0.0001 0.3109 ±\pm 0.0020 0.3110 ±\pm 0.0017 h0h_{0} 0.2879 ±\pm 0.0040 0.0598 ±\pm 0.0004 0.2981 ±\pm 0.0016 0.2839 ±\pm 0.0056 Control-only 0.3098 ±\pm 0.0024 0.0586 ±\pm 0.0001 0.3094 ±\pm 0.0013 0.3099 ±\pm 0.0028 Chemical-only 0.2239 ±\pm 0.0002 0.0616 ±\pm 0.0000 0.2998 ±\pm 0.0003 0.1832 ±\pm 0.0002 Simple concat 0.2879 ±\pm 0.0040 0.0598 ±\pm 0.0004 0.2981 ±\pm 0.0016 0.2839 ±\pm 0.0056 BBBC047 CellScientist 0.3035 ±\pm 0.0008 0.0457 ±\pm 0.0000 0.3260 ±\pm 0.0016 0.2784 ±\pm 0.0017 AIDE 0.2697 ±\pm 0.0048 0.0469 ±\pm 0.0002 0.3148 ±\pm 0.0047 0.2152 ±\pm 0.0116 CellForge 0.2727 ±\pm 0.0019 0.0467 ±\pm 0.0001 0.3175 ±\pm 0.0035 0.2190 ±\pm 0.0059 HarmonyCell 0.2983 ±\pm 0.0110 0.0459 ±\pm 0.0003 0.3271 ±\pm 0.0030 0.2654 ±\pm 0.0231 h0h_{0} 0.2697 ±\pm 0.0048 0.0469 ±\pm 0.0002 0.3148 ±\pm 0.0047 0.2152 ±\pm 0.0116 Control-only 0.3019 ±\pm 0.0015 0.0458 ±\pm 0.0001 0.3223 ±\pm 0.0011 0.2789 ±\pm 0.0023 Chemical-only 0.2159 ±\pm 0.0012 0.0480 ±\pm 0.0000 0.2935 ±\pm 0.0022 0.0924 ±\pm 0.0028 Simple concat 0.2697 ±\pm 0.0048 0.0469 ±\pm 0.0002 0.3148 ±\pm 0.0047 0.2152 ±\pm 0.0116

这段在干什么:这是附录里的一张完整数值表,列出各模型在两个数据集(BBBC036、BBBC047)上的预测指标,供审计时逐项核对。

需要解释的地方:

  • PCC / MSE:这是领域常识,PCC 是预测值与真实值的相关性(越高越好),MSE 是均方误差(越低越好)。
  • Global / CP / L1000 / BBBC036 各列:四个不同评估口径的 PCC,即全局、CP、L1000 子集、BBBC036 子集。
  • ±值:多次运行的波动范围,不是误差棒含义之外的东西。

值得留意:

  • 表中 Control-only 与 CellScientist、HarmonyCell 数值几乎持平甚至更高,而 Chemical-only 明显偏低——模型看似"预测好",可能主要来自对照组信号。这是读表时最容易漏掉的关键点。
  • Simple concat 与 h₀ 数值完全相同,两行可能是同一配置。

b160B. Input-use evidence Model EchemE_{\rm chem} EcontrolE_{\rm control} Ichem,controlI_{\rm chem,control} Q/U/I Status BBBC036 CellScientist 0.0000 ±\pm 0.0000 0.1622 ±\pm 0.0095 0.0000 ±\pm 0.0000 0/5/0 Unsup. AIDE 0.0029 ±\pm 0.0017 0.1577 ±\pm 0.0118 -0.0021 ±\pm 0.0016 0/5/0 Unsup. CellForge 0.0089 ±\pm 0.0052 0.1464 ±\pm 0.0063 -0.0000 ±\pm 0.0007 0/5/0 Unsup. HarmonyCell 0.0000 ±\pm 0.0000 0.1693 ±\pm 0.0033 -0.0000 ±\pm 0.0000 0/5/0 Unsup. h0h_{0} 0.0193 ±\pm 0.0087 0.1359 ±\pm 0.0057 -0.0003 ±\pm 0.0007 1/2/2 Inc. Control-only 0.0000 ±\pm 0.0000 0.1719 ±\pm 0.0067 0.0000 ±\pm 0.0000 – N/C Chemical-only 0.0011 ±\pm 0.0002 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0/0/5 Inc. Simple concat 0.0193 ±\pm 0.0087 0.1359 ±\pm 0.0057 -0.0003 ±\pm 0.0007 1/2/2 Inc. BBBC047 CellScientist 0.0000 ±\pm 0.0000 0.1932 ±\pm 0.0044 0.0000 ±\pm 0.0000 0/5/0 Unsup. AIDE 0.0118 ±\pm 0.0052 0.1119 ±\pm 0.0121 0.0013 ±\pm 0.0009 0/1/4 Inc. CellForge 0.0100 ±\pm 0.0045 0.1114 ±\pm 0.0059 0.0012 ±\pm 0.0008 0/2/3 Inc. HarmonyCell 0.0006 ±\pm 0.0010 0.1866 ±\pm 0.0225 0.0009 ±\pm 0.0015 0/5/0 Unsup. h0h_{0} 0.0118 ±\pm 0.0052 0.1119 ±\pm 0.0121 0.0013 ±\pm 0.0009 0/1/4 Inc. Control-only 0.0000 ±\pm 0.0000 0.2030 ±\pm 0.0054 0.0000 ±\pm 0.0000 – N/C Chemical-only 0.0111 ±\pm 0.0010 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 5/0/0 Sup. Simple concat 0.0118 ±\pm 0.0052 0.1119 ±\pm 0.0121 0.0013 ±\pm 0.0009 0/1/4 Inc.

这张表给出两个数据集(BBBC036、BBBC047)上各模型的三类输入使用证据分数,用来核对模型是否真的在用化学信息。

需要解释的地方

  • E_chem:化学输入带来的预测贡献;E_control:对照输入贡献;I_chem,control:两者的交互项。
  • Q/U/I 是计数:符合假设 / 不确定 / 违背,三数相加为 5(即 5 次随机种子或重复)。
  • Unsup./Inc./Sup.:证据不支持、不一致、支持该输入使用声明。

值得留意

  • 很多模型 E_chem≈0 且 I≈0,但 E_control 明显非零——说明预测力主要来自对照侧,而模型仍被标为 Unsup.。
  • h₀ 与 Simple concat 数值完全相同,注明两者是同一基线。

b161Aggregate status: Sup., claim-supported; Unsup., unsupported; Inc., inconclusive; N/C, not claimed.

这句是附录审计表格的表注,用来定义表格里状态列的缩写。

1. 这段在干什么:给读者一把"钥匙",说明状态列 Sup./Unsup./Inc./N/C 各代表什么意思,否则表格没法读。

2. 需要解释的地方:

  • Sup. = claim-supported,声称被支持;
  • Unsup. = unsupported,不支持;
  • Inc. = inconclusive,结论不明;
  • N/C = not claimed,压根没提这个声称。

(作者逐字给出了对应,非推测。)

3. 值得留意:它正好能解释上一段结尾的 5/0/0 Sup. 和 0/1/4 Inc.——中间那四个数字应分别对应这四类计数。但具体顺序,这段原文没有说,需回看表格表头确认。

b162A. Predictive performance Policy Model Global PCC ↑\uparrow MSE ↓\downarrow CP PCC ↑\uparrow L1000 PCC ↑\uparrow BBBC036 h0h_{0} reference h0h_{0} 0.2588 ±\pm 0.0059 0.0608 ±\pm 0.0004 0.2596 ±\pm 0.0100 0.2586 ±\pm 0.0094 Score-selected HarmonyCell 0.2932 ±\pm 0.0020 0.0590 ±\pm 0.0001 0.2909 ±\pm 0.0018 0.2942 ±\pm 0.0023 BBBC047 h0h_{0} reference h0h_{0} 0.2786 ±\pm 0.0031 0.0465 ±\pm 0.0003 0.3220 ±\pm 0.0091 0.2243 ±\pm 0.0105 Score-selected CellScientist 0.3153 ±\pm 0.0014 0.0452 ±\pm 0.0000 0.3393 ±\pm 0.0007 0.2872 ±\pm 0.0025

这段是审计面板的A 部分:预测性能,用数字支撑“选出的模型确实更好”。

关键概念:PCC 是皮尔逊相关系数,越高越好(↑);MSE 是均方误差,越低越好(↓)。h₀ 指零假设/基线,这里是“未经选择的参考模型”。两行数据分别是两个数据集(BBBC036、BBBC047),每行对比参考 h₀ 与“Score-selected”(评分选出的)模型。

值得留意:选出的模型(HarmonyCell、CellScientist)在 Global PCC、CP PCC、L1000 PCC 上都高于参考,MSE 更低,且标准差很小——说明提升稳定,不是偶然波动。但 BBBC047 参考的 L1000 PCC(0.2243)反而低于 BBBC036 的(0.2586),这处不对称表里没解释。

b163B. Input-use evidence and replication Model EchemE_{\rm chem} EcontrolE_{\rm control} Ichem,controlI_{\rm chem,control} Q/U/I Status Status replicated? BBBC036 h0h_{0} 0.0047 ±\pm 0.0031 0.1202 ±\pm 0.0085 0.0042 ±\pm 0.0011 0/5/0 Unsup. no HarmonyCell 0.0000 ±\pm 0.0000 0.1516 ±\pm 0.0035 -0.0000 ±\pm 0.0000 0/5/0 Unsup. yes BBBC047 h0h_{0} 0.0096 ±\pm 0.0037 0.1199 ±\pm 0.0109 0.0012 ±\pm 0.0007 0/1/4 Inc. yes CellScientist 0.0000 ±\pm 0.0000 0.2004 ±\pm 0.0061 0.0000 ±\pm 0.0000 0/5/0 Unsup. yes

这段是附录里的审计证据表:逐个模型列出「输入是否真被用到」的证据与复现结果。

关键列:\(E_{\rm chem}\) 是化学输入对预测的贡献(接近 0 表示没用上),\(E_{\rm control}\) 是对照量,\(I_{\rm chem,control}\) 是化学相对对照的影响;Q/U/I 是判为「存疑/未支持/不一致」的计数。Status 是审计结论,最后一列是有无复现。这是领域常识:贡献≈0 即模型没真正依赖该输入。

值得留意:HarmonyCell 与 CellScientist 的 \(E_{\rm chem}\) 全为 0.0000,却判 Unsup. 且复现 yes——「复现了」复现的是「未支持」这一结论,不是复现了模型有效。

b164Score-selected denotes prediction-score selection. Aggregate status: Unsup., unsupported; Inc., inconclusive.

这段是附录审计表格的表注/图例,用来定义前文表格里出现的缩写,属于阅读辅助而非论证内容。

需要解释的地方

  • Score-selected(预测分数选择):指按模型的预测得分高低来挑选模型,是相对其他选择标准而言的一种筛选方式。
  • Aggregate status(汇总状态):对某模型多项审计结果的总体判定。
  • Unsup. / Inc.:即 unsupported(无支撑)与 inconclusive(不确定、无法定论)。

值得留意

前一段末尾的 Inc.、Unsup. 正是靠这里才得到解释;yes 一列则未在表注中定义,含义这段没提到。

b165Task Score-selected model Selected-model PCC ↑\uparrow Control-only PCC ↑\uparrow Paired Δ\Delta [95% CI] BBBC036 HarmonyCell 0.2932 ±\pm 0.0020 0.2921 ±\pm 0.0017 +0.0011 [−-0.0008, +0.0030] BBBC047 CellScientist 0.3153 ±\pm 0.0014 0.3142 ±\pm 0.0012 +0.0011 [−-0.0005, +0.0027]

这张表是审计面板核心结果:对两个任务,比较按预测分数选出的模型与仅用对照的基线,看 PCC(预测相关性,领域常识:越大越好)差多少。

值得留意:两处 Δ 都很小(+0.0011),且 95% 置信区间都跨 0,即差距在统计上不显著——所以上一段标注为"Unsup./Inc."。这正是"发现—证伪—修正"里的证伪信号。表头未给这些符号的含义,这段没提。

b166A. BBBC036 Model Residual PCC ↑\uparrow Mean [95% CI] MSE ↓\downarrow Pert. effect ↑\uparrow h0h_{0} 0.0474 [0.0371, 0.0577] 0.0598 0.0381 CellScientist 0.0294 [0.0052, 0.0535] 0.0586 0.0000 AIDE 0.0301 [0.0236, 0.0366] 0.0593 0.0050 CellForge 0.0336 [0.0234, 0.0438] 0.0592 0.0223 HarmonyCell 0.0341 [0.0226, 0.0455] 0.0586 0.0000 Control-only Zero increment 0.0586 – Chemical-only 0.0052 [-0.0044, 0.0148] 0.0616 0.0012 Simple concat 0.0474 [0.0371, 0.0577] 0.0598 0.0381 B. BBBC047 Model Residual PCC ↑\uparrow Mean [95% CI] MSE ↓\downarrow Pert. effect ↑\uparrow h0h_{0} 0.0558 [0.0510, 0.0606] 0.0469 0.0205 CellScientist 0.0538 [0.0422, 0.0654] 0.0457 0.0000 AIDE 0.0558 [0.0510, 0.0606] 0.0469 0.0205 CellForge 0.0553 [0.0510, 0.0596] 0.0467 0.0174 HarmonyCell 0.0567 [0.0470, 0.0663] 0.0459 0.0020 Control-only Zero increment 0.0458 – Chemical-only 0.0457 [0.0382, 0.0532] 0.0480 0.0104 Simple concat 0.0558 [0.0510, 0.0606] 0.0469 0.0205

这一段是附录里的完整审计面板:把 BBBC036 和 BBBC047 两个数据集上,各模型(h₀、CellScientist、AIDE、CellForge、HarmonyCell)加几条对照基线,按三个指标逐行列出。

关键概念:Residual PCC 是残差相关性,越高越好(↑);MSE 是误差,越低越好(↓);"Pert. effect"是扰动效应,看模型是否真用上了扰动信息。Control-only、Chemical-only、Simple concat 是对照组。

值得留意:两表里 h₀ 与 Simple concat 的数值完全相同,且 Pert. effect 都偏高;而 CellScientist、HarmonyCell 的 Pert. effect 直接为 0.0000。这说明有些模型可能没真正利用扰动项,正是审计要查的疑点。

F.1 Known-truth measurement and uncertainty checks

b168The fixed-context simulator uses f⁡(z,c)=(a​z,c)f(z,c)=(az,c) and Y=(b​z+e,c)Y=(bz+e,c). Matched replacement draws an independent normal z′z^{\prime} while fixing c,Yc,Y. The loss is the sum of squared errors over both channels, so the population paired loss increment is τ=2​a​b\tau=2ab; the mean squared Euclidean prediction distance is η=2​a2\eta=2a^{2}, not RMS. Groups receive equal weight and rows equal weight within group. The experiment evaluates equal-group-weighted population contrasts; real-data row-weighted resampling is reported separately. Four classes are invariant (a,b)=(0,.2)(a,b)=(0,.2), sensitive without benefit (.2,0)(.2,0), beneficial (.2,.2)(.2,.2), and harmful (−.2,.2)(-.2,.2). Their loss truths are 0,0,.08,−.080,0,.08,-.08 and distance truths 0,.08,.08,.080,.08,.08,.08.

这段在干什么:用已知真值的模拟器校准审计指标,给出损失增量与预测距离两套"标准答案"。

需要解释的地方:

  • f(z,c)=(az,c)、Y=(bz+e,c):合成数据的生成式,z 是变量、c 是固定上下文,a 决定影响、b 决定结果侧。
  • 匹配替换:固定 c、Y,只换一个独立的 z',用来切干净"输入被替换"这一个因素。
  • τ=2ab 是配对损失增量真值,η=2a² 是均方欧氏距离真值(不是 RMS,即不做开方)。
  • 四类 (a,b) 组合分别对应:不变、敏感但无益、有益、有害。

值得留意:距离真值只依赖 a,槽位组与行都等权,等组加权与真实数据的行加权重采样是分开报告的,别混。

b169The frozen 24=162^{4}=16 scenarios cross G∈{8,32}G\in\{8,32\} independent source groups, balanced size 8 or independently sampled sizes {2,4,8,32}\{2,4,8,32\} with equal probabilities, within-group correlation ρ∈{0,.6}\rho\in\{0,.6\}, and noise SD σ∈{.25,1}\sigma\in\{.25,1\}. Each scenario has 400 independent datasets (6400 total). Maps 32/128 are nested on the same datasets, not independent cohorts. Group-percentile intervals use 999 bootstrap draws; group-tt uses G−1G-1 degrees of freedom. Source/donor, bootstrap and context seeds are 2026091401, 2026091402 and 2026091403.

讲解

这段在干什么:交代「已知真值」实验的设计清单——场景怎么组合、每组算多大、随机种子取多少,为下一节的测量结果提供可复现的设置说明。

需要解释的地方:

  • 场景网格:4 个维度相乘共 16 种(G、组大小、ρ、σ),每个场景跑 400 个独立数据集。
  • 嵌套:Maps 32 和 128 用的是同一批数据集,所以两者比较更紧,但不是独立队列。
  • bootstrap 与 G−1:分组百分位区间靠 999 次重抽样;group-t 用组数减一做自由度(领域常识:这是小样本 t 检验的常规做法)。

值得留意:前段损失真值有负值,这段的 σ 与 ρ 只描述构造方式;「not independent cohorts」是作者主动提示的局限。

b170Map-only intervals target the donor-integrated contrast of one fixed dataset; group-aware intervals target the population expectation over independent groups. Their coverage is therefore reported against both truths. Positive evidence means a strictly positive lower bound of a two-sided 95% CI: the positive one-tail nominal Type-I error is .025.025, not .05.05. Power is positive-direction evidence for benefit and negative-direction evidence for harm. Prediction distance does not test target benefit. Invariant data and intervals are identically zero and are reported as a deterministic check, not “100% calibrated” stochastic coverage.

这段在干什么

定义并区分两类区间(map-only / group-aware)各自对应的“真值”,并说明本节的覆盖率和证据判读规则。

需要解释的地方

  • map-only vs group-aware 区间:前者只针对单一固定数据集的供体整合对比,后者针对独立分组上的总体期望。领域常识:前者忽略分组不确定性,后者纳入。
  • Positive evidence:双侧 95% CI 的下界严格为正;即单侧名义 Type-I 错误是 .025 而非 .05。

值得留意

  • 覆盖是分别对两个真值报告的,不能混用。
  • 功效按方向分:正=获益证据,负=危害证据。
  • 预测距离不检验目标获益;不变量数据与区间恒为零,只作确定性检查,不等于“100% 校准”。

b171Endpoint / class CI method Population coverage Conditional coverage Directed evidence rate 32 donor maps Null loss Map only [0.1500, 0.3000][0.1500,\,0.3000] [0.9275, 0.9675][0.9275,\,0.9675] [0.2850, 0.5000][0.2850,\,0.5000] Null loss Group tt [0.9200, 0.9775][0.9200,\,0.9775] [1.0000, 1.0000][1.0000,\,1.0000] [0.0150, 0.0725][0.0150,\,0.0725] Null loss Group bootstrap [0.8175, 0.9400][0.8175,\,0.9400] [1.0000, 1.0000][1.0000,\,1.0000] [0.0350, 0.1475][0.0350,\,0.1475] Beneficial loss Map only [0.1425, 0.3350][0.1425,\,0.3350] [0.9250, 0.9650][0.9250,\,0.9650] [0.7100, 1.0000][0.7100,\,1.0000] Beneficial loss Group tt [0.9100, 0.9650][0.9100,\,0.9650] [0.9975, 1.0000][0.9975,\,1.0000] [0.1200, 1.0000][0.1200,\,1.0000] Beneficial loss Group bootstrap [0.8350, 0.9425][0.8350,\,0.9425] [0.9975, 1.0000][0.9975,\,1.0000] [0.2650, 1.0000][0.2650,\,1.0000] Harmful loss Map only [0.1050, 0.2750][0.1050,\,0.2750] [0.9375, 0.9650][0.9375,\,0.9650] [0.6775, 1.0000][0.6775,\,1.0000] Harmful loss Group tt [0.8625, 0.9725][0.8625,\,0.9725] [1.0000, 1.0000][1.0000,\,1.0000] [0.0550, 1.0000][0.0550,\,1.0000] Harmful loss Group bootstrap [0.8300, 0.9400][0.8300,\,0.9400] [1.0000, 1.0000][1.0000,\,1.0000] [0.1900, 1.0000][0.1900,\,1.0000] Squared distance Map only [0.2175, 0.5250][0.2175,\,0.5250] [0.9400, 0.9700][0.9400,\,0.9700] [1.0000, 1.0000][1.0000,\,1.0000] Squared distance Group tt [0.8375, 0.9525][0.8375,\,0.9525] [1.0000, 1.0000][1.0000,\,1.0000] [0.9975, 1.0000][0.9975,\,1.0000] Squared distance Group bootstrap [0.7800, 0.9375][0.7800,\,0.9375] [0.9975, 1.0000][0.9975,\,1.0000] [1.0000, 1.0000][1.0000,\,1.0000] 128 donor maps Null loss Map only [0.0675, 0.1625][0.0675,\,0.1625] [0.9225, 0.9725][0.9225,\,0.9725] [0.3600, 0.5375][0.3600,\,0.5375] Null loss Group tt [0.9125, 0.9775][0.9125,\,0.9775] [1.0000, 1.0000][1.0000,\,1.0000] [0.0125, 0.0775][0.0125,\,0.0775] Null loss Group bootstrap [0.8200, 0.9350][0.8200,\,0.9350] [1.0000, 1.0000][1.0000,\,1.0000] [0.0350, 0.1450][0.0350,\,0.1450] Beneficial loss Map only [0.0750, 0.1750][0.0750,\,0.1750] [0.9175, 0.9650][0.9175,\,0.9650] [0.7675, 1.0000][0.7675,\,1.0000] Beneficial loss Group tt [0.9075, 0.9725][0.9075,\,0.9725] [1.0000, 1.0000][1.0000,\,1.0000] [0.1200, 1.0000][0.1200,\,1.0000] Beneficial loss Group bootstrap [0.8400, 0.9450][0.8400,\,0.9450] [1.0000, 1.0000][1.0000,\,1.0000] [0.2475, 1.0000][0.2475,\,1.0000] Harmful loss Map only [0.0400, 0.1450][0.0400,\,0.1450] [0.9125, 0.9600][0.9125,\,0.9600] [0.7025, 1.0000][0.7025,\,1.0000] Harmful loss Group tt [0.8600, 0.9700][0.8600,\,0.9700] [1.0000, 1.0000][1.0000,\,1.0000] [0.0500, 1.0000][0.0500,\,1.0000] Harmful loss Group bootstrap [0.8200, 0.9425][0.8200,\,0.9425] [1.0000, 1.0000][1.0000,\,1.0000] [0.1875, 1.0000][0.1875,\,1.0000] Squared distance Map only [0.0925, 0.2900][0.0925,\,0.2900] [0.9350, 0.9600][0.9350,\,0.9600] [1.0000, 1.0000][1.0000,\,1.0000] Squared distance Group tt [0.8425, 0.9500][0.8425,\,0.9500] [1.0000, 1.0000][1.0000,\,1.0000] [0.9975, 1.0000][0.9975,\,1.0000] Squared distance Group bootstrap [0.8025, 0.9350][0.8025,\,0.9350] [1.0000, 1.0000][1.0000,\,1.0000] [1.0000, 1.0000][1.0000,\,1.0000]

这一段的上一句刚说完「区间恒为零时只作确定性检查、不算 100% 校准」,这段紧跟着把已知真值场景下的三种 CI 方法摆开对比。

这段在干什么:汇报已知真值实验里各端点/类别在 32 与 128 donor maps 下的群体覆盖、条件覆盖和有向证据率,验证不确定性估计是否可靠。

需要解释的地方:

  • 群体覆盖 / 条件覆盖:领域常识,指区间在总体和分组子集里真正罩住真值的比例,理想接近名义水平。
  • 有向证据率:区间能给出方向性(正/负)信号的比例。
  • Map only / Group tt / Group bootstrap:三种构造区间的方法,名字沿用前文。

值得留意:Map only 的群体覆盖明显低于另两种(如 128 maps 下 Null loss 只有约 [0.0675, 0.1625]),但它的有向证据率反而更高;Group tt 和 bootstrap 的群体覆盖接近 1,方向率却常落到 0 附近。这个权衡是作者没说破的关键。

b172Each rate uses 400 independent datasets per scenario; ranges are minima/maxima over scenarios, not pooled coverage. The last column is positive Type I for null loss (nominal .025), correctly directed power for beneficial/harmful loss, and positive-distance evidence for distance. Conditional and population targets differ. Invariant cases are excluded from stochastic coverage.

这段在干什么

交代上表结果的读法:每格数字怎么来的、代表什么,避免误读成汇总覆盖率。

需要解释的地方

  • 400 个独立数据集 / 场景:每个场景各跑 400 次独立实验。
  • 区间是各场景的最小/最大值,不是把数据汇到一起算的覆盖(常识:pooled 会把场景差异抹平)。
  • Type I(第一类错误):虚报效应,名义值 .025。
  • Conditional / population targets:条件与总体两类目标,别混着比。

值得留意

不变量(invariant)案例不进入随机覆盖率统计——做跨场景比较时,样本范围其实被缩过。

b173GG/size/ρ\rho/σ\sigma Map only Group tt Group bootstrap tt Type-I MCSE tt Type-I Wilson 95% CI 8/B/0/0.25 0.1275 0.4600 0.9500 0.0400 0.8675 0.0950 0.00980.0098 [0.0248, 0.0640][0.0248,\,0.0640] 8/B/0/1 0.1175 0.4925 0.9550 0.0325 0.8950 0.0700 0.00890.0089 [0.0191, 0.0548][0.0191,\,0.0548] 8/B/0.6/0.25 0.0875 0.5075 0.9325 0.0600 0.8200 0.1450 0.01190.0119 [0.0406, 0.0877][0.0406,\,0.0877] 8/B/0.6/1 0.0675 0.4925 0.9650 0.0225 0.8525 0.0850 0.00740.0074 [0.0119, 0.0422][0.0119,\,0.0422] 8/I/0/0.25 0.1300 0.4675 0.9325 0.0650 0.8500 0.1275 0.01230.0123 [0.0447, 0.0935][0.0447,\,0.0935] 8/I/0/1 0.1550 0.3600 0.9775 0.0125 0.8825 0.0650 0.00560.0056 [0.0054, 0.0289][0.0054,\,0.0289] 8/I/0.6/0.25 0.0725 0.5375 0.9350 0.0650 0.8400 0.1375 0.01230.0123 [0.0447, 0.0935][0.0447,\,0.0935] 8/I/0.6/1 0.0825 0.4625 0.9600 0.0250 0.8300 0.1000 0.00780.0078 [0.0136, 0.0454][0.0136,\,0.0454] 32/B/0/0.25 0.1125 0.4275 0.9325 0.0300 0.9225 0.0375 0.00850.0085 [0.0172, 0.0517][0.0172,\,0.0517] 32/B/0/1 0.1350 0.4500 0.9625 0.0225 0.9350 0.0350 0.00740.0074 [0.0119, 0.0422][0.0119,\,0.0422] 32/B/0.6/0.25 0.0850 0.4650 0.9325 0.0575 0.9175 0.0675 0.01160.0116 [0.0386, 0.0848][0.0386,\,0.0848] 32/B/0.6/1 0.0875 0.4600 0.9600 0.0250 0.9350 0.0350 0.00780.0078 [0.0136, 0.0454][0.0136,\,0.0454] 32/I/0/0.25 0.1450 0.3975 0.9450 0.0425 0.9225 0.0500 0.01010.0101 [0.0267, 0.0670][0.0267,\,0.0670] 32/I/0/1 0.1625 0.4175 0.9450 0.0375 0.9225 0.0425 0.00950.0095 [0.0229, 0.0609][0.0229,\,0.0609] 32/I/0.6/0.25 0.0975 0.4900 0.9125 0.0775 0.8925 0.0850 0.01340.0134 [0.0551, 0.1079][0.0551,\,0.1079] 32/I/0.6/1 0.0750 0.4450 0.9375 0.0450 0.9025 0.0625 0.01040.0104 [0.0287, 0.0700][0.0287,\,0.0700]

这是一张已知真值的校准结果表,用来检查度量方法本身是否可靠。

1. 这段在干什么

在已知真值的合成场景下,报告各组(不同 GG/size/ρ/σ 配置)的统计量与检验结果,验证度量流程的可靠性。

2. 需要解释的地方

  • GG/size/ρ/σ:四列配置标签,此处没展开各自含义。
  • Type-I / 95% CI / MCSE:领域常识——第一类错误率、95% 置信区间、蒙特卡洛标准误,衡量估计稳定性。
  • Group t / Group bootstrap t:两种检验得到的统计量。

3. 值得留意

表头多列在原文里没有逐一定义;Type-I 大多在 0.02–0.08 间波动,判断是否合格需对照论文正文标准,这段没明说。

b174B/I: balanced/imbalanced. Type-I nominal .025. MCSE is p⁡(1−p)/400\sqrt{p(1-p)/400} (worst possible .025); Wilson intervals are pointwise Monte Carlo intervals, not simultaneous bounds over 16 scenarios.

这段是 F.1 的图注/表注,交代上一段那串数字读法。

关键概念(领域常识):

  • B/I:平衡/不平衡数据集。
  • Type-I nominal .025:一类错误名义水平 2.5%(单侧)。
  • MCSE:蒙特卡洛标准误,公式 √(p(1−p)/400),400 是重复次数;最坏情况取 p=.025。

值得留意:Wilson 区间是逐点的,不是覆盖 16 个场景的同时置信界——别当成多重比较校正后的区间读。

b175Endpoint / class Maps Truth Bias range Bias MCSE range Null loss 32 00 [−×10−3,×10−3][-8.41\!\times\!10^{-3},\,4.95\!\times\!10^{-3}] [×10−4,×10−3][3.84\!\times\!10^{-4},\,5.21\!\times\!10^{-3}] Null loss 128 00 [−×10−3,×10−3][-8.42\!\times\!10^{-3},\,5.48\!\times\!10^{-3}] [×10−4,×10−3][3.82\!\times\!10^{-4},\,5.14\!\times\!10^{-3}] Beneficial loss 32 0.080.08 [−×10−3,×10−3][-7.32\!\times\!10^{-3},\,4.58\!\times\!10^{-3}] [×10−4,×10−3][3.50\!\times\!10^{-4},\,5.14\!\times\!10^{-3}] Beneficial loss 128 0.080.08 [−×10−3,×10−3][-7.47\!\times\!10^{-3},\,5.06\!\times\!10^{-3}] [×10−4,×10−3][3.47\!\times\!10^{-4},\,5.06\!\times\!10^{-3}] Harmful loss 32 −0.08-0.08 [−×10−3,×10−3][-4.17\!\times\!10^{-3},\,6.55\!\times\!10^{-3}] [×10−4,×10−3][5.70\!\times\!10^{-4},\,5.48\!\times\!10^{-3}] Harmful loss 128 −0.08-0.08 [−×10−3,×10−3][-4.68\!\times\!10^{-3},\,6.51\!\times\!10^{-3}] [×10−4,×10−3][5.65\!\times\!10^{-4},\,5.40\!\times\!10^{-3}] Squared distance 32 0.080.08 [−×10−4,×10−4][-4.19\!\times\!10^{-4},\,7.09\!\times\!10^{-4}] [×10−4,×10−4][1.69\!\times\!10^{-4},\,7.60\!\times\!10^{-4}] Squared distance 128 0.080.08 [−×10−4,×10−4][-3.08\!\times\!10^{-4},\,5.67\!\times\!10^{-4}] [×10−4,×10−4][1.68\!\times\!10^{-4},\,7.49\!\times\!10^{-4}]

这段是 F.1 的已知真值检验表:对 8 种损失/样本量组合,报告偏差范围与 MCSE 范围,用来证明度量本身无系统性偏差。

关键概念(领域常识):

  • Truth:该场景下已知的正确答案(0 或 ±0.08)。
  • Bias:估计值减真值。范围都跨 0,说明无系统偏。
  • MCSE:蒙特卡洛标准误,衡量重复采样的随机波动。

值得留意:Bias 范围远小于 MCSE 范围,说明误差主要来自随机性而非偏差;Squared distance 的量级整体小一个数量级——但这与损失尺度不同有关,原文未加说明。

b176All CI methods share the same point estimator. Bias MCSE is the empirical SD of estimation error divided by 400\sqrt{400}; full-precision per-scenario bias, RMSE, coverage and Wilson intervals are in the portable CSV.

讲解

1. 这段在干什么

补充说明不确定性检查的技术细节:交代置信区间(CI)方法的估计量来源和偏差误差的计算方式,并指引完整数据的位置。

2. 需要解释的地方

  • CI methods:置信区间方法,即各种给估计值划出可信范围的手段。
  • Point estimator:点估计量。意思是所有 CI 方法算出的中心值其实是同一个,差别只在区间的宽度上。
  • Bias MCSE:偏差的蒙特卡洛标准误(领域常识:MCSE 衡量模拟重复次数带来的估计波动)。这里用"估计误差的经验标准差 ÷ √400"来算,说明做了 400 次重复。
  • Wilson intervals:一种计算比例置信区间的常用公式(领域常识),比正态近似在小样本下更稳。

3. 值得留意

  • "share the same point estimator" 是易漏的关键:区间不同但中心一致,比较时才公平。
  • 400 这个重复次数只在分母里出现,作者没解释为何选它;完整数字(bias、RMSE、coverage)这段没给,只指向了 portable CSV。

b177Class Amplitude Mean effect Mean QQ Support-rate range Invariant 00 00 00 [0, 0][0,\,0] Invariant 0.20.2 00 0.087920.08792 [0, 0][0,\,0] Invariant 2.02.0 00 8.792308.79230 [0, 0][0,\,0] Sensitive/no benefit 00 −0.00067-0.00067 0.077440.07744 [0, 0.007][0,\,0.007] Sensitive/no benefit 0.20.2 −0.00067-0.00067 0.120340.12034 [0, 0.003][0,\,0.003] Sensitive/no benefit 2.02.0 −0.00067-0.00067 8.792748.79274 [0, 0][0,\,0] Beneficial 00 0.079460.07946 0.082260.08226 [0, 1.000][0,\,1.000] Beneficial 0.20.2 0.079460.07946 0.122600.12260 [0, 0.110][0,\,0.110] Beneficial 2.02.0 0.079460.07946 8.792798.79279 [0, 0][0,\,0] Harmful 00 −0.07961-0.07961 0.082180.08218 [0, 0][0,\,0] Harmful 0.20.2 −0.07961-0.07961 0.122760.12276 [0, 0][0,\,0] Harmful 2.02.0 −0.07961-0.07961 8.792488.79248 [0, 0][0,\,0]

讲解

1. 这段在干什么

这是一张「已知真值」的对照表:人工构造的四种输入类别(Invariant / Sensitive-no benefit / Beneficial / Harmful),检查审计指标能不能把它们区分开——即方法的「体检」。

2. 需要解释的地方

  • Class:该输入被预设的真实角色。
  • Amplitude:扰动强度(0、0.2、2.0),看结论是否随强度稳定。
  • Mean effect:平均预测贡献;Beneficial 为正、Harmful 为负,符合预期。
  • Mean QQ / Support-rate range:量化不确定性的指标(领域常识:QQ 即分位数-分位数对比),range 越小越自信。

3. 值得留意

Invariant 与 Sensitive 的 Mean effect 都是 0,靠生效与否区分;Amplitude 2.0 时 QQ 普遍飙到 8.79 而 support-rate 全塌成 [0,0]——高幅度下该指标基本失效,作者没点破。

b178Means average all 16 scenarios; support ranges remain scenario-specific. This replays the historical formula on negative squared-loss scores, not historical PCC. Changing orthogonal context changes its reference threshold while keeping the chemical contrast fixed. Non-support is not a false negative for the different proposition τ>0\tau>0; both map budgets and all original statuses remain in the portable CSV.

这段在干什么

交代表格数字的统计口径与前提,防止读者把状态判定误读成 PCC 上的结论。

需要解释的地方

  • 均值 vs 范围:均值把 16 个场景合起来算,support 范围仍按场景各自报。
  • replays 历史公式:领域常识——这里用的是负平方损失分数,不是 PCC,所以不能拿旧 PCC 结果直接对照。
  • 正交上下文 / 参照阈值:换上下文会挪动判定门槛,但化学对比本身不动。

值得留意

非 support 不等于命题 τ>0 的假阴性——两者不是一回事,别混读。

b179The group-aware methods still under-cover in some scenarios; the bootstrap is particularly weak with few groups, and positive null declarations can exceed the nominal one-tail rate. Independent normal donors, independent groups, and correct context matching are explicit simulation assumptions, not demonstrations of real biological conditional exchangeability or support. Finite reused donor pools and dependent biological groups require additional dependence handling. More maps reduce conditional Monte Carlo error but do not create more biological replicates.

讲解

1. 这段在干什么

这是对模拟实验局限性的自我批评:作者承认组感知方法在部分场景仍覆盖不足,并列出这些结果所依赖的假设与不能替代的东西。

2. 需要解释的地方

  • under-cover:领域常识,指置信区间/检验的实际覆盖率低于名义水平,即"低估不确定性"。
  • positive null declarations / nominal one-tail rate:把本该是零的结果判成"正",且这种误判率超过名义单尾水平。
  • bootstrap:重采样估计不确定性的方法,组数少时表现差。
  • conditional Monte Carlo error:模拟本身的随机误差,靠多加"maps"来降。

3. 值得留意

作者明确区分:加 maps 只减模拟误差,不增加生物学重复。这些独立正态、独立分组、上下文匹配都是模拟假设,不等于真实生物学的可交换性或支持度。

Appendix G Complete input-intervention results

b181The complete intervention matrix reports search performance, held-out predictive performance, compound-identity effect, control-profile effect, target-loss gain, and input-use status in separate columns. Rows summarize the full pre-specified model sets, and −⁣−-- marks contrasts that cannot be identified under the protocol. Response-block decompositions are diagnostic views of the same selected models and fitted checkpoints. Figure 6 follows the compound and control effects of the falsification-guided checkpoints across both held-out folds.

这段在干什么:交代完整干预矩阵里有哪些列、行汇总什么,并说明图表的分工——图6只看伪造引导检查点的两类效应。

需要解释的地方:

  • 干预矩阵:一张表,每行一个模型,每列一种指标。
  • compound-identity / control-profile effect:这是领域术语,指化合物身份带来的效应与对照(溶剂等)本身带来的效应。
  • −−:表示该对比在当前实验方案下无法识别,即数据不足以区分。
  • response-block decompositions:对同一批已选模型和检查点的诊断性拆解,不是新结果。

值得留意:矩阵列里包含 input-use status,这正是论文审计的对象;而分解只是"diagnostic views",别误当成独立证据。

b182A. Selection record Task Search setting Selected source Fold-3 PCC BBBC036 Prediction-score HarmonyCell 0.2985 BBBC036 Path-constrained 10-endpoint cohort 0.2925 ±\pm 0.0019 BBBC036 Falsification-guided 10-endpoint cohort 0.3006 ±\pm 0.0006 BBBC047 Prediction-score CellScientist 0.3231 BBBC047 Path-constrained 10-endpoint cohort 0.3008 ±\pm 0.0012 BBBC047 Falsification-guided 10-endpoint cohort 0.3238 ±\pm 0.0002

这段是附录里的一张选型结果表,记录每个数据集在三种任务搜索设置下选出的模型及其 Fold-3 PCC(皮尔逊相关系数,领域常识:衡量预测与真实值的相关程度,越高越好)。

关键概念:「Selection record」= 选型记录;三种设置——Prediction-score(只看预测分)、Path-constrained(路径约束,限定为 10-endpoint cohort)、Falsification-guided(用证伪引导筛选)。

值得留意:三个数据集/设置中,Falsification-guided 的 PCC 一致最高(BBBC036 为 0.3006,BBBC047 为 0.3238);且它带 ± 方差,说明跨多次运行平均,而 Prediction-score 只有单点值。表中数据集仅两个(BBBC036、BBBC047),统计口径不统一,勿过度外推。

b183B. Fold 4: held-out input-use evidence Search setting PCC EcmpdE_{\rm cmpd} EcmpdℒE^{\mathcal{L}}_{\rm cmpd} EctrlE_{\rm ctrl} Q/U/I BBBC036 Prediction-score 0.3110 ±\pm 0.0016 0.0000 ±\pm 0.0000 0.00000 ±\pm 0.00000 0.1693 ±\pm 0.0033 0/5/0 Path-constrained 0.3034 ±\pm 0.0029 0.0007 ±\pm 0.0006 0.00003 ±\pm 0.00003 0.1592 ±\pm 0.0050 0/10/0 Falsification-guided 0.3167 ±\pm 0.0013 0.0091 ±\pm 0.0010 0.00039 ±\pm 0.00004 0.1716 ±\pm 0.0042 0/10/0 BBBC047 Prediction-score 0.3035 ±\pm 0.0008 0.0000 ±\pm 0.0000 0.00000 ±\pm 0.00000 0.1932 ±\pm 0.0044 0/5/0 Path-constrained 0.2822 ±\pm 0.0017 0.0083 ±\pm 0.0016 0.00024 ±\pm 0.00005 0.1388 ±\pm 0.0052 0/6/4 Falsification-guided 0.3060 ±\pm 0.0004 0.0053 ±\pm 0.0002 0.00018 ±\pm 0.00001 0.1988 ±\pm 0.0027 0/10/0

这段是附录 G 的结果细表,把 Fold 4 在 BBBC036、BBBC047 两个数据集上的输入使用证据,按三种搜索策略逐行列出,供前面正文结论做支撑。

关键术语(领域常识):PCC 是预测与真值的相关性,越高越好;E_cmpd 等是一组“输入干预后模型输出变化”的指标,越接近 0 说明模型越不依赖该输入;Q/U/I 指被判定为合格/未用/需干预的输入个数。

值得留意:三种搜索里只有 Falsification-guided 在两个数据集上同时保住了较高 PCC 与较低 E_cmpd,且 Q/U/I 都是 0/10/0;Path-constrained 在 BBBC047 上 PCC 掉到 0.2822、I 出现 4,是三行里唯一有非零“需干预”的情况。

b184C. Fold 5: held-out input-use evidence Search setting PCC EcmpdE_{\rm cmpd} EcmpdℒE^{\mathcal{L}}_{\rm cmpd} EctrlE_{\rm ctrl} Q/U/I BBBC036 Prediction-score 0.2932 ±\pm 0.0020 0.0000 ±\pm 0.0000 0.00000 ±\pm 0.00000 0.1516 ±\pm 0.0035 0/5/0 Path-constrained 0.2886 ±\pm 0.0018 0.0039 ±\pm 0.0022 0.00016 ±\pm 0.00009 0.1423 ±\pm 0.0050 0/10/0 Falsification-guided 0.2897 ±\pm 0.0008 0.0022 ±\pm 0.0012 0.00009 ±\pm 0.00005 0.1515 ±\pm 0.0037 0/10/0 BBBC047 Prediction-score 0.3153 ±\pm 0.0014 0.0000 ±\pm 0.0000 0.00000 ±\pm 0.00000 0.2004 ±\pm 0.0061 0/5/0 Path-constrained 0.2922 ±\pm 0.0018 0.0074 ±\pm 0.0017 0.00022 ±\pm 0.00005 0.1477 ±\pm 0.0057 0/7/3 Falsification-guided 0.3173 ±\pm 0.0004 0.0048 ±\pm 0.0004 0.00017 ±\pm 0.00001 0.2051 ±\pm 0.0024 0/10/0

这一行是小节 Appendix G 表格中 Fold 5 的收尾数据。

1. 这段在干什么

汇报 5 折交叉验证里第 5 折(Fold 5)的留出输入使用证据,列出 BBBC036、BBBC047 两个数据集下三种搜索设置的表现,补齐前面各折的完整结果。

2. 需要解释的地方

  • held-out:留出集,没参与训练、用来检验的数据(领域常识)。
  • PCC:皮尔逊相关系数,衡量预测与真实值的线性相关。
  • E_cmpd / E_cmpd^L / E_ctrl:三种误差指标,数值越低通常越好。
  • Q/U/I:三类计数,原文没给全称,无法确指。
  • ± 后是标准差,表示多次运行的波动。

3. 值得留意

BBBC047 行 Path-constrained 的 Q/U/I 是 0/7/3,而其他多为 0/10/0,是唯一非零 U/I 的行,别读漏。

b185A. Predictive performance Task Fold Control-only PCC Full PCC Δ\Delta Full−-Control BBBC036 Fold 4 0.3121 ±\pm 0.0011 0.3167 ±\pm 0.0013 0.0046 ±\pm 0.0006 BBBC036 Fold 5 0.2917 ±\pm 0.0006 0.2897 ±\pm 0.0008 -0.0020 ±\pm 0.0009 BBBC047 Fold 4 0.3023 ±\pm 0.0004 0.3060 ±\pm 0.0004 0.0037 ±\pm 0.0003 BBBC047 Fold 5 0.3145 ±\pm 0.0005 0.3173 ±\pm 0.0004 0.0028 ±\pm 0.0004

这张表在汇报干预实验的预测性能:对比只用对照组(Control-only)和用完整输入(Full)的 PCC,看差值 Δ。

术语:PCC 是皮尔逊相关系数,衡量预测与真实值的线性相关,越高越好(领域常识)。Fold 是交叉验证的折。±后面是标准差。

值得留意:Δ 列才是重点,但它普遍很小(0.0028~0.0046),Fold 5 甚至出现负值 −0.0020,说明"用全输入"并不总是更好——这正是审计要 Falsify 的地方。

b186B. Input-use evidence Task Fold EcmpdE_{\rm cmpd} EcmpdℒE^{\mathcal{L}}_{\rm cmpd} EctrlE_{\rm ctrl} Model Q/U/I BBBC036 Fold 4 0.0091 ±\pm 0.0010 0.00039 ±\pm 0.00004 0.1716 ±\pm 0.0042 0/10/0 BBBC036 Fold 5 0.0022 ±\pm 0.0012 0.00009 ±\pm 0.00005 0.1515 ±\pm 0.0037 0/10/0 BBBC047 Fold 4 0.0053 ±\pm 0.0002 0.00018 ±\pm 0.00001 0.1988 ±\pm 0.0027 0/10/0 BBBC047 Fold 5 0.0048 ±\pm 0.0004 0.00017 ±\pm 0.00001 0.2051 ±\pm 0.0024 0/10/0

这段是附录 G 的输入干预结果表,逐行列出各数据集划分在干预后的误差指标,用来支撑「输入是否被模型真正使用」的审计结论。

关键概念:领域常识——E_cmpd 是干预化合物后预测变化的大小,E_cmpd^L 是它的似然/对数形式,E_ctrl 是对照误差;Q/U/I 通常指 Quiet/Used/Ignored 一类判定计数(此处未展开说明)。

值得留意:四行 Q/U/I 全是 0/10/0,即全部判定为"使用",但各折误差数值差异不小,且 E_ctrl(0.15–0.21)远大于 E_cmpd(千分位级)。这说明什么,原文这段没说,需结合上下文判断。

b187Predictor Global PCC [95% CI] Chemical PCC drop [95% CI] Target-loss gain [95% CI] BBBC036 / Fold 4 Falsification-guided 0.3167 [0.3154, 0.3181] 0.0091 [0.0081, 0.0101] 0.00039 [0.00034, 0.00043] Residual MLP 0.3056 [0.3027, 0.3085] 0.0036 [0.0028, 0.0044] 0.00015 [0.00012, 0.00019] Joint Ridge 0.0892 [0.0867, 0.0917] 0.0153 [0.0147, 0.0159] 0.00299 [0.00290, 0.00308] BBBC036 / Fold 5 Falsification-guided 0.2897 [0.2890, 0.2905] 0.0022 [0.0010, 0.0034] 0.00009 [0.00004, 0.00014] Residual MLP 0.2861 [0.2829, 0.2892] 0.0007 [0.0001, 0.0013] 0.00003 [0.00000, 0.00005] Joint Ridge 0.1114 [0.1092, 0.1136] 0.0336 [0.0328, 0.0344] 0.00569 [0.00557, 0.00582] BBBC047 / Fold 4 Falsification-guided 0.3060 [0.3056, 0.3064] 0.0053 [0.0051, 0.0055] 0.00018 [0.00017, 0.00019] Residual MLP 0.3031 [0.3007, 0.3055] 0.0119 [0.0094, 0.0145] 0.00042 [0.00033, 0.00050] Joint Ridge 0.0774 [0.0765, 0.0783] 0.0082 [0.0079, 0.0085] 0.00420 [0.00416, 0.00423] BBBC047 / Fold 5 Falsification-guided 0.3173 [0.3169, 0.3177] 0.0048 [0.0044, 0.0052] 0.00017 [0.00015, 0.00018] Residual MLP 0.3154 [0.3135, 0.3174] 0.0113 [0.0095, 0.0132] 0.00040 [0.00033, 0.00047] Joint Ridge 0.0704 [0.0693, 0.0715] 0.0052 [0.0050, 0.0053] 0.00394 [0.00393, 0.00395]

这一节把附录前面汇总表里 BBBC047 Fold 5 那两行,按四个数据集/折展开成完整三列数值。

这段在干什么:把「完整输入干预结果」以表格形式铺开,列出各模型在四个数据集/折上的三项指标及 95% 置信区间。

需要解释的地方:

  • Predictor Global PCC:预测器整体与真实值的相关性,越高越好。
  • Chemical PCC drop:换掉化学输入后相关性掉了多少,掉得多说明模型更依赖化学信息。
  • Target-loss gain:另一种衡量某输入对预测贡献的指标。
  • 三者后面方括号都是 95% 置信区间,领域常识:区间窄表示估计稳。

值得留意:Joint Ridge 的 PCC 明显低(0.07–0.11),但 Chemical PCC drop 反而最大——说明它整体预测差,却更依赖化学输入。Fold 5 的 drop 普遍比 Fold 4 小。

b188Fold Perturbation Context Interaction CP L1000 CP L1000 CP L1000 Path-constrained / BBBC036 F4 0.0012 ±\pm 0.0009 0.0004 ±\pm 0.0005 0.0250 ±\pm 0.0023 0.2164 ±\pm 0.0057 0.0003 ±\pm 0.0004 -0.0003 ±\pm 0.0006 F5 0.0008 ±\pm 0.0004 0.0053 ±\pm 0.0031 0.0290 ±\pm 0.0024 0.1935 ±\pm 0.0056 0.0005 ±\pm 0.0002 0.0013 ±\pm 0.0008 Path-constrained / BBBC047 F4 0.0130 ±\pm 0.0026 0.0020 ±\pm 0.0002 0.0911 ±\pm 0.0037 0.2178 ±\pm 0.0054 0.0026 ±\pm 0.0012 0.0007 ±\pm 0.0002 F5 0.0118 ±\pm 0.0026 0.0013 ±\pm 0.0005 0.1005 ±\pm 0.0054 0.2251 ±\pm 0.0054 0.0011 ±\pm 0.0012 -0.0001 ±\pm 0.0003 Falsification-guided / BBBC036 F4 0.0141 ±\pm 0.0023 0.0069 ±\pm 0.0007 0.0248 ±\pm 0.0015 0.2363 ±\pm 0.0041 0.0000 ±\pm 0.0000 0.0001 ±\pm 0.0000 F5 -0.0014 ±\pm 0.0016 0.0039 ±\pm 0.0013 0.0284 ±\pm 0.0016 0.2088 ±\pm 0.0035 0.0001 ±\pm 0.0000 0.0002 ±\pm 0.0001 Falsification-guided / BBBC047 F4 0.0088 ±\pm 0.0003 0.0014 ±\pm 0.0001 0.1338 ±\pm 0.0031 0.2759 ±\pm 0.0013 0.0002 ±\pm 0.0001 0.0001 ±\pm 0.0000 F5 0.0079 ±\pm 0.0007 0.0013 ±\pm 0.0002 0.1421 ±\pm 0.0025 0.2806 ±\pm 0.0014 -0.0000 ±\pm 0.0001 0.0001 ±\pm 0.0000

这段是附录 G 的完整输入干预结果表,铺列各模型/数据集组合下三种扰动(Fold、Context、Interaction)在 CP 与 L1000 两种读数上的效应值(均值±标准差),供正文结论查证。

关键概念:CP、L1000 是两种细胞表型读数(领域常识:L1000 是基因表达谱平台);Path-constrained、Falsification-guided 是两种模型发现路径。

留意:分为 BBBC036/BBBC047 两个数据集、F4/F5 两个折叠;Path-constrained 组的 Interaction 值近零甚至为负,而 Falsification-guided 组普遍更小。表格不带解读,具体论断需回正文。

b189Task Fold Δdose\Delta_{\rm dose} Compound Control Dose Compound×\times control Path-constrained BBBC036 4 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC036 5 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC047 4 – 0/6/4/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC047 5 – 0/7/3/0 10/0/0/0 0/0/0/10 0/10/0/0 Falsification-guided BBBC036 4 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC036 5 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC047 4 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0 BBBC047 5 – 0/10/0/0 10/0/0/0 0/0/0/10 0/10/0/0

这段在干什么

这是附录 G 的干预结果表,逐行比较三种建模流程(Task、Fold、Path-constrained、Falsification-guided)下,各扰动类型干预实验的命中计数。

需要解释的地方

  • 干预类型:Δdose(剂量)、Compound(化合物)、Control(对照)、Compound×control(化合物与对照交互)。
  • 冒号组数字:格式为 x/y/z/w,表示按干预类别分配的结果计数(顺序需回正文核对,这段未说明)。

值得留意

  • Path-constrained 行与 Falsification-guided 行在 BBBC047 上数字不同:前者是 0/6/4/0、0/7/3/0,后者是 0/10/0/0;作者没明说这种差异的原因。

Appendix H All-candidate source checking

b191Table 51 applies the deterministic checker to every retained prediction-score candidate and separates execution failure from checker abstention. Resolution coverage is the fraction of checked candidates whose cited source locations are all resolved. For candidates without a structured manifest declaring input use, the checker reports findings supported by pre-specified source locations. The one-key attention certificate is a code-level finding; complete-model replacement tests provide the corresponding behavioral evidence.

这段在干什么

把确定性检查器跑遍所有保留下来的预测分数候选,区分「执行失败」和「检查器弃权」两类情况。

需要解释的地方

  • 确定性检查器:按固定规则跑的源码检查工具,不看运气、每次结果一样。
  • 执行失败 vs 弃权:前者是代码跑不起来;后者是代码跑了但检查器不下判断(领域常识:静态/程序分析工具对某些情况会拒答)。
  • resolution coverage(解析覆盖率):被查候选里,引用的源码位置全部能定位到的比例。
  • 无结构化 manifest:候选没有明确声明「我用了哪些输入」,检查器只能靠预设的源码位置给结论。
  • one-key attention certificate:一个代码层面的发现;对应的行为证据来自 complete-model replacement 测试。

值得留意

代码级发现和行为证据是两回事,作者明确把二者配对——别把前者当成行为已证的结论。

b192Candidate slots Code checker Coverage Policy Total Exec. Checked Resolved Singleton Abstain Resolved/ checked Checked/ slots BBBC036 CellScientist 100 99 99 60 15 39 60.6% 99.0% AIDE 100 100 100 100 0 0 100.0% 100.0% CellForge 100 91 91 91 0 0 100.0% 91.0% HarmonyCell 100 95 95 94 0 1 98.9% 95.0% BBBC047 CellScientist 100 94 94 48 17 46 51.1% 94.0% AIDE 100 99 99 99 0 0 100.0% 99.0% CellForge 100 94 94 94 0 0 100.0% 94.0% HarmonyCell 100 95 95 94 0 1 98.9% 95.0%

好,我们就盯着这张表看,逐列捋一遍。

这段在干什么

这是附录里的数据表:列出两个数据集(BBBC036、BBBC047)下四个模型(CellScientist、AIDE、CellForge、HarmonyCell)的「全候选源码检查」结果,用覆盖率、检查数、解析数等指标评估各模型所提 input-use claim 的质量。

需要解释的地方

  • Candidate slots:候选位置总数,这里一律 100,即每模型都提了 100 条待查 claim。
  • Coverage / Checked:被实际检查到的比例与条数。
  • Resolved:能判定结论的条数;剩下的是 Singleton(只出现一次、无法比对)或 Abstain(弃权)。
  • Resolved/checked:判定成功率;Checked/slots:覆盖率。
  • 这是领域常识:AIDE 两处都 100/100,接近满分。

值得留意

CellScientist 在 BBBC047 上 Resolved/checked 只有 51.1%,且 Singleton 高达 17——说明近半数 claim 无法比对验证,是被「单例」卡住的,而非检查失败。

正文这段没提到这些数的成因,别替它归因。

b193Search Slots Exec. Resolved Singleton Abstain Resolve % Exec. % BBBC036 Prediction-score 100 99 60 15 39 60.6% 99.0% Path-constrained 100 100 100 0 0 100.0% 100.0% BBBC047 Prediction-score 100 94 48 17 46 51.1% 94.0% Path-constrained 100 100 100 0 0 100.0% 100.0%

这段在干什么

这是附录里的全候选源检查结果表,逐行给两个数据集(BBBC036、BBBC047)在两种搜索策略下的执行与解析统计,用来支撑“源检查”的审计结果。

需要解释的地方

  • Search:搜索/查找策略名。
  • Slots:候选槽位数。
  • Exec.:实际执行数。
  • Resolved:成功解析出的数。
  • Singleton:只解析出唯一结果的数。
  • Abstain:放弃/无法解析的数。
  • Resolve % / Exec. %:解析率与执行率。

值得留意

Path-constrained 两数据集都是 100/100/100、0 弃权、100%,说明该策略几乎不缺执行;而 Prediction-score 在 BBBC047 上仅解析 48、弃权 46,解析率 51.1%,差异明显。

b194A deterministic hash selects model sources within each checker category for source-interface inspection. The inspection is restricted to pre-specified source locations and public fusion interfaces; fixed behavioral replacements remain applicable across all code-check categories.

讲解

这段在干什么:交代附录 H 的检查流程——用哈希挑出各类别里的模型源码,再送到指定位置和公开融合接口做检查。

需要解释的地方:

  • 确定性哈希(deterministic hash):同一输入必得同一结果的抽选方式,用来挑样本、保证可复现。
  • checker category:先把代码检查分门别类,每类各抽一个源。
  • 公开融合接口(public fusion interfaces):领域常识——模型把多个信号合并输出的那层接口。
  • fixed behavioral replacements:固定行为替换,检查时用的一套不随类别变的替代方案。

值得留意:检查范围被明确限定(restricted)在预先指定的位置和公开接口,即不覆盖全部代码;而行为替换对所有类别一视同仁,是流程里唯一"通吃"的部分。

b195A linked sample of 48 real candidate sources extends testing beyond the one-key attention case. Of these, 47 change their predictions after compound replacement at both boundaries; 20 have positive target-loss scaffold intervals at both. An implemented multi-tower source is exactly invariant. Source consumption, fitted dependence, and target benefit thus separate in diverse generated code (Appendix M).

这段在干什么

用 48 个真实候选源码做更大范围的抽样检验,证明上一段“行为替换适用于所有代码检查类别”的说法不止于单键注意力这一个案例。

需要解释的地方

  • compound replacement at both boundaries:在两端对复合结构做替换(领域常识:边界指输入端和输出端的接口处)。
  • scaffold intervals:目标损失上“脚手架”式的区间,即呈现出结构性支撑区间。
  • multi-tower:多塔结构(推荐/检索领域的常见架构)。
  • separate:三种性质(源码消耗、拟合依赖、目标收益)互不绑定、可彼此分离。

值得留意

48 个里 47 个替换后预测会变、20 个两端都有正区间——作者用这些数字支撑“三者可分离”的结论;但只有 1 个实现的多塔源码完全不变,样本极小。数字差异的原因作者指向附录 M,本段未展开。

Appendix I Prediction-space dependence

b197The prediction-space analysis measures normalized RMS change and cosine change between f⁡(c,p,a)f(c,p,a) and predictions obtained after fixed input permutations. Table 54 reports prediction-change RMS divided by the target RMS after centering each output feature across evaluation rows, with a denominator floor of 10−1210^{-12}. Both RMS quantities average over rows and features within the reported response block. This complements target-based effects: near-zero distance indicates negligible prediction change on the tested replacements, whereas nonzero distance with near-zero target-loss gain indicates sensitivity with little measured benefit on the observed response. The calculation requires inference only over fixed checkpoints.

讲解

1. 这段在干什么

提出一种「预测空间」审计指标:看固定输入被替换后,模型输出本身变了多少,用来补足只看目标损失(target-loss)的视角。

2. 需要解释的地方

  • *normalized RMS change*:预测变化的均方根,再按目标本身的 RMS 归一化,方便跨响应块比较(领域常识)。
  • *cosine change*:用余弦距离看预测方向的偏移。
  • *fixed input permutations*:把输入做固定方式的置换/替换,再跑一遍推理。
  • *denominator floor 10⁻¹²*:防止分母为零。
  • *fixed checkpoints*:只做推理,不重训练。

3. 值得留意

近零距离≈替换没影响;非零距离但目标增益也近零,说明模型敏感却没带来实际收益——这是作者区分"敏感性"与"有用性"的关键。

b198Model Compound replacement Control replacement Global CP L1000 Global BBBC036 CellScientist 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.3112 ±\pm 0.0246 Chemical-only 0.0560 ±\pm 0.0013 0.0580 ±\pm 0.0014 0.0552 ±\pm 0.0013 0.0000 ±\pm 0.0000 Control-only 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.3390 ±\pm 0.0199 BBBC047 CellScientist 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.3593 ±\pm 0.0207 Chemical-only 0.1253 ±\pm 0.0066 0.1467 ±\pm 0.0093 0.1025 ±\pm 0.0055 0.0000 ±\pm 0.0000 Control-only 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.0000 ±\pm 0.0000 0.4011 ±\pm 0.0199

这段是一张表格式的数据对比,用来审计模型预测到底依赖哪部分输入,是附录里支撑"预测空间依赖"的证据。

在干什么:用替换实验(换掉化合物 / 换掉对照)看各模型的预测变化,检验预测究竟依赖什么。

关键概念:CP 与 L1000/BBBC036 是任务或数据集名,数值是评分(带标准差)。"Compound replacement"是把化合物换掉,"Control replacement"是换掉对照组——领域常识:把真正有用的输入换掉,分数应明显下降。

容易读漏:CellScientist 在前两列(换化合物、换对照)全是 0.0000,却在 L1000 和 BBBC036 上有 0.31、0.36 这类非零值;Chemical-only 恰好相反——前几列有值,BBBC036 一列全是 0.0000。这说明两者依赖的来源不同,作者就此没在正文里多解释。

Appendix J Representative discovery records

b200The main text follows one model from selection through source checking and held-out replication. This section complements that input-use analysis with two preselected trajectory-level run records: a representative successful search and a search containing code repair. Each record preserves the initial hypothesis, decisive score changes, component edits, failures, repair actions, prompt and source hashes, and final selection. The complete ten-slot traces show both successful and failed execution paths.

讲解

1. 这段在干什么

承上启下:正文只跟了一个模型的完整流程,附录这里补两个「轨迹级」记录——一个成功搜索、一个含代码修复,用来佐证输入使用分析。

2. 需要解释的地方

  • *held-out replication*:留出集复现,领域常识,指用没参与调参的数据重跑结果。
  • *输入使用(input-use)*:模型到底用了哪些输入特征。
  • *轨迹级记录*:整个搜索过程逐步的日志,而非只看最终模型。

3. 值得留意

记录里同时保留了哈希值(prompt/source)和失败路径——作者强调失败样本也是证据,不只是成功案例。原文未说这两个记录具体结论。

b201The two cases are selected using Fold-3 run records only: a lower-median BBBC036 trajectory and the highest-scoring BBBC047 trajectory among those with at least one recorded candidate-code repair. Their records include the initial program, complete candidate sequence, parent hashes, structured decision context, candidate source, execution and repair outcomes, provider request/response hashes, and the versioned prompt constructor that defines the generation input.

这段在干什么

交代附录 J 两个案例的挑选标准与记录内容:用 Fold-3 记录选出 BBBC036 的中位偏下轨迹和 BBBC047 的最高分轨迹。

需要解释的地方

  • Fold-3:交叉验证里的第 3 折,这里指只依据该折的运行记录来选,属领域常识。
  • 轨迹(trajectory):一次完整的候选程序搜索过程。
  • parent hashes:父候选的哈希指纹,用于追溯版本谱系。
  • provider request/response hashes:调用模型时的请求与响应哈希,作可复核凭证。

值得留意

选案例的前提是「至少有一次候选代码修复记录」,即两例都经历过失败—修复;BBBC047 是这类里的最高分。

J.1 Representative multimodal discovery trace

b203The initial candidate is a reproducible concatenation–MLP joint predictor with joint input fusion, a shared response program, a joint readout, response loss, and AdamW optimization; it obtains Fold-3 Global PCC 0.26680.2668. The recorded decision context for the retained final revision is: “Make one final diagnosis-driven revision of the best incumbent; prioritize a complete, executable perturbation-response hypothesis over a hyperparameter-only change.” The selected multi-head cross-modal-attention candidate uses separate control, compound, and dose encoders, cross-modal attention fusion, a shared response program, joint readout, block-weighted MSE, AdamW, and cosine scheduling, reaching 0.29280.2928. The trace contains 11 documented provider calls (53,096 reported tokens), each with a request hash, response hash, status, and token record.

讲解

这段在干什么:展示一次多模态发现轨迹的起点与终点——从初始的拼接-MLP 联合预测器(Fold-3 Global PCC 0.2668)到最终选定的多头跨模态注意力候选(0.2928),并交代支撑这次修订的 provider 调用记录。

需要解释的地方:

  • PCC:皮尔逊相关系数,衡量预测与真实值的线性相关,越高越好(领域常识)。
  • input fusion vs. cross-modal attention fusion:前者把不同模态输入直接拼接后送入网络;后者让不同模态通过注意力机制互相“看”对方再融合。
  • block-weighted MSE:一种按块加权的均方误差损失。
  • provider calls / tokens / hash:即调用外部模型服务的次数、消耗的 token 数,以及每次请求响应的哈希指纹,用于审计可复现性。

值得留意:作者保留的最终修订决策上下文强调的是“做一次完整的、可执行的扰动-响应假设”,而非只调超参——这暗示轨迹的筛选标准偏向机制性探索,而非单纯刷分。

J.2 Repair-containing discovery trace

b205The second trace begins from the same initial program at Fold-3 Global PCC 0.28810.2881. Its structured agenda specifies separate control, compound, and dose encoders; compound-query/control-key-value interaction with a dose residual; a shared response program; joint response prediction; CP/L1000 block-weighted MSE (0.4/0.6); AdamW; and cosine scheduling. Its selected first revision implements this design, resolves an output-shape mismatch by returning a [batch,target][\mathrm{batch},\mathrm{target}] tensor, and reaches 0.32300.3230. The complete trace contains 13 provider calls (63,572 reported tokens); a later slot exhausts all three repair attempts and remains recorded as a failed slot.

讲解

这段在干什么:展示第二条发现轨迹——从同一初始程序出发,按结构化议程做出第一次修订,性能从 0.2881 升到 0.3230,并交代整条轨迹的开销与一处失败。

需要解释的地方:

  • 结构化议程:修订前写好的设计方案清单,规定各编码器怎么分工、损失怎么加权等。
  • PCC:皮尔逊相关系数,领域常识里常用的预测拟合度量,越高越好。
  • 槽(slot)/ 修复尝试:把轨迹分成若干执行位,某个位允许重试三次,用尽即记为失败。

值得留意:性能提升只来自"第一次修订",而整条轨迹有 13 次调用、6.3 万 token,且存在一个彻底失败的槽——作者是在用这条"含修复"的轨迹说明流程并非总能成功。

Appendix K Extended related work

b207Cell Painting, L1000, Perturb-seq, and sci-Plex measure morphological and transcriptional responses across chemical and genetic interventions (Bray et al., 2016; Caicedo et al., 2017; Subramanian et al., 2017; Dixit et al., 2016; Srivatsan et al., 2020). Matched resources and harmonized collections broaden this coverage (Haghighi et al., 2022; Chandrasekaran et al., 2024; Peidli et al., 2024). Predictors use latent-state transfer, compositional models, neural transport, genetic graphs, and foundation models (Lotfollahi et al., 2019; Lotfollahi et al., 2023; Bunne et al., 2023; Roohani et al., 2024; Cui et al., 2024); cycleCDR adds cycle-consistency constraints to learn transferable perturbation representations (Huang and Liu, 2024). CIPHER combines unperturbed-cell covariance with the specified perturbation to predict responses, illustrating the predictive value of baseline cellular structure (Kuznets-Speck et al., 2025). CellAudit asks whether a selected predictor uses the particular perturbation input invoked by its design.

讲解

1. 这段在干什么

这是扩展相关工作,梳理细胞扰动数据集、现有扰动预测方法,最后落到本文的 CellAudit:追问预测器是否真用了它声称的扰动输入。

2. 需要解释的地方

  • Cell Painting / L1000 / Perturb-seq / sci-Plex:领域常识——分别测形态或转录层面对化学、遗传扰动的响应。
  • latent-state transfer、neural transport、foundation models:不同的扰动响应预测建模范式。
  • cycleCDR:加循环一致性约束来学可迁移的扰动表征。
  • CIPHER:把未扰动细胞的协方差与指定扰动结合来预测响应,说明基线细胞结构有预测价值。

3. 值得留意

作者先承认现有方法(含 CIPHER)确用扰动信息,再让 CellAudit 做"使用审计"——重点是验证而非再提一个预测器。

b208PerturbBench and PertEval-scFM standardize predictive comparisons across splits and baselines (Wu et al., 2025; Wenteler et al., 2025). Deep predictors need not outperform linear references, and common metrics can reward systematic variation shared across perturbations (Ahlmann-Eltze et al., 2025; Viñas Torné et al., 2026). Other studies show that well-calibrated predictive metrics and informative foundation-model representations reveal gains over simple baselines (Miller et al., 2025; Hasanaj et al., 2025; Cole et al., 2026). In-the-wild evaluation further examines context and perturbation shifts (Mao et al., 2026). Together, these studies establish the importance of metrics, representations, and transfer conditions. CellAudit tests a complementary property: whether registered input-use claims are supported from cited source computation through fitted dependence to target-relevant predictive contribution.

这段在干什么

做文献定位:先承认已有基准/指标/表征研究已确立重要性,再声明 CellAudit 检验的是一个互补性质——输入使用声明能否从源计算一路支持到预测贡献。

需要解释的地方

  • input-use claims(输入使用声明):模型设计声称"我用了这个扰动输入"——本段要审的就是这话是否真被支撑。
  • 互补(complementary):不是再比谁预测准,而是查"声明—实现—贡献"这条链;这是领域常识层面的区分,非论文实验结论。

值得留意

作者用"Together…"先收束前人三条线(指标、表征、迁移条件),再一句"complementary property"把 CellAudit 摘出去——别把前人的结论误当成本文结果。本段未给任何数字或结论。

b209Scientific agents span laboratory workflows, program search, and iterative machine-learning experimentation (Boiko et al., 2023; Romera-Paredes et al., 2024; Huang et al., 2024; Li et al., 2024; Lu et al., 2026; Jiang et al., 2025). CellScientist, CellForge, and HarmonyCell apply related search procedures to cellular prediction (Li et al., 2026; Tang et al., 2025; Huang et al., 2026). The existing CellScientist policy supplies candidate models in these experiments; CellAudit links their registered input-use claims to deterministic tests of source consumption, fitted dependence, and target-relevant predictive contribution.

这段在干什么

这是在铺相关工作时给本文定位:先列科学智能体的既有工作,再把它自己的 CellAudit 与 CellScientist 区分开。

需要解释的地方

  • Scientific agents:能自主做科研流程的 AI 智能体(领域常识),这里横跨实验流程、程序搜索、迭代式机器学习实验三类。
  • registered input-use claims:模型登记在案的「我用了哪些输入」的说法,是本文要审计的对象。
  • 三层测试:源码消耗 → 拟合依赖 → 目标相关预测贡献,一层比一层更贴近「这输入是否真的有用」。

值得留意

CellScientist 不是被批判的对象,而是提供候选模型的现有策略,CellAudit 是在它上面加审计环节——承接上一段那三种支撑层级。

b210Permutation reliance measures performance changes after input disruption; conditional permutations account for dependence among inputs (Fisher et al., 2019; Chamma et al., 2023). Shortcut and underspecification studies show that benchmark success can coexist with unintended or unstable decision rules (D’Amour et al., 2022; Geirhos et al., 2020; Lapuschkin et al., 2019; DeGrave et al., 2021). Attribution sanity checks, removal-based evaluation, and behavioral suites test explanations and capabilities under targeted interventions (Adebayo et al., 2018; Hooker et al., 2019; Ribeiro et al., 2020). ConceptSMILE audits concept-explanation reliability through input perturbations and local surrogate modeling (Mollapour et al., 2026); POPPER tests free-form hypotheses through agentic sequential falsification with statistical error control (Huang et al., 2025). CellAudit follows one selected cellular-response model and its registered input-use claim from the cited code to complete fitted predictions and later-data evaluation. This links source implementation, fitted dependence, and target-loss changes under registered replacements, distinguishing a blocked pathway, sensitivity without predictive benefit, and a contribution that recurs on new data.

这段在干什么

这是扩展相关工作,把 CellAudit 放进已有方法谱系里做定位:前面综述扰动/归因等已有做法,最后一句才点出本文的不同。

需要解释的地方

  • Permutation reliance:打乱某输入看性能掉多少(领域常识)。
  • Conditional permutations:考虑输入间相关性的扰动。
  • Shortcut / underspecification:基准高分可能来自非预期或不稳定的决策规则。
  • Removal-based evaluation:移除某输入再评估其作用。
  • POPPER:用智能体做序贯证伪、带统计误差控制。
  • CellAudit:把某细胞响应模型从源代码实现、拟合依赖一路追到新数据上的预测贡献。

值得留意

末尾三分类——"被阻断的通路 / 无预测收益的敏感性 / 新数据上复现的贡献"——是本文区别于上述所有方法的关键,容易读漏。

Appendix L Factorial contrasts and stability summaries

b212An explicit input route establishes a possible computation; its fitted contribution is tested with the checkpoint and observed targets held fixed. We adapt permutation reliance and behavioral testing to the registered biological inputs (Fisher et al., 2019; Ribeiro et al., 2020). For a chemical-response model, the 2×22\times 2 test combines the correct or shuffled perturbation with the correct or an alternative control profile. Let mp​cm_{pc} denote the response score with both correct inputs, mp~​cm_{\tilde{p}c} the score after a matched perturbation shuffle, mp​c~m_{p\tilde{c}} the score after a control-profile swap, and mp~​c~m_{\tilde{p}\tilde{c}} the score after both. For a higher-is-better metric mm,

这段在干什么

提出一个 2×2 因子置换检验,用扰动与控制两种输入的正确/打乱组合,测量每个输入对模型输出的实际贡献。

需要解释的地方(领域常识)

  • permutation reliance:打乱某个输入后看性能掉多少,掉得多说明模型依赖它。
  • 2×2 因子设计:把「扰动对错」和「对照对错」两个二元因素交叉成四种条件,分别算分。
  • 四个符号 $m_{pc}$、$m_{\tilde pc}$、$m_{p\tilde c}$、$m_{\tilde p\tilde c}$ 就是这四种条件的响应分数。

值得留意

原文停在「For a higher-is-better metric $m$」——公式被截断,这段没给出具体对比式。此处是在检验「输入是否真被用到」,而非只验证输入路径存在。

b213These are factorial contrasts in predictive score under the registered replacement distribution: EpertE_{\mathrm{pert}} averages perturbation-replacement effects across the two tested contexts, while Ipert,contextI_{\mathrm{pert,context}} measures their nonadditivity in score. For compound inputs we write EcmpdE_{\rm cmpd} and EctrlE_{\rm ctrl} in tables. The loss contrast EcmpdℒE^{\mathcal{L}}_{\rm cmpd} similarly averages the increase after compound replacement over both context states. Each checkpoint uses 32 fixed permutations constructed without response values or model scores. Continuous effects and their intervals are the primary quantitative evidence. For a per-model stability summary, let τ\tau be the 95th percentile of absolute pairwise differences among the corresponding reference scores. A contribution is qualified if its effect’s empirical 2.5th percentile exceeds τ\tau, unsupported if its 97.5th percentile does not exceed τ\tau, and inconclusive otherwise. These labels summarize effect stability relative to map-induced score variation; an unsupported effect may still be positive. The empirical quantiles describe replacement variability within a checkpoint; Appendix B specifies the reference scores and separate across-model uncertainty.

讲解

1. 这段在干什么

定义附录里用的因子对比记号(E、I),并给出一套把「效应大小」与「稳定性」对照的判断规则,用于给每个化合物输入打稳定性标签。

2. 需要解释的地方

  • 因子对比:领域常识,指两个因素(这里是上下文×替换类型)单独作用与交互作用的量化。
  • E vs I:E 是平均效应,I 是非可加性(交互项),即平均之外多出来的部分。
  • τ:参考分数两两绝对差的第 95 百分位,相当于一个「噪声门槛」。
  • qualified / unsupported / inconclusive:拿效应分布的分位数与 τ 比,过门槛叫 qualified,不过叫 unsupported,其余 inconclusive。

3. 值得留意

  • 判断用的是分位数而非均值,所以看的是效应分布相对于噪声的位置。
  • 作者明说 unsupported 不等于效应为负——它可能仍是正的,只是不够稳。
  • 分位数只描述单检查点内的替换波动;跨模型不确定性在 Appendix B,这段没展开。
  • 每检查点用 32 个固定排列,且构造时不看响应值和模型分数(避免泄漏)。

b214Reference-score variation can include variation induced by the other input. Holding the compound effect fixed while increasing this variation can therefore change the status. An existing synthetic check separates invariant, sensitive-without-benefit, target-relevant, and harmful predictors. All 96 target-relevant instances have positive mean effects, while 32 of 96 exceed the reference-relative criterion, identically with 32, 128, or 512 maps. The continuous measurements recover predictive benefit, whereas the categorical rule asks whether that benefit exceeds the specified reference variation. Its threshold is an operational effect scale. The additional 6,400-dataset study in Appendix F.1 tests estimator bias, interval coverage, and sensitivity to the reference scale; it distinguishes conditional replacement uncertainty from population uncertainty and retains the observed finite-group undercoverage.

这一段先给个定位:它是在说明评估标准本身的脆弱性——参考分数的波动会「污染」判定结果。

需要解释的地方:所谓「reference-score variation」指参照基准本身也在变动,而基准一动,同一个 predictor 可能就从「target-relevant」掉到别的类别里。作者举了个领域常识性的例子:把 compound effect 固定住、只抬高波动,状态就会翻。那个「96 个实例、32 个超标」的合成检验,是在验证分类规则和连续测量结果是否一致。

值得留意:两个数字要分清——「96 个全有正效应」但只有「32 个超过标准」,且 32、128、512 三种 map 数下都一模一样。作者没明说的是:连续测量能捞回真实收益,而类别规则只是在问「收益够不够超过那个操作性的阈值」。

M.1 Single-coordinate replay on observed support

b217The frozen replay changes one encoded coordinate while holding the source row’s other inputs fixed. Eligible donors satisfy the recorded metadata constraints and produce an actual input change. Table 57 reports the supported populations separately from the original context-averaged factorial summaries. Score-selected BBBC047 remains exactly invariant; the path-constrained and guided panels retain positive compound effects on both boundaries. On LKCP, dose intervals are positive and compound intervals include zero. For the replay and real-source audit below, 95% intervals resample source groups while conditioning on the selected checkpoint set, donor pools, and 128 replacement maps; BBBC and LINCS use Murcko source scaffolds, and LKCP uses registered BRD identifier groups. The matched-reference and control-bank intervals instead describe paired training-seed variation at the fixed task split. Individual intervals and positive-interval counts are descriptive and unadjusted.

讲解

1. 这段在干什么

交代单坐标重放实验的设定(只改一个编码坐标、其余输入冻结),并报告各面板的区间结果,同时说明不同区间分别刻画什么来源的不确定性。

2. 需要解释的地方

  • 单坐标重放:换掉一个输入坐标、其余不动,看输出怎么变,用来分离单个输入的作用。
  • eligible donors(合格供体):替换用的候选输入,需满足元数据约束且真的改变了输入值。
  • 95% 区间来自重采样:重放与真实源审计的区间是重采样「源组」得到的;匹配参照和控制库的区间则描述固定任务划分下训练种子的配对变异——两者不是一回事。
  • Murcko 骨架 / BRD 分组:这是领域常识,指按化学骨架、按注册化合物标识分组重采样,避免同类样本被拆散。

3. 值得留意

  • 区间是描述性的、未做校正,不能当显著性检验读。
  • BBBC047 被分数选中后完全不变,与其余面板保留正效应形成对照。
  • LKCP 上剂量区间为正、化合物区间含零——两者结论不同,别混为一谈。

b218Original-entry-weighted means: score-selected BBBC047 has 5 entries per fold; path-constrained and guided BBBC047 have 10 each; LINCS and LKCP have 50 each. LKCP’s 50 entries contain 30 unique weight hashes. Path-constrained and guided checkpoints are fixed across the two boundary evaluations. BBBC047 has no registered observed-support dose contrast.

讲解

这段在干什么:交代单坐标回放实验用的数据规模与权重构成,为后续评估设前提。

需要解释的地方:

  • "entries per fold / each":每折(或每组)拿多少条记录来打分。
  • "weight hashes":权重指纹,去重后同权重只算一个。
  • "checkpoints fixed":模型存档在两次边界评估间不变。
  • "无注册的 observed-support dose contrast":这是领域常识说法——该数据集没有登记可比的剂量对照条件。

值得留意:样本量差很大(5 vs 10 vs 50),且 LKCP 的 50 条去重后只剩 30 个唯一权重——实际独立信息比表面少。BBBC047 那句是在说明它无法做该类对比。

M.2 Complete matched input-reference family

b220The four input subsets use the same masked-MLP family, loss, fit budget, and five paired training seeds. Each selected checkpoint is evaluated on both held-out folds. Here cc denotes context, pp compound representation, and aa dose. Dose improves the LINCS reference; its increment is small on BBBC and LKCP. On LKCP Fold 5, the full reference has a positive PCC increment over g⁡(c,a)g(c,a), while its MSE interval includes zero. These trained-subset comparisons describe the fixed reference family; within-checkpoint replacement measures use in the separately selected agent models.

讲解

1. 这段在干什么

交代"匹配输入-参考族"实验的公平性设置,并汇报在固定参考族下四个输入子集的评估结果,属于方法说明兼结果过渡。

2. 需要解释的地方

  • cc/pp/aa:原文自注——上下文、化合物表示、剂量。
  • masked-MLP:领域常识,一种带掩码的多层感知机,此处与损失、训练预算、五个配对种子共用,保证四个子集可比。
  • PCC / MSE 区间:皮尔逊相关与均方误差的置信区间,用来看增量是否显著。
  • held-out folds:留出的交叉验证折。

3. 值得留意

  • 作者自己说剂量增量在 BBBC、LKCP 上"很小",只是 LINCS 上更明显——增量有限。
  • LKCP Fold 5 上 PCC 正增量、MSE 区间却含零,两个指标结论不一致,别只挑显著的看。
  • 最后一句是划界:这只是固定参考族的结论,不能直接推广到 agent 模型。

b221Paired increments and their 95% intervals appear in Table 59; the portable source preserves all per-reference and paired intervals at full precision.

这段是交代数据出处:配对增量及其 95% 区间放在表 59,并声明可移植源文件保留了全部逐参考与配对的完整精度区间。

概念:

  • *paired increments*:配对后的增量,即同一参考下两两比较的差值(领域常识)。
  • *95% intervals*:95% 置信/可信区间,用来表示估计的不确定范围。
  • *portable source*:可移植的源文件,通常指随论文发布的代码/数据仓库。

留意:原文只说「完整精度存在源文件里」,正文表格很可能是四舍五入过的;要拿精确数值得去看那个源文件,而不是表 59。

M.3 Linked source and behavior in real candidates

b223A fixed, dataset–policy-balanced and source-structure-stratified sample links 48 source records to 96 same-checkpoint boundary evaluations. Implemented source consumption coexists with distinct fitted outcomes: an implemented MultiTowerFusion source is exactly invariant on both boundaries, whereas 47 sources change predictions and 20 have positive target-loss intervals on both boundaries. Unresolved source locations also admit complete-model behavioral tests.

讲解

1. 这段在干什么

交代实证样本的构造方式,并给出「实现层面被使用」与「预测层面是否有贡献」的对照结果——即后文审计统计的取样基础。

2. 需要解释的地方

  • stratified / balanced 采样:分层、按数据集与策略配平后再抽样,避免类别偏斜(领域常识)。
  • boundary evaluation:在决策边界附近的评估,看模型输出是否翻转。
  • positive target-loss interval:去掉该来源后目标损失上升,说明它对预测有实质贡献。

3. 值得留意

「implemented」(代码里真被调用)不等于「有贡献」:文中那个 MultiTowerFusion 实现后被用,却在两个边界上都完全不变。另有 47 个改变预测、20 个在两个边界上都有正损失区间。无定位的来源也能做整模型行为测试。

b224The middle behavioral category has positive prediction distance but lacks a positive target-loss interval on at least one fold. Source status describes the cited computation. Counts describe this structurally sampled set and its scripted source classifications. Claim origin remains task_contract where no candidate-authored manifest was emitted.

讲解

1. 这段在干什么

给中间那类候选模型下定义:有正预测距离,但至少在某一折上没有正的目标损失区间。后面几句交代来源状态、计数和申报出处。

2. 需要解释的地方

  • 预测距离(prediction distance):领域常识——预测值与目标之间的偏差度量。
  • 折(fold):交叉验证的一轮划分,此处指行为测试的某个数据划分。
  • source status / claim origin:记录"该计算引用自哪里"和"声明由谁发出"的分类字段。

3. 值得留意

它是在补齐前段"两类边界"之外的中间情形,作者没明说这类算通过还是失败,但"lacks"一词暗示它是被单独标记的一档,不是简单的达标或不达标。

M.4 Control-reference reuse on sci-Plex

b226The existing control-bank comparison holds the response target Y−RY-R fixed and changes whether the context input reuses reference bank RR or uses the other bank. Reference A/B assignment is fixed before fitting and counterbalanced within cell-line, replicate, and control-layout strata (24 plate groups per orientation). Three input sets, two bank assignments, and five paired seeds give 30 fitted checkpoints. All six shared-minus-disjoint PCC intervals include zero. This comparison measures control-reference reuse on sci-Plex.

讲解

1. 这段在干什么

报告一个对照实验(control-bank comparison):固定响应目标 Y−R,只切换上下文输入用的是不是参考库 R,以测量 sci-Plex 上的"control-reference reuse"。

2. 需要解释的地方

  • control-bank / reference bank:领域常识,指用作对照的参考样本池。
  • counterbalanced(平衡分配):把 A/B 两种分配在细胞系、重复、控制布局等分层里均匀打散,避免偏向。
  • PCC:皮尔逊相关系数,衡量两个向量相似度。
  • shared-minus-disjoint PCC 区间含零:共享与不共享两种条件下的相关性差异不显著。

3. 值得留意

六个区间全含零,说明这个对照里参考库复用没有产生显著差异——但作者只说"测量了",没直接下结论。

Appendix N Matched feedback: predictive outcomes and input contribution

b228Predictor F4 PCC ↑\uparrow F4 MSE ↓\downarrow F5 PCC ↑\uparrow F5 MSE ↓\downarrow Anchor g⁡(c,a)g(c,a) 0.9664 ±\pm 0.0001 0.0308 ±\pm 0.0001 0.9728 ±\pm 0.0002 0.0251 ±\pm 0.0001 Score 0.9756 ±\pm 0.0027 0.0225 ±\pm 0.0025 0.9764 ±\pm 0.0015 0.0219 ±\pm 0.0013 Audit 0.9779 ±\pm 0.0004 0.0204 ±\pm 0.0004 0.9778 ±\pm 0.0002 0.0205 ±\pm 0.0002

讲解

这段在干什么:这是一张结果表,横向比三种模型(Anchor / Score / Audit)在 F4、F5 两个预测器上的表现,用 PCC(越高越好)和 MSE(越低越好)两组指标呈现,配合上段"共享减不共享"的对比,展示 Audit 的预测效果。

需要解释的地方:

  • PCC:皮尔逊相关系数,衡量预测值和真实值的线性吻合度,越接近 1 越好(领域常识)。
  • MSE:均方误差,预测偏离真值的平均平方,越小越好(领域常识)。
  • ±:多次运行的均值±标准差,说明结果稳定性。
  • Anchor / Score / Audit:三种模型的命名,具体含义这段没提,需回正文看。

值得留意:Audit 在两个预测器、两个指标上都优于另两者;而标准差普遍很小,说明差距不是偶然波动。

b229The anchor is shared within each seed. F4 and F5 are held-out combination boundaries of the sci-Plex development source.

上一段刚列完几组指标数字,这段是给这些数字补实验设定的。

1. 在干什么:交代这些预测结果所依据的数据划分口径,是表注式的说明。

2. 解释:anchor(锚点)——领域常识里通常指作为基准的那份参照数据或配置;这里说它「在每个 seed 内共享」,即同一随机种子的各次实验都对比同一个锚点,避免锚点差异混进结果。held-out combination boundaries(留出的组合边界)——指刻意留出、不参与训练的那部分组合条件,用于检验外推。

3. 值得留意:F4、F5 是留出的,不是训练用的;而 anchor 反而是共享的。两者性质不同,别把 F4/F5 当成普通数据点读。

b230Set Feedback Chemical PCC drop Chemical loss gain Dose PCC drop Dose loss gain F4 Score 0.0203 ±\pm 0.0046 0.0186 ±\pm 0.0042 0.0068 ±\pm 0.0026 0.0062 ±\pm 0.0024 F4 Audit 0.0240 ±\pm 0.0016 0.0220 ±\pm 0.0015 0.0090 ±\pm 0.0006 0.0082 ±\pm 0.0006 F5 Score 0.0070 ±\pm 0.0016 0.0064 ±\pm 0.0015 0.0039 ±\pm 0.0017 0.0035 ±\pm 0.0015 F5 Audit 0.0086 ±\pm 0.0005 0.0079 ±\pm 0.0005 0.0049 ±\pm 0.0004 0.0045 ±\pm 0.0004

这段在干什么

用一组对照数字,说明在 F4、F5 两个留出边界上,审计(Audit)比原始评分(Score)表现更强。这是正文结论的附表证据。

需要解释的地方

  • PCC drop / loss gain:移除某输入后预测性能的下降幅度,越大说明该输入越关键(领域常识:drop 指性能损失)。
  • ±后数字:多次运行的波动范围,不是误差绝对值。
  • F4 Score vs F4 Audit:同一模型(F4)的两种评估方式对照。

值得留意

四列里每一列,Audit 的值都高于对应的 Score,且±波动普遍更小——即更稳且更高,作者没在原文里点破。

b231Positive values favor the correct input over its legal replacement. Each displayed coordinate metric is positive in 5/5 endpoints in both arms at both boundaries.

讲解

1. 这段在干什么

给上一段那串数字(各端点、各臂的指标)下判读规则:往哪个方向读才算"对",结论一句话说的是 5/5 全为正。

2. 需要解释的地方

  • "correct input over its legal replacement":把模型的真实输入换成同等合法的替代输入,再看指标——正值表示正确输入占优。所以正号=支持原输入,负号=反过来。
  • "endpoint":评测的各个终点/端点;这里 5/5 指五个端点全部如此。
  • "both arms at both boundaries":两个实验臂、两个边界条件,四种组合都算上。

3. 值得留意

  • 这两句话是判读方向的约定,不是新证据:真正的证据是上一段那串数字,此处只是把它读成"全正"。
  • "5/5 endpoints"是一致性的说法,不等于幅度大——数值都在 0.01 量级附近,差距其实很小。

b232Set Metric Mean Δ\Delta Paired tt 95% CI Joint bootstrap 95% CI Wins F4 PCC +2.364 [-1.277, +6.005] [+0.253, +5.619] 4/5 F4 MSE -2.159 [-5.482, +1.164] [-5.136, -0.234] 4/5 F5 PCC +1.460 [-0.496, +3.416] [+0.202, +3.021] 3/5 F5 MSE -1.333 [-3.116, +0.450] [-2.748, -0.187] 3/5

这段在干什么:展示 F4、F5 两个模型组在 PCC 和 MSE 两个指标上的匹配反馈结果,用效应量、置信区间和胜出次数汇报预测层面的变化。

需要解释的地方:

  • Δ Mean:两组间指标的平均差异。
  • Paired t 95% CI:配对 t 检验给出的差异区间,跨 0 通常意味着不显著。
  • Joint bootstrap 95% CI:用重抽样法算的另一版区间。

值得留意:Paired t 区间都跨 0(如 F4 PCC 的 [-1.277, +6.005]),而 bootstrap 区间不跨 0,两者结论不一致——作者没说哪个更可信,读时别只看"Wins 4/5"。

b233Descriptive paired comparison. Paired tt: five trajectories, df=4. Joint bootstrap: 4,000 shared trajectory–compound draws. Wins favor higher PCC or lower MSE. Both uncertainty summaries retain all five trajectory pairs.

这段用三句话交代配对比较的统计口径:先给 t 检验(5 条轨迹,df=4),再用联合 bootstrap 做 4,000 次共享轨迹–化合物抽样;并说明胜负判据是 PCC 更高或 MSE 更低,两种不确定性汇总都保留全部 5 对轨迹。

概念上,df=4 是配对 t 检验的自由度(5 对减 1,领域常识);bootstrap 是重抽样估不确定性的通用方法;PCC 是皮尔逊相关系数。

留意:判据是"越高/越低越好",不是数值大小;4,000 次是抽样次数,非样本量。

b234Feedback Selected PCC Valid Failed Repairs LLM calls Reported tokens Training wall s Score 0.9757 25/25 0 11 36 ≥\geq286,995 194.93 Audit 0.9781 25/25 0 13 38 592,347 217.41

这段是附录里的对比表,比较两轮实验(Feedback 与 Audit)的预测与审计结果。

  • 关键概念:PCC 是预测值和真实值的相关性(领域常识,越接近 1 越好);Valid 指 25 次尝试全部有效;Failed Repairs 为 0,说明模型没触发返修;LLM calls 是调用大模型的次数。
  • 值得留意:两轮分数几乎持平(0.9757 vs 0.9781),但 Audit 多花了 2 次 LLM 调用、翻倍 token 和更多墙钟时间。作者可能是想说明:审计带来的性能提升很小,代价却更大。表头文字被截断,具体列名这段没完整给出。

b235Selected PCC averages the five discovery-selection scores. Calls include repairs; the Score token total is a lower bound with one missing usage receipt. Five shared anchors contribute a further 32.61 training seconds, counted once.

逐段讲解

1. 这段在干什么

给上一段表格里的数字做口头备注:说明这些分数、调用次数、token 和训练时间是怎么算出来的,属于对结果的口径澄清。

2. 需要解释的地方

  • PCC:皮尔逊相关系数(领域常识),衡量预测值与真实值的线性吻合度,越接近 1 越好。
  • Selected PCC:从候选里挑出来的模型,其 PCC 取五个"发现-选择"打分的平均。
  • Calls / Score / token:上一段表格里的几列;calls 含修复调用。
  • usage receipt:用量凭证(领域常识),这里指一条记录缺失。
  • anchors:共享锚点,五个共用,只计一次训练时间。

3. 值得留意

Score 的 token 总数是下界,因为缺了一条用量凭证——真实值应更高。32.61 秒是共享锚点额外加的训练时间,与上一段 194.93/217.41 的口径衔接。

N.1 Score–audit feedback comparison: complete records

b237We analyze Score and Audit feedback using all five paired discovery trajectories and their shared context-plus-dose anchors. The descriptive comparison holds model space, trainer, selection rule, and candidate budget fixed, and reports every selected endpoint, paired effect, and uncertainty interval. Within the sci-Plex development source, compound–cell-line roles were reassigned and frozen before response preprocessing. F4 and F5 evaluate the same selected checkpoints on held-out combinations.

讲解

1. 这段在干什么

交代 Score/Audit 反馈对比的实验设置:用五条配对轨迹,控制变量,报告全部端点与不确定性,并说明 sci-Plex 数据里角色在预处理前已冻结。

2. 需要解释的地方

  • paired discovery trajectories:成对出现的发现过程,五次配对便于直接对比。
  • shared context-plus-dose anchors:两套反馈共用同一批锚点,排除输入差异。
  • compound–cell-line roles:化合物与细胞系谁当处理、谁当背景的分配(领域常识:这样能防数据泄漏)。

3. 值得留意

控制项列得很全(模型空间、训练器、选择规则、候选预算),但共用的只有锚点——两套反馈真正的差别只剩反馈本身,这是对比能成立的关键。

b238The fit, selection, F4, and F5 packages contain 2,666, 891, 438, and 442 treatment wells. Complete compound ×\times cell-line groups, including their doses and replicate wells, remain within one role; evaluated compounds and cell lines each appear among the fit marginals. The 2,000 target genes are selected by variance among fit treatment wells, with fit-only feature scaling. Control bank A supplies inputs and bank B the reference; they contain distinct wells matched on plate, cell line, and replicate, with control pools reused across these development roles. The target is the complete gene-expression profile.

讲解

1. 这段在干什么

交代数据划分的细节:各角色(fit/selection/F4/F5)的样本量、化合物×细胞系组合归入单一角色、目标基因与对照库的选取方式。是方法可复现性的补充说明。

2. 需要解释的地方

  • treatment wells(处理孔):领域常识,多孔板上每个加样孔算一个样本。
  • marginals(边际):指各角色内可评估的化合物和细胞系仍出现在 fit 的分布中。
  • control bank A/B:两个独立孔库,A 供输入、B 供参考,按板、细胞系、重复数配对。
  • fit-only feature scaling:标准化参数只用 fit 数据估计。

3. 值得留意

"remain within one role"意味着同一组合不跨角色——这是防止数据泄漏的关键设计;作者未明说,但读者应意识到。

b239Score supplies selection metrics, sampled learning curves, model size, timing, and past designs. Audit adds grouped errors, component-source clues, and chemical/dose replacement effects. This comparison estimates the complete enhanced-feedback package. Both arms share g⁡(c,a)g(c,a) within seed, five new candidate slots, and three repair opportunities per slot. The language model may revise encoders, fusion, readout, capacity, regularization, and residual scale; a direct predictor or frozen-anchor residual is equally available. Each new component is trained from scratch. Completed selection diagnostics inform subsequent proposals; final checkpoints are frozen before F4/F5 evaluation.

讲解

1. 这段在干什么

交代两组对比(Score 与 Audit)各自提供什么信息,并列出二者共享的实验设置,为评估"完整增强反馈包"做铺垫。

2. 需要解释的地方

  • *Score / Audit*:两种反馈来源。Score 给选择指标、学习曲线、模型规模、耗时和历史设计;Audit 额外给分组错误、组件来源线索、化学/剂量替换效应。
  • *g(c,a)*、*encoders、fusion、readout* 等属领域常识:指模型各组成部分,语言模型可对这些模块做修改,也可选直接预测或冻结锚点的残差。

3. 值得留意

两臂除反馈内容外其余条件(同种子、五个候选槽、每槽三次修复机会、从头训练、最终 checkpoint 冻结后才评估)全部对齐——即差异被归因于反馈本身。

b240AdamW uses learning rate 10−310^{-3}, weight decay 10−510^{-5}, batch size 128, gradient clipping at 5, a 100-epoch maximum, and early-stopping patience 15. The shared loss is fit-standardized target MSE plus 0.02​(1−PCC)0.02(1-\mathrm{PCC}), using flattened Pearson correlation. The common limit is five million trainable parameters. Epoch and endpoint selection prioritize selection Global PCC, then lower MSE; endpoint ties additionally prefer fewer parameters and earlier slots. The anchor stays frozen in residual models.

讲解

1. 这段在干什么

交代实验的统一训练配置——优化器、损失、规模上限、选点规则,让后续所有模型在公平条件下比较。

2. 需要解释的地方

  • AdamW / weight decay / 梯度裁剪:都是领域常识的优化器设置,防过拟合、防梯度爆炸。
  • PCC:皮尔逊相关系数,衡量预测与真值的线性吻合度;损失里 1−PCC 越小越好。
  • fit-standardized MSE:在拟合数据上标准化后的均方误差。
  • early-stopping patience 15:连续 15 轮没改善就停。
  • anchor frozen:残差模型里锚点参数冻结不训练。

3. 值得留意

选点顺序是先看 Global PCC、再看 MSE,平局才比参数量和更早的 slot——这是明确但容易读漏的优先级。

b241SD uses denominator n−1n-1 across five trajectories; paired two-sided Student-tt intervals use df=4. Compound-only bootstrap fixes the fitted trajectories, trajectory-only bootstrap fixes the evaluated compounds, and joint bootstrap resamples both; each uses 4,000 shared-index draws. Appendix N.2 compares these intervals with complete leave-one-pair-out and exact sign-flip analyses. Maps are averaged within endpoints before trajectory inference. RMS measures sensitivity; PCC drop is correct-input minus replacement PCC, and loss gain is replacement minus correct-input MSE. Positive continuous effects are distinct from individual qualification or biological-mechanism evidence.

讲解

这段在干什么:这是附录的“记录说明”段,交代本研究所用统计口径(区间与重采样方案),把方法细节存档备查。

需要解释的地方:

  • SD 用 n−1:算标准差时除以样本量减一(领域常识,无偏估计),这里跨五条轨迹。
  • Student-t 区间 df=4:配对双侧 t 区间,自由度 4(因 5 条轨迹)。
  • 三种 bootstrap:只重抽轨迹、只重抽化合物、两者都重抽,各 4000 次。
  • RMS / PCC drop / loss gain:分别衡量敏感度、正确输入减替换后的 PCC、替换减正确输入的 MSE。

值得留意:三种 bootstrap 区分“固定谁、重抽谁”,决定了不确定性归因于哪一方;末句提醒正向连续效应≠个体资格或机制证据。

b242Recorded model identifier: deepseek-v41. All compared calls use this serving identity. Trajectory indices 1–5 correspond, in order, to seeds 2026091201, 2026091202, 2026091203, 2026091204, 2026091205. Full-precision endpoint, cost, and paired-effect tables accompany the renderer under results/feedback_comparison/.

讲解

这段在干什么:这是一条记录性说明,交代复核时用的模型版本、轨迹编号与随机种子的对应关系,以及配套数据放在哪。它不承担论证,属于可复现性交代。

需要解释的地方:

  • *serving identity*:实际调用模型时对外展示的版本名,领域常识是同一模型不同版本行为可能不同,所以要固定。
  • *seed*:随机数种子,控制实验可复现。
  • *paired-effect*:成对比较的效应量。

值得留意:五个种子与轨迹 1–5 是按顺序一一对应的,别读乱;完整表格没写进正文,只在 results/feedback_comparison/ 目录里,正文只给指针。

(正文没提这些数字的含义或结论。)

b243Relative MSE reduction against the anchor is computed within seed as 1−MSEendpoint,s/MSEanchor,s1-\mathrm{MSE}_{\mathrm{endpoint},s}/\mathrm{MSE}_{\mathrm{anchor},s} and then averaged. A ratio of arm mean MSEs is reported separately when used. Costs retain failed attempts and recorded interruptions. Shared anchor training totals 32.611546 seconds and is counted once, separately from each condition’s candidate-training costs. Trainer and worker wall times overlap; missing token usage produces a lower bound.

讲解

1. 这段在干什么

交代评测指标的算法与成本口径——怎么算相对 MSE 降幅、成本怎么计、锚点训练时间怎么摊。

2. 需要解释的地方

  • 相对 MSE 降幅:先在每个 seed 内算 1 − 端点MSE / 锚点MSE,再跨 seed 平均;即「比基线好了百分之多少」。
  • anchor(锚点):作为对照基线的那次训练。
  • 手臂均值 MSE 之比:另一套报告口径,仅在用到时才单列。
  • lower bound(下界):token 用量缺失时,给出的数是保守下限,不是实际上限。

3. 值得留意

  • 成本保留失败尝试与中断记录,不是只算成功的便宜账。
  • 锚点训练 32.611546 秒只计一次,与各条件的候选训练成本分开计——不重复摊到每个条件上。
  • 训练器与worker 的墙钟时间重叠,所以按各自相加会高估。

b244Tr. Model Slot PCC MSE C-R C-P C-L D-R D-P D-L Fold 4 1 Anchor 0 0.96652 30.768 0.000 0.000 0.000 28.653 1.152 1.042 1 Score 5 0.97557 22.551 141.301 20.580 18.805 84.343 6.798 6.212 1 Audit 5 0.97806 20.278 142.171 23.162 21.173 94.176 8.838 8.081 2 Anchor 0 0.96646 30.836 0.000 0.000 0.000 25.007 0.985 0.896 2 Score 3 0.97740 20.889 136.127 21.774 19.882 84.750 7.580 6.923 2 Audit 4 0.97809 20.252 155.045 25.294 23.142 101.848 9.453 8.650 3 Anchor 0 0.96647 30.813 0.000 0.000 0.000 31.861 1.311 1.189 3 Score 3 0.97570 22.432 138.732 20.644 18.823 96.003 7.982 7.278 3 Audit 3 0.97817 20.181 138.099 22.698 20.763 91.844 8.613 7.880 4 Anchor 0 0.96646 30.820 0.000 0.000 0.000 26.818 1.066 0.960 4 Score 5 0.97807 20.274 157.142 25.704 23.535 96.377 9.070 8.307 4 Audit 5 0.97725 21.023 145.497 22.784 20.858 90.603 8.193 7.502 5 Anchor 0 0.96629 30.977 0.000 0.000 0.000 17.193 0.640 0.580 5 Score 4 0.97115 26.587 117.806 13.009 11.887 48.269 2.392 2.186 5 Audit 2 0.97815 20.203 159.745 26.108 23.941 102.759 9.697 8.893 Fold 5 1 Anchor 0 0.97265 25.253 0.000 0.000 0.000 27.811 0.712 0.647 1 Score 5 0.97586 22.332 100.896 7.672 7.048 65.922 3.595 3.306 1 Audit 5 0.97785 20.495 90.349 8.720 7.978 71.763 5.036 4.612 2 Anchor 0 0.97297 24.955 0.000 0.000 0.000 24.562 0.606 0.553 2 Score 3 0.97777 20.567 80.270 7.558 6.905 71.301 4.866 4.452 2 Audit 4 0.97772 20.619 90.797 8.556 7.835 74.695 4.693 4.303 3 Anchor 0 0.97297 24.957 0.000 0.000 0.000 31.737 0.839 0.764 3 Score 3 0.97626 21.947 92.154 7.588 6.919 75.540 4.604 4.202 3 Audit 3 0.97815 20.230 86.705 8.379 7.691 71.783 4.883 4.487 4 Anchor 0 0.97270 25.230 0.000 0.000 0.000 26.502 0.688 0.621 4 Score 5 0.97771 20.623 89.824 8.217 7.526 77.179 5.209 4.776 4 Audit 5 0.97758 20.744 88.060 7.995 7.320 70.291 4.508 4.131 5 Anchor 0 0.97292 24.999 0.000 0.000 0.000 17.108 0.399 0.363 5 Score 4 0.97425 23.789 75.832 4.122 3.769 40.870 1.079 0.988 5 Audit 2 0.97786 20.504 99.006 9.405 8.635 83.693 5.505 5.060

这段在干什么:这是附录里的完整数据表,逐折、逐随机种子列出 Anchor/Score/Audit 三种模型在 PCC、MSE 和六项成本指标上的原始记录,供复核前面正文的对比结论。

需要解释的地方:PCC 是预测值和真值的相关系数(越大越准,领域常识);MSE 是均方误差(越小越好,领域常识);C-R/C-P/C-L 和 D-R/D-P/D-L 是各类训练或贡献成本指标,具体定义这段没给,需回正文查。

值得留意:Anchor 行的成本列全为 0,说明它是零成本基线,不是真实测量值;Audit 的 PCC 和 MSE 通常略优于 Score,但成本也更高,得分与代价之间是有取舍的。

b245PCC is unscaled; all other numeric outcomes are in 10−310^{-3} units. C/D: chemical/dose; R: prediction RMS; P: PCC drop; L: target-loss gain. Slot 0 is the shared anchor. RMS has no preferred direction.

上一段末尾刚给完一条完整记录(含 RMS、PCC 等数字),这段是它的读表说明——交代表里各列的单位、缩写和方向,属方法性脚注。

需要解释的地方

  • PCC:皮尔逊相关系数(领域常识),衡量预测与真值的线性相关,故「unscaled」指不缩放。
  • C/D、R、P、L 是列名缩写:化学/剂量、预测 RMS、PCC 下降、目标损失增益。
  • Slot 0 是共享锚点:各记录以此为基准比较。
  • RMS 无偏好方向:误差越小越好,但「方向」本身无意义,不像 PCC 有正负倾向。

值得留意

除 PCC 外,所有数值都乘了 10⁻³,比较数字时别忽略这个量纲。

b246Tr. Arm Slot Sel. PCC Repair Calls HTTP Tokens Train s Work s Diag. s 1 Score 5 0.97491 1 6 7 ≥\geq50,685 48.09 67.26 7.10 1 Audit 5 0.97828 2 7 7 106,239 41.24 60.03 12.10 2 Score 3 0.97757 3 8 8 62,703 41.96 60.07 7.50 2 Audit 4 0.97783 3 8 8 125,047 42.32 61.01 12.94 3 Score 3 0.97628 4 9 9 70,709 40.04 58.06 7.36 3 Audit 3 0.97825 3 8 8 136,386 44.25 61.58 12.15 4 Score 5 0.97852 3 8 8 68,757 35.40 52.75 8.73 4 Audit 5 0.97798 3 8 8 129,227 39.60 55.93 12.24 5 Score 4 0.97145 0 5 5 34,141 29.44 46.61 6.75 5 Audit 2 0.97830 2 7 7 95,448 50.00 67.66 11.96

这段是附录里的原始记录表:把 5 组模型各自的 Score 与 Audit 两次运行逐行列出,供读者核对上一段摘要背后的完整数据。

需要解释的地方

  • Tr.:试验编号,1–5 对应五组模型;每组有 Score、Audit 两行,即同一模型在两种评分设定下的复现。
  • Arm/Slot:被检验的"输入用途声明"编号;Slot 0 是共享锚点(上段已说明)。
  • PCC:预测相关系数,越大越准。
  • Repair/Calls:修复次数与该次审计的模型调用次数。
  • HTTP Tokens:接口传输的 token 量,反映算力/成本。
  • Train/Work/Diag. s:训练、运行、诊断耗时(秒)。

值得留意

Audit 行普遍 PCC 更高,但 Calls、Tokens、耗时也大幅上升(如第 5 组 Token 从 34,141 涨到 95,448)——精度提升伴随成本膨胀。表中只给数据,未作评价。

b247Every trajectory completed 5/5 new candidates and had zero fully failed slots. Train is known completed trainer wall time; Work is candidate worker wall time; Diag. is postfit diagnostic wall time. Logical calls and physical HTTP attempts are distinct. Complete token breakdowns and missing-cost flags are in the CSV.

讲解

1. 这段在干什么

它是 N.1 小节的数据说明,交代这张表里各列的定义,以及样本完整性。

2. 需要解释的地方

  • 5/5 new candidates / 零 fully failed slots:每条轨迹都成功产出了全部 5 个新候选,没有哪个槽位彻底失败。
  • Train / Work / Diag.:分别是训练器、候选工作进程、postfit 诊断三段的墙钟时间(领域常识:wall time 指真实流逝时间,不是 CPU 时间)。
  • Logical calls vs physical HTTP attempts:逻辑调用 ≠ 实际发出的 HTTP 请求,重试会让后者变多。

3. 值得留意

作者先声明"零完全失败",但下一句就点出逻辑调用与 HTTP 尝试不同——即可能有重试或部分失败被藏在这两个计数差里。完整 token 拆分与 missing-cost 标记不在正文,在 CSV。

b248Set Metric Mean Δ\Delta Paired 95% CI +/–/0 F4 PCC +2.364 [-1.277, +6.005] 4/1/0 F4 MSE -2.159 [-5.482, +1.164] 1/4/0 F4 Chem. RMS +9.890 [-16.189, +35.969] 3/2/0 F4 Chem. PCC drop +3.667 [-3.577, +10.912] 4/1/0 F4 Chem. loss gain +3.389 [-3.269, +10.048] 4/1/0 F4 Dose RMS +14.298 [-16.033, +44.628] 3/2/0 F4 Dose PCC drop +2.194 [-1.639, +6.028] 4/1/0 F4 Dose loss gain +2.020 [-1.496, +5.536] 4/1/0 F5 PCC +1.460 [-0.496, +3.416] 3/2/0 F5 MSE -1.333 [-3.116, +0.450] 2/3/0 F5 Chem. RMS +3.189 [-13.712, +20.089] 2/3/0 F5 Chem. PCC drop +1.580 [-1.069, +4.228] 4/1/0 F5 Chem. loss gain +1.458 [-0.979, +3.896] 4/1/0 F5 Dose RMS +8.283 [-16.535, +33.100] 3/2/0 F5 Dose PCC drop +1.054 [-1.483, +3.592] 3/2/0 F5 Dose loss gain +0.974 [-1.355, +3.303] 3/2/0

这段在干什么:这是 N.1 节的完整数据表,逐行记录两组模型(F4、F5)在 8 个指标上「score–audit」反馈对比的均值差、配对 95% 置信区间和 +/-/0 计数。

需要解释的地方:Δ 是两次评分之差;Paired 95% CI 是配对样本的置信区间,跨 0 说明差异不稳健;「+/–/0」是各次配对里改善/变差/持平的出现次数,三者之和即样本量(F4 多为 5,F5 各行为 5)。

值得留意:多数 CI 都跨 0,看着均值大(如 F4 Dose RMS +14.298)其实不稳定;F4 PCC、F5 各 PCC 系列几乎全为 4/1/0,方向一致但样本极小,别把「4:1」当成显著。

b249Descriptive effect estimates over all five paired trajectories. Signs count positive, negative, and exactly zero differences, not wins. Lower MSE is favorable; RMS has no favorable direction.

这段是 N.1 小节的总说明,交代下面五条配对轨迹的数字该怎么读。

关键概念:描述性效应估计——只汇总差值、不做显著性检验;3/2/0 是正/负/恰好为零的计数,不是胜负场次。MSE 越小越好,所以方向可判;RMS 没有"好"的方向(领域常识:它是误差量纲的量,无所谓正负优劣)。

值得留意:上段那些 +1.054、+0.974 都是"越大越好"的正向指标,与 MSE 相反,别混着看。

b250Set Metric Paired tt CI Compound-only CI Joint CI F4 PCC [-1.277, +6.005] [+1.300, +3.704] [+0.253, +5.619] F4 MSE [-5.482, +1.164] [-3.389, -1.197] [-5.136, -0.234] F5 PCC [-0.496, +3.416] [+1.006, +1.975] [+0.202, +3.021] F5 MSE [-3.116, +0.450] [-1.800, -0.924] [-2.748, -0.187]

你正在看的这段,其实是上一句的「数据正文」。

这段在干什么

它把 F4、F5 两个模型在 PCC、MSE 两个指标下的配对 t 检验置信区间列全,供读者与正文图表对照。

需要解释的地方

  • PCC:皮尔逊相关系数,领域常识里通常越大越好。
  • MSE:均方误差,上一段已说「越小越好」。
  • 三种 CI:Compound-only、Joint、Paired t,是三种不同的置信区间口径,区别来自统计对象,这段没细说。
  • F4/F5:两个模型编号。

值得留意

区间跨没跨 0 才是重点:跨 0 说明效应不稳健,别只盯着区间宽窄。具体哪些跨 0 需要你逐行核对,这段没替你说。

b251Bootstrap intervals use 4,000 draws. No coordinate-effect bootstrap interval is substituted from predictive statistics, and no new confirmatory tests are introduced.

紧接上面那串区间,这段是给整套审计数字定规矩的,不是新结果。

1. 这段在干什么:交代区间怎么算、边界在哪——只报 bootstrap 区间,不拿预测统计量去顶替坐标效应区间,也不加新的确证检验。

2. 需要解释的地方:

  • Bootstrap(自助法):领域常识,反复从原数据里重抽样、算统计量,用这些重抽样结果的分布来估区间;「4,000 draws」就是抽了 4,000 次。
  • 坐标效应:指模型里单个坐标/变量的效应量,和「预测统计量」(如预测误差类指标)是两码事。
  • 确证检验:预先设定、用来坐实结论的正式检验。

3. 值得留意:「is substituted(不被替换)」是刻意划界——作者在声明坐标效应的区间一律来自坐标效应本身,绝不拿预测侧的数字来凑。这句是防读者误读上面那三列区间怎么来的。

b252Set Coordinate Eligible wells Maps Unique First 32 unique F4 Chemical 438/438 128 128 32 F4 Dose 428/438 128 1 1 F5 Chemical 434/442 128 128 32 F5 Dose 436/442 128 1 1

讲解

1. 这段在干什么

这是一张审计汇总表,逐条列出「集合—坐标—合格孔数—映射数—唯一数」,用来完整记录分数审计的反馈对比结果。

2. 需要解释的地方

这是领域常识:表中「合格孔数」写作 438/438 这类分数,前者是实际合格数,后者是总数;128 疑似每组映射的固定规模;F4/F5 是两个实验集合,Chemical 与 Dose 是同一集合下的两种坐标口径。

3. 值得留意

Chemical 行映射为 128,Dose 行却只有 1,两者差异极大——原文未解释原因。另外 Dose 的 432 与 1 两个数字含义也不明,作者没有交代。

b253Replacements match plate, cell line, replicate, and treatment time. Chemical replacement holds dose fixed; dose replacement holds compound fixed and requires at least 0.1 log-dose difference. No cross-group fallback is used. Discovery used 32 maps; final evaluation used 128, with only one distinct dose map. Unsupported wells are excluded, not assigned zero effect.

讲解

1. 这段在干什么

交代「替换」的配对规则和实验设置,是上段那张统计表的配套说明。

2. 需要解释的地方

  • Replacements(替换):指用于配对比较的替代样本,须在板、细胞系、重复、处理时间上与原样本对齐。
  • Chemical / Dose replacement:化学替换固定剂量;剂量替换固定化合物,且要求至少 0.1 log 剂量差。
  • No cross-group fallback:找不到合规替换就作罢,不跨组凑数。
  • map:此处指所用映射数量。

3. 值得留意

最终评估 128 个 map,却只有 1 个剂量 map(对应上表 Dose 行);不支持的孔是剔除而非记 0,避免引入假信号。

N.2 Robustness across the complete five trajectory pairs

b255This retrospective sensitivity analysis retains seeds 2026091201–2026091205 and all eight metrics at both boundaries. Audit minus Score is computed within each selected-trajectory pair; F4 and F5 reuse the same selected checkpoints and do not make N=10N=10. Means and descriptive two-sided 95% t4t_{4} intervals use all five pairs. All 32 sign assignments are enumerated for the absolute unstudentized mean under a pair-exchangeability/sign-symmetry null, not a randomized-assignment guarantee. The smallest attainable two-sided tail probability is 2/32=.06252/32=.0625; no multiplicity-adjusted confirmatory claim is made.

这一段在做稳健性/敏感性分析:用固定5个种子、8个指标,检验“Audit−Score”结果是否只是特定选点造成的。

关键概念:

  • 配对内的 Audit minus Score:只在同一对选中轨迹里做差,避免跨对混算。
  • F4/F5 复用同一批 checkpoint:所以样本量不翻倍成 N=10,避免了重复计数。
  • 符号枚举 2/32:把32种正负号全排一遍,看均值多极端;这是可交换性/符号对称零假设下的检验,不是随机分配保证(领域常识:置换检验常这样近似)。

值得留意:最小双侧尾部概率只有 .0625,即5对样本下无法做到 p<.05;作者明确不做多重校正后的确证声明,结论只是探索性。

b256Fold Metric Mean Paired-tt 95% CI Median +/−/0+/-/0 Tail F4 Global PCC 2.3642.364 [−1.277, 6.005][-1.277,\,6.005] 2.4652.465 4/1/0 6/32 F4 MSE −2.159-2.159 [−5.482, 1.164][-5.482,\,1.164] −2.251-2.251 1/4/0 6/32 F4 Chemical RMS 9.8909.890 [−16.189, 35.969][-16.189,\,35.969] 0.8700.870 3/2/0 12/32 F4 Chemical loss gain 3.3893.389 [−3.269, 10.048][-3.269,\,10.048] 2.3682.368 4/1/0 8/32 F4 Chemical PCC drop 3.6673.667 [−3.577, 10.912][-3.577,\,10.912] 2.5822.582 4/1/0 8/32 F4 Dose RMS 14.29814.298 [−16.033, 44.628][-16.033,\,44.628] 9.8339.833 3/2/0 10/32 F4 Dose loss gain 2.0202.020 [−1.496, 5.536][-1.496,\,5.536] 1.7271.727 4/1/0 6/32 F4 Dose PCC drop 2.1942.194 [−1.639, 6.028][-1.639,\,6.028] 1.8731.873 4/1/0 6/32 F5 Global PCC 1.4601.460 [−0.496, 3.416][-0.496,\,3.416] 1.8881.888 3/2/0 8/32 F5 MSE −1.333-1.333 [−3.116, 0.450][-3.116,\,0.450] −1.716-1.716 2/3/0 8/32 F5 Chemical RMS 3.1893.189 [−13.712, 20.089][-13.712,\,20.089] −1.763-1.763 2/3/0 24/32 F5 Chemical loss gain 1.4581.458 [−0.979, 3.896][-0.979,\,3.896] 0.9300.930 4/1/0 4/32 F5 Chemical PCC drop 1.5801.580 [−1.069, 4.228][-1.069,\,4.228] 0.9990.999 4/1/0 4/32 F5 Dose RMS 8.2838.283 [−16.535, 33.100][-16.535,\,33.100] 3.3943.394 3/2/0 20/32 F5 Dose loss gain 0.9740.974 [−1.355, 3.303][-1.355,\,3.303] 0.2850.285 3/2/0 12/32 F5 Dose PCC drop 1.0541.054 [−1.483, 3.592][-1.483,\,3.592] 0.2790.279 3/2/0 12/32

这段是逐层核对稳健性:把 F4、F5 两对轨迹在各指标上的配对差异摆出来,看结论是否普遍成立。

关键术语:*Mean* 是配对差值均值;*Paired-t* 是配对 t 统计量;*95% CI* 是置信区间;*Median* 是中位差;+/−/0 是正向/负向/无变化计数;*Tail* 是列尾概率(分母 32)。

值得留意:多个指标的 95% CI 跨零(如 F5 Chemical RMS 的 [−13.7, 20.1]),说明方向不稳;Tail 列分母多为 32,与上段 2/32 的离散下限一致——效应再大也难达显著,故不作确认性声明。

b257Tail is the exact sign-flip count out of 32. All 16 paired-tt intervals contain zero and all exact tails are at least .125. RMS measures sensitivity; target-loss gain and PCC drop measure changes in loss and correlation.

这段是五组轨迹对稳健性检验的收尾说明,给上表数据做注解。

关键概念:

  • Tail:32 次里符号翻转的精确次数。
  • **paired-*t* 区间**:配对 *t* 检验的置信区间,全含 0 → 差异不显著(领域常识)。
  • RMS:衡量敏感度;target-loss gain / PCC drop:衡量损失与相关性变化。

值得留意:作者其实在说没有稳健的差异——16 个区间全含 0、所有 tail 至少 .125,即翻不出显著结果。但这段只提"度量什么",没解释表里 3/2/0 的含义。

b258Fold Metric Compound only Trajectory only Joint F4 Global PCC [1.300, 3.704][1.300,\,3.704] [0.144, 4.828][0.144,\,4.828] [0.253, 5.619][0.253,\,5.619] F4 MSE [−3.389,−1.197][-3.389,\,-1.197] [−4.408,−0.132][-4.408,\,-0.132] [−5.136,−0.234][-5.136,\,-0.234] F4 Chemical RMS [6.425, 13.160][6.425,\,13.160] [−4.738, 28.126][-4.738,\,28.126] [−5.652, 27.694][-5.652,\,27.694] F4 Chemical loss gain [2.032, 5.060][2.032,\,5.060] [−0.482, 8.094][-0.482,\,8.094] [−0.456, 8.402][-0.456,\,8.402] F4 Chemical PCC drop [2.201, 5.452][2.201,\,5.452] [−0.534, 8.787][-0.534,\,8.787] [−0.508, 9.067][-0.508,\,9.067] F4 Dose RMS [4.676, 22.279][4.676,\,22.279] [−2.007, 34.959][-2.007,\,34.959] [−3.228, 38.642][-3.228,\,38.642] F4 Dose loss gain [0.689, 3.753][0.689,\,3.753] [0.039, 4.490][0.039,\,4.490] [0.066, 5.631][0.066,\,5.631] F4 Dose PCC drop [0.750, 4.083][0.750,\,4.083] [0.028, 4.884][0.028,\,4.884] [0.068, 6.187][0.068,\,6.187] F5 Global PCC [1.006, 1.975][1.006,\,1.975] [0.302, 2.618][0.302,\,2.618] [0.202, 3.021][0.202,\,3.021] F5 MSE [−1.800,−0.924][-1.800,\,-0.924] [−2.392,−0.274][-2.392,\,-0.274] [−2.748,−0.187][-2.748,\,-0.187] F5 Chemical RMS [−1.073, 7.471][-1.073,\,7.471] [−6.751, 14.920][-6.751,\,14.920] [−7.875, 15.416][-7.875,\,15.416] F5 Chemical loss gain [0.824, 2.180][0.824,\,2.180] [0.248, 3.260][0.248,\,3.260] [0.143, 3.537][0.143,\,3.537] F5 Chemical PCC drop [0.892, 2.364][0.892,\,2.364] [0.266, 3.528][0.266,\,3.528] [0.146, 3.843][0.146,\,3.843] F5 Dose RMS [1.196, 15.103][1.196,\,15.103] [−3.716, 25.621][-3.716,\,25.621] [−5.151, 28.995][-5.151,\,28.995] F5 Dose loss gain [0.324, 1.866][0.324,\,1.866] [−0.261, 2.557][-0.261,\,2.557] [−0.253, 3.122][-0.253,\,3.122] F5 Dose PCC drop [0.349, 2.021][0.349,\,2.021] [−0.294, 2.768][-0.294,\,2.768] [−0.284, 3.397][-0.284,\,3.397]

这段在干什么:把稳健性检验从之前的两条轨迹扩到全部五对轨迹,用区间汇总 F4、F5 六项指标的波动范围。

需要解释的地方:

  • 每个单元格是一个区间(下界, 上界),不是单点值,表示跨轨迹对的波动范围。
  • 三列是对照设置:Compound only(只换化合物)、Trajectory only(只换轨迹)、Joint(两者都换)。
  • MSE 区间全为负、PCC 区间全为正,是领域常识:MSE 越低越好、PCC 越高越好,故含义相反。

值得留意:三列区间明显逐列变宽,Joint 最宽;Chemical/Dose RMS 在 Trajectory only 和 Joint 下下界转负,而 Compound only 保持正。

b2594000 draws, seed 69313. Compound weights are shared across all paired arms and trajectories; trajectory draws resample the five paired indices. Joint draws resample both units. The 128 donor maps are fixed. Sufficient statistics are pooled before PCC; compound PCCs are not averaged. Compound-only intervals condition on five selected trajectories; trajectory-only intervals condition on the observed compound set, while joint intervals vary both empirical units.

这段在交代重采样方案与区间口径,确保五对轨迹下的稳健性结果可比。

关键概念:

  • compound / trajectory / joint draws:分别指只重抽化合物、只重抽轨迹、两者都抽,对应三类置信区间。
  • sufficient statistics pooled before PCC:先合并统计量再算相关性,不是把各化合物的 PCC 平均。(PCC=皮尔逊相关系数,领域常识)
  • condition on:指抽样时固定住另一方不变。

值得留意:三类区间的条件对象不同——compound-only 固定五条轨迹,trajectory-only 固定观测到的化合物集,joint 两者都变,所以数值不可直接横向比较。128 张 donor map 是固定的,不参与抽样。

b260Trajectory-only percentile bootstrap and paired-tt inference target the same conditional mean paired contrast; their different conclusions reflect the five-point empirical bootstrap distribution versus Student-tt sampling assumptions, not different estimands. Compound-only and joint intervals additionally have different resampling scopes. All predictive bootstrap intervals exclude zero, whereas the paired-tt intervals and exact sign-flip tails do not provide the same evidence. Increasing bootstrap draws does not add independent trajectories or establish broad generalization.

这段在干什么:对比几种统计推断方法为何结论不一致,指出差异来自假设与重采样范围,而非估计目标不同。

需要解释的地方:

  • 百分位自助法 vs 配对 t:领域常识——前者直接看重采样分布,后者假设正态;两者目标相同(条件均值配对差异),结论却可能不同。
  • 重采样范围:trajectory-only 固定化合物集,joint 则连化合物也变。

值得留意:作者明说「增加自助抽样次数并不能带来独立轨迹或广泛泛化」——别把区间变窄误当成证据变强。

b261Fold Metric All-five mean Fifth-pair d5/5d_{5}/5 Omit-fifth mean All five LOO means F4 Global PCC 2.3642.364 1.3991.399 1.2061.206 [1.206, 3.160][1.206,\,3.160] F4 MSE −2.159-2.159 −1.277-1.277 −1.103-1.103 [−2.886,−1.103][-2.886,\,-1.103] F4 Chemical RMS 9.8909.890 8.3888.388 1.8771.877 [1.877, 15.273][1.877,\,15.273] F4 Chemical loss gain 3.3893.389 2.4112.411 1.2231.223 [1.223, 4.906][1.223,\,4.906] F4 Chemical PCC drop 3.6673.667 2.6202.620 1.3091.309 [1.309, 5.314][1.309,\,5.314] F4 Dose RMS 14.29814.298 10.89810.898 4.2494.249 [4.249, 19.315][4.249,\,19.315] F4 Dose loss gain 2.0202.020 1.3411.341 0.8480.848 [0.848, 2.726][0.848,\,2.726] F4 Dose PCC drop 2.1942.194 1.4611.461 0.9160.916 [0.916, 2.962][0.916,\,2.962] F5 Global PCC 1.4601.460 0.7220.722 0.9220.922 [0.922, 1.859][0.922,\,1.859] F5 MSE −1.333-1.333 −0.657-0.657 −0.845-0.845 [−1.697,−0.845][-1.697,\,-0.845] F5 Chemical RMS 3.1893.189 4.6354.635 −1.808-1.808 [−1.808, 6.622][-1.808,\,6.622] F5 Chemical loss gain 1.4581.458 0.9730.973 0.6060.606 [0.606, 1.875][0.606,\,1.875] F5 Chemical PCC drop 1.5801.580 1.0571.057 0.6540.654 [0.654, 2.030][0.654,\,2.030] F5 Dose RMS 8.2838.283 8.5658.565 −0.352-0.352 [−0.352, 12.075][-0.352,\,12.075] F5 Dose loss gain 0.9740.974 0.8150.815 0.1990.199 [0.199, 1.379][0.199,\,1.379] F5 Dose PCC drop 1.0541.054 0.8850.885 0.2110.211 [0.211, 1.493][0.211,\,1.493]

这段是五组轨迹对的稳健性汇总表:每行一个「模型×指标」,看五对全纳入和只留第五对时结论是否一致。

需要解释的

  • d5/5:第五对单独贡献占五对均值多少,>1 说明它比平均更大。
  • LOO means:留一法,逐对剔除后重算,看区间是否跨 0。

值得留意

  • F5 的 Chemical RMS、Dose RMS 的 d5/5 为负(−1.808、−0.352),且 LOO 区间跨 0——这两项在第五对上方向相反、不稳定。
  • 其余多数指标 d5/5 为正且 LOO 区间不含 0,说明主要结论不靠单对支撑。

b262The fifth-pair column is its contribution to the five-pair mean, not its whole paired difference. Every leave-one-out estimate and t3t_{3} interval is retained in the portable CSV. F5 chemical and dose RMS have positive all-five means but negative omit-fifth means; the F5 chemical-RMS median is negative, whereas the dose-RMS median is positive. All five pairs remain in the primary estimate.

这段在干什么:解释第五对轨迹在五对均值里的权重口径,并说明留一法结果的符号分歧。

需要解释的地方:

  • 第五对列:表里 F5 那列是它对五对均值的贡献,不是它自己成对的差值——两者别混。
  • omit-fifth:去掉第五对后重算的均值。
  • t₃ 区间:仅三对时的区间估计,样本极少。

值得留意:F5 化学与剂量的 RMS 出现"全五对均值为正、去掉第五对为负"的翻转,且两者中位数符号相反——说明结论对是否包含 F5 敏感。但作者仍把五对全留在主估计里。

b263Retained full predictions independently verify global PCC/MSE and eligible-source correct statistics. Full counterfactual prediction arrays were not retained for this legacy feedback package, so coordinate reconstruction verifies stored sufficient statistics but is not a new independent full-array counterfactual check. These additional descriptions do not reselect checkpoints, retrain models, or revise historical qualification labels.

这段在给结论加限定:说明保留的是完整预测,但反事实数组没留。

需要解释的地方:

  • *PCC/MSE*:领域常识,相关系数与均方误差,衡量预测准不准。
  • *充分统计量*:领域常识,能概括数据的汇总值;重算坐标只能验证这些汇总值,不算独立的完整数组复核。
  • *legacy feedback package*:旧版反馈包,即历史遗留数据。

值得留意:作者主动承认证据层级不同——全局统计可信,但没做新的完整反事实检验。末句强调这些描述不改变既有结论,防止误读为重新筛选。

Appendix O Recorded feedback, source revisions, and frozen endpoints

b265The two cases are the first two registered trajectory seeds (2026091201 and 2026091202), selected by registration order. Each retains the common initial predictor and all five new audit-feedback proposals. Parent links denote design ancestry: every new candidate is trained from scratch, while the shared context-plus-dose predictor g⁡(c,a)g(c,a) is frozen. This anchor concatenates 2,000 context features with dose and uses a 2001–256–256–2000 MLP with 1,092,816 trainable parameters. Local metrics describe the development-selection partition. The public case package contains the hash-verified anchor source, actual system/user prompts, proposals, candidate source, diffs, and received feedback.

这段在干什么

登记两个轨迹种子案例(2026091201/02),交代它们的设计来源、冻结锚点与公开包内容。

需要解释的地方

  • 种子/登记顺序:这两例按登记先后入选,是"最早两个"。
  • Parent links / 祖先:只表示设计血缘;所有新候选都从头训练,不继承权重。
  • 锚点 g(c,a):上下文+剂量的预测器,被冻结,全程不动。这是领域常见的"固定参照"设定。
  • 2001–256–256–2000 MLP:输入2000+剂量共2001维,两层隐藏各256,输出2000。参数1,092,816个。
  • Local metrics:只在开发选择分区上统计。

值得留意

承上句强调"不改历史标签",这里进一步点明锚点冻结、候选从头训练——即审计改动只发生在候选侧,参照系始终不变。

b266In the first trajectory, slot 2 is the incumbent after its child slot 3 fails to improve selection PCC. Before slot 4, the actual prompt contains slot 2’s best epoch (5), sampled learning curve, group-error diagnostics, source-check findings, and positive compound and dose replacement effects. The model hypothesizes overfitting and proposes reducing encoder capacity; compound memorization is a proposed explanation, not a separately measured biological or learning mechanism. Inspection of the emitted source verifies that each chemical/context encoder changes from two Linear layers to one, with dropout 0.1 added in the encoders and fusion. Dose-conditioned FiLM and the rank-128 bilinear interaction remain. Trainable parameters decrease from 4,056,400 to 2,754,896; total parameters including the fixed anchor are 5,149,216 and 3,847,712. Slot 2 versus slot 4 selection PCC is 0.97808647 versus 0.97818047; MSE is 0.02028296 versus 0.02019429; compound target-loss gain is 0.01846193 versus 0.01949683. The static checker reports syntactic return-dependency hints while leaving functional usefulness unresolved; fixed-model input replacement supplies the separate behavioral evidence.

这段在干什么:它是第一条轨迹的实验记录——报告 slot 3 输给 slot 2、slot 4 提出“过拟合、缩减编码器容量”的假设,然后核对改后源码、参数量和指标。

需要解释的地方:

  • “PCC”指预测与真实值的相关性,越高越好;“MSE”是均方误差,越低越好,这两个是领域常识。
  • “incumbent”指当前保留的胜出版本。
  • “compound memorization”只是模型提出的解释,并未被单独测量——原文特意强调了这点。
  • “static checker / 行为证据”是说静态检查只给语法线索,真正有用与否靠替换输入实验来判。

值得留意:slot 4 的 PCC 和 MSE 都略好于 slot 2,但幅度很小;参数量从 4,056,400 降到 2,754,896(含固定锚点则是 5,149,216→3,847,712)。

b267Both first-case sources zero-initialize their output readout. Under the fixed residual wrapper this initializes the newly fitted predictor at g⁡(c,a)g(c,a). Slot 5 adds a dose-conditioned scalar interaction gate and an extra projected chemical feature before fusion, then wins by selection PCC. The emitted chem_direct feature enters the nonlinear fusion, whereas its proposal calls it a direct additive target-space path. Slot 5 has higher selection PCC but lower compound target-loss gain than slot 4; selection and input contribution therefore remain distinct. The same frozen slot-5 checkpoint is subsequently evaluated on the two held-out sci-Plex partitions.

这段在干什么

回答上一段遗留的问题:固定模型下换掉输入后,slot 5 的行为证据是什么——即它赢在选中指标,而非输入贡献本身。

需要解释的地方

  • zero-initialize 输出读出:输出层初值设为零(领域常识做法)。
  • 残差包装器:把新预测器接在残差结构上。
  • selection PCC:按选中标准算的相关系数;target-loss gain:对目标损失的改善。两者在此不一致。

值得留意

slot 5 的提案称 chem_direct 是「直接加性路径」,实际它进入非线性融合——提案与实现的描述对不上,这正是「审计」的落点。另外,选中指标高≠输入贡献大,两者须分开看。

b268For seed 2026091202, slot 2 decreases selection PCC from its slot-1 parent (0.977817 to 0.977703). Slot 3 improves on slot 2 but remains below slot 1. Slot 4 explicitly returns to slot 1 as its design parent and is selected at PCC 0.977826; its MSE (0.02053288) is slightly worse than slot 1’s (0.02052894), consistent with PCC-first selection. Slot 5 adds explicit pairwise projections but falls to PCC 0.977506 and is not selected. Its rationale claims that zero-initialized new terms preserve the parent, but the shared from-scratch training contract does not inherit fitted parent weights: this claim is preserved as an agent statement in the public proposal, not endorsed as a measured guarantee.

讲解

这段在干什么:记录 seed 2026091202 下各 slot 的选择结果,展示选代轨迹,并揭穿 slot 5 的一个说法。

需要解释的地方:

  • PCC:预测值与真实值的相关性指标(领域常识),越大越好。
  • MSE:均方误差(领域常识),越小越好。
  • slot:这里指迭代产生的候选设计版本,后一个以前一个为父本。

值得留意:slot 1→2 下降、3 回升但仍不如 1、4 直接回到 slot 1;选 4 时 PCC 更好但 MSE 更差,作者点明这是「PCC-first」取向。slot 5 的理由说零初始化能保留父本,作者反驳:共享训练契约并不继承已拟合权重——这只是 agent 声明,不是被验证的保证。

b269The first case uses one JSON-contract repair at each of slots 4 and 5; the second uses one at slot 4 and two at slot 5. These logged failures are completion-token-limit/JSON-contract failures. Lower-scoring, successfully trained candidates are retained without score-driven repair. The case package preserves these outcomes and the actual branch structure.

这段在对比两个案例的修复开销,并交代这些失败被如实记录、未做分数驱动的修补。

关键概念

  • JSON-contract 修复:领域常识,指输出不符合预定 JSON 结构时触发的修补动作。
  • completion-token-limit:生成被输出长度上限截断,导致 JSON 不完整。
  • slot 4 / slot 5:本例中发生修复的位置,具体含义这段没展开。

值得留意

  • 两个案例修复次数不同,说明修复量因案例而异。
  • 「低分但训练成功的候选不做分数修复」——分数不是修复依据,隐含修复只针对结构失败,而非性能。