b003The Enzyme Commission (EC) numbering scheme provides a hierarchical way to classify enzymes according to their catalytic functions. While recent protein language model (PLM) based approaches like CLEAN and ProteInter have improved sequence-based EC number prediction, they struggle with fine-grained classification at the deepest hierarchical level. Structure-based approaches for grouping similar proteins using alignment tools excel at finding proteins that share overall global structure, but suffer from high false positive rates when classifying proteins that are globally structurally similar but functional differentiation depends on a localized region. This problem is particularly relevant to EC number prediction, as enzymatic function depends on its catalytic domain, which is a relatively small, specific region of the protein. We introduce Deep Enzyme Function Transfer (DEFT) that harmonizes sequence- and structure-based approaches through the key insight that PLM based annotations of the first two EC number hierarchy levels vastly reduce false positives that are likely to show in purely structure-based EC number prediction. Given an enzyme of interest, DEFT first uses a PLM based method to assign the first two levels of the enzyme’s EC number, and then uses a structure-based method to predict the remaining two levels of the EC number. Using benchmarking datasets, we demonstrate that DEFT achieves superior accuracy compared with current state-of-the-art tools for EC number prediction. Furthermore we show that DEFT’s computational efficiency enables high-throughput, genome-wide annotations of total enzyme repertoires in organisms. We illustrate this capability by experimentally validating DEFT predicted glycoside hydrolase (GH) profiles of intestinal mucus associated bacteria.
1. 这段在干什么
这是摘要,先指出「序列法」和「结构法」各自的短板,再提出 DEFT:用序列法定前两层 EC,用结构法定后两层,并声称在基准测试中更准。
2. 需要解释的地方
3. 值得留意
作者的洞见是「先粗后细」——前两层靠序列法压下假阳性,而非纯靠结构。全文的落点在最后一句:不只是比准确率,还强调能高通量、全基因组地注释,并做了实验验证。
b005Enzymes are ubiquitous proteins that catalyze chemical reactions of living cells. Enzymes are classified using a hierarchical numbering system called Enzyme Commission (EC) numbers that describe the chemical reactions the enzymes catalyze, from a general reaction type (e.g., breaking bonds, transferring chemical groups, etc.) to more specific aspects such as chemical bonds and substrates involved in the reaction. We present a new machine learning method for predicting EC numbers called Deep Enzyme Function Transfer (DEFT). This method improves on previous methods that use either protein sequence- or three-dimensional (3D) structure-based comparisons between enzymes of known and unknown classification. DEFT combines the strengths of both approaches by first using a protein sequence-based model to predict the general enzyme category and then using protein structure comparisons to predict the finer subcategories. We demonstrate that DEFT achieves superior accuracy compared with current state-of-the-art tools for EC number prediction. We next demonstrate how DEFT’s computational efficiency enables us to perform high-throughput, genome-wide annotations of organisms’ enzyme repertoires. We illustrate this capability by experimentally validating DEFT predicted sugar metabolizing enzyme profiles of intestinal mucus associated bacteria.
1. 这段在干什么
这是 Author summary 的收尾段,向非专业读者概述全文:提出新方法 DEFT 用于预测酶的 EC 编号,并宣称它比现有工具更准、还能做全基因组规模注释。
2. 需要解释的地方
3. 值得留意
DEFT 的核心卖点是把两类老方法的优点拼起来(序列 + 结构),而不是单靠一种。注意"superior accuracy""computational efficiency"都是作者自述的结论,原文未给具体数字。
b006Enzymes are ubiquitous proteins that catalyze chemical reactions of living cells. Enzymes are classified using a hierarchical numbering system called Enzyme Commission (EC) numbers that describe the chemical reactions the enzymes catalyze, from a general reaction type (e.g., breaking bonds, transferring chemical groups, etc.) to more specific aspects such as chemical bonds and substrates involved in the reaction. We present a new machine learning method for predicting EC numbers called Deep Enzyme Function Transfer (DEFT). This method improves on previous methods that use either protein sequence- or three-dimensional (3D) structure-based comparisons between enzymes of known and unknown classification. DEFT combines the strengths of both approaches by first using a protein sequence-based model to predict the general enzyme category and then using protein structure comparisons to predict the finer subcategories. We demonstrate that DEFT achieves superior accuracy compared with current state-of-the-art tools for EC number prediction. We next demonstrate how DEFT’s computational efficiency enables us to perform high-throughput, genome-wide annotations of organisms’ enzyme repertoires. We illustrate this capability by experimentally validating DEFT predicted sugar metabolizing enzyme profiles of intestinal mucus associated bacteria.
1. 这段在干什么
这是作者自述的方法概述段:先交代酶与 EC 编号是什么,再抛出本文新方法 DEFT,说明它比旧方法强在哪、强到什么程度,最后点出它能做全基因组规模化注释。
2. 需要解释的地方
3. 值得留意
作者称 DEFT 精度"优于当前最先进工具",并强调其计算效率带来高通量能力——但具体数字这段没给,要看正文。
b007Citation: Erden M, Schult T, Yanagi K, Sahoo JK, Kaplan DL, Cowen LJ, et al. (2026) Fast structural search for classification of gut bacterial mucin O-glycan degrading enzymes. PLoS Comput Biol 22(9): e1014034. https://doi.org/10.1371/journal.pcbi.1014034
这一段本身只是Citation(引用信息),不是正文——它列出作者、期刊、年份、DOI 等,作用是被引用/溯源,不含论证或过渡。
值得留意:上一段结尾讲"通过实验验证 DEFT 预测的……酶谱",而这段没有任何内容承接那句话,所以别指望从这里读出结论——它只是文献条目。
需要解释:DOI 是数字对象标识符,相当于论文的永久链接;这是领域常识,不是本文观点。
b008Editor: Boyang Ji, BioInnovation Institute, DENMARK
这一行不是论文正文,而是期刊排版中的编辑信息行:
1. 这段在干什么:标明处理这篇稿件的责任编辑是 Boyang Ji(任职于丹麦 BioInnovation Institute),属于期刊元数据,不承载论文的科学内容。
2. 需要解释的地方:PLoS 系列期刊会在文中列出 Academic Editor,"Editor: 姓名, 单位, 国家"就是这条格式(领域常识,非本文观点)。
3. 值得留意:它与上下文的正文、Author summary 无逻辑关系,别误读成作者、合作者或机构署名;真正的作者信息在标题下方。
b009Received: February 16, 2026; Accepted: September 15, 2026; Published: September 25, 2026
好,看这一段。
1. 这段在干什么
这不是正文,是论文的投稿记录栏,列出收稿、接收、见刊三个日期。
2. 需要解释的地方
3. 值得留意
三个日期都在 2026 年,且从投稿到见刊仅约 7 个月,说明审稿周期较快。但这只是时间信息,不含任何研究内容,别当成摘要来读。
b010Copyright: © 2026 Erden et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
这是论文的版权与开放获取声明,不是正文内容。它说明文章版权归作者、以 CC 署名许可发布,任何人可自由使用。
这段完全没提研究方法、菌群或酶的任何信息。它属于期刊固定模板文字,别当作论文观点来读;引用或转载时唯一要遵守的是"署名"这一条件。
b011Data Availability: The code for DEFT is available at https://github.com/merterden98/DEFT Model training weights and data are archived at https://zenodo.org/records/17858733. All other data are included in the manuscript.
这段在干什么:这是数据可用性声明,交代代码、模型权重和数据的存放位置,方便他人复现。
需要解释的地方:
值得留意:它明确区分了两类材料——代码放在 GitHub,训练权重和数据归档在 Zenodo,其余数据则"随文收录",没给外部链接。这句话承接上行授权声明,但内容上与之无关,是独立的信息块。
b012Funding: This work was supported in part by a grant from the Army Research Office (W911NF2210185) (to K.L. and D.K.) and the Karol Family Professorship (to K.L.). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
这段在干什么:这是论文的基金声明段,交代研究经费来源及资助方角色,属于投稿必需的声明性内容。
需要解释的地方:
值得留意:两个资助号只挂在 K.L. 和 D.K. 名下,暗示其余作者由其他经费支持,但这段没说明。
b014Enzymes are the main catalytic units in metabolism. Collectively, the enzymes of an organism describe the biochemical reactions that are possible for the organism since almost all cellular reactions are enzyme-catalyzed. The standard [1] way of describing the catalytic function of an enzyme is to assign an Enzyme Commission (EC) number to the enzyme, which, by construction, is hierarchical. Unlike an accession number in UniProt, which is unique to the corresponding protein, the same EC number can be assigned to two or more proteins if the corresponding enzymes catalyze the same reaction. In addition, a single enzyme can be assigned more than one EC number if it is multifunctional and catalyzes multiple reactions. In some cases, a single EC number represents multiple reactions of a similar type. In other cases, a single EC number can represent two or more distinct reactions because it is associated with an enzyme that is pleiotropic or exhibits substrate promiscuity. The complexity of relationships between enzymes, EC numbers, and reactions [2] has presented challenges in developing efficient algorithms to systematically classify enzymes based on their catalytic functions.
这段在干什么:交代背景——酶功能的标准化描述是 EC 号,而酶、EC 号、反应三者关系复杂,为后文做铺垫。
需要解释的地方:
值得留意:作者反复强调「多对多」的复杂关系,正是为了引出——现有分类算法难做,这正是本文要解决的痛点。
b015Enzymes are the main catalytic units in metabolism. Collectively, the enzymes of an organism describe the biochemical reactions that are possible for the organism since almost all cellular reactions are enzyme-catalyzed. The standard [1] way of describing the catalytic function of an enzyme is to assign an Enzyme Commission (EC) number to the enzyme, which, by construction, is hierarchical. Unlike an accession number in UniProt, which is unique to the corresponding protein, the same EC number can be assigned to two or more proteins if the corresponding enzymes catalyze the same reaction. In addition, a single enzyme can be assigned more than one EC number if it is multifunctional and catalyzes multiple reactions. In some cases, a single EC number represents multiple reactions of a similar type. In other cases, a single EC number can represent two or more distinct reactions because it is associated with an enzyme that is pleiotropic or exhibits substrate promiscuity. The complexity of relationships between enzymes, EC numbers, and reactions [2] has presented challenges in developing efficient algorithms to systematically classify enzymes based on their catalytic functions.
这段在干什么:为全文做背景铺垫——先立起"酶的功能靠 EC 号描述"这一标准做法,再指出酶、EC 号、反应三者关系复杂,从而引出"系统分类酶功能很难"这一问题。
需要解释的地方:
值得留意:作者反复强调"同一 EC 号可能代表不同反应",这正是在暗示:只靠 EC 号分类不可靠,为后文提出结构搜索方法埋伏笔。
b016Recently, several deep learning methods have been developed to predict the EC number of an enzyme directly from its amino acid sequence [3–7] but even the best of these methods struggle to correctly assign EC numbers, especially down to the lowest level (fourth digit). When predicting the function of an enzyme, its protein structure adds valuable information because the catalytic mechanism depends on the shape and physicochemical properties of the enzyme’s substrate binding site(s) [8]. Given this, it follows that structure-based methods could help with the EC number prediction task. However, even when an enzyme’s full 3D structure is available from Alphafold2 [9], Alphafold3 [10] or a related protein structure prediction tool (ESMfold [11], OmegaFold [12], etc.), it remains an open question how these structures can be used to effectively predict the enzyme’s EC number without resorting to computationally expensive docking methods. A naive approach is to find an annotated enzyme that is the closest 3D structure match to the enzyme of interest using some global structural alignment method (e.g., TMalign [13]), and then transfer the annotated enzyme’s EC number to the enzyme of interest. However, this approach can produce false positive annotations because enzymes having similar global structures can have distinct catalytic regions due to divergent evolution.
1. 这段在干什么
承接上段"酶功能分类很难",先说明深度学习直接预测 EC 号的不足,再解释为何结构信息有帮助,最后指出朴素结构比对会因"全局相似、局部催化区不同"而误判——为下文提出自己的方法做铺垫。
2. 需要解释的地方
3. 值得留意
作者埋了一个关键矛盾:结构有助于预测功能,但"全局结构相似"≠"催化区相似",这正是他后文要绕开的坑。
b017We introduce a new EC number prediction method, Deep Enzyme Function Transfer (DEFT), that combines the advantages of sequence-based deep learning and structure-based search to achieve superior enzyme classification performance. This new method leverages computationally predicted enzyme 3D structure in two ways. First, DEFT uses SaProt [14], a structure-informed protein language model (PLM), to predict the first two EC number levels (class and subclass), for an enzyme of interest. Then, DEFT uses the Foldseek [15] structure-based 3Di string to find the closest annotated structural match having the same first two EC number levels as the enzyme of interest. That matching enzyme’s full EC number (class, subclass, sub-subclass, and serial number) is then transferred to the enzyme of interest to complete the EC number prediction. We benchmark DEFT against several other EC number prediction tools and show that it achieves superior recall and precision scores on the benchmark datasets.
这段在干什么:提出新方法 DEFT——用「结构感知语言模型 + 结构搜索」两招组合来预测酶的 EC 号,并用基准测试说它更强。
需要解释的地方:
值得留意:它只借用结构来找匹配,而非序列;前两级靠模型,后两级靠“抄”最近邻——分工是这方法的关键。
b020Enzyme annotation by DEFT can be split into three distinct steps: coarse prediction, alignment, and filtering (Fig 1). We outline each step below.
这段是典型的方法总览段——承接上句对DEFT的夸赞后,立刻把它的内部流程拆成三步,为下文逐条展开铺路。
关键概念
值得留意
三步是"由粗到精"的漏斗结构,作者特意点明(Fig 1),说明后面会配图讲。这段只给流程名,具体做法全在下面,别在此处找细节。
b021Enzyme annotation by DEFT can be split into three distinct steps: coarse prediction, alignment, and filtering (Fig 1). We outline each step below.
这段在干什么:紧接上段介绍完 DEFT 后,这里交代它的流程结构——把酶注释拆成三步,并预告下文将逐步展开。
需要解释的地方:
(三步的具体算法属领域常识层面,这段本身没展开。)
值得留意:作者用"can be split into three"强调这是可分解的流水线,三步有先后顺序——这一点决定了后文的叙述结构,Fig 1 是它的可视化对照。
b022In step 1, the primary amino acid sequence (AA) and structural representation (3Di) of the enzyme of interest are merged and passed to a fine-tuned protein language model (PLM) to predict the first two levels of the EC number hierarchy. In step 2, the enzyme’s structure is aligned against a reference database. In step 3, the predicted first two digits of the EC number (’prefix’) from the PLM is used to filter out the alignments, where an alignment is kept if the prefix matches the first two digits of the predicted EC number. The final output is a ranked list of EC number predictions for the enzyme of interest sorted by E-values, where the E-value correlates with the probability that the EC number match occurred by chance. See the methods section for details of the E-value calculation.
这段就是按上一句说的三步,把 DEFT 的流程逐步骤讲了一遍。
b023https://doi.org/10.1371/journal.pcbi.1014034.g001
这一段实际只有一条图注链接(指向论文图1),没有正文文字。
1. 这段在干什么:它本身不承担论述,而是给出图1的链接,正文对图1的解读要靠你自己点开看图。
2. 需要解释的地方:doi.org/...g001 是期刊给每张图分配的独立DOI,点开即图1,属领域常识,不是这段文字讲的。
3. 值得留意:你说的「上一段结尾」在这段里没有出现,这里也没有任何关于E-value、EC号的文字。别把上下文读成这段的内容。
如果方便,把图1的图注文字贴给我,我再按同样要求讲。
b024The EC number system classifies enzymes using a four-level hierarchy based on the reaction they catalyze, with each digit providing increasing specificity. Given the protein sequence and 3D structure of an enzyme of interest, we use the SaProt [14] PLM embedding to learn how to predict the first two EC number digits. The Fine Prediction step, described next, completes the EC number prediction. SaProt’s input alphabet pairs each amino acid with a structural state , representing a protein of length m as a sequence . The amino acid alphabet contains the 20 standard residues plus a mask token (21 tokens), and the structural alphabet contains the 20 Foldseek 3Di states plus a mask token (21 tokens). The combined paired vocabulary therefore has unique tokens, each representing one (residue, structure) pair that can appear at a given position. Fine-tuning uses LoRA [21] and a standard multilayer perceptron classifier trained via cross-entropy loss.
这段在干什么:交代作者用 SaProt 蛋白语言模型从序列+3D结构预测 EC 号前两位,为下一步 Fine Prediction 做铺垫。
需要解释的地方:
值得留意:配对词表规模原文没给具体数字,只说"unique tokens"。
b025In our coarse classification step, we ignore the hierarchy in the first two EC number levels and instead treat each EC number prefix (through level 2) as a separate class. Thus the coarse classification task assigns each enzyme to one of 78 classes. Freezing the granularity at level 2 and not attempting to predict the entire 4-level EC number as separate stand-alone labels at this stage both reduces the variance that would emerge from training a model with a large number of classes, and allows us to take advantage of the hierarchical EC number structure in the next, fine classification step.
这段在干什么:交代粗分类的具体做法——把前两级 EC 号当独立类别(共 78 类),并说明为何不一次性预测完整 4 级。
需要解释的地方:EC 号是酶的功能编号,分 4 级(领域常识)。这里被"截断"到第 2 级来分类;每个二级前缀算一类,于是有 78 类。作者只做粗分类,细分类留到下一步。
值得留意:作者给出的两个理由——类少则方差小、能利用层级结构——是设计动机,不是实验结论。数字 78 是类别数,不是样本数。
b026Taking the coarse prediction above, DEFT next leverages the hierarchical structure of EC numbers and completes the final two levels using structural alignment. Specifically, DEFT performs a local alignment against a reference database on the 3Di representation of the enzyme returning all sufficiently strong matches, to be filtered in the final step. We utilize Foldseek’s optimized implementation of the Smith-Waterman algorithm [22] for structure matching. In practice, the reference database can be any corpus of 3D structures, for DEFT we utilize the AlphaFold predicted structures of the CLEAN 50% sequence homology training set as to limit leakage of data during evaluations.
这段在干什么:承接上一步的粗分类,说明 DEFT 如何利用 EC 编号的层级结构,用结构比对补齐最后两级。
需要解释的地方:
值得留意:参考库用 AlphaFold 预测结构而非实验结构,且特意选 CLEAN 50% 同源训练集,目的是减少评测数据泄漏——这点容易被读漏。
b027After the coarse prediction and alignment step, entries with the lowest E-values are selected and their annotations are transferred as the final prediction. In the Results section showing enzyme classification performance on benchmark data, we compare the final prediction to the baseline structural alignment method that predicts all four levels of the EC number solely based on the closest Foldseek match.
1. 这段在干什么
交代方法流程的最后一步:粗预测+比对后,挑 E-value 最低的条目,把它们的注释作为最终预测;并说明在结果部分会拿它和只用最近 Foldseek 匹配的基线方法做对比。
2. 需要解释的地方
3. 值得留意
"solely based on the closest Foldseek match"——基线只取一个最近匹配,而本方法先粗筛再挑最低 E-value,走的不是同一条路径,这个对比框架是理解性能差异的关键。
b028To rank DEFT’s predictions we rely on Foldseek E-values, which are analogous to standard sequence alignment E-values but differs in several important ways. Foldseek E-values are defined as follows.
1. 这段在干什么
承接上段(DEFT 用 Foldseek 最近邻预测 EC 号),这里说明:给 DEFT 预测排序,靠的是 Foldseek 的 E-value,并指出它和序列比对的 E-value 相似但不完全相同,随后引出其定义。
2. 需要解释的地方
3. 值得留意
作者特意强调"类似但有重要差异",但差异是什么,这段没展开——显然留给下面的定义。读时要盯住"differs in several important ways"这个伏笔。
b029where S is a raw alignment score as reported by the alignment algorithm, and K are statistical parameters that control decay rate and densities respectively, and m and n are the lengths of the aligned sequences. Foldseek [15] dynamically controls the hyperparameters K and and provides a predetermined substitution matrix S that is calibrated for the structural alignment task. The resulting Foldseek alignment scores (which we utilize unchanged) are similar to E-values seen in primary sequence alignment, where lower values imply a statistically more likely match, i.e., the probability a random pair of sequences could achieve this alignment is lower. However, because the alignment scores from Foldseek are based on different hyperparameters, they do not have the same distribution as sequence alignment E-values, and a direct numeric comparison cannot be made. Thus, to estimate E-values for each match, Foldseek trains a neural network to predict the mean and scale parameter of the extreme value distribution for each query.
这段在干什么:交代 Foldseek 的 E-value 是怎么来的,并说明它和序列比对 E-value 为什么不能直接比大小。
需要解释的地方:
值得留意:作者明说 Foldseek 分数「我们原样使用」,但反复强调它和序列 E-value 分布不同、不能直接数值比较——这是理解全文方法的前提。
b031All chemicals and reagents were purchased from MilliporeSigma (Burlington, MA) or Thermo Fisher Scientific (Waltham, MA) unless otherwise specified. Seven anaerobic bacterial strains were selected for in vitro cell culture to experimentally evaluate DEFT’s predictions of mucin O-glycan utilization. The selected strains comprised two known mucin grazers, four non-grazers, and a subspecies that can conditionally degrade mucin O-glycans. The mucin grazers were Akkermansia muciniphila ATCC BAA-835 (Am) and Bacteroides thetaiotaomicron ATCC 29148 (Bt). The non-grazers were Lactobacillus plantarum ATCC 8014 (Lp), Lactobacillus reuteri ATCC 23272 (Lr), Bifidobacterium dentium ATCC 27678 (Bd), and Escherichia coli Nissle (EcN). The conditional mucin-grazer was a strain (ATCC 15707) of Bifidobacterium longum subsp. longum (Bl), a subspecies that has a large repertoire of GHs [23] and encodes a core-1 O-glycan degradation pathway [24–26]. Except EcN (Creative Biolabs, Shirley, NY), all strains were sourced from ATCC (Manassas, VA).
1. 这段在干什么
交代实验材料来源,并列出被选来做体外培养验证的七株菌,说明它们各自在黏液 O-聚糖利用上的预期角色(吃 / 不吃 / 条件性吃)。
2. 需要解释的地方
3. 值得留意
这七株菌是拿来验证 DEFT 预测的,不是训练集;"条件性"那株的划分依据是它带 core-1 降解通路,原文给了引文支撑。
b032All chemicals and reagents were purchased from MilliporeSigma (Burlington, MA) or Thermo Fisher Scientific (Waltham, MA) unless otherwise specified. Seven anaerobic bacterial strains were selected for in vitro cell culture to experimentally evaluate DEFT’s predictions of mucin O-glycan utilization. The selected strains comprised two known mucin grazers, four non-grazers, and a subspecies that can conditionally degrade mucin O-glycans. The mucin grazers were Akkermansia muciniphila ATCC BAA-835 (Am) and Bacteroides thetaiotaomicron ATCC 29148 (Bt). The non-grazers were Lactobacillus plantarum ATCC 8014 (Lp), Lactobacillus reuteri ATCC 23272 (Lr), Bifidobacterium dentium ATCC 27678 (Bd), and Escherichia coli Nissle (EcN). The conditional mucin-grazer was a strain (ATCC 15707) of Bifidobacterium longum subsp. longum (Bl), a subspecies that has a large repertoire of GHs [23] and encodes a core-1 O-glycan degradation pathway [24–26]. Except EcN (Creative Biolabs, Shirley, NY), all strains were sourced from ATCC (Manassas, VA).
1. 这段在干什么
交代实验材料来源,并说明选了哪七株菌——用来实验验证 DEFT 对黏蛋白 O-聚糖利用的预测。
2. 需要解释的地方
3. 值得留意
作者按"觅食/非觅食/条件性"三类挑菌,正好对应上一段的预测,是为了后面做对照验证。EcN 不是 ATCC 来源,其余都来自 ATCC。
b034For all experiments, bacterial strains were outgrown from glycerol stocks onto pre-reduced agar plates and into 1 ml starter cultures in the appropriate recovery medium (100% MRS for Bifidobacteria and Lactobacilli, and 100% YCFAC for all other strains), and incubated for 24 h at 37° C. This initial growth stage was used to revive cells from cryostorage. Cultures were then pelleted (2,900, 10 min), washed in sterile 1 PBS, and transferred into 1 ml of 100% YCFAC for a second 24 h incubation to promote outgrowth and media adaptation. During these two stages, Am was further supplemented with 0.1% (w/v) autoclaved PGM to allow adequate growth. All washing and transfers were performed within the anaerobic chamber, while centrifugation was performed externally in anaerobically-sealed plates. Biological replicates were generated from independent agar plate colonies when feasible, or from separately inoculated tubes prepared from distinct glycerol stocks in the case of Am, which does not show robust growth on YCFAC agar plates. Cell-free media controls were included at each stage to monitor for contamination, and base YCFAC medium inoculated with each strain served as a substrate-independent growth control for cell viability.
这段在干什么:交代厌氧培养的完整操作流程,是方法学细节,为后续酶活/生长实验提供可复现的菌体来源。
需要解释的地方:
值得留意:洗涤、转管在厌氧舱内做,离心却在舱外——靠"厌氧密封板"维持无氧,这是容易被读漏的操作细节。
b035For all experiments, bacterial strains were outgrown from glycerol stocks onto pre-reduced agar plates and into 1 ml starter cultures in the appropriate recovery medium (100% MRS for Bifidobacteria and Lactobacilli, and 100% YCFAC for all other strains), and incubated for 24 h at 37° C. This initial growth stage was used to revive cells from cryostorage. Cultures were then pelleted (2,900, 10 min), washed in sterile 1 PBS, and transferred into 1 ml of 100% YCFAC for a second 24 h incubation to promote outgrowth and media adaptation. During these two stages, Am was further supplemented with 0.1% (w/v) autoclaved PGM to allow adequate growth. All washing and transfers were performed within the anaerobic chamber, while centrifugation was performed externally in anaerobically-sealed plates. Biological replicates were generated from independent agar plate colonies when feasible, or from separately inoculated tubes prepared from distinct glycerol stocks in the case of Am, which does not show robust growth on YCFAC agar plates. Cell-free media controls were included at each stage to monitor for contamination, and base YCFAC medium inoculated with each strain served as a substrate-independent growth control for cell viability.
1. 这段在干什么
交代整套厌氧培养流程:从甘油管复苏 → 两次 24 h 培养 → 洗涤转接,为后续酶活/生长实验准备标准化细胞。
2. 需要解释的地方
3. 值得留意
离心在厌氧室外进行,但用「厌氧密封板」;洗涤转接则全在室内——这是本段唯一的操作折中,别读漏。
b037Prior to LC-MS analysis, sugars were extracted from supernatant samples and derivatized using 3-Methyl-1-phenyl-2-pyrazoline-5-one (PMP). After thawing on ice, 25 L of each culture supernatant sample was mixed with 5 L of internal standard (D-Glucose-, 25 g/mL), followed by the addition of 60 L of ice-chilled methanol. After vortexing and centrifuging at 4000 rpm for 10 minutes, 60 L of each supernatant mixture was transferred to a well of a new plate. Then, 20 L of freshly prepared 0.5M PMP solution (87 mg/mL in methanol) and 16.6 L of ammonium hydroxide solution (20% in water) were added, followed by vortexing. The mixture was incubated at 60°C for 30 min to facilitate the derivatization reaction. The reaction was quenched with 16.6 L of formic acid, after which the mixture was vortexed and briefly spun down. After sequential extractions with chloroform, 40 L of the final aqueous phase was collected and diluted with 120 L of HPLC-grade water prior to LC-MS injection.
这一段是方法部分的样品前处理步骤:把培养上清里的糖提取出来,并用 PMP 做衍生化,让后续 LC-MS 能检测到。
要解释的
值得留意
b039Chromatographic separation of PMP-derivatized sugars was performed on a Gemini 5 m C18 110 Å column (250 x 2 mm; Phenomenex, Torrance, CA) using a gradient elution method. The LC-MS system was an 1260 Infinity III LC System (Agilent, Santa Clara, CA) coupled to a TripleTOF 5600 quadrupole-time-of-flight (QToF) mass spectrometer (AB Sciex, Framingham, MA). The mobile phases consisted of solvent A (acetonitrile:water, 5:95 v/v) with 20 mM ammonium formate (pH adjusted to 9.45) and solvent B (100% acetonitrile). The flow rate was set to 450 L/min, with an injection volume of 10 L. The column oven temperature was maintained at . The elution gradient started at 90% A, decreased to 80% by 2 min, to 75% A at 10 min, and to 5% A at 12 min. The solvent composition was held at 5% A until 20 min, then returned to 90% A by 20.5 min and maintained until 28 min. The mass spectrometer was operated in negative electrospray ionization mode. Data acquisition was performed using multiple product ion experiments (Table 1).
这一段的开头是仪器参数,你对照着看:
这段在干什么:交代 PMP 衍生化糖的 LC-MS 分离与检测条件,属于方法学细节,为后面糖谱数据提供可复现的凭证。上一段刚把样品处理好准备进样,这里接的就是"进样之后怎么跑"。
需要解释的地方:
值得留意:
b040Chromatographic separation of PMP-derivatized sugars was performed on a Gemini 5 m C18 110 Å column (250 x 2 mm; Phenomenex, Torrance, CA) using a gradient elution method. The LC-MS system was an 1260 Infinity III LC System (Agilent, Santa Clara, CA) coupled to a TripleTOF 5600 quadrupole-time-of-flight (QToF) mass spectrometer (AB Sciex, Framingham, MA). The mobile phases consisted of solvent A (acetonitrile:water, 5:95 v/v) with 20 mM ammonium formate (pH adjusted to 9.45) and solvent B (100% acetonitrile). The flow rate was set to 450 L/min, with an injection volume of 10 L. The column oven temperature was maintained at . The elution gradient started at 90% A, decreased to 80% by 2 min, to 75% A at 10 min, and to 5% A at 12 min. The solvent composition was held at 5% A until 20 min, then returned to 90% A by 20.5 min and maintained until 28 min. The mass spectrometer was operated in negative electrospray ionization mode. Data acquisition was performed using multiple product ion experiments (Table 1).
1. 这段在干什么
交代糖类 LC-MS 分析的具体仪器与色谱条件——接上一段"样品已备好、准备进样",说明进样后到底怎么跑。
2. 需要解释的地方
3. 值得留意
原文里流动相流速和进样体积的数字(450、10 等)单位丢失,柱温数值也是空的——这几处明显是排版/转录残缺,读时别当成作者省略。
b043We benchmarked DEFT against five state-of-the-art sequence-based methods for EC number prediction: ECPred [7], DeepEC [5], ProteInfer [4], DeepECTransformer [28], and CLEAN [6]. ECPred takes an ensemble approach to EC number prediction by combining multiple feature types. It integrates subsequence-level features, homology-based information, and biochemical properties into a weighted scoring scheme for final predictions. While effective for well-characterized enzyme families, ECPred’s reliance on pre-computed features limits its ability to generalize to novel enzyme architectures. DeepEC pioneered the application of deep learning to enzyme classification. The method represents protein sequences as one-hot encoded matrices of size (where m is sequence length and 21 accounts for the 20 amino acids plus unknown residues) and applies convolutional neural networks. DeepEC employs a multi-task architecture with three separate convolutional layers: one for binary enzyme/non-enzyme classification, and two others for predicting EC levels 1–3 and level 4, respectively, in a multilabel fashion. ProtInfer builds upon DeepEC’s convolutional approach but incorporates dilated convolutions to capture long-range sequence dependencies. By gradually increasing the receptive field size, ProtInfer can better model the relationship between distant residues that may be critical for enzymatic function. DeepECTransformer relies on a transformer based architecture that directly predicts EC labels without relying on a pretrained protein language model. CLEAN employs a contrastive learning framework that embeds protein sequences using ESM-1b followed by LayerNorm. The method learns representations where enzymes with similar EC annotations are brought closer together in embedding space while dissimilar enzymes are pushed apart. This approach has shown superior performance on standard benchmarks and provides well-curated evaluation datasets with controlled sequence similarity.
这段在干什么
逐一介绍五个基准对手,为 DEFT 的分类性能做铺垫对比。
需要解释的地方
值得留意
每介绍一个方法都顺带点了它的局限(ECPred 依赖预计算特征、难泛化到新结构),这是为后文 DEFT 的"superior"埋伏笔。
b044We used Foldseek [15] to detect candidate enzymes for structure-based alignment and EC number annotation. Briefly, the enzyme of interest’s 3D structure is first discretized with Foldseek’s tokenizer yielding a 3Di string, which describes the geometric conformation of each residue i in the protein backbone with its spatially closest residue j. The 3Di string is then aligned against our reference set, and the best match (as reported by E-values) has its EC number transferred to the unknown enzyme. In our benchmarking experiments we chose our reference set to be the enzymes in the training split.
介绍 Foldseek 做候选酶检测与 EC 号注释的流程,为后续 benchmark 铺垫方法基础。
原文指出参照集用的是训练集里的酶;被注释的是"未知酶",即测试对象——注意别把它误当成全体数据。
b045The benchmarking experiments followed the framework of the CLEAN study by Yu et al. [6]. The same two datasets, New-392 and Price-149, were used, with the same training/testing splits. The datasets comprise, respectively, 392 new additions to UniProt introduced after the curation of the SwissProt datasets and 149 proteins released by ProteInfer [4] as historically difficult to annotate. To afford comparisons with the previous study by Yu et al., we also used the original SwissProt splits from CLEAN for hyperparameter selection, taking care to ensure that there is no overlap between the training and evaluation sets. In addition to the five sequence-based methods, we also compared DEFT’s performance against a naive structure-based approach that transfers the entire EC number from an already annotated enzyme that is structurally most similar to the enzyme of interest (referred to as Foldseek in Table 2). Additionally, we compared DEFT against a hybrid approach that takes the CLEAN prediction, retains only the first two levels of EC number prediction, and then finds the annotated enzyme that is structurally most similar among those whose first two EC numbers match CLEAN’s prediction (referred to as CLEAN Prefix + Foldseek in Table 2). An alternative structural approach that fine-tunes the PLM head of SaProt [29] to predict all four EC number levels was also tested. Finally, we tested GraphEC [30], a recent method that operates on 3D graph representations of enzymes, as an example of a sophisticated structure-based approach designed specifically for the task at hand.
交代基准测试的设计:沿用什么数据集和划分,以及 DEFT 要对比哪几类竞争方法。
作者特意强调训练/评估集无重叠;对比对象既有序列法也有结构法,说明是要公平地证明结构法(DEFT)更强。
b046https://doi.org/10.1371/journal.pcbi.1014034.t002
你给的「原文」只有一条表 2 的 DOI 链接,没有正文文字,所以我没法逐句带读。就这段内容写:
1. 这段在干什么:这里只给出表 2 的链接,作用是引出支撑「DEFT 分类性能更优」这一论点的数据表;正文论述没有出现。
2. 需要解释的地方:
3. 值得留意:表本身不在这里,所有具体数字、指标、对比结论都这段没提到,要判断「superior」成立与否,必须点开表 2 再看,别凭标题下结论。
b047On both New-392 and Price-149 datasets, DEFT showed superior precision and recall performance compared with all other methods, whether they are sequence- or structure-based. On Price-149, DEFT achieved an F1 score of 0.72, compared with 0.48 for CLEAN and 0.61 for GraphEC. On New-392, DEFT achieved an F1 score of 0.84 compared with 0.59 on GraphEC. Overall, the performance by DEFT gave double digit improvements over both the best sequence- and structure-based methods on both benchmark datasets.
用两个数据集上的 F1 分数,论证 DEFT 的分类性能优于所有序列和结构方法。
b048We further analyzed the cross-validation results to characterize the dependence of annotation accuracy on the frequency of an EC number in the training dataset. Table 3 displays the precision and recall of DEFT stratified by number of times a particular EC number occurs in the training dataset. The precision and recall metrics shown in Table 3 compare DEFT’s cross-validation performance on the same SwissProt dataset with at most 50% sequence similarity between any two proteins. As expected, precision and recall of both DEFT and CLEAN declined as the frequency of a particular EC number in the training dataset decreased. However, the performance of CLEAN declined more sharply. Even for rare EC numbers (less than 5 occurrences), DEFT’s recall was , whereas CLEAN’s recall declined to .
这段在干什么:上一段刚夸完 DEFT 总体大幅领先,这一段用交叉验证按 EC 号出现频次分层,检查优势是否稳健——结果是:稀有 EC 号上 CLEAN 掉得更狠。
需要解释的地方:
值得留意:作者先说两者都随频次下降而变差(符合预期),真正想强调的是"CLEAN 下降更陡"。但原文两个空缺处没给出具体数值,所以"DEFT 稀有类召回是多少"这段无法回答。
b049https://doi.org/10.1371/journal.pcbi.1014034.t003
这一段的原文只有一个表格链接(DOI 指向 Table 3),没有可讲解的正文文字——我看不到表格里的数据、标题、行列或注释。
1. 这段在干什么
这段没提到。仅凭一个指向 t003 的 DOI 链接,无法判断它在文中承担什么论述功能;需要打开表格才能确认。
2. 需要解释的地方
3. 值得留意
只有链接、没有正文,说明表格是论据本体。真正要读的是表里的指标(如 recall、precision 之类)和它对应的酶类/EC 分组,别跳过表格只看正文。
建议把 Table 3 的标题和表头也贴出来,我再逐行带读。
b051A key problem in biology is searching for homologous proteins. Typical techniques to search for homologous proteins involve BLAST like queries on the primary amino acid sequence against a reference database. Foldseek intuitively extends a BLAST like search to also utilize structural information, while maintaining the computational benefits of utilizing character searching algorithms. Often, the purpose of such a search is to find similarly functioning proteins in other species. The challenge with such a search in the context of enzymes is that through evolutionary pressure catalytic sites tend to be more conserved than the rest of the protein. In such a case, the regions of two proteins that need to be aligned tend to be small portions of the proteins. As a result, the signal of finding similar enzymes on the basis of global alignment tends to be drowned out by noise.
1. 这段在干什么
为后文引出 DEFT 做铺垫:先讲同源蛋白搜索的常规做法和 Foldseek 的思路,再指出酶搜索的特殊困难——催化位点比整体更保守,导致全局比对信号被噪声淹没。
2. 需要解释的地方
3. 值得留意
"催化位点比别处更保守"意味着真正该比的是局部小片段,而不是整条蛋白——这正是后面方法要解决的问题。
b052We tested DEFT’s capability to perform genome-wide enzyme classification by characterizing the GH profiles of several gut bacteria. The bacteria were selected to represent anaerobes that inhabit the intestinal mucus in mammals, and included both mucin grazers (Am and Bt) and non-grazers (Lp, Lr, Bd, and EcN). The analysis also included a Bl strain that has a large repertoire of GHs [23] but has been shown to poorly degrade mucin O-glycans [31,32]. The inputs to the DEFT predictions were whole genomes of the bacteria and an expertly curated list of EC numbers representative of GH activities required for mucin O-glycan metabolism (Fig 2). For these predictions, we utilized publicly available AlphaFold2 predicted structures from EMBL-EBI [33]. Genome-wide scans using DEFT found high-probability (low E-value) matches for all of the key GHs in the mucin grazers, including alpha-fucosidase (EC number 3.2.1.51), alpha-N-acetylgalactosaminidase (3.2.1.49), beta-N-acetylhexosaminidase (3.2.1.52), alpha-acetylglucosaminidase (3.2.1.50) and neuraminidase (3.2.1.18). In contrast, these enzymatic functions were either not detected in the non-grazers or matched with weaker (orders of magnitude higher) E-values. Interestingly, the predicted GH profile varied among the mucin grazers, with Am having stronger (orders of magnitude lower) E-values than Bt for hydrolysis of terminal N-acetylhexosamines from glycan chains (EC number 3.2.1.52).
用多个肠道菌的全基因组扫描,验证 DEFT 能否做全基因组范围的酶分类。
b053The probability DEFT found a matching enzyme for an EC number of interest by chance is expressed by an E-value. See Methods for E-value definition. A scaling of is applied to emphasize the strength of the match. A gray square indicates that DEFT did not find any matching enzymes in the organism for the EC number (see text for abbreviations of organism names). E-values less than were truncated to - values of 100.
1. 这段在干什么
给图配图例:说明图里 E-value、灰色方块、截断值分别代表什么,好让读者看懂 DEFT 的匹配强度。
2. 需要解释的地方
3. 值得留意
原文里几个具体数值(scaling 倍数、截断阈值)在给出的文本中是空缺的,被句首"less than"后直接断掉,别当成作者没写——很可能是排版或摘录丢失,读全文时需回原文核对。
b054https://doi.org/10.1371/journal.pcbi.1014034.g002
这段原文实际上没有正文,只有一张图的 DOI 链接(指向 PLOS 官网的 Figure 2),所以能讲的非常有限。
1. 这段在干什么:它只是把小节标题「DEFT facilitates genome-wide profiling of enzyme repertoires」对应的图 2 挂出来,本身不含文字论述。
2. 需要解释的地方:doi.org/...g002 是出版方给图片分配的永久链接标识(领域常识),点开才是真正的图,图里通常画的是全基因组层面的酶谱分析结果——但具体内容这段没提到,要看原图才知道。
3. 值得留意:你贴的「上一段结尾」里有一句 E-values less than were truncated...,句子本身不完整(缺了数字),说明那是转录/排版截断,不是原文全貌。另外,仅凭一个图链接,无法判断作者在本段想论证什么;需要结合图注和图本身,才能衔接上下文。这点要留意,别把图的含义当成正文说出来的。
b056We experimentally evaluated the predicted ability of selected gut bacteria to utilize mucin O-glycans as metabolic substrates by measuring their growth on various mucin supplemented culture media. The substrates used for the experiments were porcine gastric mucin (PGM), porcine mucin 2 (Muc2), and two mucin mimetics that have simpler, defined O-glycan profiles. The mimetics were Bombyx mori silk proteins modified with GalNAc (SA) or GlcNAc (SU) at the hydroxyl groups of serine and threonine (S/T) residue derivatives. Without mucin supplementation, bacterial growth did not stratify [27] by the ability to metabolize mucin O-glycans (Fig 3). Growth was slowest for Am, reaching an OD600 value less than 0.1. The fastest growing group comprised Bt, Lp, Lr, and Bd, which reached OD600 values between 0.35-0.45. Mucin supplementation had the largest growth promoting effect on Am. This is consistent with the predicted GH profiles (Fig 2) showing Am with the strongest E-values for enzyme functions (EC numbers) needed to hydrolyze common glycosidic bonds in mucin O-glycans [34]. With PGM or Muc2 supplementation, Am growth was 4-fold higher compared with the base YCFAC medium. Supplementation with SA or SU had a weaker, but still significant (36 to 41%) growth-promoting effect for Am. Mucin supplementation also enhanced the growth of Bt, which was predicted to encode a GH profile comparable to Am, albeit with a weaker E-value for EC number 3.2.1.52. The increases in Bt growth compared with the base medium ranged from 15 (SU) to 24% (Muc2). Mucin supplementation had no significant effect on growth of Bl, which had weaker E-values than Am or Bt for EC numbers (3.2.1.49, 3.2.1.52, and 3.2.1.50) corresponding to release of N-acetylhexosamines (Fig 2). As expected, the other non-grazers (Bd, Lp, Lr, and EcN) did not respond significantly to mucin supplementation.
这段在干什么:用生长实验验证 DEFT 的预测——能降解黏蛋白 O-聚糖的菌,在加黏蛋白的培养基里应该长得更好。
需要解释的地方:
值得留意:
b057All cultures were inoculated at an approximate OD600 value of 0.05. Final OD600 values were recorded after 24 h of culture in base (YCFAC) or mucin-supplemented medium (PGM, Muc2, SA, or SU). Data shown are mean SD of n = 4 biological replicates. Statistical significance was determined using two-way ANOVA followed by multiple comparisons with Dunnett’s adjustment relative to the YCFAC, non-mucin control. * p < 0.05; ** p < 0.01.
好的,我们来看这一段。它紧接着上一段对非黏液降解菌的描述,转向交代实验测量方法。
1. 这段在干什么:说明生长实验如何做、数据怎么算、显著性怎么检验。属于方法交代,为后面展示各菌对黏液补充的响应结果提供依据。
2. 需要解释的地方:
3. 值得留意:统计比较的基准是YCFAC非黏液对照,所以所有显著性都是"相对对照"而言,不是菌与菌之间比。另外终点是24小时的最终OD,反映的是生长结果而非速率。
b058https://doi.org/10.1371/journal.pcbi.1014034.g003
你给我的这段「原文」其实只有一个图片链接(PLOS Comput Biol 的图3),没有正文文字,所以我只能就这个链接本身来讲。
1. 这段在干什么:它指向论文的图3,是该小节展示实验数据的配图,正文靠它撑起「实验生长速率和糖测量验证了 DEFT 预测的酶谱」这一论证。
2. 需要解释的地方:链接里的 g003 是「Figure 3」的编号(这是 PLOS 期刊图片 URL 的命名惯例,属领域常识);doi.org/10.1371/journal.pcbi.1014034 是这篇论文的 DOI。
3. 值得留意:图注里那句「YCFAC, non-mucin control」「* p<0.05」你已在上段结尾看到——说明这幅图用的是相对 YCFAC 对照的多重比较(Dunnett 校正)。具体画了什么、有哪些菌和糖,这段没提到,得打开图本身才知道。
b059To further investigate the differential utilization of mucins by the bacteria, we measured the medium concentrations of sugars comprising the bulk of O-glycan core structures and termini. Targeted LC-MS assays showed significant increases in GalNAc, GlcNAc, N-acetylneuraminic acid (Neu5Ac), fucose, and galactose in Muc2-supplemented Am cultures (Fig 4A-4E). Except for fucose, the other four sugars were also significantly increased in Am cultures. However, the increases in GlcNAc and Neu5Ac were lower compared with Am cultures with Muc2 supplementation. Incubation of Am in SA- or SU-supplemented medium selectively increased the concentrations of GalNAc and GlcNAc, respectively (Fig 4A and 4B), consistent with predicted hydrolysis of the sugars from the mucin mimetics (Fig 4F). Compared with Am, Bt cultures showed only limited increases in sugar concentrations when incubated with mucins. Incubation of Bt with PGM significantly increased GalNAc compared with the cell-free control, while incubation with Muc2 did not increase any measured sugars. As was the case for Am, incubation of Bt in SA- or SU-supplemented medium selectively increased the concentrations of GalNAc and GlcNAc, respectively (Fig 4A and 4B). Besides Am and Bt, no other bacteria increased any measured sugars.
1. 这段在干什么
用糖浓度实测数据,验证 DEFT 预测的酶谱是否真实反映了细菌对黏蛋白的降解能力。
2. 需要解释的地方
3. 值得留意
b060(A–E) Changes in O-glycan core and terminal sugar concentrations following incubation of mucin-grazing and non-grazing gut bacteria in base YCFAC medium or YCFAC supplemented with mucin substrates (PGM, Muc2, SA, or SU). See text for sugar name abbreviations. Bars represent medium (cell-free) control subtracted concentrations. Positive values indicate net accumulation of monosaccharide; negative values indicate net consumption. Statistical significance was determined using two-way ANOVA followed by multiple comparisons with Dunnett’s adjustment relative to the YCFAC, non-mucin control. Data shown are mean ± SD of n = 2 biological replicates from N = 2 independent experiments. * p < 0.05; ** p < 0.01. (F) Putative top three O-glycan structures present in each mucin substrate (PGM [35], Muc2 [36], SA, or SU) and the potential EC numbers of GHs required to cleave the bonds.
这是图4的图注,交代A–E各柱状图的实验设计与读法,并预告F图给出各黏蛋白底物的O-聚糖结构与对应糖苷酶EC号。
正负号是读图关键:正值=单糖净积累,负值=净消耗;对照是YCFAC无黏蛋白组。重复数很小(n=2,N=2)。
b061https://doi.org/10.1371/journal.pcbi.1014034.g004
你给的“原文”其实只有一个图4的链接,没有正文文字,所以下面只能就这个引用本身讲。
1. 这段在干什么:这里插入了图4(PLOS Comput Biol 的 g004),按上下文,它应该是在把上一段预测出的酶谱(针对 PGM、Muc2、SA、SU 的 GH/EC)与实验测得的生长速率和糖测量结果对照,用来验证 DEFT 的预测。但具体图表内容这段没给,无法替它下结论。
2. 需要解释的地方:
3. 值得留意:图题和正文都没在这里出现,不要凭图号反推结论;真要用,得先打开图4看它的坐标轴、分组和统计。
b062Interestingly, Neu5Ac levels were significantly reduced relative to the cell-free control when Bd, Lr, or EcN were incubated in PGM-supplemented medium (Fig 4C). Except for Am, all other bacteria exhibited significant net consumption of fucose when incubated in PGM-supplemented medium (Fig 4D). As these trends were not observed when the bacteria were incubated in Muc2-supplemented medium, we investigated if PGM- and Muc2-supplemented fresh media had different sugar profiles. Targeted LC-MS analysis revealed that the PGM-supplemented medium, without exposure to cells, had significantly elevated levels of Neu5Ac and fucose compared with the base YCFAC medium (S1 Fig A). This suggested that fresh PGM-supplemented medium may contain free sugars from non-enzymatic degradation of glycan termini during sterilization or sample preparation. To test this possibility, we measured sugar concentrations in different dilutions of PGM (0.05-0.8% w/v). This analysis found that the sugar concentrations inversely correlated with PGM dilution (S1 Fig B). Fucose and Neu5Ac were detected at the highest levels (60 M at 0.8% w/v PGM). These results confirmed that PGM supplementation introduced significant levels of free sugars to the base medium, independent of bacterial activity. However, at the level of PGM supplementation used in this study (0.2% w/v), only fucose and Neu5Ac were present as free sugars at sufficient concentrations to support significant consumption by the cells (Fig 4C and 4D).
这段在解释一个反常现象:为什么在 PGM 培养基里,Neu5Ac 和 fucose 会被显著消耗(甚至降到低于无细胞对照)?作者的做法是——先怀疑培养基本身。他们用 LC-MS 测了没接触过细菌的新鲜 PGM 培养基,发现其中的 Neu5Ac 和 fucose 本来就比基础培养基 YCFAC 高,且随 PGM 浓度升高而升高,说明这些"游离糖"来自灭菌或样品制备时糖链末端的非酶降解,并非细菌的功劳。
关键术语:Neu5Ac(唾液酸的一种)、fucose(岩藻糖)是黏蛋白糖链末端的单糖;PGM 是猪胃黏蛋白,Muc2 是小鼠黏蛋白;LC-MS 是液相色谱-质谱联用,用来定量小分子。
值得留意:作者最后提醒,在本研究实际用的 0.2% PGM 浓度下,只有 fucose 和 Neu5Ac 的游离量足以支撑"显著消耗"——也就是说,观察到的糖消耗现象有一部分被培养基背景污染"放大"了,解释时要扣住这个浓度限定。
b063The profiles of non-glycan sugars qualitatively differed from those of glycan sugars (Fig 5). Significant increases in ManNAc were measured in SU- and SA-supplemented Am cultures (1.6- to 2.1-fold higher, respectively, than Am culture in base medium without mucin) (Fig 5A). However, ManNAc was not used to modify either silk protein. Mannose levels of mucin-supplemented Am and Bt cultures did not change significantly compared with the corresponding base medium controls (Fig 5B). Both Bl and Lr cultures produced mannose when incubated in SU- or SA-supplemented medium, whereas Bd cultures consumed mannose when incubated in PGM-, Muc2-, or SU-supplemented medium. Glucose, present as a component of the base medium (S1 Fig A), was substantially depleted by Bt, Bl, Bd, Lp, Lr, and EcN under all medium conditions, indicating a shared reliance on free glucose as a carbon source. By comparison, Am only depleted glucose in PGM- or Muc2-supplemented medium. Together with the lower OD600 value of Am in YCFAC and SA- or SU-supplemented medium, these results suggest that basal glucose consumption correlates with cell growth. Overall, the measured bacterial profiles of non-glycan sugars, unlike glycan sugars, did not cluster according to the predicted GH repertoire.
1. 这段在干什么
验证 DEFT 预测:非糖类单糖(如甘露糖、GlcNAc 类)的实测变化能否对上预测的酶谱——结论是不能。最后一句是落脚点。
2. 需要解释的地方
3. 值得留意
b064(a–c) Changes in non-glycan sugar concentrations following incubation of mucin-grazing and non-grazing gut bacteria in base YCFAC medium or YCFAC supplemented with mucin substrates (PGM, Muc2, SA, or SU). See text for sugar name abbreviations. Bars represent medium (cell-free) control subtracted concentrations. Positive values indicate net accumulation of monosaccharide; negative values indicate net consumption. Statistical significance was determined using two-way ANOVA followed by multiple comparisons with Dunnett’s adjustment relative to the YCFAC, non-mucin control. Data shown are mean ± SD of n = 2 biological replicates from N = 2 independent experiments. * p < 0.05; ** p < 0.01.
这是图 (a–c) 的图注,说明这几张图测的是什么、怎么算、怎么判显著——即用糖浓度变化来验证 DEFT 预测的酶谱。
上一段说非糖苷糖不按预测的 GH 谱聚类,这里却仍在用糖浓度"验证"预测酶谱——两组结果的张力值得留意。原文未点明。
b065https://doi.org/10.1371/journal.pcbi.1014034.g005
你贴的其实只有图注链接和一句图注文字,没有正文段落。我先按已有的这点内容讲:
1. 这段在干什么:这是图5的图注结尾,交代数据来源与统计显著性标注方式,本身不提出新论点,只是为图5的结果做说明。
2. 需要解释的地方:
3. 值得留意:n=2 的重复数偏少;且此段没有给出任何具体数值或结论,别把它当结果来读。
b067The present study introduces a novel enzyme classification method, DEFT, which takes a hybrid of two different strategies on coarse and fine prediction of EC number to substantially outperform existing methods. Applying the same benchmarking approach as the study by Yu et al. [6], we demonstrate that DEFT achieves substantial improvements in precision and recall, achieving F1 scores 1.7- and 1.5-fold higher than the next best performing method (CLEAN) on New-392 and Price-149 datasets, respectively. The hybrid strategy takes advantage of the fast structural search capabilities of Foldseek [15] to perform the fine classification of the last two EC number digits, while avoiding the pitfall of using structural alignment to assign the most general EC number levels (first two digits), instead using the SaProt PLM representation to learn the correct portions of the structure that are important for enzymatic activity.
1. 这段在干什么
这是 Discussion 的开头,总结本文提出的新方法 DEFT 的分类性能,并解释它为什么能赢——混合策略。
2. 需要解释的地方
3. 值得留意
作者特意说,不用结构比对去定"最笼统的"前两位,是为了避开一个坑(pitfall)——原文没具体说是什么坑,别自行脑补。另外对比基准和 F1 倍数都来自和 Yu et al. 相同的评测流程。
b068To prevent data leakage from influencing the performance evaluations, we screened out all proteins from the benchmarking data sets that have a 50% or greater sequence similarity with proteins in the training data set. It is possible to have high structural similarity without sequence similarity. However, if our method was only learning structural similarity, then our ablations that predict EC numbers solely based on structural similarity or the hybrid CLEAN+Foldseek method would outperform our method. This was not the case, as DEFT outperformed the ablations and hybrid method (Table 2).
回应审稿人可能的质疑:DEFT 的表现提升是否来自结构相似性而非酶功能学习,作者用消融实验反驳这一点。
作者先承认"序列不像也可能结构像",这是让步;但随即指出:若 DEFT 只是在学结构相似,那两个纯结构消融应更强——事实相反。这种"先承认、再反驳"是典型的防守式论证,容易被读漏。
b069SaProt-based assignment of the first two EC number levels, followed by Foldseek structural search-based assignment of the last two EC number levels, is very fast. Annotating 5,000 proteins takes under 5 minutes on a single NVIDIA H200 machine. This enables DEFT to be used for genome-wide profiling of an organism’s entire enzyme repertoire. As an illustrative use case, we used DEFT to predict the ability of representative gut bacteria to hydrolyze mucin O-glycans. The predictions were validated experimentally by analyzing bacterial growth and glycan sugar level changes in mucin-supplemented media. The experimental results independently verified the predicted mucin-grazing enzyme profiles of Am and Bt. DEFT also correctly predicted the non-grazing enzyme profiles of the remaining species. These results are in good agreement with previous studies reporting that Am and Bt can degrade mucin O-glycans and utilize the glycan sugars as substrates for growth [37,38].
1. 这段在干什么
总结 DEFT 的速度优势,并用肠道菌 mucin O-glycan 降解预测 + 实验验证作为应用范例,证明其预测准确。
2. 需要解释的地方
3. 值得留意
b070A key finding of our study is that mucin-grazing bacteria encode a more diverse set of GH functions compared with non-grazing bacteria. This insight would not have been obtained by focusing on a single GH and determining if an organism encodes the enzyme. We show that the breadth and diversity of the enzymatic repertoire characterize the mucin grazing phenotype. By capturing the genome-wide functional landscape, DEFT generates a robust biological signal capable of reliably differentiating mucin grazers from non-grazers. To test whether existing methods could recover these findings, we repeated the genome-wide scans for GHs using CLEAN and Foldseek, where the latter served as a representative method for structural homology search. Both CLEAN (S2 Fig) and Foldseek (S3 Fig) incorrectly predicted GH activity for E. coli, whereas DEFT correctly avoided these false positives. These results suggest that existing methods lack the necessary specificity to reliably characterize the functional repertoire diversity of glycohydrolases across different bacteria reported in our study.
总结本文核心发现——「食黏蛋白菌」的糖苷酶(GH)功能谱更广更多样,并用与 CLEAN、Foldseek 的对比证明 DEFT 才能准确区分。
在 *E. coli* 上,两个对照方法都误报了 GH 活性,DEFT 没误报——这是作者用来支撑「既有方法特异性不足」的关键证据。
b071In this study, the experimental validation focused on the organismic phenotypes predicted by our tool as the ability to rapidly and accurately scan entire genomes for enzymatic functions of interest is the key novel feature. For the seven gut bacteria tested in the study, we found good agreement between the predicted O-glycan degradation enzyme profiles (Fig 2) and experimental O-glycan sugar profiles (Fig 4). At the enzyme level, we performed computational validation on benchmark data sets (Tables 2 and 3). We note that the benchmark data may not fully capture the sequence/structural diversity of the mucin-degrading GHs identified here; nonetheless, based on DEFT’s benchmark precision/recall, we expect a broadly comparable error rate. Experimental validation for specific, mucin metabolism relevant GHs, which warrants a future study, could be done by performing enzyme activity assays using purified proteins under appropriate reaction conditions.
交代本研究的验证策略:主验证是整基因组扫描这一新功能,用7株肠道菌的预测表型与实验糖谱比对来支持。
酶层面只做了计算验证;特定酶的实验验证(酶活测定)作者明确留作未来工作——这是本文的局限,别误读成已验证。
b072Although Bl is a potential mucin grazer because it has an endo--N-acetylgalactosaminidase (engBF) [25] and an intracellular degradation pathway for core-1 structure (Gal1–3GalNAc) [24,26], the engBF encoded endo--N-acetylgalactosaminidase is substrate-restricted to act only on the core-1 structure [25], which is rare in natural mucins. A previous study reported that among four Bifidobacterium longum subsp. longum strains encoding engBF genes, only one strain (NCIMB8809) exhibited appreciable mucin O-glycan degradation activity [32]. Combined with the lack of a comprehensive O-glycan degradation enzyme repertoire in Bl—as revealed by DEFT’s genome—wide analysis-this likely explains why Bifidobacterium longum subsp. longum ATCC 15707 is unable to hydrolyze PGM or Muc2 glycans into monosaccharides. Using gut bacterial mucin O-glycan degradation as a case study, we demonstrate DEFT’s strong ability to infer complex organismic functions by predicting multiple catalytic activities (EC numbers) in a computationally efficient manner.
这段在干什么:解释为什么 *Bl* 菌株降解不了黏蛋白糖链——酶谱不全、关键酶底物受限,顺带用这个案例证明 DEFT 能推复杂功能。
需要解释的地方:
值得留意:engBF 虽是"潜在"降解酶,但底物范围窄,加上 *Bl* 缺整套酶,两者叠加才解释了 ATCC 15707 降解不了 PGM/Muc2——单一原因不够。
b073We note that a correct EC number is sometimes not a sufficiently fine specification of enzyme activity. Taking glycan chain degradation as an example, current EC number assignments do not clearly distinguish among functionally distinct activities. A representative case is EC number 3.2.1.52 [39,40], assigned to -N-acetylhexosaminidase. This enzyme has been reported to exhibit both endo- and exo-acting GH activities [41], as well as activity toward both GalNAc and GlcNAc. However, endo-acting GHs can, in some contexts, degrade mucins more effectively because they initiate cleavage within the oligosaccharide chain rather than acting only at the termini. Whether DEFT, provided with appropriate training examples, can capture such subtle yet critical activity distinctions warrants further study. Prospectively, users may provide a curated enzyme training set with pseudo–last (e.g., fifth) EC number digits encoding user-defined, context-dependent activity subclasses. Without retraining the entire enzyme classification model, DEFT should in principle be able to learn structural similarities and make predictions using the user-provided protein sequences along with structural information and custom subclass definitions. In this way, DEFT has the potential to bridge the gap between rigid EC number classification and the functional complexity inherent in biological systems, a challenge commonly encountered in studying biological degradation of both natural and synthetic polymers.
1. 这段在干什么
承认 EC 号有时太粗,提出 DEFT 未来可用「自定义亚类」来补足,展望而非已证实。
2. 需要解释的地方
3. 值得留意
通篇是 *warrants further study / in principle / has the potential*——全是展望,不是本文已验证的结果。别把它读成结论。
b074We note that current EC number assignments do not always provide a sufficiently fine-grained description of enzyme activity. Furthermore, experimentally validated enzymes may exist without any EC assignment. For example, a recently characterized sulfoglycosidase has been reported to release 6-O-sulfated -GlcNAc from sulfated mucins [42], but this enzyme has not yet received an official EC classification. As a result, newly identified enzymatic activities may not be represented within an EC-based prediction framework despite having demonstrated biochemical function.
1. 这段在干什么
指出基于 EC 编号的预测框架有个硬伤:EC 号本身描述不够细,且有些已实验验证的酶压根没 EC 号,所以新发现的酶活可能被漏掉。
2. 需要解释的地方
3. 值得留意
作者举的「已证实功能却无 EC 号」的例子,正是要论证:不是酶没功能,而是数据库没收录——这为全文用结构而非 EC 来搜索埋下理由。
b075Even when EC numbers are available, they may not fully distinguish among functionally important activity subclasses. Taking glycan degradation as an example, EC number 3.2.1.52 [40], assigned to -N-acetylhexosaminidase, encompasses enzymes reported to exhibit activity toward both GalNAc and GlcNAc. Differences in substrate preference or cleavage specificity may influence whether an enzyme can effectively access and depolymerize densely glycosylated mucin structures, thereby altering the rate and extent of mucin utilization by microorganisms. Similar limitations may arise in other enzyme classes. For example, cytochrome P450 enzymes [43] and signaling kinases (EC 2.7.11.1) often share EC classifications despite exhibiting substantial differences in substrate specificity or reaction selectivity. Consequently, EC-number-based annotations may not always capture the full functional diversity of these enzyme families.
1. 这段在干什么
紧接上句,用具体例子论证:EC 编号本身不够用,无法区分功能上重要的活性亚类。
2. 需要解释的地方
3. 值得留意
作者承认"即便有 EC 号也不够",语气比上句更退一步——不是没注释,而是注释粒度太粗。
b076An additional limitation of the current DEFT framework is that it focuses on the presence of enzyme functions and does not consider other differences. Enzymes with the same predicted function can vary in properties such as catalytic activity, substrate affinity, and expression level, which affect the enzymes’ biological function. Consequently, the predicted functional repertoire should be interpreted as an indicator of potential metabolic capability rather than a quantitative measure of enzymatic performance or phenotypic strength. Future studies may explore incorporating enzyme activity information, such as kinetic parameters or activity-based subclasses, to further refine functional predictions. For example, users may provide curated enzyme training sets with pseudo–last, i.e., fifth EC number digits encoding user-defined, context-dependent activity subclasses. However, such approaches would require dedicated training data and independent validation before their predictive value can be established.
这段在干什么:承接上段对 EC 编号的质疑,坦白 DEFT 框架本身的另一处短板——只看"有没有这个酶",不看酶的表现强弱。
需要解释的地方:
值得留意:作者把"有酶"和"酶好用"明确切开,提醒别把预测结果当作表型强弱的定量指标;提出的改进方向也只是"未来可探索",并没有验证过。
b078The code for DEFT is available at https://github.com/merterden98/DEFT Model training weights and data are archived at https://zenodo.org/records/17858733
这段是论文的数据与代码可用性声明,属于全文末尾的补充说明,交代 DEFT 的代码和训练数据去哪找。
需要解释的地方:
值得留意:代码和数据分开放在两个平台(GitHub 放代码、Zenodo 放权重和数据),这是常见做法,目的是让数据有长期稳定的引用链接。这段没提使用许可或版本号。
b081Free sugar concentrations in (a) mucin-supplemented fresh (cell-free) media and (b) different PGM dilutions in 1x PBS (0.05-0.8% w/v). The concentrations were determined using targeted LC-MS analysis as described in Methods.
1. 这段在干什么
这是 S1 Fig 的图注:说明图中两条曲线分别测的是「加黏蛋白的新鲜无细胞培养基」和「不同稀释度的 PGM 溶液」里的游离糖浓度,数据来自 LC-MS。
2. 需要解释的地方
3. 值得留意
(a) 特意标注「cell-free(无细胞)」,说明测的是培养基本身的糖背景,不是细菌代谢后的结果——这点容易读漏。
b082https://doi.org/10.1371/journal.pcbi.1014034.s001
这一条其实没给出正文,只有一个补充材料的链接(S1 Fig 的 DOI)。所以严格按铁律,这段没有可讲的正文内容。
b085The probability CLEAN found a matching enzyme for an EC number of interest by chance is expressed by an E-value. See Methods for E-value definition. A scaling of is applied to emphasize the strength of the match. A gray square indicates that CLEAN did not find any matching enzymes in the organism for the EC number (see text for abbreviations of organism names). E-values less than were truncated to - values of 100.
1. 这段在干什么
这是在给 S2 Fig. 的图例(legend)写读图说明,交代图中颜色、数值和符号的含义,让读者能看懂该热图/矩阵。
2. 需要解释的地方
3. 值得留意
原文两处数值被抹掉了(“A scaling of ___”“E-values less than ___”),所以具体取什么标度、截到多少,这段里看不到,别乱补。
(约 150 字)
b086https://doi.org/10.1371/journal.pcbi.1014034.s002
这段其实只有一个 DOI 链接,指向 S2 Fig. 的补充材料(图本身)。它本身不展开论证,只是把上一段文字里提到的 CLEAN 匹配结果落到一张补充图上。
读者容易把这段当成结论,其实它只是"证据在这张图里"的指路,真正内容要点开链接看图,文字本身没给任何数字或判断。
b089The probability Foldseek found a matching enzyme for an EC number of interest by chance is expressed by an E-value. See Methods for E-value definition. A scaling of is applied to emphasize the strength of the match. A gray square indicates that Foldseek did not find any matching enzymes in the organism for the EC number (see text for abbreviations of organism names). E-values less than were truncated to - values of 100.
1. 这段在干什么
这是 S3 Fig. 的图注,说明该图用什么指标、什么符号来展示 Foldseek 的匹配结果。
2. 需要解释的地方
3. 值得留意
原文中"applied to emphasize"前的缩放系数、以及"less than"后的阈值都缺了数字,引用时别自行补。
b090https://doi.org/10.1371/journal.pcbi.1014034.s003
1. 这段在干什么
这段本身只是一条补充材料的链接(S3 Fig.),指向一张图,展示用 Foldseek 在代表性肠道厌氧菌里做结构搜索得到的匹配结果。
2. 需要解释的地方
3. 值得留意
原文这里只有一行 URL,没有任何图注、数值或结论。上一段提到的 E-value=100 截断、菌名缩写等,是上一段的内容,不能算到这段头上。
注意:这段没提到具体匹配数、物种名单或任何分析结论,别脑补。