b002Statistical couplings between protein sites have been enormously useful for protein structure prediction, but there is still debate on the selective forces that generate them. Here we explicitly predict the couplings considering selection for protein folding stability, both against unfolding and misfolding, which we model through the Stability constrained model of protein evolution (SCPE). The SCPE is based on contact interactions, and it predicts stability against misfolding through the Random Energy Model in the space of contact matrices. We adopt the minimum selection principle, assuming that the multivariate amino acid distribution has minimum Kullback-Leibler divergence with respect to a site-unspecific distribution that is not influenced by protein folding stability. We predict the couplings at first order, assuming that they are small (small coupling approximation, SCA), and obtain an explicit formula that relates them with the contact interaction matrix and with the structural properties of the pairs of sites: presence of native contact, distance along the sequence and buriedness of the sites. We test these predictions on a representative set of the Protein Data Bank. For pairs of sites in contact, both predicted and observed couplings are inversely related with the contact energies of the amino acid pairs, while for pairs not in contact this relation is positive, i.e. negative design destabilizes pairs of amino acids for short-range pairs of sites that can form wrong contacts with high probability. Our results suggest that the strongest couplings correspond to native contacts, but an even larger number of couplings are influenced by negative design. Intriguingly, the most informative pairs are formed by surface sites, suggesting a role for protein-protein interactions and functional motions. We interpret the site-unspecific couplings, which are not influenced by selection for protein folding stability, as the result of variations of global mutational or selective forces across the lineages where the proteins evolve. Consistently, the strongest global couplings are self-couplings between equal amino acids. Cys-Cys, which form disulfide bridges, show the strongest global coupling, followed by pairs of metabolically costly amino acids (Trp-Trp, Tyr-Tyr, Trp-Tyr) and positively charged amino acids that tend to interact with nucleic acids (Lys-Lys, Arg-Arg). Ser, Thr, Asn and Gln form a cluster of globally correlated polar amino acids.
1. 这段在干什么
这是摘要全文:先立论——用「折叠稳定性选择」来解释蛋白位点间的统计耦合,再给出模型(SCPE)、近似(SCA)、预测公式和 PDB 验证结果。
2. 需要解释的地方
3. 值得留意
b009The evolution of protein sequences is commonly modelled mathematically as a substitution process, i.e. a stochastic process that describes the probability that an amino acid at a given site in a protein is substituted by another one as the function of time. This formalism is very powerful for inferring phylogenetic trees and other evolutionary properties (Yang 2006). Substitution processes typically assume that different sites of the protein evolve independently of each other, which allows feasible computation of the likelihood function required for evolutionary inferences. However, it is known that the pairwise correlations that are neglected in this approach are key for predicting structural interactions between protein sites (Göbel et al. 1994; De Juan et al. 2013). The inference of these evolutionary correlations through neural networks lead to spectacular success in protein structure prediction (Jumper et al. 2021). Large part of these evolutionary correlations are thought to arise from the selective pressure for maintaining the functional structure of the protein and its folding stability (Zeldovich et al. 2007; Goldstein 2008; Liberles et al. 2012; Bastolla et al. 2017). Within this framework, one of us and co-workers developed a model that predicts the stationary amino acid distribution of a protein family with known native structures (Bastolla et al. 2006), from which we derived the stability-constrained protein evolution (SCPE) model, a substitution process for phylogenetic inference based on selection on protein folding stability (Arenas et al. 2015), which we explicitly predict based on a statistical physics model of the misfolded state (Minning et al. 2013). We later extended the SCPE model to enforce structure conservation, since several lines of evidence suggest that mutations that change the native structure are selected against more strongly than mutations that maintain the structure and only affect its folding stability. Following (Echave 2008), we predict the effect of mutations on the protein structure as the linear response of the protein represented as an elastic network model. We obtain in this way the structure and stability constrained protein evolution (SSCPE) models (Lorca-Alonso et al. 2025).
这段在干什么:交代本文的理论背景与来路——从「替换过程」假设位点独立,引出被忽略的相关性对结构预测很关键,再说明作者团队如何一步步做出 SCPE 和 SSCPE 模型。
需要解释的地方:
值得留意:相关性被认为主要来自「维持功能结构与折叠稳定性」的选择压力,这正是全文标题的由来。
b010For computational treatability, the substitution processes derived from the SSCPE models assumed that protein sites evolve independently of each other. The model takes into account structural interactions effectively, through the effects of mutations on folding stability and structural changes. The SSCPE models produce site-specific amino acid distributions that present qualitative differences at sites buried in the protein core, which tend to be more hydrophobic, more conserved and evolve more slowly, with respect to exposed sites that tend to be polar, have higher sequence variation and evolve more rapidly. These differences arise from the different constraints imposed by natural selection for folding stability, and they are consistent with observations from multiple sequence alignments (MSA) (Echave et al. 2016).
1. 这段在干什么
为方便计算,SSCPE 模型先假设各位点独立演化,但通过突变对折叠稳定性的影响,把结构相互作用间接纳入进来——为后文讨论位点协变铺垫。
2. 需要解释的地方
3. 值得留意
作者承认"独立演化"只是计算上的权宜,但坚称结构效应已被有效吸收,这是全文立论的关键前提。最后一句用 MSA 观测佐证,说明预测与实证一致。
b011Adopting the SCPE model, we relaxed the assumption that protein sites evolve independently and assume that the stationary amino acid distribution presents pairwise couplings described by the Potts model (Wu 1982) pioneered by Lapedes et al. (2002); Weigt et al. (2009); Marks et al. (2011) for inferring evolutionary correlations and studied by several groups (for instance, Morcos et al. 2011; Skwark et al. 2013; Ekeberg et al. 2013; Kamisetty et al. 2013).
1. 这段在干什么
承接上一段的 SCPE 模型,作者在此放宽"位点独立演化"的假设,引入 Potts 模型来刻画位点间的成对耦合。
2. 需要解释的地方
3. 值得留意
作者一口气堆了多篇文献,都是支撑点:Potts 模型并非本文首创,而是被多个团队用过的既有方法。
b012The study of the pairwise stability constrained model in the framework of the "small-couplings" approximation (SCA) described below was one of the subjects of the PhD thesis of Jonas Minning (Minning 2012), directed by Markus Porto, at that time at the University of Darmstadt, with the external advice of Ugo Bastolla. Unfortunately, despite reiterate attempts, we have not been able to get in contact with Markus Porto, which is the reason why he does not appear among the authors although he is clearly one of the main contributors to this work. Should he read these lines, we will be more than happy if he gets in touch with us, and he is included in the authors list, as he deserves.
1. 这段在干什么
这是一段作者说明:交代该研究的来源(Jonas Minning 的博士论文,导师 Markus Porto),并解释为何 Porto 未列入作者名单。
2. 需要解释的地方
3. 值得留意
作者坦承 Porto 是"主要贡献者之一"却因联系不上而缺席作者名单——这在学术署名中不常见,也透露出本文与那篇博士论文的继承关系。
b013Here we consider the pairwise SCPE model within the SCA approximation, which provides an analytic formula that we test here, leaving the test of the general theory for future work. Despite its simplicity, the SCA constructively supports the hypothesis that site-specific selection for protein folding stability, together with global mutational and selective forces that act at the level of the whole protein, may produce the observed covariation among protein positions.
1. 这段在干什么
交代本文的研究范围(只做 SCA 近似下的成对 SCPE 模型,并检验其解析公式),并点出它支持的假设。
2. 需要解释的地方
3. 值得留意
b016Following (Weigt et al. 2009), we model the L-dimensional probability distribution that sites \(A_i\), \(i\in [0,L-1]\) in a protein family are occupied by the amino-acids \(a_i\) as the exponential distribution
这段在干什么:紧接上段提到的“选择压力可能造成位置间共变”,这里提出用指数分布这种概率模型来刻画一个蛋白家族中各位置氨基酸的联合分布。
需要解释的地方:
值得留意:作者明确说“Following (Weigt et al. 2009)”,即这套建模是沿用前人框架,不是本文首创;真正的重点在下一段才会展开。
b017We use the notation \(x_i= \nu i +a_i\), which has \(L\nu\) possible values where L is the number of sites and \(\nu\) is the number of amino acid types (21, since we consider gaps as an additional symbol). In the following, we indicate by sum over \(a_i\) the sum over the \(\nu\) amino acids at position i, and by sum over \(x_i\) the sum over all L sites and their possible amino acids \(a_i\). The single site parameter vector \(h_{x_i}\) has \(L\nu\) components and the pairwise coupling matrix \(g_{x_i,x_j}\) is an \(L\nu \times L\nu\) symmetric matrix whose elements with \(j=i\) vanish.
1. 这段在干什么
给能量模型(看起来像 Potts/Ising 模型)定义变量和参数的符号体系,为后面写分布函数做铺垫。
2. 需要解释的地方
3. 值得留意
\(g\) 是 \(L\nu \times L\nu\) 的大矩阵;对角项被设为 0,说明自耦合不计入——这点容易读漏。
b018In the physical literature, Eq.(1) is called "Potts model" (Wu 1982), and it has been extensively studied in the context of disordered systems before being recognized as the exponential pairwise distribution that describes a protein family.
这段是承接上文:上一段刚给完耦合矩阵 \(g\) 的定义,这里补一句它的物理出处。
需要解释的:
值得留意:作者强调这是先在物理中被研究、后来才被认作描述蛋白家族的分布——说明这套数学不是为蛋白发明的,而是借用。
b019In general, we can not normalize the Potts distribution either analytically or numerically, because computing the normalization factor \(Z=\sum _{a_0\cdots a_{L-1}} P(a_0,\cdots a_{L-1})\) requires an exponential number of operations. Thus, its parameters are usually inferred adopting approximated methods, such as the message passing algorithm originally proposed by Weigt et al. (2009), the regularized matrix inversion proposed by Jones et al. (2012), or the pseudo-likelihood approximation (Besag 1975), applied to this problem by Aurell and coworkers (Ekeberg et al. 2013) and (Kamisetty et al. 2013), among others.
1. 这段在干什么
承接上句把 Potts 分布认定为描述蛋白家族的成对分布后,指出它无法归一化,所以参数只能靠近似方法推断,并列出三条代表性路线。
2. 需要解释的地方
3. 值得留意
作者明说"解析和数值都不行",是指数级操作导致——这是整段论证的根。三个方法名不展开,只作引用定位。
b021For notational convenience, we define the propensity matrix
这段在干什么:这是一句过渡句,作者开始引入「propensity matrix(倾向性矩阵)」这个记号,为下文定义它做铺垫。
需要解释的地方:「notational convenience(记号的便利)」是论文常见说法,意思是"接下来为了书写方便先约定一个符号",本身没有科学含义。「propensity matrix」这段还没给出定义,只是一个名字,具体内容要看下一段。
值得留意:注意括号里是"(Ekeberg et al. 2013)",而上文写的是"Aurell and coworkers"——领域常识:Aurell 是通讯作者,Ekeberg 是第一作者,指同一批工作。但这段原文本身没解释这点。
b022that represents the ratio between the conditional probability of \(a_i\) conditioned to \(a_j\) and the unconditioned probability, minus 1. It is zero if the amino aids are independent, positive if amino acid \(a_j\) at site j favors \(a_i\) at site i, and negative if it affects it negatively. We call this quantity “propensity” between the two variables \(a_i\) and \(a_j\). This name is used in bioinformatics to describe the propensity of a given amino acid to take any secondary structure (Koehl and Levitt 1999). We also call \(\log (1+Q)\) “logarithmic propensity”. At first order in Q, the propensity and the logarithmic propensity coincide: \(\log (1+Q)\approx Q\).
定义「propensity(倾向性)」这个量:它怎么算、正负号怎么读、为什么借用「倾向性」这个名字。
「propensity」这名字是从二级结构倾向性研究借来的(同词不同义),别和本文的定义混。
b023The mutual information between sites i and j, \(M_{ij}\), is related to the propensities through
这句话是一个过渡句:它引出互信息和倾向性(propensity)之间的数学关系,为下文推导做铺垫。
需要解释的地方
值得留意
原文只说二者"相关(is related)",但具体关系式这里没给出,得看下文。
b024In the limit of small propensities the mutual information is approximated by the weighted mean square propensity.
1. 这段在干什么
接着上一段 \(M_{ij}\) 与 propensity 的关系式,给一个近似:当 propensity 很小时,互信息约等于加权的均方 propensity。
2. 需要解释的地方
3. 值得留意
关键词是「in the limit of small propensities」——这个近似只在弱关联时成立,作者没展开说误差有多大。
b026The distribution Eq.(1) presents more parameters than necessary, in the sense that different parameter sets produce exactly the same distribution. These extra parameters must be fixed through normalization conditions or "gauge" (Weigt et al. 2009; Ekeberg et al. 2013). We change parameters as \(p_{x_i}\equiv \exp (h_{x_i})\), \(q_{xi,xj}\equiv \exp (g_{x_i,x_j})-1\) and write the multivariate distribution as
这段在干什么:指出 Eq.(1) 的参数多于必要的,说明要用归一化条件(即"规范/gauge")把多余参数固定住,并给出两组参数替换式。
需要解释的地方:
值得留意:替换式只给了定义,后面的多元分布表达式在本段并未写出。
b027We adopt the following gauge conditions:
这段在干什么
一句话过渡:宣布接下来要设定一组规范条件(gauge conditions),为后面的参数化简铺路。
需要解释的地方
值得留意
"the following"指的是下一段,这里只有一句引子;真正的规范条件内容不在此段。
b028The sum over \(a_i\) denotes the sum over all possible \(\nu\) amino acids at position i. We call these conditions the "zero-mean" gauge because it sets to zero the weighted mean value of the couplings over the amino acids at one of the site, and we adopt it throughout this paper.
本段紧接「We adopt the following gauge conditions:」,具体给出这条规范条件的内容。
这段在干什么:定义并命名论文全篇采用的「零均值」规范条件。
需要解释的地方:
值得留意:原文只说把该位点氨基酸上的加权均值「置零」,权重是什么、零点从哪来,这段都没提。「全文沿用」意味着后面所有耦合值都相对于这个基准。
b029Note that the propensities satisfy the zero mean condition \(\sum _{a_j}p_{x_j}Q_{x_i,x_j}=\sum _{a_i}p_{x_i}Q_{x_i,x_j}=0\).
紧接上文提出的"加权平均"处理,用公式明确倾向性(propensities)满足零均值条件——这是全文采用的一项约定。
两个连等号分别是对 \(a_j\) 与 \(a_i\) 求和,说明任一位点都满足该条件,而非只对某一侧成立。这是领域里常见的"gauge 选择"手法,作者靠它把参量标准化。
b031We model the complete pairwise distribution through the combination of site specific parameters that represent natural selection for folding stability and global parameters \(p^\textrm{glob}_a\), \(q^\textrm{glob}_{ab}\) that are the same for all sites of the protein. These global parameters reflect the genetic code and the mutational process, which may be different in different genomes and different genomic regions. Indeed, considering the codon level instead of the amino acid level improves the inference of the couplings (Jacob et al. 2015). The global parameters also represent selective pressures that are uniform across the protein (see Discussion).
这段在干什么:提出一个建模框架——把完整的成对分布拆成“每个位点自己的选择参数”加“全蛋白共用的全局参数”。
需要解释的地方:
值得留意:全局参数不只管突变,还代表“全蛋白均匀的选择压力”,作者把它留到 Discussion 展开,这里只点名。
b032We represent the global pairwise distribution as \(P^\textrm{glob}_{a,b} = p^\textrm{glob}_{a} p^\textrm{glob}_{b} \left( 1 +q^\textrm{glob}_{a,b}\right)\) for all pairs of sites. It is easy to see that the global propensities \(q^\textrm{glob}_{a,b}\) obey the zero-mean gauge with respect to \(p^\textrm{glob}_{a}\).
1. 这段在干什么
上一段刚定义了“全局选择压力”,这段接着给它一个数学表达:把任意两个位点的联合分布写成各自边缘分布的乘积,再乘一个修正项。
2. 需要解释的地方
3. 值得留意
作者只说“容易看出”成立,但没推导。这里的 \(q\) 是全局参数,后面与位点特异的协变对比时,正是靠它当基准。
b033We determine the \(\nu +\nu ^2\) global parameters \(\{p^\textrm{glob}_a\},\{q^\textrm{glob}_{a,b}\}\) through a regularized fit, as explained in Methods.
这段在干什么
交代全局参数(线性项 \(p^\textrm{glob}_a\) 与二次项 \(q^\textrm{glob}_{a,b}\))的确定方式:通过正则化拟合得到,细节见 Methods。
需要解释的地方
值得留意
接上一段,这些 \(q^\textrm{glob}\) 满足关于 \(p^\textrm{glob}\) 的零均值规范——即参数并非独立自由,读时别把它们当作可任意取值的量。
b035The site-specific components of the modelled distribution reflect selection for maintaining the native structure of the protein and its folding stability. We determine them through the minimum selection principle (Arenas et al. 2015). This principle assumes that the complete distribution is as similar as possible to the global distribution, which is not influenced by selection for correct folding, in the sense that it has minimal Kullback-Leibler (KL) (Kullback and Leibler 1951) divergence from the global distribution, while constraining the average protein folding free energy \(\Delta G\).
从全局分布过渡到「位点特异性分布」,提出用最小选择原则来反推每个位点受折叠稳定性选择影响下的分布。
「minimal selection」是假设,不是推导出的结论;文中只说是为了体现折叠稳定性选择,并未给出求解细节——那应该在后文/Methods里。
b036The minimum selection principle generalizes the commonly adopted maximum entropy principle (see for instance Lapedes et al. 2002; Bastolla et al. 2006) if the amino acids are not equivalent. In fact, maximum entropy is equivalent to minimum KL divergence from the uniform distribution. Since the genetic code, the mutational process and the unspecific selective pressures differentiate the amino acids, the minimum KL divergence is preferable.
给上一段提出的最小选择原则做定位:说明它是更常见的最大熵原则的推广,并论证为何该用前者。
作者的关键理由:正因氨基酸不等价,用「相对均匀分布」的最小 KL 比用「最大熵」更合适——这是逻辑推理,不是实验结论。
b037This principle can be justified based on the analogy between molecular evolution and the Boltzmann distribution of statistical physics (Sella and Hirsh 2005; Mustonen and Lässig 2010; Puente-Sánchez et al. 2024). In particular, the evolutionary process, modelled as a stochastic process in genotype space (Fisher 1930), converges to a stationary exponential distribution analogous to the Boltzmann distribution in statistical physics, with the logarithm of the fitness playing the role of minus the energy and the inverse of the effective population size playing the role of temperature, meaning that small populations are more tolerant towards sequences with low fitness.
给上一段的结论(为何取最小 KL 散度)补一个理论依据:把分子演化类比成统计物理的玻尔兹曼分布。
作者只是"预设"了这个类比成立("can be justified"),并引用前人文献,并未在此推导。
b038In the present context, the inverse of the selective temperature, or selective strength, is represented by the Lagrange multiplier \(\Lambda\) that enforces the mean value of the protein folding free energy \(\Delta G\). This is the only free parameter of the stability constrained model. We fit \(\Lambda\) by maximizing the log-likelihood of the observed protein sequences with respect to the model.
好,我们来看这一段。
1. 这段在干什么
定义稳定性约束模型里唯一的自由参数 Λ,并说明它怎么定出来——用观测序列做最大似然拟合。
2. 需要解释的地方
3. 值得留意
ΔG 均值被"强制"(enforce),不是自然涌现;Λ 是倒数关系,别把方向和温度搞混。
b039Additionally, we impose the gauge conditions Eq.(5,6,7) through the Lagrange multipliers \(\mu _i\), \(\zeta _{x_i,j}\) and \(\eta _{x_j,i}\). These are not free parameters, since they are determined by imposing the constraints.
这段在干什么:紧接上一段拟合 \(\Lambda\) 的模型,补充说明还通过拉格朗日乘子 \(\mu_i\)、\(\zeta_{x_i,j}\)、\(\eta_{x_j,i}\) 施加上规范条件 Eq.(5,6,7)。
需要解释的地方:
值得留意:作者特意强调这些乘子"不是自由参数"——它们是约束的产物,不是可调的模型参数。读者别把它们误当成需要拟合的量。
b040We determine the site-specific Potts parameters \(\{p_{x_i}\},\{q_{x_i,x_j}\}\) of the stability constrained model by minimizing the objective function \(\Phi\), which is the sum of the KL divergence plus the functions that describe the constraints.
1. 这段在干什么
承接上段:说明那些非自由参数到底怎么定——通过最小化目标函数 \(\Phi\),反解出位点特异的 Potts 参数。
2. 需要解释的地方
3. 值得留意
约束不是硬性扣在参数上,而是作为 \(\Phi\) 的加项一起最小化,所以叫"稳定性约束模型"。
b041The matrix \(G_{x_i,x_j}\) represents the effective free energy of the pair \(x_i\) and \(x_j\) with respect to unfolding and incorrect folding (misfolding). These are not free parameters: We predict them with our model of protein folding stability, which considers the native contact matrix and the statistics of non-native contacts (Minning et al. 2013), as we describe in section 3.4. This computation is approximate since we approximate the free energy of the misfolded ensemble as a sum of pairwise terms, while the REM prediction involves terms that depend on the amino acid at four protein sites.
这段在干什么:给上一段目标函数里出现的耦合矩阵 \(G_{x_i,x_j}\) 定性——它不是自由拟合参数,而是由作者的折叠稳定性模型预测出来的。
需要解释的地方:
值得留意:作者主动承认这是近似——把错误折叠的自由能简化成两两配对之和,而 REM 预测涉及四个位点的氨基酸,精度上是有代价的。
b042We compute the optimal values of the site-specific parameters \(p_{x_i}\) and \(q_{x_i,x_j}\) by minimizing the score \(\Phi\), where \(\Lambda\), \(\{p_a^\textrm{glob}\},\{q^\textrm{glob}_{a,b}\}\) are pre-set global parameters and the external information comes from the predicted values of the effective free energies \(G_{x_i,x_j}\).
交代怎么定出位点参数:固定全局参数,通过最小化打分 \(\Phi\) 来求 \(p\)、\(q\),外部信息来自预测的自由能 \(G\)。
\(p\)、\(q\) 是“最优值”,但不是说它们有唯一解——是靠最小化 \(\Phi\) 得到的。另外要分清谁是固定、谁是待求:glob 参数固定,只有 \(p_{x_i}\)、\(q_{x_i,x_j}\) 在优化。
b043We impose the constraints Eq.(5,6,7) through the Lagrange multipliers \(\mu _i\), \(\zeta _{j,x_i}\) and \(\eta _{i,x_j}\), whose values are determined by the corresponding constraints. In this way, we determine the site-specific couplings and the single-site parameters as the function of the selective strength \(\Lambda\).
说明如何用拉格朗日乘子把约束(式5-7)落实,从而把耦合项与单点参数解成选择强度 \(\Lambda\) 的函数。
这段只讲"怎么解",没写解出的具体形式或结果;\(\mu_i,\zeta_{j,x_i},\eta_{i,x_j}\) 三个乘子各对应式5、6、7,原文未逐一说明谁对应哪条式子。
b044Finally, we determine the optimal value of the parameter \(\Lambda\) by maximizing the likelihood of the empirical pairwise distributions observed in the PDB for all types of pairs of sites.
前面已把耦合与单点参数写成 \(\Lambda\) 的函数,这段收尾:用最大似然定 \(\Lambda\) 的最优值,拟合对象是 PDB 中各类位点对的经验成对分布。
留意:它只提"似然最大化"这一准则,没给具体数值或优化算法——别自行脑补。
b045In our previous work (Arenas et al. 2015), we determined the single site parameters \(p_{x_i}\) assuming that all sites are independent, i.e. assuming that \(q_{x_i,x_j}\equiv q^\textrm{glob}_{a,b}\equiv 0\). Here we relax this assumption, as we also did in Minning (2012) within the small couplings approximation (SCA) discussed below.
承接上文「逐对位点似然」的语境,这段在交代模型假设的松紧变化:此前假设位点彼此独立,现在放弃这个假设。
留意:两个「先例」要分清——Arenas et al. 2015 是令耦合为零的旧做法,Minning 2012 则和本文一样放松了假设,这是本文方法上的直接来源。
b047In the numerical tests presented below, we adopt the small-coupling approximation (SCA) that neglects second and higher order terms in the couplings in the formula of the pairwise probability, i.e. we only consider the first order terms in \(q_{x_i,x_j}\). Because of the zero-mean gauge, the resulting values of the single site probabilities are \(P_i(a_i)=p_{x_i}\). The propensities Eq.(2) are
1. 这段在干什么
定义正文数值测试所采用的 SCA 近似:在配对概率公式里只保留耦合的一阶项。
2. 需要解释的地方
3. 值得留意
"only first order" 是全文近似的核心代价——高阶耦合被直接丢弃,后面结论都建立在此之上。
b048i.e. at the first order in q the propensities are equal to the direct couplings, and we neglect the correlations that arise from indirect couplings, although they can be recovered a-posteriori as we do when we predict the propensities of indirect contacts.
这段在干什么:承接上一段 propensity 的定义,说明在一阶近似下,propensity 就等于直接耦合,并交代间接耦合的处理方式。
需要解释的地方:
值得留意:作者坦白“间接耦合带来的关联被忽略了”,但强调可以事后再捞回来(a-posteriori),这正是后文预测间接接触要做的事。
b049The simplicity of the SCA approximation stems from the fact that the couplings of the site of pairs ij only depend on its properties of the pair, and not on the other pairs. This approximation is reminiscent of the first approach that was adopted to infer the couplings \(q_{x_i,x_j}\) through the mutual information, which in the small coupling limit is equal to the weighted mean square \(Q_{x_i,x_j}\). The seminal work by Weigt et al. (2009) and its sequels clearly showed that the approximation based on the mutual information is much less accurate for inferring the couplings than the approach based on the Potts model that also considers the indirect couplings. Nevertheless, we consider here the SCA because of two reasons: (1) The SCA provides an analytic formula that is easy to interpret and sheds light on the selective forces that are behind the couplings. (2) The resulting pairwise distributions depend only on the “local” properties of each pair ij, which allows grouping pairs with similar properties and performing the simple and exhaustive numerical tests that we describe below.
1. 这段在干什么
为 SCA 近似做定位:说明它简单在哪(耦合只依赖配对 ij 自身)、与互信息法的渊源,再给出作者仍采用它的两条理由。
2. 需要解释的地方
3. 值得留意
作者明知 SCA 精度不如 Potts,仍选它——用"可解释 + 可做穷举数值检验"换精度。这不是最好的方法,而是最好用的工具。
b050Using the approximation (9) and computing the derivatives of Eq.(8) with respect to the parameters \(\{p_i\}\) and \(\{q_{ij}\}\), we can explicitly determine the values of the parameters that minimize \(\Phi\) as
1. 这段在干什么
承接上段的近似思路,说明作者接下来要对目标函数 \(\Phi\) 求极值,从而显式解出最优参数。
2. 需要解释的地方
3. 值得留意
b051In the above equations, we denote with \(\textrm{ZM}\left( G_{x_i,x_j}\right)\) the zero-mean of the matrix \(G_{x_i,x_j}\), whose weighted mean value across the amino acids \(a_i\) and \(a_j\) is zero. Eq.(10) is the same that we obtained in our previous mean-field approximation in which we neglected the couplings and considered sites as independent, i.e. \(q_{x_i,x_j}\equiv 0\). The same formula continues to hold also at first order in the couplings, which does not change the single-site probabilities.
1. 这段在干什么
紧接上文给出的参数解,这里补充公式符号约定,并说明 Eq.(10) 与之前平均场近似的联系。
2. 需要解释的地方
3. 值得留意
作者强调:在一阶耦合下同一公式仍成立,且单点概率不变——说明耦合只影响两点项,不影响 \(\{p_i\}\),这是后面用协方差反推耦合的关键前提。
b054For assessing Eq. 11, we sampled pairs of sites from a representative subset of the protein data bank, which we grouped in various groups according to their properties.
这段在干什么:为验证上一节末尾的公式 Eq. 11,交代数据来源和抽样方式——从 PDB 的代表性子集中抽取位点对,并按性质分组。
需要解释的地方:
值得留意:"representative subset"(代表性子集)暗示作者没有用整个 PDB,而是做了筛选,但筛选标准这段没说;分组的具体依据同样留到后文。
b055Firstly, we distinguished four types of proteins according to sequence length (short, \(<150\) residues, and long) and average contact energy, related with hydrophobicity (hydrophobic, \(\overline{U}<0\) and polar). These groups allow us analyzing how selection for stability depends on the properties of proteins. Alternatively, we lumped together all proteins in the same group.
这段在干什么:交代分组方式——先按序列长度(短/长)和平均接触能(疏水/极性)分出四种蛋白,也提到另一种做法:把所有蛋白并成一组。
需要解释的地方:
值得留意:原文用了 "Firstly",说明分组方案不止一种;末尾 "Alternatively" 正是备选。分组依据是长度和接触能两个维度,别读成四类各自独立。
b056For every group of proteins, we grouped sites in two classes (buried or exposed) depending on their number of native contacts \(n_i=\sum _j C^\textrm{nat}_{ij}\), thus obtaining three types of pairs of sites: exposed-exposed, exposed-buried and buried-buried. Alternatively, we considered only one type of sites.
1. 这段在干什么
交代分析方案的分类方式:按溶剂暴露程度给位点分两类,从而产生三种位点对,或只取一类。
2. 需要解释的地方
3. 值得留意
分类阈值怎么定、接触数怎么算,这段都没说。
(约 150 字)
b057We then classified pairs of sites i, j according to their native contact \(C^\textrm{nat}_{ij}\): either pairs that form a native contact (C1), or pairs that form an indirect contact through another site (C2) or pairs that do not form direct or indirect contacts.
按分类维度分层:先按位点类型分(上一段),再按空间接触类型分(本段)。
1. 这段在干什么
在上一段按位点暴露程度分类之后,这一段换一个维度——按两个位点在天然结构里是否接触——给位点对重新分类。
2. 需要解释的地方
3. 值得留意
作者用「either … or … or」把三类并列,说明 C1、C2、无接触是互斥且穷尽的划分,后面做统计时应该是三组分开比较,而不是只分"接触/不接触"两类。
b058Finally, we grouped the distance of the sites along the sequence \(|i-j|\) (contact range). We found that five groups of contact ranges are sufficient to obtain good results, and we show this value in most of the figures.
交代方法收尾:把位点对沿序列的距离 \(|i-j|\) 分组(contact range),并说明为什么选 5 组——足够拿到好结果,后续多数图都用这个划分。
上一段刚把接触分成 C1 直接、C2 间接、无接触三类;这里换了个维度,按序列间隔分组——两组分类是并行使用的,别混为一谈。"5 组足够好"的依据这段没提。
b059As a result, we obtained 3 pairs of sites times 3 types of contacts times 5 contact ranges, yielding 45 types of pairs for each group of proteins.
这段是方法交代:告诉读者后续分析里“位点对”的类型总数是怎么算出来的。
关键概念:3 pairs of sites(每组的位点配对)× 3 种接触类型 × 5 个接触范围距离档,相乘得 45,即每组蛋白有 45 种位点对类型。这里的 contacts、contact ranges 具体指什么,这段没展开。
值得留意:三种因素是被当成独立维度做笛卡尔积——这个乘法结构意味着后文分析是按“位点对×接触类型×距离档”分层展开的,不是一个笼统的 45。至于这些类型最终如何用,得看下文。
b060For each group of proteins and for each group of pair of sites p, we sampled the pairwise amino acid distribution from protein sequences in the PDB and computed the observed propensities \(Q^\textrm{obs}_{p}(a,b)=F_p(a,b)/(F_i(a)F_j(b))-1\) where \(F_p(a,b)\) is the observed frequency of the pair of amino acids a, b at pairs of sites that belong to class p. We obtained the single-site amino acid frequencies \(F_{i}(a), F_j(b)\) by marginalizing the observed pairwise distribution either symmetrically, if the sites are equivalent (core-core and surface-surface) or asymmetrically if they are different (core-surface).
1. 这段在干什么
给上一段列出的 45 类位点对,逐一定义怎么从 PDB 序列里算出可观测的氨基酸配对倾向 \(Q^\textrm{obs}_p(a,b)\)。
2. 需要解释的地方
3. 值得留意
「减 1」让随机情形落在 0,不是 1。\(i\)、\(j\) 的具体含义这段没展开,要靠上一段的分类来对号。
b062We determine the global parameters \(\{p^\textrm{glob}_a\},\{q^\textrm{glob}_{a,b}\}\) as follows.
这段在干什么:承接上文对位点类型的讨论,说明下文要确定全局层面的参数 \(\{p^\textrm{glob}_a\}\)、\(\{q^\textrm{glob}_{a,b}\}\)。
需要解释的地方:\(p^\textrm{glob}_a\) 是单个位置取某氨基酸 \(a\) 的全局频率,\(q^\textrm{glob}_{a,b}\) 是位置对 \(\{a,b\}\) 的全局频率——即"背景"分布(这是蛋白统计领域的常见记号)。
值得留意:作者用"as follows"引出后文,说明这些参数的具体确定方法在紧接着的下一段,本段只是过渡,没有给出任何公式或数据。
b063First, we sampled the pairwise empirical distributions at all sites, without considering site specificity, i.e. we averaged the site-specific frequencies over all classes of pairs that belong to the same type of proteins. We thus get what we call \(F_{a,b}^\textrm{glob,obs}\), from which we obtain the single-site frequencies \(F^{\textrm{glob,obs}}_{a}\) by marginalization and the propensities as \(q^{\textrm{glob,obs}}_{a,b}=F^{\textrm{glob,obs}}_{a,b}/\left( F^{\textrm{glob,obs}}_{a}F^{\textrm{glob,obs}}_{b}\right) -1\).
承接上句,具体定义全局参数怎么从数据里算出来:先把所有位点的配对频率对同类蛋白求平均,得到全局观测频率,再算出边缘频率和倾向性 \(q\)。
b064This data is affected by the reduced sample size, which we treat as an overfitting problem. We fit the global parameters \(p^\textrm{glob}_a\) and \(q^\textrm{glob}_{a,b}\) from the observed parameters adopting a variant of the “rescaled ridge regression” method proposed by one of us and a coworker (Dehouck and Bastolla 2017). Precisely, we minimize the weighted mean of the errors of the fits. For propensities, we adopt the weights \(w_{a,b}=F^\textrm{glob,obs}_a F^\textrm{glob,obs}_b\) that satisfies the zero mean condition \(\sum _{a,b} w_{a,b}q^\textrm{glob,obs}_{a,b}=0\). We then add the L2 regularization term with regularization parameter R.
交代如何处理上面算出的观测值,转而拟合一套全局参数,并说明所用回归方法与正则化。
作者是最小化误差的加权平均,权重 $w_{a,b}=F^\textrm{glob,obs}_a F^\textrm{glob,obs}_b$,文中特意点出这个权重满足零均值条件——这是为了让拟合基准合理。具体怎么最小化、$R$ 取多少,这段都没说。
b065In the spirit of the rescaled ridge regression, we impose an additional condition on the average parameter values, \(\sum _{a,b} w_{a,b} q^\textrm{glob}_{a,b}\), for preventing that the scale of the fitted parameters vanish for infinite R. The additional term is formally similar to an L1 regularization with the signed value of the fitted parameters \(q^\textrm{glob}_{a,b}\) instead of the absolute value. For notational convenience, we indicate with c any pair of amino acids a, b. The scoring function that we minimize is
这段在干什么:接着上一段的 L2 正则化,作者再补一个约束项——给参数均值加条件,防止 R 变大时拟合参数整体缩到零;并预告下一步写出要最小化的打分函数。
需要解释的地方:
值得留意:上一段已让观测均值等于 0;这段是给拟合参数的均值再加约束,两者对象不同。作者只说"形式相似",并非真是 L1。
b066We determine the value of the regularization parameter \(\mu\) by minimizing the mean square error of the fit for given R. This yields \(\mu =R\sum _{c} w_{c} q_{c}^\textrm{glob,obs}\equiv R\langle q^\textrm{glob,obs}\rangle\) (see Appendix). In other words, the rescaled ridge regression sets the scale of the fitted parameters that minimizes the error of the fit, and it only imposes the regularization on the relative values of the parameters, leaving the fit unregularized when there is only one parameter.
确定正则化参数 μ 的取值:通过最小化拟合的均方误差,把它解出来。
作者特意说"只有一个参数时不施加正则化"——因为此时正则化对唯一参数的单值无相对意义。
b067Finally, we apply the specific heat criterion (Dehouck and Bastolla 2017) and we determine the regularization parameter R by maximizing the derivative of the error with respect to R, which represents a specific heat, in the statistical mechanics analogy proposed by Dehouck and Bastolla. This yields \(R=0.5\) (see Appendix).
好,我们来看这一段。
1. 这段在干什么:给正则化参数 \(R\) 定值。它用一种“比热”准则选出最优的 \(R\),最终得到 \(R=0.5\),供后续计算使用。
2. 需要解释的地方:
3. 值得留意:这里只是“选参数”这一步,\(R=0.5\) 是结果,不是结论性的科学发现。具体推导在 Appendix,正文只给结果。
b068The final result is \(q_c=\frac{2}{3}q^\textrm{obs}_c+\frac{1}{3}\langle q^\textrm{obs}\rangle\). For the propensities, because of the zero mean gauge, it holds \(\langle q^\textrm{obs}\rangle =0\). For the single-site probabilities \(p^\textrm{glob}_a\), we choose weights \(w_a= p^\textrm{obs}_a\), and \(\langle p^\textrm{obs}\rangle =\sum _a (p^\textrm{obs}_a)^2\). We finally obtain
1. 这段在干什么
这是在推一套把观测统计量"重新加权"成全局分布的公式,属于方法推导的收尾(\(\langle\cdot\rangle\) 表示按权重的平均)。上一段刚给出 \(R=0.5\),这里转入具体公式。
2. 需要解释的地方
3. 值得留意
\(\langle p^{\rm obs}\rangle=\sum_a (p_a^{\rm obs})^2\) 这一步——"平均"在这里不是除以 \(a\) 的数目,而是按 \(p_a^{\rm obs}\) 自身加权,所以平均值恰好是概率平方和。这是容易读漏的地方;为什么这么选,原文未进一步说明。
b070We then predict the pairwise distribution for each group of pairs of sites adopting Eq.(11). We use the average value of \(\left\langle C_{ij} \right\rangle\) over all pairs in the group and the observed single-site frequencies, since we want to focus on the prediction of the pairwise couplings.
1. 这段在干什么:交代预测 pairwise 分布的具体做法——对每组位点对套用 Eq.(11),是方法落地的一步(承接上一段刚定好的权重,转入实际预测)。
2. 需要解释的地方:
3. 值得留意:作者明说用平均值+观测单点频率,是为了「专注预测耦合」,即故意去掉其他来源的差异,只测 Eq.(11) 对 coupling 的预测力。
b071We considered three scores that quantify the agreement between the observed and predicted pairwise amino acid distributions. These scores are related with the log-likelihood scores that we used for optimizing the parameter \(\Lambda\). The additional scores allow a more transparent assessment, and we show them in the figures.
1. 这段在干什么
交代评价方法:声明后面用什么指标来判断"预测的氨基酸配对分布"和"实际观测"吻合得好不好,并解释为什么换用这三个分数——为了展示更直观。
2. 需要解释的地方
3. 值得留意
作者明说新分数和优化用的对数似然"相关"(related with),说明它们是同一回事的不同呈现,不是独立证据——这点容易被读成"双重验证"。
b073where \(F_{p}^\textrm{obs}(a,b)\) is the frequency of the amino acid pairs a, b observed at the pair of type p, \(Q_{p}^\textrm{pred}(a,b)\) is the predicted propensity, and \(w_{i,j}\) is the weight of pair p (see below). The log-likelihood is related with the Kullback-Leibler divergence, except for a term that does not depend on \(Q_{p}^\textrm{pred}\), so that optimizing either score yields the same results:
承接上句,给出评价预测所用的打分公式中各符号的含义,并说明该 log-likelihood 与 KL 散度只差一个与 $Q_p^{\text{pred}}$ 无关的项。
$w_{i,j}$ 原文只说"see below",含义未在这段给出;"optimizing either score yields the same results"是指优化 log-likelihood 和优化 KL 散度结果一致,不是指这两个量相等。
b075This score is quadratic, and it is more sensitive to outliers. Its value allows us to assess the relative error of the predicted logarithmic propensities.
这段在干什么:紧接上句,说明这个二次型评分的特点,并交代它的用途——用来评估预测的对数倾向值的相对误差。
需要解释的地方:
值得留意:这一段只讲了该评分"更敏感"和"能评估相对误差",但没提具体数值或阈值。
b076weighted Pearson correlation coefficient \(\textrm{wr}\) between the set observed and predicted logarithmic propensities, \(\log (1+Q_{p}^\textrm{obs}(a,b))\) and \(\log (1+Q_{p}^\textrm{pred}(a,b))\), in which each pair of amino acids a, b is again weighted proportionally to their observed frequency \(F_{p}^\textrm{obs}(a,b)\). We show this quantity because it is easy to interpret.
提出用加权 Pearson 相关系数 wr 衡量预测与观测的对数倾向(log propensity)的吻合度,并说明选它因为它好解读。
注意是「对数倾向」之间的相关,不是原始倾向;且「好解读」是作者给出的选它的理由,而非该指标的固有属性。
b077In all of these scores, we weight each pair of amino acids proportionally to its observed number of instances \(F_{p}^\textrm{obs}(a,b)\), which down weights infrequent pairs that present large relative error, improving the reliability.
这段在干什么:交代这几项评分里对氨基酸配对的加权方式——按观测实例数 \(F_p^{\text{obs}}(a,b)\) 加权,顺带说明为什么这么做。
需要解释的地方:
值得留意:作者给的理由是"降权出现少的配对,提高可靠性"——这是对噪声的处理,不是生物学结论。另外,加权用的是观测频率,而预测本身想预测的就是共变,所以这里已经隐含了"数据说话"的立场。
b078Finally, we fitted the selection parameter \(\Lambda\) separately from each group of proteins by maximizing the weighted sum of the log-likelihood of all pairs p, using as weight of each pair its mutual information, \(w_{p}=M_{p}\equiv \sum _{a,b}F_{p}^\textrm{obs}(a,b)\) \(\log (F_{p}^\textrm{obs}(a,b)/F_i^\textrm{obs}(a)F_j^\textrm{obs}(b))\), so that we weight more the more informative pairs.
交代作者最终怎么给每组蛋白单独拟合选择参数 \(\Lambda\)——按信息量加权求和所有配对的对数似然。
用互信息当权重,等于让"信息量大"的配对主导 \(\Lambda\) 的估计;它与上一段的降权思路方向相反——一个防噪声、一个重信号,共同决定每对的实际影响力。
b079We performed the maximization through an iterative quadratic interpolation algorithm that evaluates the score for three values of the parameter, interpolates the score with a quadratic curve and determines the maximum, then it substitutes the candidate maximum to one of the three previous values and iterates until convergence.
这段在交代:搜出评分最大值是怎么算的——即对上一段那个带权重的评分函数做迭代求极大。
关键概念:只给了三个参数值试算评分,用二次曲线拟合出峰值位置,再把峰值的参数替换掉三个旧值之一,如此反复到收敛。属于数值优化的常规做法(领域常识),不是本文独创。
留意:它只保证收敛,没说收敛到全局最大;也没给迭代次数或收敛判据。
b080After fitting the selection parameter, we compute the total score for each group of proteins using as weight \(w_{p}\) of each group of pair the number of sampled pairs. In this way, the final scores are not affected by the way in which the pairs are grouped, and their comparison is meaningful.
这段在干什么
承接上一步的拟合,说明如何把选择参数换算成每个蛋白组的总分,并为跨组比较做铺垫。
需要解释的地方
值得留意
这里用采样对数当权重,目的是让最终分数不依赖 pair 怎么分组——即分组方式不影响结果,这样不同组的分数才能放在一起比。这是作者为后面比对预测与实际协变做的公平性处理。
b082The strength of the stability constrained model consists in the fact that we predict the folding stability \(\Delta G\) without any free parameter by adopting a simple contact free energy model. In this model, protein structures are represented as binary contact matrices \(C_{ij}\in \{0,1\}\), with residues i and j considered in contact if any of their heavy atoms are closer than 4.5Å. For a given sequence \(a_i\), we estimate the effective free energy of the native state as \(G^\textrm{nat}=\sum _{i<j}C^\textrm{nat}_{ij}U(a_i,a_j)\,,\) where \(C^\textrm{nat}_{ij}\) is the contact matrix of the average native structure, \(a_i, a_j\) are the amino acids at positions i and j in the sequence, and U(a, b) are the effective contact interaction parameters determined by Bastolla et al. (2000).
交代稳定性约束模型的核心——用无自由参数的接触能模型去预测折叠稳定性 \(\Delta G\)(即天然态自由能 \(G^\textrm{nat}\))。
作者说 Gⁿᵃᵗ 是"有效自由能",且用的是平均天然结构的接触矩阵——不是某一条具体结构,容易读漏。
b083Importantly, the free energy parameters U(a, b) where not determined from the statistics of amino acids in the PDB as in the quasi-Boltzmann approximation (Sippl 1995), which would be circular with the approach outlined here, but they were determined by imposing that the native state of a representative set of PDB structures are maximally stable against non-native decoys.
1. 这段在干什么
交代 U(a,b) 的来源:不是从 PDB 统计拟合,而是用"天然态对 decoy 最稳定"来定参。
2. 需要解释的地方
3. 值得留意
作者特意声明这里不循环,是在防"你拿数据又造数据"的质疑;U(a,b) 对应上段 Bastolla 等 2000 的接触作用参数。语法上 where 应为 were。
b084We estimate the effective free energy of the unfolded state as \(G^\textrm{unf}=L S_U\), where \(S_U\) is the average configurational entropy of an unfolded residue. For the ensemble of wrongly folded (misfolded) compact conformations, we adopt the Random Energy Model (Derrida 1981) in the space of compact contact matrices, which approximate the free energy of the ensemble of misfolded conformation with the mean and the variance of their contact free energies. Although corrections to the REM that involve the third moment of the energies are not negligible (Minning et al. 2013), here we do not consider them for simplicity.
这段在干什么:设定模型里"未折叠态"和"错误折叠态"两部分能量的算法,以便后面算出某个序列的折叠稳定性。
需要解释的地方:
值得留意:作者自己承认 REM 忽略了三阶矩修正(Minning 等 2013 指出不可忽略),只是"为简单起见"跳过——这是模型已知的粗糙处。
(约 160 字,可略删)
b085In the above formula, the brackets \(\left\langle \cdot \right\rangle\) represent the mean value over the set of the possible compact contact matrices of proteins with L residues sampled from native structures in the PDB. Thus, the first term represents the mean contact free energy of misfolded conformations and the second term represents their variance.
给上面那个公式里的两个项"翻译"含义:第一项是错折叠构象的平均接触自由能,第二项是它们的方差。
b086Because of the contact correlation terms \(U(A_i,A_j)U(A_k,A_l)\), the free energy of the misfolded ensemble depends on sets of four sites, not just pairs of sites. We therefore average these terms in order to obtain an approximated expression that only depends on pairs of sites:
1. 这段在干什么
提出一个技术困难并说明应对:错折叠系综的自由能涉及四体项,作者要把它近似成只含成对项的表达式,好进入后续计算。
2. 需要解释的地方
3. 值得留意
作者明说这是 approximation——四体信息被丢掉了,这是全段最该记住的让步。「contact correlation」是领域常识:指不同接触之间非独立。
b087To evaluate the above expression, we sample the moments of the distributions of contacts from the PDB as follows.
上一段刚把公式近似成只看「位点对」的形式,这一段是接着交代怎么实际算出那个近似式:不解析推导,而是直接从 PDB 里抽样。
b088The contact frequency \(\left\langle C_{ij} \right\rangle\) represents the mean value of the contact between residues i and j in proteins of L residues. In order to sample it, we assume that this is the product of a function c(L) that only depends on the protein length and a function \(f(|i-j|)\) that only depends on the distance of the two residues along the chain, which we estimate from contact matrices in the protein data bank (PDB). \(\left\langle C_{ij} \right\rangle\) increases with L (because the surface to volume ratio decreases with L, decreasing the influence of surface residues that have fewer contacts) and it decreases with \(|i-j|\) (because short range contacts form more easily than long range ones). We sample these functions from observed contact matrices in the PDB, accelerating the subsequent computations.
构造一个可计算的接触频率近似公式 \(\langle C_{ij}\rangle \approx c(L)\cdot f(|i-j|)\),参数直接从 PDB 实际接触矩阵拟合。
作者没证明可分离性假设,只说“我们假定”;两大趋势(随 L 增、随 |i−j| 减)都只给了定性理由,无数据。
b089In the second term, \(N_C = \sum _{i<j}C_{ij}\) is the total number of contacts. The average number of contacts per residue, \(\left\langle N_C \right\rangle /L\), proportional to c(L), increases with L due to the decrease of the surface to volume ratio. We compute numerically its correlation with the contact frequencies as \(\left\langle C_{ij} N_C \right\rangle \approx Lc(L)^2 f1(|i-j|)\), sampling the factors \(f1(|i-j|)\) prior to the computation.
1. 这段在干什么
承接上句引入的接触统计,给出把「接触数 \(N_C\)」与「接触频率」关联起来的近似计算式,为后续能量项服务。
2. 需要解释的地方
3. 值得留意
注意 \(\langle C_{ij}N_C\rangle\) 是近似(\(\approx\)),不是恒等式;\(c(L)\) 是 \(L\) 的函数,别当成常数;"sampling prior"这句正呼应上句所说的加速计算。
b090We define \(n_i=\sum _j C_{ij}\) as the number of contacts formed by residue i. Furthermore, we assume that its average value is the same for all residues (disregarding the important but limited end effect: residues at the chain termini form fewer contacts), and we numerically estimate the correlations between residues at distance \(|i-j|\) along the sequence as \(\left\langle n_i n_j \right\rangle - \left\langle n_i \right\rangle \left\langle n_j \right\rangle \approx c(L)^2 f2(|i-j|)\), again sampling the factors \(f2(|i-j|)\) prior to the computation.
把上一段对“角度”相关性的处理,平行搬到“接触数”上,给出变异相关性的估计式。
末端接触少被作者点名为“重要但有限”,却没在公式里体现,等于把链两端和中间一视同仁了。
b091Finally, in the spirit of the mean-field approximation, we approximate the terms in Eq.() that depend on the sequence with their average value, except the pairwise terms \(U(A_i,A_j)\). The factor \(\overline{U}\) is the average energy of non-native contacts, which we estimate with the global distribution \(P^\textrm{glob}\):
这段在干什么:在平均场近似下,把方程里依赖序列的项用平均值替代,只保留成对项 \(U(A_i,A_j)\),好把非天然接触的平均能量 \(\overline{U}\) 写出来。
需要解释的地方:
值得留意:被特殊对待的只有成对项 \(U(A_i,A_j)\),说明它被认为携带了关键的序列信息,其余项粗略平均掉。
b093is the average energy of residue \(a_i\) sitting at position i (and same for residue \(a_j\) sitting at position j), which we compute at the beginning of the computation based on the site-specific single site distributions \(p_{x_j}\) and the global amino correlations \(q^\textrm{glob}_{a,b}\), but neglecting the site-specific couplings \(q_{x_i,x_j}\).
1. 这段在干什么
接着上一段定义 \(\bar{e}_U\),这段给出另一项——单个残基坐在某个位置上的平均能量,指明它是怎么算出来的。
2. 需要解释的地方
3. 值得留意
能量只用单点分布 + 全局相关算,故意丢掉位置耦合——即这项不刻画两个位置相互依赖。
b094If \(\overline{U}>0\), the free energy of the misfolded ensemble is larger than the free energy of the unfolded ensemble, and we consider the latter for computing the folding free energy \(\Delta G\):
1. 这段在干什么
提出一个判断条件:当 \(\overline{U}>0\) 时,改选未折叠态(而非错误折叠态)作为计算折叠自由能 \(\Delta G\) 的参照基准。
2. 需要解释的地方
3. 值得留意
这里的"用哪个做参照"不是随便选的——它决定了 \(\Delta G\) 数值的含义。作者是按自由能高低"挑更低的那个"来当基准,而非固定用未折叠态。\(\overline{U}\) 的具体定义在这段没提到,需回前文找。
b095As the result of these approximations, we can decompose the folding free energy as the sum of pairwise terms,
1. 这段在干什么
这是一句承上启下的过渡:在上一段说明用未折叠态能量来算折叠自由能 ΔG 之后,这里交代——经过一系列近似,ΔG 可以拆成成对项(pairwise terms)之和。这个"可拆成对"是后续分析位置协变的前提。
2. 需要解释的地方
3. 值得留意
作者用"approximations"(近似)一词,说明这里的分解是简化结果、并非严格成立;但原文没说具体是哪些近似、误差多大。
b097We analyzed proteins in the PDB grouped at the 70% sequence identity level (list available in the PDB web server), retaining only structures with more than 40 residues studied by X-ray crystallography, which amounts to 37108 protein chains. In order to accelerate the computations, we produced files that contain the sequences and the contact matrices
1. 这段在干什么
交代数据集怎么筛出来的:从 PDB 按 70% 序列一致性分组,只留 X 射线晶体学解析、长度 >40 残基的结构,最终得 37108 条链;并预先算好序列和接触矩阵文件以便加速。
2. 需要解释的地方
3. 值得留意
b100First, we optimized the temperature parameter T of the Random Energy Model (REM) applied to contact matrices, Eq.(18), which sets the scale of the free energy parameters U(a, b). We plot in Fig. 1 the scores summed over all pairs and all classes of proteins as the function of the parameter T. Fig. 1A and C show the log-likelihood score and Fig. 1B and D show the mean squared error of the logarithmic propensities. The top plots A,B do not consider the global propensities (\(q^\textrm{glob}_{ab}=0)\), which focuses more on folding stability, and the bottom plots C,D are obtained adopting \(q^\textrm{glob}\).
作者在调参:先把随机能量模型(REM)里的温度参数 T 优化好,再看打分随 T 怎么变。
b101In all cases there is an optimum at intermediate T (\(T=0.8\) without \(q^\textrm{glob}\), \(T=0.7\) with \(q^\textrm{glob}\) for log-likelihood, \(T=0.9\) without \(q^\textrm{glob}\), \(T=0.8\) with \(q^\textrm{glob}\) for wMSE, in the arbitrary units of the contact interaction parameters).
1. 这段在干什么
汇报结果:不管用哪种打分方式,折叠稳定性选择都在某个中间温度 T 处最优——太冷太热都不行。
2. 需要解释的地方
3. 值得留意
四种组合的最优 T 全在 0.7–0.9 这个窄区间,说明这个"最优"很稳,不挑指标也不挑模型。括号里的单位是"任意单位",别当物理温度读。
b102The continuous lines in Fig. 1 show the scores computed considering only the native free energy and disregarding the average free energy of the misfolded ensemble, \(\sum _{ij}\left\langle C_{ij} \right\rangle U(A_i,A_j)\). The dashed line is the limit of infinite T, which is obtained neglecting the variance of the energy of the misfolded ensemble, U2.
这段在干什么:解释 Fig. 1 里两种曲线的算法差异——实线只算天然态自由能,虚线则是温度趋于无穷时的极限。
需要解释的地方:
值得留意:实线=忽略错误折叠平均能,虚线=再额外忽略方差,两条线的差别层层递进,是后文对比的基准。
b103The data show that it is important to consider selection against misfolding, which is usually neglected when modelling the folding free energy.
承接上一句对 U2(错折叠系综能量方差)的忽略,这一段给出结论。
1. 这段在干什么:收束上文——说明建模折叠自由能时,考虑"抗错折叠的选择"很重要,而这一点常被忽略。
2. 需要解释的地方:
3. 值得留意:作者指出"通常被忽略"——这正是本文要补的缺口,也是它主张 U2 不可省的理由。
b104Agreement between observed and predicted pairwise propensities as the function of the parameter T of the REM. A, C: log-likelihood score. B, D: Mean Square Error. In plots A, B, the only fitted parameter was the selection strength \(\Lambda\). In plots C,D we also adopt the global propensity matrix \(q^\textrm{glob}\), which we obtain from the data
这段是结果小节的图注开头:交代下面四张图在比较什么、横轴是什么、两种误差指标分别是什么。
关键概念:T 是 REM(一种把折叠稳定性写成可调参数的模型)的参数;A/C 用对数似然,B/D 用均方误差——都是衡量“预测的位点配对倾向”与“观测值”差多远。Λ 是选择强度,A、B 里唯一拟合的参数。C、D 额外引入全局倾向矩阵 q^glob,从数据本身得到。
值得留意:原文在 q^glob 处戛然而止,句子没写完,别以为有省略的结论;这里只描述图的设置,尚未给出结果。
b105In all the following figures, we show data obtained at \(T=1.0\) and considering \(q^\textrm{glob}\), unless otherwise stated.
先接着上一段看:上一段刚说完"图 C、D 用全局倾向矩阵 \(q^\textrm{glob}\)",这一段就是在给后面所有图定默认条件。
1. 这段在干什么:定个"默认设置"——以下(之后的)所有图,除非专门说明,都是在 \(T=1.0\)、用 \(q^\textrm{glob}\) 的情况下得到的数据。
2. 需要解释的地方:
3. 值得留意:这是个"除非另有说明"的声明(unless otherwise stated),意味着后面的图只要不特别标注,都套用这套设置;看后续图时别默认它们换了参数。
b107From Eq.(21) and (11), we see that the predicted propensities of amino acid pairs at pairs of sites p that form native contacts (\(C^\textrm{nat}_{p}=1\)) are negatively correlated with their contact energy U(a, b), i.e. amino acids that interact more favorably are predicted to be more frequent for pairs of sites in contact. On the contrary, for pairs of sites that do not form native contacts, the propensities are predicted to be positively correlated with U(a, b), i.e. amino acid pairs that repel each other are predicted to be more frequent, particularly so for short contact range for which \(\left\langle C_{p} \right\rangle\) is large (i.e. they can form contacts more easily).
从公式推出:有天然接触的位点对,倾向性与接触能负相关;无天然接触的则正相关。
有接触时"越合得来越常见",无接触时反过来——"越互相排斥越常见",尤其在短程接触范围。这是作者从公式预测的,不是观测数据。
b108From Eq.(21), we predict that the propensities are correlated with the zero-mean value of the energy parameters, defined as
1. 这段在干什么
承接上句关于「短程接触」的讨论,作者从公式 (21) 出发,提出一个预测:倾向性(propensities)会与能量参数的零均值版本相关联,为下文对比 native / non-native 接触对做铺垫。
2. 需要解释的地方
3. 值得留意
这里只说「预测相关」(correlated),没说正相关还是负相关;方向问题留到标题里的「opposite ways」才展开。读时别提前代入结论。
b109We assess this expectation in Figure 2, which shows the scatter plots of the observed logarithmic propensities versus the zero-mean energy. We group together all proteins, and we distinguish pairs with direct, indirect and no native contacts and with short and long contact ranges.
承接上句的预测,这段说明用 Figure 2 的散点图来检验它:横轴是零均值能量,纵轴是观测到的对数倾向性。这是领域常识:倾向性指某对位置共变的偏好程度,零均值能量是能量参数去中心化后的值——上句刚下了预测,这段就给出检验方案。
分组有两套维度:按天然接触分直接、间接、无三类,按接触距离分长短。
值得留意:此段只交代了图怎么分组,还没给任何结论或数值,别把它读成"已经验证了预测"。
b110Comparing the different types of contacts, we see that the propensities of native contacts present strong negative correlation with the contact energies, as predicted. This means that stabilizing pairs of amino acids are more frequent. The corresponding slope is much stronger for long range native contacts (Fig. 2D) than for short range native contacts (Fig. 2A), as expected since short range contacts are subject to “negative design” that tends to make them less favorable.
对比不同接触类型,验证「天然接触的倾向与接触能量负相关」,并解释长程比短程斜率更陡。
长程斜率更强,是因为短程接触受「负设计」压制而变得不利——这是作者给出的解释,不是数据本身。
b111In agreement with negative design, short range pairs without native contacts (Fig. 2B) and with indirect contacts (Fig. 2C) are positively correlated with the contact energy, as predicted by our formula. This means that destabilizing pairs of amino acids are more frequent. However, this correlation disappears for long range pairs without contact, which are less subject to negative design (Fig. 2E) and it becomes negative for long range indirect contacts, which are more similar to native contacts than to no contacts (Fig. 2F).
承接上一段短程接触受负设计影响,这段把图2B–F的分类逐一说清楚。
这段在干什么:按“有无天然接触 × 短程/长程”四类,分别报告它们与接触能的相关性方向。
需要解释的地方:
值得留意:
b112We can see that the amino acid pair Cysteine-Cysteine (CC) has one of the largest propensities in all plots. This pair, which can form disulfide bridges, has the most stabilizing contact energy. However, its nature of outlier is not due to its effect on stability but it is due to the fact that CC has the largest global propensity, which we attribute to the mutation process and to global selective pressures (see below).
用 CC(半胱氨酸对)做例子,说明它倾向值大但不是因为稳定作用,而是突变过程和全局选择压。
作者特意把"稳定作用最强"和"倾向值最大"拆开:CC 虽接触能最稳定,但它的异常高倾向被归因于突变与全局选择,原因留到下文。
b113Scatter plots of the observed logarithmic propensities of amino acid pairs versus the zero-mean energy of the same amino acids, distinguishing pairs in contact (A, D), pairs without contacts (B, E) and pairs with indirect contacts (C, F). The first line represents short range contacts (A, B, C), the second line represents long range contacts (D, E, F)
这是图注,用散点图展示氨基酸对的观测对数倾向(propensity)与能量的关系,并按接触类型分组。
正文段落讲"两类接触中倾向与能量的关系方向相反",但这条图注本身只描述图形,没有复述这个结论——别把正负方向归到图注里。
b114Figure 3 plots the scatter plots of observed versus predicted propensities. They represent the same type of pairs as Fig. 2.
这段是过渡句,把读者的注意力从上一段的图 2 转到图 3,交代图 3 画的是什么。
需要解释的地方
值得留意
b115We can see that the SCA approximation allows reasonable predictions of the logarithmic propensities. The quality of the prediction is higher for long range native contacts (Fig. 3D) than for short range native contacts (Fig. 3A) that are subject to the contrasting forces of negative and positive design.
拿 SCA 预测的对数倾向性跟观测值对比,报告拟合质量:长程天然接触预测得比短程天然接触好。
短程天然接触预测差,作者归因于它同时受"负设计和正设计"两种相反力量拉扯——这句是解释原因的关键,别只看到"预测差"就跳过。
b116For sites without contact, the prediction is much better for long range pairs that are not subject to negative design (Fig. 3E). This is due to the fact that the prediction in this case is exclusively based on the global propensities \(q^\textrm{glob}\), which may suffer from overfitting despite the regularization that we adopted.
1. 这段在干什么
解释一个现象:对于没有接触的位点,长程配对的预测反而更好;并说明原因是此时预测只依赖\(q^\textrm{glob}\),可能过拟合。
2. 需要解释的地方
3. 值得留意
作者把"预测更好"归因于只用了全局倾向,暗示这种"好"可能是假象——因为没有局部细节可依赖,也就没有过拟合的来源被摊薄。这一段是在为前面的对比做补充说明,不是新论点。
b117Indirect contacts show lower agreement (Fig. 3C and F). In all cases, the amino acid pair CC presents one of the highest propensities.
接着上句比较不同接触类型的倾向性:指出间接接触的吻合度更低,并强调 CC 这种氨基酸对在各类情形下倾向性都偏高。
b118Observed versus predicted logarithmic propensities of amino acid pairs for pairs of sites in contact (A, D), without contacts (B, E) and with indirect contacts (C, F). The first line represents short range contacts (A, B, C), the second line represents long range contacts (D, E, F)
这是图3的图注,说明六个子图(A–F)分别对应哪类位点对和哪种作用距离,为正文比较“观测 vs 预测倾向性”提供坐标。
本段只说图的布局(A–F 对应关系),不含任何结论或数值;结论在上文与后文。
b120In Fig. 4, we show how the contact range \(|i-j|\) influences the weighted correlation coefficients between the logarithmic propensities of each amino acid pair, \(\log (1+Q^\textrm{obs}(a,b))\), and the zero-mean of the free energy matrix, Eq.(22), on which our theoretical prediction is based. The figure depicts pairs of sites in contact (top row), not in contact (middle row) and in indirect contact (bottom row). Each column represents one of the four classes of proteins (Short hydrophobic, short polar, Long hydrophobic and Long polar) and each symbol represent a type of pairs of sites (surface-surface, surface-core and core-core).
这段在干什么:介绍 Fig. 4 的坐标含义——看接触距离 \(|i-j|\) 如何影响氨基酸对的加权相关系数(观测倾向 vs. 理论自由能矩阵)。
需要解释的地方:
值得留意:图分三行——接触、不接触、间接接触;课文上一段提的"short/long range"对应这里的列分类。
b121The correlations are negative for pairs of sites in contact (top row, Fig. 4A, B, C, D), as expected based on positive design, but they are less negative for short range contacts as expected based on negative design, since short range contacts have large contact frequency \(\langle C_{ij}\rangle\), i.e. they are frequently formed in misfolded conformations. The strength of the correlations and their dependence on contact range is larger for pairs of sites both in the core (triangles) than pairs on the surface (circles). Correlations are relatively weaker for short polar proteins (Fig. 4B).
1. 这段在干什么
解读 Fig. 4A–D 的结果:位点间的相关性为负,并随接触距离、埋藏程度和蛋白类型而变化,用以支持“正设计/负设计”共变的解释。
2. 需要解释的地方
3. 值得留意
短程接触负相关“较弱”,作者归因于它们易在错误折叠中出现——这是一个解释,不是直接观测。相关性对距离的依赖,在核心内比表面更强。
b122For pairs of sites that do not form native contacts (middle row, Fig. 4E, F, G, H) the correlation between propensities and contact energies is positive for short range contacts, as expected based on negative design, and it becomes negative for long range contacts. However, these negative correlations at long ranges are not attributable to positive design, as for pairs of sites in contact, but to the negative correlation between the contact energies and the global propensities \(q^\textrm{glob}\), which present positive propensities between pairs of hydrophobic amino acids. In fact, \(q^\textrm{glob}\) determines almost completely the propensities of long range pairs without contacts (Fig. 3E). In long proteins (Fig. 4G,H), the strongest correlations for non-contact pairs happen for pairs of site in the core (triangles), as it is the case for contact pairs, while for short proteins (Fig. 4E,F) they happen for surface-surface sites (the core is not so well defined for short proteins, and many of these proteins are forming complexes, but quaternary contacts are not taken into account in our analysis).
解释非接触位点对的相关性随距离的变化:短程为正、长程为负,且负相关另有原因。
长程负相关不是正设计造成的,而是被 \(q^\textrm{glob}\) 主导——这是本段的关键转折。另外短蛋白的核(core)本身定义就模糊,且文中忽略了四级接触。
b123Similar to no-contact pairs, indirect contact pairs (Fig. 4I,J,K,L) show positive correlations at short range, as expected based on negative design, and negative at long range, where correlations are more negative than for no-contact pairs, i.e. long range pairs that form indirect contacts tend to interact more favourably than pairs that do not form contacts, which suggests that some kind of positive design is taking place for indirect contacts.
这段在干什么:报告间接接触对的相关性随距离变化的规律,并据此推测间接接触存在正设计。
需要解释的地方:
值得留意:长程间接接触对的负相关比无接触对更强,作者据此推断它们相互作用更有利——这是相关性推断,非直接实验证据。
b124weighted Pearson correlation coefficients between the propensities of each amino acid pair, \(Q^\textrm{obs}(a,b)\), and the zero-mean of the free energy matrix, Eq.(22), on which our theoretical prediction is based, as a function of the contact range \(|i-j|\), for pairs in contact (first row, plots A,B,C,D), pairs not in contact (middle row, E,F,G,H) and pairs with indirect contact (bottom row,I,J,K,L). The three curves show three different pairs of sites. The four columns show the different classes of proteins
1. 这段在干什么
段首用一整句长句交代这张图的"画法":横轴是接触距离 \(|i-j|\),纵轴是观测倾向 \(Q^\textrm{obs}(a,b)\) 与自由能矩阵零均值之间的加权 Pearson 相关系数。承接上一段"间接接触存在正设计"的推断,这段把相关性按接触类型分层画出来。
2. 需要解释的地方
3. 值得留意
三行 × 四列共 12 张图,但第二句才交代曲线的身份——三条曲线代表三对位点;四个列对应不同蛋白类别。原文没给出任何趋势结论,只是"说明书",结论要看图本身。
b125We show in Fig. 5 the quality of the predictions, assessed through the weighted Pearson correlation coefficients between observed and predicted propensities, which are based on four global propensities matrices, one for each group of proteins. As in the previous figure, the three rows represent the three types of contacts (native contact, no contact and indirect contacts) and the four columns represent the four groups of proteins.
这段在干什么:交代图5要展示预测质量,并说明图5的行列布局。
需要解释的地方:
值得留意:图和上一段图共用同一套行列结构——三行=三类接触,四列=四组蛋白;上一段说的"三条曲线""不同位点对"是指另一张图,别混。
b126Comparing the contacts, one can see that the predictions are generally worse for indirect contacts (bottom row), which are not well represented in the SCA approximation. Comparing pairs of sites, we see that the predictions tend to be comparatively slightly worse for core-surface pairs (green curves). Comparing the different contact ranges, we see that the predictions are generally worse for short range contacts, which are more strongly affected by negative design, than for long range contacts, for essentially all types of contacts. Finally, comparing the four types of proteins, we can see that the worst predictions are achieved for short polar proteins, which are less affected by selection for negative design, which is an important component of our model. The peculiarity of short polar proteins could also be related with the fact that many of them occur in multichain complexes, so that the distinction between core and surface sites is not so clear-cut.
逐条比较不同类别的预测效果,总结哪类接触/位点/蛋白预测更差,作为对上文图表的解读收尾。
b127Weighted Pearson correlation coefficients between the observed and predicted propensities of each amino acid pair, \(Q^\textrm{obs}(a,b)\), and \(Q^{\textrm{pred}}(a,b)\) as a function of the contact range \(|i-j|\), for pairs in contact (first row, A, B, C, D), pairs not in contact (middle row, E, F, G, H) and pairs with indirect contact (bottom row, I, J, K, L). The three curves show three different pairs of sites. The four columns show the different classes of proteins. The dashed line is a reference for the eye at \(r=0.5\)
这段是图注,交代一张按接触范围分层展示预测效果的大图怎么读。
关键概念:\(Q^\textrm{obs}\) 与 \(Q^\textrm{pred}\) 是某氨基酸对在真实数据与模型预测下的共变倾向(领域常识:打分/倾向性);接触范围 \(|i-j|\) 指两个位点在序列上隔多远;indirect contact 指不直接接触但通过第三方间接关联。横轴即 \(|i-j|\),纵轴是加权 Pearson 相关系数,衡量预测与观测的吻合度。
值得留意:图分三类(接触/不接触/间接接触)×四列蛋白类别,每条曲线是"三对位点",容易误读成三类各一条。虚线 \(r=0.5\) 只是眼睛的参考线,不是判定阈值。
b129As the result of the different selective pressures that act on native contacts, no-contacts and indirect contacts and on short and long contact ranges, the information content varies for different types of pairs.
这段话是过渡句:把前面「不同选择压力」和后面「信息量因配对类型而异」串起来,为下文分类讨论做铺垫。
需要解释的地方:
值得留意:作者把「接触类型」和「范围长短」当作两个独立维度同时影响信息量,说明后面分析会交叉分类,而不是只看一种。
b130To illustrate this fact, we computed the observed mutual information for every class of pairs, removing the mutual information obtained from the global propensities \(q^\textrm{glob}\) in order to focus only on the site-specific contributions to mutual information.
接上一段"信息量因配对类型而异",用具体计算来展示这一事实。
扣掉 q^glob 是为了聚焦位点本身的贡献,不是否定全局倾向的存在——作者只是想让信号更干净。
b131We show the results in Fig. 6, lumping all proteins together and distinguishing contact types and pairs of sites (surface-surface in Fig. 6A, core-surface in Fig. 6B and core-core in Fig. 6C). As expected, the highest mutual information occurs for pairs in contact, followed by indirect contacts, while pairs that do not form contacts present almost vanishing values (i.e., most of their mutual information arises from the global propensities), except at short range where they are influenced by negative design.
这段把 Fig. 6 的结果按接触类型拆开讲:作者把上一段处理过的 MI(扣掉全局倾向、只留位点特异贡献)拿来看不同接触方式的差异。
这段在干什么:交代 Fig. 6 的核心观察——接触对的 MI 最高,间接接触次之,不接触的几乎为零。
需要解释的地方:
值得留意:不接触对"几乎消失"是扣掉全局倾向后的结论;但短程处它们仍受 negative design 影响,这个例外容易被略过。
b132As expected, the mutual information of sites in contact tends to increase for longer ranges. In fact, we expect that short range contacts present small mutual information because they are subject to the contrasting forces of positive and negative design.
顺着上一段往下推进:给出"接触位点随距离增大互信息更高"这个观察,并预告短程接触为何互信息偏低。
"contrasting forces"是作者对短程偏低原因的预期解释,不是已证实的结论;原文只给方向,没给数字。
b133The contrary holds for sites without contacts and with indirect contacts, which present the strongest mutual information at short range, as expected based on negative design.
用「相反的情况」对上一段做对照:无接触/间接接触位点的模式恰好相反,并给出解释。
b134Nevertheless, unexpectedly, pairs of surface sites in contact (Fig. 6A circles) present higher mutual information than pairs of core sites in contact (Fig. 6C circles). Selection for protein folding stability cannot explain this difference, since surface sites contribute less to protein stability. It is intriguing to speculate that this result may be due to selection for function and/or for proper protein interfaces.
1. 这段在干什么
抛出一个反常结果:表面接触对(Fig. 6A 圆圈)的互信息反而比核心接触对(Fig. 6C 圆圈)高,并给出推测性解释。
2. 需要解释的地方
3. 值得留意
作者明确说稳定性选择解释不了这个差异——因为表面位点对稳定性贡献本来就小。所以结尾只是"猜测(speculate)"可能源于功能选择或界面选择,并非结论。
b135Mutual information of each class of pairs of sites as a function of the contact range \(|i-j|\), for pairs in contact (cicles), pairs not in contact (squares) and pairs with indirect contact (triangles), distinguishing surface-surface (plot A), core-surface (plot B) and core-core sites (plot C)
这段在干什么:这是图注,说明用互信息(MI)衡量位点对随接触距离 \(|i-j|\) 的变化趋势,并把位点分为接触/不接触/间接接触三类,再按表层-表层、核心-表层、核心-核心分三张图。
需要解释的地方(领域常识):互信息是衡量两个位点氨基酸组成是否协同变化的量;接触距离 \(|i-j|\) 指序列上相隔多远;cicles/squares/triangles 是图上的圆形、方形、三角标记。
值得留意:标题说接触对 MI 随距离升高、不接触对降低,但图注本身只描述坐标轴与分组,未给趋势结论,别把标题当图注内容。
b137Figure 4 suggests that long and hydrophobic proteins are subject to stronger selection for protein folding stability than short and polar proteins. To clarify this aspect, we plotted in Fig. 7A the selection parameter \(\Lambda\) fitted for the four different classes of proteins, considering one global propensity matrix \(q^\textrm{glob}\) for each protein class (Qglob4, black bars), only one \(q^\textrm{glob}\) for all protein class (Qglob1, green bars) and no \(q^\textrm{glob}\) at all (Qglob0, orange bars). We also compared models that consider core and surface sites (NC2, solid bars) with models that consider only one type of site, abolishing the distinction between core and surface sites (NC1, dotted bars).
承接上文「长蛋白、疏水蛋白受到更强的折叠稳定性选择」的推测,用 Fig. 7A 做定量对比:比较不同蛋白分类下拟合出的选择参数 Λ。
图中用黑色、绿色、橙色、实心/虚线区分不同模型,读图时要先分清「条形颜色管 q^glob 设定,实虚管位点区分」,别混在一起。
(注:本段只描述图的设计,未给出任何结论或数值。)
b138As expected, the more detailed a model is, the lower the selection coefficient is, which indicates that it complies well with the mimimum selection principle. The lowest selection coefficients \(\Lambda\) corresponds to the most accurate global model NC2 Qglob4 (two sites and one global propensity matrix for each protein class).
解释模型越细、选择系数越低,说明结果符合「最小选择原则」;并点名最低系数来自最精细的模型 NC2 Qglob4。
「模型越细、系数越低」是作者说「符合预期」的,不是意外的发现;且最低值恰好给了最精细的模型,这个对应关系是这段的落点。是否唯一最低,原文未明说。
b139In almost all cases, we find larger selection coefficients for long proteins, which suggests that longer proteins are subject to stronger selection for protein folding stability. This is consistent with our previous conclusion that long proteins are more affected by misfolding problems, and they are subject to stronger negative design (Minning et al. 2013).
这段在干什么:顺着上一段的模型结果,报告一个发现——长蛋白的 selection coefficient 更大,并把它和之前"长蛋白更容易受错折叠困扰、受到更强的 negative design"的结论接上。
需要解释的地方:
值得留意:作者用"almost all cases"而非"全部",说明这一趋势是普遍的但非绝对;且结论是"支持/一致于"旧结论,而非新证据本身。
b140Nevertheless, when we compare polar and hydrophobic proteins, we find results contrary to our expectations. Misfolding problems are expected to be more serious for hydrophobic proteins, which are subject to stronger negative design (Minning et al. 2013), but the value of \(\Lambda\) is slightly larger for long polar proteins than for long hydrophobic proteins.
1. 这段在干什么
报告一个与预期相反的结果:疏水蛋白本该问题更大,但长极性蛋白的 \(\Lambda\) 反而略大于长疏水蛋白。属于转折/反例段落。
2. 需要解释的地方
3. 值得留意
b141To further investigate this issue, we plot in Fig. 7C the maximum correlation between propensities and contact energies across no-contact pairs. This correlation allows quantifying negative design without any model. It tends to be larger for hydrophobic than for polar proteins and for long than for short proteins, as expected. Thus, we conclude that our qualitative expectation on negative design is confirmed, but the fitted value of \(\Lambda\) may not be accurate enough to show subtle differences.
这段在干什么:用一张新图(Fig. 7C)换个角度验证前面关于负设计的猜测,并承认拟合出的 \(\Lambda\) 可能不够灵敏。
需要解释的地方:
值得留意:作者明说定性预期被证实,但 \(\Lambda\) 的拟合值精度不足——这是主动承认局限;图里的相关只是"跨所有无接触对取最大值",不是平均值。
b142Similarly, we expect that short proteins experience stronger positive design for stabilizing the native state against unfolding, since the number of contacts per residues increases with chain length due to the surface to volume scaling, and they present fewer native interactions for compensating the loss of conformational entropy upon folding (Bastolla and Demetrius 2005). We quantify positive design for stabilizing the native state against unfolding through the maximum negative correlation between propensities and contact energies across all contact pairs (Fig. 7B). We expect that positive design is stronger for polar proteins, but we find that it is weaker both when assessed through the maximum negative correlation and when assessed through the selection coefficient \(\Lambda\) of short proteins (Fig. 7A), contrary to our expectation.
1. 这段在干什么
提出一个预期(短蛋白应受更强的正设计),然后用两种指标去检验,结果发现与预期相反,属于「预期—反证」的过渡段。
2. 需要解释的地方
3. 值得留意
作者明确说结果「contrary to our expectation」——短蛋白正设计反而更弱。这与上一段「负设计已确认」形成对照,暗示两类选择信号强弱不同。
b143This discrepancy may be due to the fact that many proteins in our data set, in particularly short proteins, tend to occur in protein complexes, and they do not need to be stable on their own. To test this hypothesis, we restricted our data set only to monomeric proteins determined by Xray crystallography, and we tested two thresholds of length, 150 amino acids and 90 amino acids that is often considered the threshold for proteins that fold through two-states folding mechanisms without intermediate states (Barrick 2009). Indeed, for these short lengths the negative design score is very reduced (Fig. 7F orange bars), consistent with the absence of metastable intermediates, and the selection coefficient \(\Lambda\) is larger for short polar proteins than for short hydrophobic proteins (Fig. 7D orange bars), as we expected. However, the positive design score is again larger for short hydrophobic proteins than for short polar proteins (Fig. 7E).
针对上段"短蛋白结果与预期相反"的疑问,作者提出假说:这可能是因短蛋白多以复合物形式存在、无需独立稳定,并只用单体蛋白重新检验来验证。
作者假设"复合物中的短蛋白不需独立稳定",但结论并不完全干净利落:负设计分如预期降低,但正设计分反而在短疏水蛋白里更高,作者也照实写出,没掩饰。
b144Selective parameters \(\Lambda\) fitted for the four different classes of proteins (plots A and D) and maximum correlations between propensities and contact energies for native contacts (B and E) and no contacts (C and F) as a function of different global propensities \(q^\textrm{glob}\) and number of types of sites NC1 and NC2. Plots D,E,F refer to monomeric proteins determined by X ray
这是图注,说明四类蛋白质拟合出的选择参数 Λ,以及倾向性与接触能之间的最大相关性随 \(q^\textrm{glob}\)、NC1、NC2 变化的情况;D–F 特指 X 射线测定的单体蛋白。
提到"最大相关性"和"native contacts / no contacts"两组对比——A–C 与 D–F 应是不同蛋白集合,但这段没说 A–C 指什么。
b146By definition, the global propensities are not site-specific. Therefore, they are not influenced by selection for protein folding stability. They mainly reflect variation of the mutational process and of global (not site-specific) selective forces across the lineages where the proteins evolve. Even if here we do not consider multiple sequence alignments (MSA) or phylogenetic trees from which we can derive evolutionary information, we expect that the global propensities are related with the phylogenetic propensities that can be derived from the MSA and do not depend on site-specific structural or functional constraints.
1. 这段在干什么
过渡段:说明全局倾向(global propensities)为什么与折叠稳定性选择无关,并预告它与系统发育倾向有关。
2. 需要解释的地方
3. 值得留意
作者用"Even if…we expect…"作了让步:即便不用 MSA/树,仍预期 global propensities 与可从 MSA 推出的 phylogenetic propensities 相关,且不依赖位点特异的约束——这是推测(expect),不是已证结论。
b147We show in Fig. 8 the ranked values of the global propensities \(q^\textrm{glob}_{a,b}\), focusing on the amino acid pairs that present the most positive (Fig. 8A) and most negative (Fig. 8B) propensities.
这段在干什么
接着上段引出 Fig. 8:把全局倾向 \(q^\textrm{glob}_{a,b}\) 按数值排名,只挑正、负倾向最强的氨基酸对来展示。
需要解释的地方
值得留意
图只呈现"最极端"的两端,而非全部氨基酸对;正负分开写成 Fig. 8A、8B。这段没交代具体数值或排行结果,那些在图中。
b148The first observation is that pairs of the same amino acids have the strongest positive propensities. This is expected since, if the mutational preference for amino acid a is higher in some lineages and lower in other lineages, this variation will create positive self-correlations of the amino acid frequency. More in general, the mutational preferences between amino acids a and b are similar if they are coded by similar codons, predicting positive global propensities in this case. Conversely, if the codons that code two amino acids overlap less than expected by chance, their mutational preferences will be anticorrelated, possibly leading to negative propensities.
解释上段图中那些正/负倾向(propensity)从何而来——用突变偏好和密码子相似性说明全局氨基酸相关性的成因。
b149A similar mechanism holds for variations of the selective preference for any given amino acid. These global preferences may include several selective pressures, which act globally at the level of the whole protein and modulate the mutational propensities:
这段在干什么:承接上文关于"两个位点偏好相反会导致负倾向"的说法,把机制从"成对"推广到"整体"——任一氨基酸的选择偏好变动,同样受全局性选择压力调节。
需要解释的地方:
值得留意:这段是从"两个位点相互影响"过渡到"整条序列层面的全局因素",为后文讨论遗传密码和一般选择压力如何影响相关性做铺垫。真正的内容(具体是哪些压力、如何调节)这段并未展开。
b150The pressure towards disulfide bridges in short proteins that exist in an oxidative environment (Bastolla and Demetrius 2005). Consistently, Cysteine-Cysteine presents the strongest global propensities \(q^\textrm{glob}_{\textrm{CC}}=0.312\) (Fig. 8).
紧接上句「全局选择压力调节突变倾向」,举二硫键为例:氧化环境中的短蛋白倾向于形成二硫键,而数据里 Cys-Cys 的全局倾向最高(0.312)。
作者用「Consistently」把机制预期(该成键)和统计结果(最高值)对上,这是旁证式的呼应,属相关而非因果证明;文献引用 (Bastolla and Demetrius 2005) 支持的是前半句机制说法,不是 0.312 这个数字。
b151The pressure to avoid biosynthetically costly amino acids, which is particularly relevant in bacterial community, which present hints of cross-feeding where only some community members synthesize some specific amino acids and leak them to the environment (Mustonen and Lässig 2010; Puente-Sánchez et al. 2024), and is lower in organisms that are less limited by resources. Consistently, tryptophane, considered the most biosynthetically costly amino acid, present the second-largest value of the global propensity \(q^\textrm{glob}_{\textrm{WW}}=0.152\), tyrosine presents the 7th highest propensity \(q^\textrm{glob}_{\textrm{YY}}=0.064\) and \(q^\textrm{glob}_{\textrm{WY}}=0.042\). Phenylalanine belongs to the same cluster: \(q^\textrm{glob}_{\textrm{FF}}=0.038\), \(q^\textrm{glob}_{\textrm{FW}}=0.018\) and \(q^\textrm{glob}_{\textrm{FY}}=0.027\).
承接上句「半胱氨酸-半胱氨酸 propensity 最高」,这段在论证全局氨基酸相关性还受合成成本和遗传密码影响,为后面解释共变来源做铺垫。
需解释
值得留意
作者把高成本氨基酸(W、Y、F,同属芳香簇)与高 propensity 挂钩,但 W 只说"第二大",没直接说因果。
b152The pressure to avoid codons that frequently mutate to stop codons and can interrupt the translated protein, which mainly applies to the same amino acids discussed above (Tryptophane, Tyrosine and Cysteine), whose codons are close to stop codons.
这段紧接着上一段的 \(q^\textrm{glob}\) 数值,转而提出一个选择压力来源:密码子若靠近终止密码子,易突变成终止子,中断翻译。
需要解释的地方:
值得留意:
作者用"mainly applies to the same amino acids discussed above"把它和上文的三种氨基酸绑定,说明这可能是解释它们协变的一个候选机制——但这段本身还没说它如何影响协变,别读成结论。
b153The pressure favoring positively charged amino acids in proteins that interact with nucleic acids that are negatively charged. Consistently, Lysine and Arginine have the third and the sixth-largest global propensity, \(q^\textrm{glob}_{\textrm{KK}}=0.100\) and \(q^\textrm{glob}_{\textrm{RR}}=0.064\). However, \(\textrm{RK}\) is disfavored, \(q^\textrm{glob}_{\textrm{RK}}=-0.058\), possibly because there are specific selective pressures that differentiate Arg and Lys besides their common charge.
承接上段"密码子邻近终止子"的偏好话题,换角度说明:氨基酸的全局关联还同时受电荷等一般性选择压力影响。
b154The pressure towards hydrophobic amino acids in membrane proteins. However, pairs of different hydrophobic residues are not favored, except Methionine and Isoleucine (\(q^\textrm{glob}_{\textrm{MI}}=0.015\)) that are coded by very similar codons.
1. 这段在干什么
承接上文对氨基酸配对偏好的讨论,指出膜蛋白里虽偏好疏水残基,但不同疏水残基之间并不互相偏好——是个例外说明。
2. 需要解释的地方
3. 值得留意
作者用"码子相似"来解释 MI 的偏号,暗示这种相关性未必来自选择压力,而可能只是密码子结构的副产品——这正是本小节标题所谓"遗传密码影响相关性"的一个例证。
b155Interestingly, the four polar amino acids Serine, Threonine, Asparagine and Glutamine form a cluster, since they are positively correlated with each other (the highest value is \(q^\textrm{glob}_{\textrm{ST}}=0.030\) and the lowest is \(q^\textrm{glob}_{\textrm{TQ}}=0.003\)). These correlations may also be caused by the similarity of the codons that code for them.
这段紧接上句,继续补充「全局氨基酸相关性」的成因。
这段在干什么:指出 Ser、Thr、Asn、Gln 四种极性氨基酸彼此正相关,形成一个簇,并推测原因可能是密码子相似。
需要解释的地方:
值得留意:这段只是"may also be caused"的推测,并未给出证据;最低值 TQ=0.003 说明相关性其实很弱,别误读成强关联。
b156Global propensities \(q^\textrm{glob}_{ab }\). Plot A shows the 25 most positive propensities and plot B shows the 25 most negative ones
这段在干什么:这是图注的开头。上一段末尾提到相关系数可能来自密码子相似性,本段引入「全局倾向 \(q^\textrm{glob}_{ab}\)」,并说明两张子图分别展示最正/最负的 25 个倾向。
需要解释的地方:\(q^\textrm{glob}_{ab}\) 是某对氨基酸 \(a,b\) 在全局范围内一起出现的倾向(领域常识:下标 \(a,b\) 通常指两个氨基酸,glob 指全局统计,而非某蛋白内部)。
值得留意:原文只说 A 图给最正的 25 个、B 图给最负的 25 个,没有给具体数值或结论;上一段末句的因果猜测,这里也没正面回答。
b158In this paper, we introduce the stability constrained model of protein evolution (SCPE) in the pairwise approximation, which generalizes our previous SCPE model that assumes that the protein sites are independent (Arenas et al. 2015). We model the site-specific amino acid distribution of a protein family based on minimal selection for folding stability, requiring that the probability distribution of the amino acids has minimal Kullback-Leibler (KL) divergence from the global (not site-specific) amino acid distribution, while constraining the protein folding stability \(\Delta G\) with respect to both the unfolded state and compact but uncorrectly folded misfolded conformations.
好的,我们来看这一段。它出现在讨论部分,是作者对自己工作的定位与概括。
1. 这段在干什么
交代本文提出了什么模型:把此前"位点相互独立"的 SCPE 推广到成对近似版本。
2. 需要解释的地方
3. 值得留意
模型不只约束"折叠态 vs 去折叠态",还额外约束了错误折叠构象——这点容易被一扫而过。
b159Minimum KL divergence coincides with the more familiar maximum entropy principle when amino acids are equiprobable, since maximum entropy is the same as minimum KL divergence from the uniform distribution, but it is more suitable for describing real situations in which the amino acids are not equally distributed. The main assumption of the model is that protein fitness can be approximated with the stability of the native state.
这一段先替模型“选尺子”,再交代核心假设。
1. 这段在干什么:说明为什么用最小 KL 散度来描述氨基酸分布,并点出模型的中心假设。
2. 需要解释的地方:
3. 值得留意:第二句“蛋白质适应度可用天然态稳定性近似”是全文的关键假设,不是已证结论。
b160Importantly, we do not fit stability constraints from the data, but we predict them from the protein sequence and the experimentally known protein structure with a simple contact model that also takes into account the stability of the misfolded state modelled with the Random Energy Model (Minning et al. 2013). In this way, we predict the parameters of a stability constrained Potts model that describes the amino acid distribution of the protein family in the pairwise approximation (Lapedes et al. 2002; Weigt et al. 2009; Marks et al. 2012; Morcos et al. 2011) only based on global amino acid propensities and site-specific stability constraints. Within the REM, the free energy of the misfolded ensemble depends on groups of four sites, while we approximate it here as a sum of pairwise terms. This approximation may explain in part the reduced accuracy of our predictions for short range pairs, which are dominated by selection on misfolding.
承接上段「fitness≈native稳定性」的假设,说明预测 Potts 模型参数的来源:不拟合数据,而是从序列+实验结构出发预测。
b162We adopt here the small-couplings approximation, i.e. we consider only first order terms in the Potts couplings, which is related with the mutual information approximation originally adopted for inferring the statistical couplings. We obtain an explicit formula that predicts the couplings between amino acid \(a_i\) at site i and amino acid \(a_j\) at site j, and depends on:
引入小耦合近似,为下一句「预测任意两个位点上氨基酸间耦合的显式公式」做铺垫。
原文最后停在「depends on:」,冒号后内容被截断,公式依赖什么并未出现在本段。
b163The global (non-site-specific) couplings \(q^\textrm{glob}_{a_i,a_j}\) between the amino acids \(a_i\) and \(a_j\). In our interpretation, global couplings are the result of variations across organismic lineages of the mutation process and the selective forces that favor or disfavor \(a_i\) and \(a_j\). The global couplings are generally small, and they are strongest when \(a_i=a_j\), since in this case the mutational and selective forces that favor them are perfectly correlated. In particular, we find the strongest self-coupling for Cysteine, which is subject to global selective force for the formation of disulfide bridges. The second one is for Triptophane, which is subject to global selective force for sparing this metabolically costly amino acid. The same applies to Tyrosine and, to a limited extent, Phenilanine, and the three form a cluster of positively correlated amino acids. Another cluster is formed by the polar amino acids Serine, Threonine, Glutamine and Asparagine.
这一段在介绍 SCPE 模型公式里的第一个输入项——氨基酸之间的全局耦合 \(q^\textrm{glob}\),也就是上一段公式冒号后面列出的第一类参数。
b164Site-specific selective forces due to positive design that stabilize the native state with respect to the unfolded state. They mainly act at pairs of sites that form native contacts. The predicted amino acid propensities of native contacts are approximately equal to minus the interaction free energy of the two amino acids, i.e. amino acid pairs that stabilize the native state are predicted to be more frequently observed at these pairs of sites, with contrasting influence from negative design on short range contacts.
1. 这段在干什么
介绍 SCPE 模型里的一类选择压力——来自 positive design、稳定天然态的位点特异性作用力。
2. 需要解释的地方
3. 值得留意
末尾一句提到 short range contacts 还受 negative design 的相反影响——这句话是补充,别漏读,它说明近距离接触的规律比天然接触更复杂。
b165Site-specific selective forces due to negative design that destabilize misfolded states. They mainly act at pairs of sites that are at short range along the sequence and do not form native contacts, and they predict amino acid propensities that are approximately equal to the interaction free energy of the two amino acids, i.e. amino acid pairs that destabilize misfolded conformations are more frequently observed at these pairs of sites.
承接上文,具体说明 SCPE 模型里“负设计”这种选择压力:它怎么作用、作用在哪些位点对、对氨基酸偏好做出什么预测。
负设计专门盯“短程、非天然接触”的位点对,而非天然接触对——这是它区别于正设计的关键。
b166We test these predictions on a representative subset of 37108 non-redundant protein chains from the PDB, grouping together 45 classes of pairs of sites with similar properties (native contacts versus no contacts and indirect contacts, 5 different distances along the sequence, core or surface sites). We also consider four types of proteins (long or short, hydrophobic or polar and their combinations).
这段在干什么:交代验证上一段预测所用的数据规模和分组方式——把 PDB 中 37108 条非冗余蛋白链按位点对性质分成 45 类,并另按蛋白类型分四类。
需要解释的地方:
值得留意:45 = 3(接触类型)× 5(序列距离)× 3(core/surface 之类),三种属性交叉分组;而"long or short, hydrophobic or polar 及其组合"是另一套独立的四类划分,两套分组别混。
b168We examine the observed propensities of each class of pairs of sites p, \(Q_{p}(a,b)=P_{p}(a,b)/\left( P_{i}(a)P_{j}(b)\right) -1\), which in the small coupling approximation is proportional to the corresponding coupling.
提出用观测倾向 \(Q_p(a,b)\) 来量化每类位点对,并说明它在弱耦合近似下正比于耦合强度,为后文检验 SCPE 模型提供可观测的代理量。
这里说的是"正比",不是"等于",中间的比例系数未给出,因此 \(Q_p\) 只能做相对比较,不能直接读出耦合绝对值。
b169The qualitative dependence of the observed couplings on the contact free energy is for each type of pair in line with what our model predicts (see Fig. 4), and it is modulated by native contacts, the buriedness of the sites, their distance along the sequences, and the length of the proteins as the SCPE model predicts. In particular, the correlations between the couplings and the contact energy are negative for pairs that form native contacts, as expected based on positive design, and they are less negative for short range pairs, as expected based on negative design. They are more negative for core-core sites, as expected based on their importance for protein stability, and they are more negative for longer proteins. For pairs that do not form native contacts, the correlation is positive for short range pairs, as expected based on negative design, and they are stronger for longer proteins, which are expected to be subject to stronger selection against misfolding. This also holds for pairs that form indirect contacts through another residue. On the other hand, the couplings of long range pairs that do not form native contacts are essentially determined by the global couplings, with which they correlate very strongly.
检验上文模型预测的耦合强度与接触自由能的关系是否在真实数据中成立,逐类配对逐一对照。
作者不说“完全吻合”,只说“in line with / as expected”——是定性趋势一致,不是定量证明。
b170The agreement with our predictions is reasonably good, with correlation coefficients r larger than 0.5 for most pairs of sites (see Fig. 3, where the dashed line represents \(r=0.5\)).
1. 这段在干什么
给出 SCPE 模型预测与实际数据吻合程度的定量结论:大部分位点对的相关系数 r > 0.5,算「相当好」。
2. 需要解释的地方
3. 值得留意
(约 150 字)
b171However, one can see that there are some regimes in which our predictions do not perform very well. In particular, predictions are poor for pairs that form indirect contacts (bottom line in Fig. 3), which are neglected in the small coupling approximation and can be computed only a posteriori. Predictions are also poor for short polar proteins (second column in Fig. 3). This is in part due to the fact that most of these chains (59%) appear in multimeric proteins, so that they may be unstable as monomers. Also, 49% of these chains have been determined by NMR and may be different from average crystal structures that dominate for long proteins. Finally, predictions are poor for short range contacts. There are two likely explanations of this fact: one is the breakdown of the small couplings approximation for short range pairs that do not form native contacts, whose couplings are large (see Fig. 6), and the other one is the approximation of the free energy of the misfolded ensemble predicted through the REM as a sum of pairwise terms, Eq.(), which is necessary for applying the pairwise model but may be not sufficiently accurate.
1. 这段在干什么
紧接上一段的乐观结论(多数位点对 \(r>0.5\)),这段转折,列出预测效果差的几类情形,并逐一交代可能原因。
2. 需要解释的地方
3. 值得留意
作者只是罗列"可能原因"(likely explanations, may be not accurate),并未给出确证;数据(59%、49%)只针对短极性蛋白这一类。
b173These results suggest that a big part of the couplings in protein evolution originate from the variation of global mutational and selective forces, which act on most pairs of sites, plus couplings due to positive selection for enhancing the stability of the native state, which mainly act at pairs of sites that form native contacts, plus couplings due to negative design for destabilizing pairs that are close along the sequence and do not form native contacts.
这段在干什么:论文观察到蛋白质演化中的耦合(couplings)不只一种来源,这里把它们归纳成三部分叠加。
需要解释的地方:
值得留意:作者说耦合是「三部分之和」,正好呼应上一段「pairwise 之和」的说法——但这里说的是来源,不是模型本身。
b174Perhaps unexpectedly, a larger number of couplings are influenced by negative design against misfolding (positive correlations between observed propensities and contact energies) than those influenced by positive design for stabilizing the native state against unfolding.
抛出一个"意料之外"的观察结论:受负设计影响的耦合,比受正设计影响的更多。
作者用 "Perhaps unexpectedly" 点明这结果与直觉相悖——通常以为稳定天然态才是主导。且这是观察量多寡的比较,不代表负设计作用更强。
b175Another interesting and perhaps unexpected result is that pairs of surface-surface sites in contact possess higher mutual information than pairs of core-core suites in contact (compare circles in Fig. 6A and C) despite the fact that the latters contribute more to protein folding stability. This suggests that surface-surface contacts are subject to other selective pressures besides protein folding stability. Strong candidates are functional movements and interface properties.
1. 这段在干什么
提出一个反直觉的观察:表面-表面接触位点对的互信息,反而比核心-核心接触位点对更高。
2. 需要解释的地方
3. 值得留意
(原文只给"强候选"是功能运动和界面性质,没说已证实。)