Selection for Folding Stability Predicts Observed Covariation Between Protein Positions in the PDB

全文来源:generic  · 全文共 141 段,已全部带读  · 单段均价 ¥0.002013

Abstract

b002Statistical couplings between protein sites have been enormously useful for protein structure prediction, but there is still debate on the selective forces that generate them. Here we explicitly predict the couplings considering selection for protein folding stability, both against unfolding and misfolding, which we model through the Stability constrained model of protein evolution (SCPE). The SCPE is based on contact interactions, and it predicts stability against misfolding through the Random Energy Model in the space of contact matrices. We adopt the minimum selection principle, assuming that the multivariate amino acid distribution has minimum Kullback-Leibler divergence with respect to a site-unspecific distribution that is not influenced by protein folding stability. We predict the couplings at first order, assuming that they are small (small coupling approximation, SCA), and obtain an explicit formula that relates them with the contact interaction matrix and with the structural properties of the pairs of sites: presence of native contact, distance along the sequence and buriedness of the sites. We test these predictions on a representative set of the Protein Data Bank. For pairs of sites in contact, both predicted and observed couplings are inversely related with the contact energies of the amino acid pairs, while for pairs not in contact this relation is positive, i.e. negative design destabilizes pairs of amino acids for short-range pairs of sites that can form wrong contacts with high probability. Our results suggest that the strongest couplings correspond to native contacts, but an even larger number of couplings are influenced by negative design. Intriguingly, the most informative pairs are formed by surface sites, suggesting a role for protein-protein interactions and functional motions. We interpret the site-unspecific couplings, which are not influenced by selection for protein folding stability, as the result of variations of global mutational or selective forces across the lineages where the proteins evolve. Consistently, the strongest global couplings are self-couplings between equal amino acids. Cys-Cys, which form disulfide bridges, show the strongest global coupling, followed by pairs of metabolically costly amino acids (Trp-Trp, Tyr-Tyr, Trp-Tyr) and positively charged amino acids that tend to interact with nucleic acids (Lys-Lys, Arg-Arg). Ser, Thr, Asn and Gln form a cluster of globally correlated polar amino acids.

讲解

1. 这段在干什么

这是摘要全文:先立论——用「折叠稳定性选择」来解释蛋白位点间的统计耦合,再给出模型(SCPE)、近似(SCA)、预测公式和 PDB 验证结果。

2. 需要解释的地方

  • 统计耦合:两个位点氨基酸"同变"的强度,常用来推结构。
  • SCPE:把稳定性选择写进进化模型的框架(领域常识:稳定性约束模型)。
  • 随机能量模型:用来估算错误折叠概率的工具(领域常识)。
  • 最小选择原则 / KL 散度:选最"接近无选择背景"的那个分布。
  • SCA:假定耦合很小,一阶展开,得到显式公式。
  • 负设计:不是"选对接触",而是"避免形成错误接触"。

3. 值得留意

  • 预测公式只依赖三样结构特征:是否有天然接触、序列距离、埋藏程度。
  • 接触对与不接触对的耦合-能量关系符号相反(负 vs 正),这是负设计的关键证据。
  • 最有信息量的竟是表面位点——作者自己都说"intriguingly",暗示与蛋白互作、功能运动有关,但这只是推测。
  • 全局耦合最强的 Cys-Cys 等,被解释为与折叠稳定性无关的因素(突变/选择随谱系变化)。

Introduction

b009The evolution of protein sequences is commonly modelled mathematically as a substitution process, i.e. a stochastic process that describes the probability that an amino acid at a given site in a protein is substituted by another one as the function of time. This formalism is very powerful for inferring phylogenetic trees and other evolutionary properties (Yang 2006). Substitution processes typically assume that different sites of the protein evolve independently of each other, which allows feasible computation of the likelihood function required for evolutionary inferences. However, it is known that the pairwise correlations that are neglected in this approach are key for predicting structural interactions between protein sites (Göbel et al. 1994; De Juan et al. 2013). The inference of these evolutionary correlations through neural networks lead to spectacular success in protein structure prediction (Jumper et al. 2021). Large part of these evolutionary correlations are thought to arise from the selective pressure for maintaining the functional structure of the protein and its folding stability (Zeldovich et al. 2007; Goldstein 2008; Liberles et al. 2012; Bastolla et al. 2017). Within this framework, one of us and co-workers developed a model that predicts the stationary amino acid distribution of a protein family with known native structures (Bastolla et al. 2006), from which we derived the stability-constrained protein evolution (SCPE) model, a substitution process for phylogenetic inference based on selection on protein folding stability (Arenas et al. 2015), which we explicitly predict based on a statistical physics model of the misfolded state (Minning et al. 2013). We later extended the SCPE model to enforce structure conservation, since several lines of evidence suggest that mutations that change the native structure are selected against more strongly than mutations that maintain the structure and only affect its folding stability. Following (Echave 2008), we predict the effect of mutations on the protein structure as the linear response of the protein represented as an elastic network model. We obtain in this way the structure and stability constrained protein evolution (SSCPE) models (Lorca-Alonso et al. 2025).

讲解

这段在干什么:交代本文的理论背景与来路——从「替换过程」假设位点独立,引出被忽略的相关性对结构预测很关键,再说明作者团队如何一步步做出 SCPE 和 SSCPE 模型。

需要解释的地方:

  • 替换过程:把氨基酸随时间被换成另一个的概率,写成随机过程(领域常识)。
  • 位点独立假设:假定每个位点各变各的,好处是似然函数好算,坏处是丢了成对相关。
  • SCPE / SSCPE:作者团队的两个模型,前者基于折叠稳定性选择,后者再加「结构保守」约束,用弹性网络模型预测突变对结构的影响。

值得留意:相关性被认为主要来自「维持功能结构与折叠稳定性」的选择压力,这正是全文标题的由来。

b010For computational treatability, the substitution processes derived from the SSCPE models assumed that protein sites evolve independently of each other. The model takes into account structural interactions effectively, through the effects of mutations on folding stability and structural changes. The SSCPE models produce site-specific amino acid distributions that present qualitative differences at sites buried in the protein core, which tend to be more hydrophobic, more conserved and evolve more slowly, with respect to exposed sites that tend to be polar, have higher sequence variation and evolve more rapidly. These differences arise from the different constraints imposed by natural selection for folding stability, and they are consistent with observations from multiple sequence alignments (MSA) (Echave et al. 2016).

讲解

1. 这段在干什么

为方便计算,SSCPE 模型先假设各位点独立演化,但通过突变对折叠稳定性的影响,把结构相互作用间接纳入进来——为后文讨论位点协变铺垫。

2. 需要解释的地方

  • 独立演化:位点各变各的,互不影响,是简化假设,不是真实情况。
  • SSCPE:结构+稳定性约束的蛋白演化模型,这里强调它能给出每位点的氨基酸分布。
  • 埋藏位点 vs 暴露位点:这是领域常识——核心内部偏疏水、保守、变得慢;表面偏极性、多变、变得快。

3. 值得留意

作者承认"独立演化"只是计算上的权宜,但坚称结构效应已被有效吸收,这是全文立论的关键前提。最后一句用 MSA 观测佐证,说明预测与实证一致。

b011Adopting the SCPE model, we relaxed the assumption that protein sites evolve independently and assume that the stationary amino acid distribution presents pairwise couplings described by the Potts model (Wu 1982) pioneered by Lapedes et al. (2002); Weigt et al. (2009); Marks et al. (2011) for inferring evolutionary correlations and studied by several groups (for instance, Morcos et al. 2011; Skwark et al. 2013; Ekeberg et al. 2013; Kamisetty et al. 2013).

讲解

1. 这段在干什么

承接上一段的 SCPE 模型,作者在此放宽"位点独立演化"的假设,引入 Potts 模型来刻画位点间的成对耦合。

2. 需要解释的地方

  • SCPE:上一段提出的模型,此处作为起点,本段对它做扩展。
  • 位点独立演化:领域常识,指假定蛋白质各位置互不影响地各自演化——这里被放松了。
  • Potts 模型:用来描述位点两两之间相互作用的统计模型,原文说它由 Lapedes 等(2002)开创、用于推断演化相关性。
  • pairwise couplings(成对耦合):平稳氨基酸分布中,两个位点之间的关联项。

3. 值得留意

作者一口气堆了多篇文献,都是支撑点:Potts 模型并非本文首创,而是被多个团队用过的既有方法。

b012The study of the pairwise stability constrained model in the framework of the "small-couplings" approximation (SCA) described below was one of the subjects of the PhD thesis of Jonas Minning (Minning 2012), directed by Markus Porto, at that time at the University of Darmstadt, with the external advice of Ugo Bastolla. Unfortunately, despite reiterate attempts, we have not been able to get in contact with Markus Porto, which is the reason why he does not appear among the authors although he is clearly one of the main contributors to this work. Should he read these lines, we will be more than happy if he gets in touch with us, and he is included in the authors list, as he deserves.

逐段讲解

1. 这段在干什么

这是一段作者说明:交代该研究的来源(Jonas Minning 的博士论文,导师 Markus Porto),并解释为何 Porto 未列入作者名单。

2. 需要解释的地方

  • "small-couplings" approximation (SCA):本文下面要展开的核心方法。领域常识:它假设蛋白位点间耦合很弱,据此从序列协变反推稳定性约束。这里只是点名,细节在"described below"。
  • PhD thesis:这篇论文的部分工作源自 Minning 的博士论文。

3. 值得留意

作者坦承 Porto 是"主要贡献者之一"却因联系不上而缺席作者名单——这在学术署名中不常见,也透露出本文与那篇博士论文的继承关系。

b013Here we consider the pairwise SCPE model within the SCA approximation, which provides an analytic formula that we test here, leaving the test of the general theory for future work. Despite its simplicity, the SCA constructively supports the hypothesis that site-specific selection for protein folding stability, together with global mutational and selective forces that act at the level of the whole protein, may produce the observed covariation among protein positions.

逐段讲解

1. 这段在干什么

交代本文的研究范围(只做 SCA 近似下的成对 SCPE 模型,并检验其解析公式),并点出它支持的假设。

2. 需要解释的地方

  • SCPE / SCA:领域常识——两类从序列数据推断位点间共变(耦合)的统计模型;SCA 是其中一种近似,能给出解析公式。
  • covariation:指不同蛋白位点在进化中"同变"、彼此关联的现象。
  • folding stability:蛋白折叠稳定性,受选择压力影响。

3. 值得留意

  • "leaving the test of the general theory for future work":作者自承只验证了特例,一般理论没测——范围有限。
  • "may produce":只是可能产生,是假设性措辞,不是结论。
  • 本段与上一段(一封致谢/收录来信者的段落)内容上无承接关系。

Multivariate Amino Acid Distribution

b016Following (Weigt et al. 2009), we model the L-dimensional probability distribution that sites \(A_i\), \(i\in [0,L-1]\) in a protein family are occupied by the amino-acids \(a_i\) as the exponential distribution

这段在干什么:紧接上段提到的“选择压力可能造成位置间共变”,这里提出用指数分布这种概率模型来刻画一个蛋白家族中各位置氨基酸的联合分布。

需要解释的地方:

  • “L 维概率分布”:这里有 L 个位点,每个位点取哪种氨基酸是一维,合起来就是要描述整条序列出现的概率。
  • “指数分布(此处指玻尔兹曼式形式)”:领域常识——把序列概率写成指数形式,是直接耦合分析(DCA)一类方法的起点。原文只给了名字,没展开公式。

值得留意:作者明确说“Following (Weigt et al. 2009)”,即这套建模是沿用前人框架,不是本文首创;真正的重点在下一段才会展开。

b017We use the notation \(x_i= \nu i +a_i\), which has \(L\nu\) possible values where L is the number of sites and \(\nu\) is the number of amino acid types (21, since we consider gaps as an additional symbol). In the following, we indicate by sum over \(a_i\) the sum over the \(\nu\) amino acids at position i, and by sum over \(x_i\) the sum over all L sites and their possible amino acids \(a_i\). The single site parameter vector \(h_{x_i}\) has \(L\nu\) components and the pairwise coupling matrix \(g_{x_i,x_j}\) is an \(L\nu \times L\nu\) symmetric matrix whose elements with \(j=i\) vanish.

讲解

1. 这段在干什么

给能量模型(看起来像 Potts/Ising 模型)定义变量和参数的符号体系,为后面写分布函数做铺垫。

2. 需要解释的地方

  • \(x_i = \nu i + a_i\):把「第 i 位 + 第 a 种氨基酸」两个下标压缩成一个整数,避免写双下标。
  • \(\nu = 21\):20 种氨基酸加 gap,共 21 种。
  • \(h\) 是单体项,\(g\) 是两两耦合项,矩阵对称且对角为 0(\(j=i\) 时元素消失)。

3. 值得留意

\(g\) 是 \(L\nu \times L\nu\) 的大矩阵;对角项被设为 0,说明自耦合不计入——这点容易读漏。

b018In the physical literature, Eq.(1) is called "Potts model" (Wu 1982), and it has been extensively studied in the context of disordered systems before being recognized as the exponential pairwise distribution that describes a protein family.

这段是承接上文:上一段刚给完耦合矩阵 \(g\) 的定义,这里补一句它的物理出处。

需要解释的:

  • Potts model(波茨模型):物理里描述多个“格点”各取多种状态、且状态间有相互作用能的模型(领域常识)。作者说式(1)就是这个模型。
  • disordered systems(无序系统):物理中相互作用随机、非均匀的系统(领域常识)。
  • exponential pairwise distribution:即式(1)那种“概率随两两耦合指数衰减”的形式。

值得留意:作者强调这是先在物理中被研究、后来才被认作描述蛋白家族的分布——说明这套数学不是为蛋白发明的,而是借用。

b019In general, we can not normalize the Potts distribution either analytically or numerically, because computing the normalization factor \(Z=\sum _{a_0\cdots a_{L-1}} P(a_0,\cdots a_{L-1})\) requires an exponential number of operations. Thus, its parameters are usually inferred adopting approximated methods, such as the message passing algorithm originally proposed by Weigt et al. (2009), the regularized matrix inversion proposed by Jones et al. (2012), or the pseudo-likelihood approximation (Besag 1975), applied to this problem by Aurell and coworkers (Ekeberg et al. 2013) and (Kamisetty et al. 2013), among others.

讲解

1. 这段在干什么

承接上句把 Potts 分布认定为描述蛋白家族的成对分布后,指出它无法归一化,所以参数只能靠近似方法推断,并列出三条代表性路线。

2. 需要解释的地方

  • 归一化因子 Z:把分布所有可能取值加起来,让它总和为 1。这里字母组合数随序列长度指数增长,算不动。
  • message passing / 正则化矩阵求逆 / 伪似然:三种绕开算 Z 的近似推断手段(领域常识,非本文提出)。

3. 值得留意

作者明说"解析和数值都不行",是指数级操作导致——这是整段论证的根。三个方法名不展开,只作引用定位。

Propensities

b021For notational convenience, we define the propensity matrix

这段在干什么:这是一句过渡句,作者开始引入「propensity matrix(倾向性矩阵)」这个记号,为下文定义它做铺垫。

需要解释的地方:「notational convenience(记号的便利)」是论文常见说法,意思是"接下来为了书写方便先约定一个符号",本身没有科学含义。「propensity matrix」这段还没给出定义,只是一个名字,具体内容要看下一段。

值得留意:注意括号里是"(Ekeberg et al. 2013)",而上文写的是"Aurell and coworkers"——领域常识:Aurell 是通讯作者,Ekeberg 是第一作者,指同一批工作。但这段原文本身没解释这点。

b022that represents the ratio between the conditional probability of \(a_i\) conditioned to \(a_j\) and the unconditioned probability, minus 1. It is zero if the amino aids are independent, positive if amino acid \(a_j\) at site j favors \(a_i\) at site i, and negative if it affects it negatively. We call this quantity “propensity” between the two variables \(a_i\) and \(a_j\). This name is used in bioinformatics to describe the propensity of a given amino acid to take any secondary structure (Koehl and Levitt 1999). We also call \(\log (1+Q)\) “logarithmic propensity”. At first order in Q, the propensity and the logarithmic propensity coincide: \(\log (1+Q)\approx Q\).

这段在干什么

定义「propensity(倾向性)」这个量:它怎么算、正负号怎么读、为什么借用「倾向性」这个名字。

需要解释的地方

  • 条件概率比无条件概率再减 1:衡量「已知 j 位是 a_j 后,i 位出现 a_i 的概率」相对「不管 j、i 位本来是 a_i 的概率」变了多少,即一种关联强度。
  • 零/正/负:分别对应两个位点独立、正相关、负相关。
  • logarithmic propensity:取 log(1+Q);Q 很小时 log(1+Q)≈Q,两者近似相等。

值得留意

「propensity」这名字是从二级结构倾向性研究借来的(同词不同义),别和本文的定义混。

b023The mutual information between sites i and j, \(M_{ij}\), is related to the propensities through

这句话是一个过渡句:它引出互信息和倾向性(propensity)之间的数学关系,为下文推导做铺垫。

需要解释的地方

  • 互信息 $M_{ij}$:领域常识里,它衡量两个位点 $i$、$j$ 的关联程度——一个位点变化能否提供另一个位点的信息。
  • propensity 倾向性:结合上一段,指每个位点倾向于某种状态的统计量。

值得留意

原文只说二者"相关(is related)",但具体关系式这里没给出,得看下文。

b024In the limit of small propensities the mutual information is approximated by the weighted mean square propensity.

讲解

1. 这段在干什么

接着上一段 \(M_{ij}\) 与 propensity 的关系式,给一个近似:当 propensity 很小时,互信息约等于加权的均方 propensity。

2. 需要解释的地方

  • propensity(倾向性):这里指两个位点共同变化的统计倾向强度。
  • mutual information(互信息):领域常识,衡量两个位点的变化是否关联;关联越强值越大。
  • 小 propensity 极限:假设这种关联很弱,才允许用近似式替代原式。

3. 值得留意

关键词是「in the limit of small propensities」——这个近似只在弱关联时成立,作者没展开说误差有多大。

Zero-Mean Gauge

b026The distribution Eq.(1) presents more parameters than necessary, in the sense that different parameter sets produce exactly the same distribution. These extra parameters must be fixed through normalization conditions or "gauge" (Weigt et al. 2009; Ekeberg et al. 2013). We change parameters as \(p_{x_i}\equiv \exp (h_{x_i})\), \(q_{xi,xj}\equiv \exp (g_{x_i,x_j})-1\) and write the multivariate distribution as

这段在干什么:指出 Eq.(1) 的参数多于必要的,说明要用归一化条件(即"规范/gauge")把多余参数固定住,并给出两组参数替换式。

需要解释的地方:

  • 参数多于必要 / 规范(gauge):不同的参数组合能给出完全相同的分布,多出来的自由度得靠额外条件锁死。这是领域常识,类似"同一物理状态多种写法"。
  • 参数替换:把 \(p\)、\(q\) 分别写成 \(\exp(h)\) 和 \(\exp(g)-1\),是换一组参数来写同一个分布。

值得留意:替换式只给了定义,后面的多元分布表达式在本段并未写出。

b027We adopt the following gauge conditions:

这段在干什么

一句话过渡:宣布接下来要设定一组规范条件(gauge conditions),为后面的参数化简铺路。

需要解释的地方

  • gauge(规范):物理/统计模型里的领域常识——模型本身有冗余自由度(同一分布可由多组参数表示),固定某些参数取值来消除冗余、让解唯一。作者选的可能是零均值之类的约定,但具体选了什么,这段没提。

值得留意

"the following"指的是下一段,这里只有一句引子;真正的规范条件内容不在此段。

b028The sum over \(a_i\) denotes the sum over all possible \(\nu\) amino acids at position i. We call these conditions the "zero-mean" gauge because it sets to zero the weighted mean value of the couplings over the amino acids at one of the site, and we adopt it throughout this paper.

本段紧接「We adopt the following gauge conditions:」,具体给出这条规范条件的内容。

这段在干什么:定义并命名论文全篇采用的「零均值」规范条件。

需要解释的地方:

  • \(\sum_{a_i}\):对位点 i 上所有 \(\nu\) 种氨基酸求和。
  • 零均值规范:通过选取参照,使某一位点上耦合强度的加权平均为零(领域常识:规范即人为选定的参照基准,此处作者说要全文沿用)。
  • couplings(耦合):刻画两个位点间关联强度的量。

值得留意:原文只说把该位点氨基酸上的加权均值「置零」,权重是什么、零点从哪来,这段都没提。「全文沿用」意味着后面所有耦合值都相对于这个基准。

b029Note that the propensities satisfy the zero mean condition \(\sum _{a_j}p_{x_j}Q_{x_i,x_j}=\sum _{a_i}p_{x_i}Q_{x_i,x_j}=0\).

这段在干什么

紧接上文提出的"加权平均"处理,用公式明确倾向性(propensities)满足零均值条件——这是全文采用的一项约定。

需要解释的地方

  • propensities / couplings:描述氨基酸之间耦合强弱的量。
  • 零均值条件:对某个位点上的各种氨基酸按其频率加权求和,结果恒为 0。相当于"把基准线平移一下",让平均值为零。

值得留意

两个连等号分别是对 \(a_j\) 与 \(a_i\) 求和,说明任一位点都满足该条件,而非只对某一侧成立。这是领域里常见的"gauge 选择"手法,作者靠它把参量标准化。

Global Distribution

b031We model the complete pairwise distribution through the combination of site specific parameters that represent natural selection for folding stability and global parameters \(p^\textrm{glob}_a\), \(q^\textrm{glob}_{ab}\) that are the same for all sites of the protein. These global parameters reflect the genetic code and the mutational process, which may be different in different genomes and different genomic regions. Indeed, considering the codon level instead of the amino acid level improves the inference of the couplings (Jacob et al. 2015). The global parameters also represent selective pressures that are uniform across the protein (see Discussion).

这段在干什么:提出一个建模框架——把完整的成对分布拆成“每个位点自己的选择参数”加“全蛋白共用的全局参数”。

需要解释的地方:

  • 「全局参数 \(p^\textrm{glob}_a\)、\(q^\textrm{glob}_{ab}\)」:对所有位点都一样的部分,反映遗传密码和突变过程(不同基因组可能不同)。
  • 「site specific parameters」:各位点独有的部分,代表对折叠稳定性的自然选择。
  • 从密码子层面而非氨基酸层面建模能改善耦合推断——这是引用 Jacob et al. 2015 的领域结论。

值得留意:全局参数不只管突变,还代表“全蛋白均匀的选择压力”,作者把它留到 Discussion 展开,这里只点名。

b032We represent the global pairwise distribution as \(P^\textrm{glob}_{a,b} = p^\textrm{glob}_{a} p^\textrm{glob}_{b} \left( 1 +q^\textrm{glob}_{a,b}\right)\) for all pairs of sites. It is easy to see that the global propensities \(q^\textrm{glob}_{a,b}\) obey the zero-mean gauge with respect to \(p^\textrm{glob}_{a}\).

讲解

1. 这段在干什么

上一段刚定义了“全局选择压力”,这段接着给它一个数学表达:把任意两个位点的联合分布写成各自边缘分布的乘积,再乘一个修正项。

2. 需要解释的地方

  • \(P^\textrm{glob}_{a,b}\):全蛋白范围内,位点1是氨基酸a、位点2是氨基酸b的联合概率。
  • \(p^\textrm{glob}_{a}\):单看一个位点,它是a的概率(边缘分布)。
  • 若两位点独立,联合分布就等于两者相乘;\(q^\textrm{glob}_{a,b}\) 就是衡量“偏离独立”程度的耦合项。
  • 零均值规范(gauge):领域常识——指对 \(p^\textrm{glob}_a\) 加权求和时 \(q\) 的均值为零,保证这个改写在数学上自洽。

3. 值得留意

作者只说“容易看出”成立,但没推导。这里的 \(q\) 是全局参数,后面与位点特异的协变对比时,正是靠它当基准。

b033We determine the \(\nu +\nu ^2\) global parameters \(\{p^\textrm{glob}_a\},\{q^\textrm{glob}_{a,b}\}\) through a regularized fit, as explained in Methods.

这段在干什么

交代全局参数(线性项 \(p^\textrm{glob}_a\) 与二次项 \(q^\textrm{glob}_{a,b}\))的确定方式:通过正则化拟合得到,细节见 Methods。

需要解释的地方

  • \(p^\textrm{glob}_a\):单个位置 \(a\) 的全局倾向,可理解为该位点的“基准偏好”。
  • \(q^\textrm{glob}_{a,b}\):位置对 \(a,b\) 之间的全局耦合项(\(\nu+\nu^2\) 指参数总数:\(\nu\) 个线性 + 约 \(\nu^2\) 个二次)。
  • 正则化拟合:在拟合时加惩罚项以抑制过拟合,是统计建模的领域常识。

值得留意

接上一段,这些 \(q^\textrm{glob}\) 满足关于 \(p^\textrm{glob}\) 的零均值规范——即参数并非独立自由,读时别把它们当作可任意取值的量。

Site-Specific Couplings: Stability Constrains and Minimal Selection

b035The site-specific components of the modelled distribution reflect selection for maintaining the native structure of the protein and its folding stability. We determine them through the minimum selection principle (Arenas et al. 2015). This principle assumes that the complete distribution is as similar as possible to the global distribution, which is not influenced by selection for correct folding, in the sense that it has minimal Kullback-Leibler (KL) (Kullback and Leibler 1951) divergence from the global distribution, while constraining the average protein folding free energy \(\Delta G\).

这段在干什么

从全局分布过渡到「位点特异性分布」,提出用最小选择原则来反推每个位点受折叠稳定性选择影响下的分布。

需要解释的地方

  • KL 散度:衡量两个概率分布差多远(领域常识)。文中要它「最小」,即新分布尽量贴近不受折叠选择影响的全局分布。
  • 平均折叠自由能 ΔG:稳定性的度量,作为「约束」加进去——即让分布贴近全局,同时保证平均稳定性达标。

值得留意

「minimal selection」是假设,不是推导出的结论;文中只说是为了体现折叠稳定性选择,并未给出求解细节——那应该在后文/Methods里。

b036The minimum selection principle generalizes the commonly adopted maximum entropy principle (see for instance Lapedes et al. 2002; Bastolla et al. 2006) if the amino acids are not equivalent. In fact, maximum entropy is equivalent to minimum KL divergence from the uniform distribution. Since the genetic code, the mutational process and the unspecific selective pressures differentiate the amino acids, the minimum KL divergence is preferable.

这段在干什么

给上一段提出的最小选择原则做定位:说明它是更常见的最大熵原则的推广,并论证为何该用前者。

需要解释的地方

  • 最大熵原则:在所有满足给定约束的分布里,挑「最无偏见、最不额外假设」的那个(领域常识)。
  • KL 散度:衡量两个分布差多远的量;它最小 ⟺ 熵最大(等价的两种说法)。这里说「从均匀分布出发的最小 KL」,就是相对于完全无偏好状态的偏离最小。
  • 氨基酸不等价:因为遗传密码、突变过程、非特异选择压力,20 种氨基酸地位并不对等。

值得留意

作者的关键理由:正因氨基酸不等价,用「相对均匀分布」的最小 KL 比用「最大熵」更合适——这是逻辑推理,不是实验结论。

b037This principle can be justified based on the analogy between molecular evolution and the Boltzmann distribution of statistical physics (Sella and Hirsh 2005; Mustonen and Lässig 2010; Puente-Sánchez et al. 2024). In particular, the evolutionary process, modelled as a stochastic process in genotype space (Fisher 1930), converges to a stationary exponential distribution analogous to the Boltzmann distribution in statistical physics, with the logarithm of the fitness playing the role of minus the energy and the inverse of the effective population size playing the role of temperature, meaning that small populations are more tolerant towards sequences with low fitness.

这段在干什么

给上一段的结论(为何取最小 KL 散度)补一个理论依据:把分子演化类比成统计物理的玻尔兹曼分布。

需要解释的地方

  • 基因型空间里的随机过程:把序列突变+选择当成随机游走,这是领域常识(Fisher 1930)。
  • 稳态指数分布:该过程最终收敛到的分布形式,对应物理里的玻尔兹曼分布。
  • 三个对应:fitness 取对数 ↔ 负能量;有效种群大小倒数 ↔ 温度。
  • 推论:小种群 ≈ 高温,更能容忍低适应度序列。

值得留意

作者只是"预设"了这个类比成立("can be justified"),并引用前人文献,并未在此推导。

b038In the present context, the inverse of the selective temperature, or selective strength, is represented by the Lagrange multiplier \(\Lambda\) that enforces the mean value of the protein folding free energy \(\Delta G\). This is the only free parameter of the stability constrained model. We fit \(\Lambda\) by maximizing the log-likelihood of the observed protein sequences with respect to the model.

好,我们来看这一段。

1. 这段在干什么

定义稳定性约束模型里唯一的自由参数 Λ,并说明它怎么定出来——用观测序列做最大似然拟合。

2. 需要解释的地方

  • 选择温度/选择强度:承接上一段"温度"的比喻,这里换成它的倒数 Λ,控制选择有多严。
  • 拉格朗日乘子 Λ:领域常识,一种用来"强制"某个平均值满足设定值的数学工具。这里它强制折叠自由能 ΔG 的均值。
  • 唯一自由参数:整个模型只有 Λ 可调。
  • 最大似然拟合:调 Λ,让模型生成观测序列的概率最大。

3. 值得留意

ΔG 均值被"强制"(enforce),不是自然涌现;Λ 是倒数关系,别把方向和温度搞混。

b039Additionally, we impose the gauge conditions Eq.(5,6,7) through the Lagrange multipliers \(\mu _i\), \(\zeta _{x_i,j}\) and \(\eta _{x_j,i}\). These are not free parameters, since they are determined by imposing the constraints.

讲解

这段在干什么:紧接上一段拟合 \(\Lambda\) 的模型,补充说明还通过拉格朗日乘子 \(\mu_i\)、\(\zeta_{x_i,j}\)、\(\eta_{x_j,i}\) 施加上规范条件 Eq.(5,6,7)。

需要解释的地方:

  • 规范条件(gauge conditions):领域常识,指模型存在冗余自由度时人为固定的约束,避免同一物理状态有多种参数写法。
  • 拉格朗日乘子:领域常识,把约束"塞进"优化里的数学工具,乘子本身随约束被满足而自动确定。

值得留意:作者特意强调这些乘子"不是自由参数"——它们是约束的产物,不是可调的模型参数。读者别把它们误当成需要拟合的量。

b040We determine the site-specific Potts parameters \(\{p_{x_i}\},\{q_{x_i,x_j}\}\) of the stability constrained model by minimizing the objective function \(\Phi\), which is the sum of the KL divergence plus the functions that describe the constraints.

逐段讲解

1. 这段在干什么

承接上段:说明那些非自由参数到底怎么定——通过最小化目标函数 \(\Phi\),反解出位点特异的 Potts 参数。

2. 需要解释的地方

  • Potts 参数 \(\{p_{x_i}\},\{q_{x_i,x_j}\}\):描述每个位点取某氨基酸的偏好,以及两个位点之间的耦合(领域常识)。
  • KL 散度:衡量模型分布与观测分布差多少(领域常识)。
  • \(\Phi\) = KL + 约束项:即在"满足约束"的前提下,让模型尽量贴近数据。

3. 值得留意

约束不是硬性扣在参数上,而是作为 \(\Phi\) 的加项一起最小化,所以叫"稳定性约束模型"。

b041The matrix \(G_{x_i,x_j}\) represents the effective free energy of the pair \(x_i\) and \(x_j\) with respect to unfolding and incorrect folding (misfolding). These are not free parameters: We predict them with our model of protein folding stability, which considers the native contact matrix and the statistics of non-native contacts (Minning et al. 2013), as we describe in section 3.4. This computation is approximate since we approximate the free energy of the misfolded ensemble as a sum of pairwise terms, while the REM prediction involves terms that depend on the amino acid at four protein sites.

这段在干什么:给上一段目标函数里出现的耦合矩阵 \(G_{x_i,x_j}\) 定性——它不是自由拟合参数,而是由作者的折叠稳定性模型预测出来的。

需要解释的地方:

  • \(G_{x_i,x_j}\):位置 \(i\)、\(j\) 上氨基酸组合 \(x_i,x_j\) 对「去折叠 + 错误折叠」的有效自由能,即这个配对有多"不利/有利"。
  • native contact matrix:天然结构中相互接触的残基对,是领域常识。
  • misfolded ensemble:所有错误折叠构象的集合。
  • REM:此处只被提及作对比,原文未展开,这段没解释它是什么。

值得留意:作者主动承认这是近似——把错误折叠的自由能简化成两两配对之和,而 REM 预测涉及四个位点的氨基酸,精度上是有代价的。

b042We compute the optimal values of the site-specific parameters \(p_{x_i}\) and \(q_{x_i,x_j}\) by minimizing the score \(\Phi\), where \(\Lambda\), \(\{p_a^\textrm{glob}\},\{q^\textrm{glob}_{a,b}\}\) are pre-set global parameters and the external information comes from the predicted values of the effective free energies \(G_{x_i,x_j}\).

这段在干什么

交代怎么定出位点参数:固定全局参数,通过最小化打分 \(\Phi\) 来求 \(p\)、\(q\),外部信息来自预测的自由能 \(G\)。

需要解释的地方

  • 位点参数 \(p_{x_i}\)、\(q_{x_i,x_j}\):单点和成对的位置相关参数,就是要拟合的对象。
  • \(\Phi\)、\(\Lambda\)、glob 参数:\(\Phi\) 是要最小化的目标函数;\(\Lambda\) 和带 glob 上标的量是预设好的全局参数,不参与优化。
  • \(G_{x_i,x_j}\):“外部信息”的来源,指有效自由能的预测值。

值得留意

\(p\)、\(q\) 是“最优值”,但不是说它们有唯一解——是靠最小化 \(\Phi\) 得到的。另外要分清谁是固定、谁是待求:glob 参数固定,只有 \(p_{x_i}\)、\(q_{x_i,x_j}\) 在优化。

b043We impose the constraints Eq.(5,6,7) through the Lagrange multipliers \(\mu _i\), \(\zeta _{j,x_i}\) and \(\eta _{i,x_j}\), whose values are determined by the corresponding constraints. In this way, we determine the site-specific couplings and the single-site parameters as the function of the selective strength \(\Lambda\).

这段在干什么

说明如何用拉格朗日乘子把约束(式5-7)落实,从而把耦合项与单点参数解成选择强度 \(\Lambda\) 的函数。

需要解释的地方

  • 拉格朗日乘子(领域常识):一种带约束求极值的数学工具,乘子值本身由约束条件反解出来。
  • site-specific couplings:位点间的耦合,刻画两个位置取值的关联。
  • 单点参数 / 选择强度 \(\Lambda\):前者是各位置自身的项,后者衡量选择压力大小。

值得留意

这段只讲"怎么解",没写解出的具体形式或结果;\(\mu_i,\zeta_{j,x_i},\eta_{i,x_j}\) 三个乘子各对应式5、6、7,原文未逐一说明谁对应哪条式子。

b044Finally, we determine the optimal value of the parameter \(\Lambda\) by maximizing the likelihood of the empirical pairwise distributions observed in the PDB for all types of pairs of sites.

前面已把耦合与单点参数写成 \(\Lambda\) 的函数,这段收尾:用最大似然定 \(\Lambda\) 的最优值,拟合对象是 PDB 中各类位点对的经验成对分布。

  • \(\Lambda\):选择强度参数,前文当作自变量。
  • 最大似然:领域常识,即挑一个让观测数据出现概率最大的参数值。
  • 经验成对分布:直接从 PDB 数出来的位点对共现频率。

留意:它只提"似然最大化"这一准则,没给具体数值或优化算法——别自行脑补。

b045In our previous work (Arenas et al. 2015), we determined the single site parameters \(p_{x_i}\) assuming that all sites are independent, i.e. assuming that \(q_{x_i,x_j}\equiv q^\textrm{glob}_{a,b}\equiv 0\). Here we relax this assumption, as we also did in Minning (2012) within the small couplings approximation (SCA) discussed below.

承接上文「逐对位点似然」的语境,这段在交代模型假设的松紧变化:此前假设位点彼此独立,现在放弃这个假设。

  • \(p_{x_i}\) / \(q_{x_i,x_j}\):这是领域常识——前者是单个位点的氨基酸分布参数,后者是两位点间的耦合参数。原文说清楚了两件事:之前的工作令耦合项 \(q\equiv 0\);本文不再这样设。
  • SCA(小耦合近似):原文只给了名字和「讨论见下文」,具体含义这段没展开,读到后面再看。

留意:两个「先例」要分清——Arenas et al. 2015 是令耦合为零的旧做法,Minning 2012 则和本文一样放松了假设,这是本文方法上的直接来源。

The Small-Coupling Approximation (SCA)

b047In the numerical tests presented below, we adopt the small-coupling approximation (SCA) that neglects second and higher order terms in the couplings in the formula of the pairwise probability, i.e. we only consider the first order terms in \(q_{x_i,x_j}\). Because of the zero-mean gauge, the resulting values of the single site probabilities are \(P_i(a_i)=p_{x_i}\). The propensities Eq.(2) are

逐段带读

1. 这段在干什么

定义正文数值测试所采用的 SCA 近似:在配对概率公式里只保留耦合的一阶项。

2. 需要解释的地方

  • small-coupling approximation:假设位置间耦合很弱,把公式按耦合做泰勒展开后只留线性项,高阶项忽略。
  • zero-mean gauge:领域常识,指把耦合参数约束成零均值,从而单点概率退化回 \(p_{x_i}\)。
  • propensity:氨基酸在某位点出现的倾向。

3. 值得留意

"only first order" 是全文近似的核心代价——高阶耦合被直接丢弃,后面结论都建立在此之上。

b048i.e. at the first order in q the propensities are equal to the direct couplings, and we neglect the correlations that arise from indirect couplings, although they can be recovered a-posteriori as we do when we predict the propensities of indirect contacts.

这段在干什么:承接上一段 propensity 的定义,说明在一阶近似下,propensity 就等于直接耦合,并交代间接耦合的处理方式。

需要解释的地方:

  • 一阶(first order in q):领域常识,指把小参数 q 按泰勒展开只保留一次项,高阶项忽略。
  • propensity(倾向):这里指单点概率推出的量,Eq.(2) 给出的东西。
  • direct / indirect couplings:前者是位点间直接相互作用,后者是经中间位点间接产生的关联。

值得留意:作者坦白“间接耦合带来的关联被忽略了”,但强调可以事后再捞回来(a-posteriori),这正是后文预测间接接触要做的事。

b049The simplicity of the SCA approximation stems from the fact that the couplings of the site of pairs ij only depend on its properties of the pair, and not on the other pairs. This approximation is reminiscent of the first approach that was adopted to infer the couplings \(q_{x_i,x_j}\) through the mutual information, which in the small coupling limit is equal to the weighted mean square \(Q_{x_i,x_j}\). The seminal work by Weigt et al. (2009) and its sequels clearly showed that the approximation based on the mutual information is much less accurate for inferring the couplings than the approach based on the Potts model that also considers the indirect couplings. Nevertheless, we consider here the SCA because of two reasons: (1) The SCA provides an analytic formula that is easy to interpret and sheds light on the selective forces that are behind the couplings. (2) The resulting pairwise distributions depend only on the “local” properties of each pair ij, which allows grouping pairs with similar properties and performing the simple and exhaustive numerical tests that we describe below.

逐段讲解

1. 这段在干什么

为 SCA 近似做定位:说明它简单在哪(耦合只依赖配对 ij 自身)、与互信息法的渊源,再给出作者仍采用它的两条理由。

2. 需要解释的地方

  • SCA:把每个位置对的耦合单独近似,忽略其他对的间接影响,所以式子是解析可写的。
  • 互信息 vs Potts 模型(领域常识):互信息只看两个位点直接相关,容易把间接关联当成直接耦合;Potts 模型能扣除间接效应,因此更准。

3. 值得留意

作者明知 SCA 精度不如 Potts,仍选它——用"可解释 + 可做穷举数值检验"换精度。这不是最好的方法,而是最好用的工具。

b050Using the approximation (9) and computing the derivatives of Eq.(8) with respect to the parameters \(\{p_i\}\) and \(\{q_{ij}\}\), we can explicitly determine the values of the parameters that minimize \(\Phi\) as

逐段带读

1. 这段在干什么

承接上段的近似思路,说明作者接下来要对目标函数 \(\Phi\) 求极值,从而显式解出最优参数。

2. 需要解释的地方

  • \(\Phi\):需要被最小化的目标函数(通常是某种误差或自由能近似)。
  • 式(9):上一段引入的「小耦合近似」,这里直接拿来用。
  • 求导=找极值:对 \(\{p_i\}\)、\(\{q_{ij}\}\) 求偏导并令其为零(领域常识:这是求最小值点的标准做法),解出使 \(\Phi\) 最小的参数值。

3. 值得留意

  • \(\{p_i\}\) 是单点参数,\(\{q_{ij\}\) 是成对耦合参数——注意是两组参数一起解。
  • 「explicitly determine」是承诺:后面应给出参数的解析表达式,而非数值迭代。

b051In the above equations, we denote with \(\textrm{ZM}\left( G_{x_i,x_j}\right)\) the zero-mean of the matrix \(G_{x_i,x_j}\), whose weighted mean value across the amino acids \(a_i\) and \(a_j\) is zero. Eq.(10) is the same that we obtained in our previous mean-field approximation in which we neglected the couplings and considered sites as independent, i.e. \(q_{x_i,x_j}\equiv 0\). The same formula continues to hold also at first order in the couplings, which does not change the single-site probabilities.

讲解

1. 这段在干什么

紧接上文给出的参数解,这里补充公式符号约定,并说明 Eq.(10) 与之前平均场近似的联系。

2. 需要解释的地方

  • 零均值 ZM:把矩阵 \(G_{x_i,x_j}\) 按氨基酸对加权平均后平移,使其均值为零(领域常识:便于后续协方差分析)。
  • 平均场近似:忽略位点间耦合,把各位点当独立处理,此时 \(q_{x_i,x_j}\equiv 0\)。

3. 值得留意

作者强调:在一阶耦合下同一公式仍成立,且单点概率不变——说明耦合只影响两点项,不影响 \(\{p_i\}\),这是后面用协方差反推耦合的关键前提。

Site-Specific Distributions

b054For assessing Eq. 11, we sampled pairs of sites from a representative subset of the protein data bank, which we grouped in various groups according to their properties.

这段在干什么:为验证上一节末尾的公式 Eq. 11,交代数据来源和抽样方式——从 PDB 的代表性子集中抽取位点对,并按性质分组。

需要解释的地方:

  • PDB(蛋白质数据库):存放蛋白质三维结构的公共数据库,这是领域常识。
  • 位点对(pairs of sites):指蛋白质序列上两个位置的组合。
  • 分组(grouped in various groups):按蛋白质的某种属性归类,具体按什么属性,这段没提。

值得留意:"representative subset"(代表性子集)暗示作者没有用整个 PDB,而是做了筛选,但筛选标准这段没说;分组的具体依据同样留到后文。

b055Firstly, we distinguished four types of proteins according to sequence length (short, \(<150\) residues, and long) and average contact energy, related with hydrophobicity (hydrophobic, \(\overline{U}<0\) and polar). These groups allow us analyzing how selection for stability depends on the properties of proteins. Alternatively, we lumped together all proteins in the same group.

这段在干什么:交代分组方式——先按序列长度(短/长)和平均接触能(疏水/极性)分出四种蛋白,也提到另一种做法:把所有蛋白并成一组。

需要解释的地方:

  • 平均接触能 \(\overline{U}\):领域常识,粗略衡量残基间相互作用强弱,越负通常越偏疏水。
  • selection for stability:选择压力偏好维持折叠稳定性的序列。

值得留意:原文用了 "Firstly",说明分组方案不止一种;末尾 "Alternatively" 正是备选。分组依据是长度和接触能两个维度,别读成四类各自独立。

b056For every group of proteins, we grouped sites in two classes (buried or exposed) depending on their number of native contacts \(n_i=\sum _j C^\textrm{nat}_{ij}\), thus obtaining three types of pairs of sites: exposed-exposed, exposed-buried and buried-buried. Alternatively, we considered only one type of sites.

讲解

1. 这段在干什么

交代分析方案的分类方式:按溶剂暴露程度给位点分两类,从而产生三种位点对,或只取一类。

2. 需要解释的地方

  • \(n_i=\sum_j C^{\textrm{nat}}_{ij}\):位点 \(i\) 的天然接触数,即它在天然结构里"挨着"多少个其他残基(这是领域常识,接触常用来近似埋藏程度)。
  • buried / exposed:埋藏 / 暴露,按接触数多少来划。
  • 三种配对:暴露-暴露、暴露-埋藏、埋藏-埋藏。
  • "Alternatively…only one type":作为对照,作者也试过不分类。

3. 值得留意

分类阈值怎么定、接触数怎么算,这段都没说。

(约 150 字)

b057We then classified pairs of sites i, j according to their native contact \(C^\textrm{nat}_{ij}\): either pairs that form a native contact (C1), or pairs that form an indirect contact through another site (C2) or pairs that do not form direct or indirect contacts.

按分类维度分层:先按位点类型分(上一段),再按空间接触类型分(本段)。

1. 这段在干什么

在上一段按位点暴露程度分类之后,这一段换一个维度——按两个位点在天然结构里是否接触——给位点对重新分类。

2. 需要解释的地方

  • native contact(天然接触):领域常识,指蛋白质在天然折叠状态下,两个残基在空间上靠得足够近、有直接相互作用。
  • C1 / C2 / 无接触:C1 是 i、j 直接接触;C2 是它俩不直接接触,但通过第三个位点间接连上;剩下的是两者都不沾。

3. 值得留意

作者用「either … or … or」把三类并列,说明 C1、C2、无接触是互斥且穷尽的划分,后面做统计时应该是三组分开比较,而不是只分"接触/不接触"两类。

b058Finally, we grouped the distance of the sites along the sequence \(|i-j|\) (contact range). We found that five groups of contact ranges are sufficient to obtain good results, and we show this value in most of the figures.

这段在干什么

交代方法收尾:把位点对沿序列的距离 \(|i-j|\) 分组(contact range),并说明为什么选 5 组——足够拿到好结果,后续多数图都用这个划分。

需要解释的地方

  • \(|i-j|\):两个位点在序列上的间隔数,如第 5 位和第 12 位,\(|i-j|=7\)。注意是序列距离,不是空间距离(领域常识)。
  • contact range:按序列间隔划分的档位/区间。
  • sufficient:作者自己判定的经验标准,没给量化依据。

值得留意

上一段刚把接触分成 C1 直接、C2 间接、无接触三类;这里换了个维度,按序列间隔分组——两组分类是并行使用的,别混为一谈。"5 组足够好"的依据这段没提。

b059As a result, we obtained 3 pairs of sites times 3 types of contacts times 5 contact ranges, yielding 45 types of pairs for each group of proteins.

这段是方法交代:告诉读者后续分析里“位点对”的类型总数是怎么算出来的。

关键概念:3 pairs of sites(每组的位点配对)× 3 种接触类型 × 5 个接触范围距离档,相乘得 45,即每组蛋白有 45 种位点对类型。这里的 contacts、contact ranges 具体指什么,这段没展开。

值得留意:三种因素是被当成独立维度做笛卡尔积——这个乘法结构意味着后文分析是按“位点对×接触类型×距离档”分层展开的,不是一个笼统的 45。至于这些类型最终如何用,得看下文。

b060For each group of proteins and for each group of pair of sites p, we sampled the pairwise amino acid distribution from protein sequences in the PDB and computed the observed propensities \(Q^\textrm{obs}_{p}(a,b)=F_p(a,b)/(F_i(a)F_j(b))-1\) where \(F_p(a,b)\) is the observed frequency of the pair of amino acids a, b at pairs of sites that belong to class p. We obtained the single-site amino acid frequencies \(F_{i}(a), F_j(b)\) by marginalizing the observed pairwise distribution either symmetrically, if the sites are equivalent (core-core and surface-surface) or asymmetrically if they are different (core-surface).

讲解

1. 这段在干什么

给上一段列出的 45 类位点对,逐一定义怎么从 PDB 序列里算出可观测的氨基酸配对倾向 \(Q^\textrm{obs}_p(a,b)\)。

2. 需要解释的地方

  • 边际化(marginalize):把二维配对频率 \(F_p(a,b)\) 对另一个位置求和,压成单点频率 \(F_i(a)\)、\(F_j(b)\)。
  • 倾向公式:观测频率除以两个单点频率的乘积再减 1,即「实际比随机预期多/少多少」。
  • 核心-核心 / 表面-表面 vs 核心-表面:前两类两个位点地位对等,按对称方式压;核心-表面不对称,要按不对称方式压。

3. 值得留意

「减 1」让随机情形落在 0,不是 1。\(i\)、\(j\) 的具体含义这段没展开,要靠上一段的分类来对号。

Global Distribution

b062We determine the global parameters \(\{p^\textrm{glob}_a\},\{q^\textrm{glob}_{a,b}\}\) as follows.

这段在干什么:承接上文对位点类型的讨论,说明下文要确定全局层面的参数 \(\{p^\textrm{glob}_a\}\)、\(\{q^\textrm{glob}_{a,b}\}\)。

需要解释的地方:\(p^\textrm{glob}_a\) 是单个位置取某氨基酸 \(a\) 的全局频率,\(q^\textrm{glob}_{a,b}\) 是位置对 \(\{a,b\}\) 的全局频率——即"背景"分布(这是蛋白统计领域的常见记号)。

值得留意:作者用"as follows"引出后文,说明这些参数的具体确定方法在紧接着的下一段,本段只是过渡,没有给出任何公式或数据。

b063First, we sampled the pairwise empirical distributions at all sites, without considering site specificity, i.e. we averaged the site-specific frequencies over all classes of pairs that belong to the same type of proteins. We thus get what we call \(F_{a,b}^\textrm{glob,obs}\), from which we obtain the single-site frequencies \(F^{\textrm{glob,obs}}_{a}\) by marginalization and the propensities as \(q^{\textrm{glob,obs}}_{a,b}=F^{\textrm{glob,obs}}_{a,b}/\left( F^{\textrm{glob,obs}}_{a}F^{\textrm{glob,obs}}_{b}\right) -1\).

这段在干什么

承接上句,具体定义全局参数怎么从数据里算出来:先把所有位点的配对频率对同类蛋白求平均,得到全局观测频率,再算出边缘频率和倾向性 \(q\)。

需要解释的地方

  • 不考虑位点特异性:不区分是哪个位点,把所有位点混在一起平均。
  • 边缘化(marginalization):把配对频率对其中一个位点求和,得到单一位点的频率——领域常识。
  • 倾向性 \(q\):观测的配对频率,比上“若两位点独立时的期望频率”,再减 1。等于 0 就说明无耦合。

值得留意

  • \(q\) 的定义是 \(F_{a,b}/(F_a F_b)-1\),那个 −1 容易看漏,它把“独立”归零。
  • 这是观测值(obs),不是模型预测值——后文才会拿来比对。

b064This data is affected by the reduced sample size, which we treat as an overfitting problem. We fit the global parameters \(p^\textrm{glob}_a\) and \(q^\textrm{glob}_{a,b}\) from the observed parameters adopting a variant of the “rescaled ridge regression” method proposed by one of us and a coworker (Dehouck and Bastolla 2017). Precisely, we minimize the weighted mean of the errors of the fits. For propensities, we adopt the weights \(w_{a,b}=F^\textrm{glob,obs}_a F^\textrm{glob,obs}_b\) that satisfies the zero mean condition \(\sum _{a,b} w_{a,b}q^\textrm{glob,obs}_{a,b}=0\). We then add the L2 regularization term with regularization parameter R.

这段在干什么

交代如何处理上面算出的观测值,转而拟合一套全局参数,并说明所用回归方法与正则化。

需要解释的地方

  • overfitting(过拟合):样本太少时,模型会把噪声也当成规律学进去。
  • 全局参数 $p^\textrm{glob}_a$、$q^\textrm{glob}_{a,b}$:相对于前面逐条算出的观测值,这里的"glob"表示一套统一的、被拟合出来的参数。
  • rescaled ridge regression:岭回归(ridge regression)是领域常识——在最小化拟合误差时额外加一个惩罚项(L2),防止参数跑得太野;"rescaled"指这个变体带了权重/缩放。
  • L2 正则化项:即那个惩罚项,用参数 $R$ 控制惩罚强度。

值得留意

作者是最小化误差的加权平均,权重 $w_{a,b}=F^\textrm{glob,obs}_a F^\textrm{glob,obs}_b$,文中特意点出这个权重满足零均值条件——这是为了让拟合基准合理。具体怎么最小化、$R$ 取多少,这段都没说。

b065In the spirit of the rescaled ridge regression, we impose an additional condition on the average parameter values, \(\sum _{a,b} w_{a,b} q^\textrm{glob}_{a,b}\), for preventing that the scale of the fitted parameters vanish for infinite R. The additional term is formally similar to an L1 regularization with the signed value of the fitted parameters \(q^\textrm{glob}_{a,b}\) instead of the absolute value. For notational convenience, we indicate with c any pair of amino acids a, b. The scoring function that we minimize is

讲解

这段在干什么:接着上一段的 L2 正则化,作者再补一个约束项——给参数均值加条件,防止 R 变大时拟合参数整体缩到零;并预告下一步写出要最小化的打分函数。

需要解释的地方:

  • rescaled ridge regression(重标度岭回归):领域常识,岭回归是带 L2 惩罚的回归;"重标度"指因变量/参数被重新缩放。
  • L1 正则化:常识,惩罚参数绝对值之和。这里形式像 L1,但用的是带符号的 \(q^\textrm{glob}_{a,b}\) 而非绝对值。
  • \(c\):作者为省事,用 c 代表任意一对氨基酸 (a,b)。

值得留意:上一段已让观测均值等于 0;这段是给拟合参数的均值再加约束,两者对象不同。作者只说"形式相似",并非真是 L1。

b066We determine the value of the regularization parameter \(\mu\) by minimizing the mean square error of the fit for given R. This yields \(\mu =R\sum _{c} w_{c} q_{c}^\textrm{glob,obs}\equiv R\langle q^\textrm{glob,obs}\rangle\) (see Appendix). In other words, the rescaled ridge regression sets the scale of the fitted parameters that minimizes the error of the fit, and it only imposes the regularization on the relative values of the parameters, leaving the fit unregularized when there is only one parameter.

这段在干什么

确定正则化参数 μ 的取值:通过最小化拟合的均方误差,把它解出来。

需要解释的地方

  • ridge regression(岭回归):领域常识,一种带正则项的回归,正则项惩罚参数过大。
  • μ:这里就是正则化强度的缩放系数,不是随便设的,而是由数据(R、权重 w_c、观测值)算出来的。
  • $\langle q^\textrm{glob,obs}\rangle$:字面是加权平均,符号简写。

值得留意

作者特意说"只有一个参数时不施加正则化"——因为此时正则化对唯一参数的单值无相对意义。

b067Finally, we apply the specific heat criterion (Dehouck and Bastolla 2017) and we determine the regularization parameter R by maximizing the derivative of the error with respect to R, which represents a specific heat, in the statistical mechanics analogy proposed by Dehouck and Bastolla. This yields \(R=0.5\) (see Appendix).

好,我们来看这一段。

1. 这段在干什么:给正则化参数 \(R\) 定值。它用一种“比热”准则选出最优的 \(R\),最终得到 \(R=0.5\),供后续计算使用。

2. 需要解释的地方:

  • 正则化参数 \(R\):控制拟合“松紧”的旋钮,防止过拟合。
  • 比热准则(领域常识):借统计物理的类比,把“误差对 \(R\) 的导数”当作比热,取它最大处来选 \(R\)。

3. 值得留意:这里只是“选参数”这一步,\(R=0.5\) 是结果,不是结论性的科学发现。具体推导在 Appendix,正文只给结果。

b068The final result is \(q_c=\frac{2}{3}q^\textrm{obs}_c+\frac{1}{3}\langle q^\textrm{obs}\rangle\). For the propensities, because of the zero mean gauge, it holds \(\langle q^\textrm{obs}\rangle =0\). For the single-site probabilities \(p^\textrm{glob}_a\), we choose weights \(w_a= p^\textrm{obs}_a\), and \(\langle p^\textrm{obs}\rangle =\sum _a (p^\textrm{obs}_a)^2\). We finally obtain

讲解

1. 这段在干什么

这是在推一套把观测统计量"重新加权"成全局分布的公式,属于方法推导的收尾(\(\langle\cdot\rangle\) 表示按权重的平均)。上一段刚给出 \(R=0.5\),这里转入具体公式。

2. 需要解释的地方

  • 规范/规范固定(gauge):这是领域常识,指变量整体平移不影响物理结果,故可约定平均为零,于是 \(\langle q^{\rm obs}\rangle=0\)——作者正是借这一点去掉第一式里的第二项。
  • propensity(倾向性):每个位点氨基酸出现的偏好程度。
  • \(q_c^{\rm obs}\) / \(p_a^{\rm obs}\):上标 obs 是"观测到的"值,即直接从 PDB 数出来的频率。
  • 权重 \(w_a=p_a^{\rm obs}\):用观测单点概率当权重,是人为选择,非推导出来的。

3. 值得留意

\(\langle p^{\rm obs}\rangle=\sum_a (p_a^{\rm obs})^2\) 这一步——"平均"在这里不是除以 \(a\) 的数目,而是按 \(p_a^{\rm obs}\) 自身加权,所以平均值恰好是概率平方和。这是容易读漏的地方;为什么这么选,原文未进一步说明。

Assessment of the Predictions

b070We then predict the pairwise distribution for each group of pairs of sites adopting Eq.(11). We use the average value of \(\left\langle C_{ij} \right\rangle\) over all pairs in the group and the observed single-site frequencies, since we want to focus on the prediction of the pairwise couplings.

1. 这段在干什么:交代预测 pairwise 分布的具体做法——对每组位点对套用 Eq.(11),是方法落地的一步(承接上一段刚定好的权重,转入实际预测)。

2. 需要解释的地方:

  • \(C_{ij}\):位点 i、j 之间的耦合量(协变强度),下标 ij 指某一对位点。
  • 「用组内平均值」:把同一组所有位点对的 \(\langle C_{ij}\rangle\) 取平均,用来算该组的 pairwise 分布。
  • 单点频率 \(p_a\) 用观测值:沿用上一段定的 \(w_a=p^{\text{obs}}_a\)。

3. 值得留意:作者明说用平均值+观测单点频率,是为了「专注预测耦合」,即故意去掉其他来源的差异,只测 Eq.(11) 对 coupling 的预测力。

b071We considered three scores that quantify the agreement between the observed and predicted pairwise amino acid distributions. These scores are related with the log-likelihood scores that we used for optimizing the parameter \(\Lambda\). The additional scores allow a more transparent assessment, and we show them in the figures.

讲解

1. 这段在干什么

交代评价方法:声明后面用什么指标来判断"预测的氨基酸配对分布"和"实际观测"吻合得好不好,并解释为什么换用这三个分数——为了展示更直观。

2. 需要解释的地方

  • pairwise amino acid distributions(配对氨基酸分布):指两个位点上氨基酸组合出现的概率分布,是预测和观测比较的对象。
  • log-likelihood scores:对数似然分数,用来衡量模型拟合好坏,也是优化参数 Λ 时用的目标。这是领域常识。
  • Λ:模型里的参数集合(原文用大写 Lambda 表示),具体是什么这段没展开。
  • more transparent assessment:作者的意思是这三个分数比之前优化用的对数似然更"看得懂",所以图里用它们。

3. 值得留意

作者明说新分数和优化用的对数似然"相关"(related with),说明它们是同一回事的不同呈现,不是独立证据——这点容易被读成"双重验证"。

b073where \(F_{p}^\textrm{obs}(a,b)\) is the frequency of the amino acid pairs a, b observed at the pair of type p, \(Q_{p}^\textrm{pred}(a,b)\) is the predicted propensity, and \(w_{i,j}\) is the weight of pair p (see below). The log-likelihood is related with the Kullback-Leibler divergence, except for a term that does not depend on \(Q_{p}^\textrm{pred}\), so that optimizing either score yields the same results:

这段在干什么

承接上句,给出评价预测所用的打分公式中各符号的含义,并说明该 log-likelihood 与 KL 散度只差一个与 $Q_p^{\text{pred}}$ 无关的项。

需要解释的地方

  • $F_p^{\text{obs}}$:在 $p$ 型位置对上实际观察到的氨基酸对 $a,b$ 的频率。
  • $Q_p^{\text{pred}}$:模型预测的倾向性(predicted propensity)。
  • $w_{i,j}$:该对的权重,原文说"见下文",此处未展开。
  • KL 散度:衡量两个概率分布差异的常用量(领域常识)。因差的项不含 $Q_p^{\text{pred}}$,所以优化两个分数等价。

值得留意

$w_{i,j}$ 原文只说"see below",含义未在这段给出;"optimizing either score yields the same results"是指优化 log-likelihood 和优化 KL 散度结果一致,不是指这两个量相等。

b075This score is quadratic, and it is more sensitive to outliers. Its value allows us to assess the relative error of the predicted logarithmic propensities.

这段在干什么:紧接上句,说明这个二次型评分的特点,并交代它的用途——用来评估预测的对数倾向值的相对误差。

需要解释的地方:

  • 二次(quadratic):评分里含平方项。领域常识:平方会把大偏差放得更大。
  • 对离群值更敏感:正因为平方,个别偏离很远的点会显著拉高评分。
  • 对数倾向(logarithmic propensities):模型对每个位置预测出的倾向取对数后的值。

值得留意:这一段只讲了该评分"更敏感"和"能评估相对误差",但没提具体数值或阈值。

b076weighted Pearson correlation coefficient \(\textrm{wr}\) between the set observed and predicted logarithmic propensities, \(\log (1+Q_{p}^\textrm{obs}(a,b))\) and \(\log (1+Q_{p}^\textrm{pred}(a,b))\), in which each pair of amino acids a, b is again weighted proportionally to their observed frequency \(F_{p}^\textrm{obs}(a,b)\). We show this quantity because it is easy to interpret.

1. 这段在干什么

提出用加权 Pearson 相关系数 wr 衡量预测与观测的对数倾向(log propensity)的吻合度,并说明选它因为它好解读。

2. 需要解释的地方

  • 对数倾向:对观测/预测的 \(Q_p(a,b)\) 取 \(\log(1+Q)\),压缩量级、缓和极端值,这是领域常识。
  • 加权:每对氨基酸 a,b 按观测频率 \(F_p^{\text{obs}}(a,b)\) 加权——常见对多的权重更大。
  • wr:即加权后的 Pearson 相关系数。

3. 值得留意

注意是「对数倾向」之间的相关,不是原始倾向;且「好解读」是作者给出的选它的理由,而非该指标的固有属性。

b077In all of these scores, we weight each pair of amino acids proportionally to its observed number of instances \(F_{p}^\textrm{obs}(a,b)\), which down weights infrequent pairs that present large relative error, improving the reliability.

讲解

这段在干什么:交代这几项评分里对氨基酸配对的加权方式——按观测实例数 \(F_p^{\text{obs}}(a,b)\) 加权,顺带说明为什么这么做。

需要解释的地方:

  • \(F_p^{\text{obs}}(a,b)\):位置 p 上氨基酸 a 和 b 这对组合实际被观察到多少次。
  • 按实例数加权:出现得多的配对,在总分里占的份量更重。
  • 相对误差:领域常识,指误差相对于数值本身的大小;出现次数少的配对,这个比值容易偏大。

值得留意:作者给的理由是"降权出现少的配对,提高可靠性"——这是对噪声的处理,不是生物学结论。另外,加权用的是观测频率,而预测本身想预测的就是共变,所以这里已经隐含了"数据说话"的立场。

b078Finally, we fitted the selection parameter \(\Lambda\) separately from each group of proteins by maximizing the weighted sum of the log-likelihood of all pairs p, using as weight of each pair its mutual information, \(w_{p}=M_{p}\equiv \sum _{a,b}F_{p}^\textrm{obs}(a,b)\) \(\log (F_{p}^\textrm{obs}(a,b)/F_i^\textrm{obs}(a)F_j^\textrm{obs}(b))\), so that we weight more the more informative pairs.

这段在干什么

交代作者最终怎么给每组蛋白单独拟合选择参数 \(\Lambda\)——按信息量加权求和所有配对的对数似然。

需要解释的地方

  • \(\Lambda\) 单独拟合:不是全体蛋白共用一个值,而是分组各估一个。
  • 对数似然:衡量当前 \(\Lambda\) 下模型对数据的拟合好坏,越大越好。
  • 权重 \(w_p=\) 互信息 \(M_p\):互信息(领域常识)衡量两个位点关联的强弱;关联越强,这对位点在拟合里分量越重。

值得留意

用互信息当权重,等于让"信息量大"的配对主导 \(\Lambda\) 的估计;它与上一段的降权思路方向相反——一个防噪声、一个重信号,共同决定每对的实际影响力。

b079We performed the maximization through an iterative quadratic interpolation algorithm that evaluates the score for three values of the parameter, interpolates the score with a quadratic curve and determines the maximum, then it substitutes the candidate maximum to one of the three previous values and iterates until convergence.

这段在交代:搜出评分最大值是怎么算的——即对上一段那个带权重的评分函数做迭代求极大。

关键概念:只给了三个参数值试算评分,用二次曲线拟合出峰值位置,再把峰值的参数替换掉三个旧值之一,如此反复到收敛。属于数值优化的常规做法(领域常识),不是本文独创。

留意:它只保证收敛,没说收敛到全局最大;也没给迭代次数或收敛判据。

b080After fitting the selection parameter, we compute the total score for each group of proteins using as weight \(w_{p}\) of each group of pair the number of sampled pairs. In this way, the final scores are not affected by the way in which the pairs are grouped, and their comparison is meaningful.

这段在干什么

承接上一步的拟合,说明如何把选择参数换算成每个蛋白组的总分,并为跨组比较做铺垫。

需要解释的地方

  • 选择参数:上一步刚拟合出的那个参数,这里直接拿来用。
  • 以采样对数为权重:每组按采样到的 pair 数量加权,而不是简单平均。

值得留意

这里用采样对数当权重,目的是让最终分数不依赖 pair 怎么分组——即分组方式不影响结果,这样不同组的分数才能放在一起比。这是作者为后面比对预测与实际协变做的公平性处理。

Stability with Respect to Misfolding

b082The strength of the stability constrained model consists in the fact that we predict the folding stability \(\Delta G\) without any free parameter by adopting a simple contact free energy model. In this model, protein structures are represented as binary contact matrices \(C_{ij}\in \{0,1\}\), with residues i and j considered in contact if any of their heavy atoms are closer than 4.5Å. For a given sequence \(a_i\), we estimate the effective free energy of the native state as \(G^\textrm{nat}=\sum _{i<j}C^\textrm{nat}_{ij}U(a_i,a_j)\,,\) where \(C^\textrm{nat}_{ij}\) is the contact matrix of the average native structure, \(a_i, a_j\) are the amino acids at positions i and j in the sequence, and U(a, b) are the effective contact interaction parameters determined by Bastolla et al. (2000).

这段在干什么

交代稳定性约束模型的核心——用无自由参数的接触能模型去预测折叠稳定性 \(\Delta G\)(即天然态自由能 \(G^\textrm{nat}\))。

需要解释的地方

  • 接触矩阵 \(C_{ij}\):把结构简化成 0/1 表格,两残基任意重原子间距 < 4.5Å 就算"接触"。
  • 无自由参数:模型参数 U(a,b) 直接取自 Bastolla et al. (2000),不用再拟合,这点是作者强调的优势。

值得留意

作者说 Gⁿᵃᵗ 是"有效自由能",且用的是平均天然结构的接触矩阵——不是某一条具体结构,容易读漏。

b083Importantly, the free energy parameters U(a, b) where not determined from the statistics of amino acids in the PDB as in the quasi-Boltzmann approximation (Sippl 1995), which would be circular with the approach outlined here, but they were determined by imposing that the native state of a representative set of PDB structures are maximally stable against non-native decoys.

讲解

1. 这段在干什么

交代 U(a,b) 的来源:不是从 PDB 统计拟合,而是用"天然态对 decoy 最稳定"来定参。

2. 需要解释的地方

  • quasi-Boltzmann 近似:领域常识,指直接用 PDB 里氨基酸出现频率反推能量参数(Sippl 1995)。
  • circular(循环论证):本文已从 PDB 统计协变,若参数又来自同一统计,就自证自。
  • decoy:领域常识,人为构造的非天然结构,用作对照。

3. 值得留意

作者特意声明这里不循环,是在防"你拿数据又造数据"的质疑;U(a,b) 对应上段 Bastolla 等 2000 的接触作用参数。语法上 where 应为 were。

b084We estimate the effective free energy of the unfolded state as \(G^\textrm{unf}=L S_U\), where \(S_U\) is the average configurational entropy of an unfolded residue. For the ensemble of wrongly folded (misfolded) compact conformations, we adopt the Random Energy Model (Derrida 1981) in the space of compact contact matrices, which approximate the free energy of the ensemble of misfolded conformation with the mean and the variance of their contact free energies. Although corrections to the REM that involve the third moment of the energies are not negligible (Minning et al. 2013), here we do not consider them for simplicity.

这段在干什么:设定模型里"未折叠态"和"错误折叠态"两部分能量的算法,以便后面算出某个序列的折叠稳定性。

需要解释的地方:

  • \(G^\textrm{unf}=LS_U\):未折叠态自由能 = 链长 \(L\) × 每个残基的平均构象熵,很粗的近似。
  • Random Energy Model(REM):领域常识,一种把大量随机能量当独立同分布来估计系综自由能的简化模型,这里用来近似错误折叠构象的能量。
  • 它只用接触自由能的均值和方差来刻画整个错误折叠系综。

值得留意:作者自己承认 REM 忽略了三阶矩修正(Minning 等 2013 指出不可忽略),只是"为简单起见"跳过——这是模型已知的粗糙处。

(约 160 字,可略删)

b085In the above formula, the brackets \(\left\langle \cdot \right\rangle\) represent the mean value over the set of the possible compact contact matrices of proteins with L residues sampled from native structures in the PDB. Thus, the first term represents the mean contact free energy of misfolded conformations and the second term represents their variance.

这段在干什么

给上面那个公式里的两个项"翻译"含义:第一项是错折叠构象的平均接触自由能,第二项是它们的方差。

需要解释的地方

  • 方括号 $\langle\cdot\rangle$:对"一组可能构象"取平均。这里取的是从 PDB 天然结构采样出的、L 个残基的紧密接触矩阵。(领域常识:PDB 是蛋白质结构数据库。)
  • 接触自由能:把蛋白质结构抽象成"哪些残基互相接触"后,用这套接触算出的能量。

值得留意

  • 平均和方差是分开的两项——第一项管"错折叠构象平均有多不稳",第二项管"不同错折叠构象之间差多少"。上一段刚说三阶矩被略去,这里正好只剩一、二阶。

b086Because of the contact correlation terms \(U(A_i,A_j)U(A_k,A_l)\), the free energy of the misfolded ensemble depends on sets of four sites, not just pairs of sites. We therefore average these terms in order to obtain an approximated expression that only depends on pairs of sites:

讲解

1. 这段在干什么

提出一个技术困难并说明应对:错折叠系综的自由能涉及四体项,作者要把它近似成只含成对项的表达式,好进入后续计算。

2. 需要解释的地方

  • \(U(A_i,A_j)U(A_k,l)\):两个接触项相乘,牵涉 \(i,j,k,l\) 四个位点,所以是「四体」而非「两两」。
  • 平均这些项:把四体位点的贡献粗暴地压回成对形式(近似,非精确等价)。

3. 值得留意

作者明说这是 approximation——四体信息被丢掉了,这是全段最该记住的让步。「contact correlation」是领域常识:指不同接触之间非独立。

b087To evaluate the above expression, we sample the moments of the distributions of contacts from the PDB as follows.

上一段刚把公式近似成只看「位点对」的形式,这一段是接着交代怎么实际算出那个近似式:不解析推导,而是直接从 PDB 里抽样。

  • 这段在干什么:从公式推导过渡到数值实现——声明用 PDB 抽样来求接触分布的各阶矩。
  • 需要解释的地方:
  • 「moments(矩)」:领域常识,指均值、方差这类描述分布形状的量;这里指接触数的分布特征。
  • 「sample the moments…from the PDB」:即用 PDB 里真实结构统计出这些量,而非理论假设。
  • 值得留意:作者说的是「抽样」求矩,意味着后面是近似/估计值,不是精确解。具体怎么抽、抽什么,这段没提。

b088The contact frequency \(\left\langle C_{ij} \right\rangle\) represents the mean value of the contact between residues i and j in proteins of L residues. In order to sample it, we assume that this is the product of a function c(L) that only depends on the protein length and a function \(f(|i-j|)\) that only depends on the distance of the two residues along the chain, which we estimate from contact matrices in the protein data bank (PDB). \(\left\langle C_{ij} \right\rangle\) increases with L (because the surface to volume ratio decreases with L, decreasing the influence of surface residues that have fewer contacts) and it decreases with \(|i-j|\) (because short range contacts form more easily than long range ones). We sample these functions from observed contact matrices in the PDB, accelerating the subsequent computations.

这段在干什么

构造一个可计算的接触频率近似公式 \(\langle C_{ij}\rangle \approx c(L)\cdot f(|i-j|)\),参数直接从 PDB 实际接触矩阵拟合。

需要解释的地方

  • 接触频率:残基 i 和 j 在空间中靠近的平均次数。
  • 可分离假设:把 \(\langle C_{ij}\rangle\) 拆成长度函数 \(c(L)\) 与链上距离函数 \(f(|i-j|)\) 的乘积,简化采样。
  • 表面/体积比:领域常识——蛋白越大,表面占比越小,表面残基接触少,故 \(c(L)\) 随 L 增。

值得留意

作者没证明可分离性假设,只说“我们假定”;两大趋势(随 L 增、随 |i−j| 减)都只给了定性理由,无数据。

b089In the second term, \(N_C = \sum _{i<j}C_{ij}\) is the total number of contacts. The average number of contacts per residue, \(\left\langle N_C \right\rangle /L\), proportional to c(L), increases with L due to the decrease of the surface to volume ratio. We compute numerically its correlation with the contact frequencies as \(\left\langle C_{ij} N_C \right\rangle \approx Lc(L)^2 f1(|i-j|)\), sampling the factors \(f1(|i-j|)\) prior to the computation.

讲解

1. 这段在干什么

承接上句引入的接触统计,给出把「接触数 \(N_C\)」与「接触频率」关联起来的近似计算式,为后续能量项服务。

2. 需要解释的地方

  • \(N_C=\sum_{i<j}C_{ij}\):所有残基对的接触总数。
  • \(\langle N_C\rangle/L\):每个残基平均接触数;因表面积/体积比随链长 \(L\) 下降而增大(领域常识)。
  • \(f1(|i-j|)\):按序列间隔采样的因子,先采样再算,图的是省算力。

3. 值得留意

注意 \(\langle C_{ij}N_C\rangle\) 是近似(\(\approx\)),不是恒等式;\(c(L)\) 是 \(L\) 的函数,别当成常数;"sampling prior"这句正呼应上句所说的加速计算。

b090We define \(n_i=\sum _j C_{ij}\) as the number of contacts formed by residue i. Furthermore, we assume that its average value is the same for all residues (disregarding the important but limited end effect: residues at the chain termini form fewer contacts), and we numerically estimate the correlations between residues at distance \(|i-j|\) along the sequence as \(\left\langle n_i n_j \right\rangle - \left\langle n_i \right\rangle \left\langle n_j \right\rangle \approx c(L)^2 f2(|i-j|)\), again sampling the factors \(f2(|i-j|)\) prior to the computation.

这段在干什么

把上一段对“角度”相关性的处理,平行搬到“接触数”上,给出变异相关性的估计式。

需要解释的地方

  • \(n_i=\sum_j C_{ij}\):第 i 个残基与别人形成的接触总数;\(C_{ij}\) 只在与 j 有接触时非零,所以求和就是数它有几个接触伙伴。
  • \(\langle n_i\rangle\) 对所有残基取同一值:把末端效应忽略掉,好让公式简化。这是论文自己说明的假设。
  • 协方差 \(\langle n_i n_j\rangle-\langle n_i\rangle\langle n_j\rangle\):领域常识,衡量两个位置接触数是否同涨同落,为零就是不相关。
  • \(f2(|i-j|)\):作者不直接算,先按序列距离采样,与上一段的 \(f1\) 同一套做法。

值得留意

末端接触少被作者点名为“重要但有限”,却没在公式里体现,等于把链两端和中间一视同仁了。

b091Finally, in the spirit of the mean-field approximation, we approximate the terms in Eq.() that depend on the sequence with their average value, except the pairwise terms \(U(A_i,A_j)\). The factor \(\overline{U}\) is the average energy of non-native contacts, which we estimate with the global distribution \(P^\textrm{glob}\):

这段在干什么:在平均场近似下,把方程里依赖序列的项用平均值替代,只保留成对项 \(U(A_i,A_j)\),好把非天然接触的平均能量 \(\overline{U}\) 写出来。

需要解释的地方:

  • 平均场近似(领域常识):把复杂多体相互作用中每个个体替换成"平均水平",简化计算。
  • 非天然接触:序列里本不该配对的残基之间形成的接触。
  • \(\overline{U}\) 就是用全局分布 \(P^\textrm{glob}\) 估计出的这种接触的平均能量。

值得留意:被特殊对待的只有成对项 \(U(A_i,A_j)\),说明它被认为携带了关键的序列信息,其余项粗略平均掉。

b093is the average energy of residue \(a_i\) sitting at position i (and same for residue \(a_j\) sitting at position j), which we compute at the beginning of the computation based on the site-specific single site distributions \(p_{x_j}\) and the global amino correlations \(q^\textrm{glob}_{a,b}\), but neglecting the site-specific couplings \(q_{x_i,x_j}\).

讲解

1. 这段在干什么

接着上一段定义 \(\bar{e}_U\),这段给出另一项——单个残基坐在某个位置上的平均能量,指明它是怎么算出来的。

2. 需要解释的地方

  • \(a_i\) 坐在位置 \(i\):残基就是氨基酸,像「某个座位(位置)上坐着谁(残基)」。
  • 单点分布 \(p_{x_j}\):每个位置单独看,氨基酸出现的概率。
  • 全局相关 \(q^\textrm{glob}_{a,b}\):不区分位置,两个氨基酸一起出现的倾向(领域常识:共进化/协变的核心量)。
  • 在开头就算好,并且忽略位置特异性耦合 \(q_{x_i,x_j}\)。

3. 值得留意

能量只用单点分布 + 全局相关算,故意丢掉位置耦合——即这项不刻画两个位置相互依赖。

b094If \(\overline{U}>0\), the free energy of the misfolded ensemble is larger than the free energy of the unfolded ensemble, and we consider the latter for computing the folding free energy \(\Delta G\):

讲解

1. 这段在干什么

提出一个判断条件:当 \(\overline{U}>0\) 时,改选未折叠态(而非错误折叠态)作为计算折叠自由能 \(\Delta G\) 的参照基准。

2. 需要解释的地方

  • 错误折叠态 / 未折叠态(领域常识):蛋白质可能塌成"错结构",也可能松散没折叠。二者都是相对天然态的竞争状态。
  • \(\overline{U}\):这段没给定义,从上下文看是两类状态自由能高低的比较量。
  • \(\Delta G\):折叠自由能,即天然态相对参照态的稳定性。

3. 值得留意

这里的"用哪个做参照"不是随便选的——它决定了 \(\Delta G\) 数值的含义。作者是按自由能高低"挑更低的那个"来当基准,而非固定用未折叠态。\(\overline{U}\) 的具体定义在这段没提到,需回前文找。

b095As the result of these approximations, we can decompose the folding free energy as the sum of pairwise terms,

逐段讲解

1. 这段在干什么

这是一句承上启下的过渡:在上一段说明用未折叠态能量来算折叠自由能 ΔG 之后,这里交代——经过一系列近似,ΔG 可以拆成成对项(pairwise terms)之和。这个"可拆成对"是后续分析位置协变的前提。

2. 需要解释的地方

  • folding free energy(折叠自由能):蛋白质折叠前后自由能之差,衡量折叠是否有利,是领域常识。
  • pairwise terms:指把总能量写成"每两个残基之间"的相互作用项之和,而不是高阶的多体项。这是本文模型的核心简化假设。

3. 值得留意

作者用"approximations"(近似)一词,说明这里的分解是简化结果、并非严格成立;但原文没说具体是哪些近似、误差多大。

Data Sets

b097We analyzed proteins in the PDB grouped at the 70% sequence identity level (list available in the PDB web server), retaining only structures with more than 40 residues studied by X-ray crystallography, which amounts to 37108 protein chains. In order to accelerate the computations, we produced files that contain the sequences and the contact matrices

讲解

1. 这段在干什么

交代数据集怎么筛出来的:从 PDB 按 70% 序列一致性分组,只留 X 射线晶体学解析、长度 >40 残基的结构,最终得 37108 条链;并预先算好序列和接触矩阵文件以便加速。

2. 需要解释的地方

  • PDB:蛋白质结构数据库(领域常识)。
  • 70% 序列一致性分组:把彼此相似度超过 70% 的序列归为一组,避免近亲序列重复占权重(领域常识)。
  • 残基:氨基酸单元。
  • 接触矩阵:记录哪些位置在空间上相互靠近的表格。

3. 值得留意

  • 筛掉 <40 残基的结构,说明只关注能稳定折叠的小型以上蛋白。
  • 限定 X 射线晶体学,是数据来源的取舍,原因这段没说。
  • 接触矩阵是为后续计算预先准备的,本段没提具体怎么算。

Influence of the Misfolding Model

b100First, we optimized the temperature parameter T of the Random Energy Model (REM) applied to contact matrices, Eq.(18), which sets the scale of the free energy parameters U(a, b). We plot in Fig. 1 the scores summed over all pairs and all classes of proteins as the function of the parameter T. Fig. 1A and C show the log-likelihood score and Fig. 1B and D show the mean squared error of the logarithmic propensities. The top plots A,B do not consider the global propensities (\(q^\textrm{glob}_{ab}=0)\), which focuses more on folding stability, and the bottom plots C,D are obtained adopting \(q^\textrm{glob}\).

这段在干什么

作者在调参:先把随机能量模型(REM)里的温度参数 T 优化好,再看打分随 T 怎么变。

需要解释的地方

  • T(温度参数):这里是拟合用的自由参数,不是物理温度,它决定能量参数 U(a,b) 的整体尺度(原文即此意)。
  • REM + 接触矩阵式(18):领域常识——把蛋白质折叠稳定性用接触矩阵上的随机能量模型来近似。
  • log-likelihood score / MSE:一个是似然分,一个是倾向对数值的均方误差,都是拿来看拟合好坏的。
  • q^glob:全局倾向项;A、B 图设它为零,更聚焦折叠稳定性;C、D 图保留它。

值得留意

  • 四张图分两组:上排(A、B)不考虑 q^glob,下排(C、D)考虑,对比就在这里。原文只展示结果,具体结论这段没提。

b101In all cases there is an optimum at intermediate T (\(T=0.8\) without \(q^\textrm{glob}\), \(T=0.7\) with \(q^\textrm{glob}\) for log-likelihood, \(T=0.9\) without \(q^\textrm{glob}\), \(T=0.8\) with \(q^\textrm{glob}\) for wMSE, in the arbitrary units of the contact interaction parameters).

逐段带读

1. 这段在干什么

汇报结果:不管用哪种打分方式,折叠稳定性选择都在某个中间温度 T 处最优——太冷太热都不行。

2. 需要解释的地方

  • T:不是摄氏温度,是接触相互作用参数的单位,一个可调的"选择压力强度"刻度。
  • 有无 \(q^\textrm{glob}\):对应上段说的两条路线——不用它更偏折叠稳定性,用上它则额外约束整体构象。
  • log-likelihood / wMSE:两种衡量拟合好坏的指标,各自的最优 T 不同。

3. 值得留意

四种组合的最优 T 全在 0.7–0.9 这个窄区间,说明这个"最优"很稳,不挑指标也不挑模型。括号里的单位是"任意单位",别当物理温度读。

b102The continuous lines in Fig. 1 show the scores computed considering only the native free energy and disregarding the average free energy of the misfolded ensemble, \(\sum _{ij}\left\langle C_{ij} \right\rangle U(A_i,A_j)\). The dashed line is the limit of infinite T, which is obtained neglecting the variance of the energy of the misfolded ensemble, U2.

这段在干什么:解释 Fig. 1 里两种曲线的算法差异——实线只算天然态自由能,虚线则是温度趋于无穷时的极限。

需要解释的地方:

  • 那个求和项 \(\sum_{ij}\langle C_{ij}\rangle U(A_i,A_j)\) 是错误折叠系综的平均自由能,实线把它丢掉了。
  • "无限 T 极限"指忽略错误折叠能量方差 U2 后的结果,属于领域常识:温度越高,能量涨落的影响越被抹平。

值得留意:实线=忽略错误折叠平均能,虚线=再额外忽略方差,两条线的差别层层递进,是后文对比的基准。

b103The data show that it is important to consider selection against misfolding, which is usually neglected when modelling the folding free energy.

承接上一句对 U2(错折叠系综能量方差)的忽略,这一段给出结论。

1. 这段在干什么:收束上文——说明建模折叠自由能时,考虑"抗错折叠的选择"很重要,而这一点常被忽略。

2. 需要解释的地方:

  • selection against misfolding:自然选择会淘汰那些容易错误折叠的蛋白序列,即"抗错折叠的选择压力"。
  • folding free energy(折叠自由能):领域常识,指蛋白折叠态与去折叠态之间的自由能差,是稳定性的度量。

3. 值得留意:作者指出"通常被忽略"——这正是本文要补的缺口,也是它主张 U2 不可省的理由。

b104Agreement between observed and predicted pairwise propensities as the function of the parameter T of the REM. A, C: log-likelihood score. B, D: Mean Square Error. In plots A, B, the only fitted parameter was the selection strength \(\Lambda\). In plots C,D we also adopt the global propensity matrix \(q^\textrm{glob}\), which we obtain from the data

这段是结果小节的图注开头:交代下面四张图在比较什么、横轴是什么、两种误差指标分别是什么。

关键概念:T 是 REM(一种把折叠稳定性写成可调参数的模型)的参数;A/C 用对数似然,B/D 用均方误差——都是衡量“预测的位点配对倾向”与“观测值”差多远。Λ 是选择强度,A、B 里唯一拟合的参数。C、D 额外引入全局倾向矩阵 q^glob,从数据本身得到。

值得留意:原文在 q^glob 处戛然而止,句子没写完,别以为有省略的结论;这里只描述图的设置,尚未给出结果。

b105In all the following figures, we show data obtained at \(T=1.0\) and considering \(q^\textrm{glob}\), unless otherwise stated.

先接着上一段看:上一段刚说完"图 C、D 用全局倾向矩阵 \(q^\textrm{glob}\)",这一段就是在给后面所有图定默认条件。

1. 这段在干什么:定个"默认设置"——以下(之后的)所有图,除非专门说明,都是在 \(T=1.0\)、用 \(q^\textrm{glob}\) 的情况下得到的数据。

2. 需要解释的地方:

  • \(T=1.0\):这里的 \(T\) 是模拟里的温度参数,属于领域常识——这类能量模型通常用无量纲温度,\(1.0\) 常作为基准值,不一定等于室温。
  • 原文没解释 \(q^\textrm{glob}\) 是什么,指回上一段那个"全局倾向矩阵"。

3. 值得留意:这是个"除非另有说明"的声明(unless otherwise stated),意味着后面的图只要不特别标注,都套用这套设置;看后续图时别默认它们换了参数。

The Propensities are Related with the Contact Energy in Opposite Ways for Pairs with Native Contacts and Non-native Contacts

b107From Eq.(21) and (11), we see that the predicted propensities of amino acid pairs at pairs of sites p that form native contacts (\(C^\textrm{nat}_{p}=1\)) are negatively correlated with their contact energy U(a, b), i.e. amino acids that interact more favorably are predicted to be more frequent for pairs of sites in contact. On the contrary, for pairs of sites that do not form native contacts, the propensities are predicted to be positively correlated with U(a, b), i.e. amino acid pairs that repel each other are predicted to be more frequent, particularly so for short contact range for which \(\left\langle C_{p} \right\rangle\) is large (i.e. they can form contacts more easily).

这段在干什么

从公式推出:有天然接触的位点对,倾向性与接触能负相关;无天然接触的则正相关。

关键概念

  • 倾向性(propensity):某氨基酸对出现在该位点的频率倾向。
  • 天然接触(native contact):蛋白真实结构里实际贴在一起的两个位点,这是领域常识。
  • 接触能 U(a,b):越低越"合得来"。
  • ⟨C_p⟩ 大:该位点对容易形成接触。

值得留意

有接触时"越合得来越常见",无接触时反过来——"越互相排斥越常见",尤其在短程接触范围。这是作者从公式预测的,不是观测数据。

b108From Eq.(21), we predict that the propensities are correlated with the zero-mean value of the energy parameters, defined as

讲解

1. 这段在干什么

承接上句关于「短程接触」的讨论,作者从公式 (21) 出发,提出一个预测:倾向性(propensities)会与能量参数的零均值版本相关联,为下文对比 native / non-native 接触对做铺垫。

2. 需要解释的地方

  • Eq.(21):前文推导出的关系式,这里直接引用,具体形式这段没给。
  • propensities(倾向性):指某两个位置一起发生协变(covariation)的统计偏好强度。
  • zero-mean value(零均值化):领域常识——把能量参数减去其平均值,使均值为 0,便于只关心「偏离平均」的部分,正负号才有可比意义。

3. 值得留意

这里只说「预测相关」(correlated),没说正相关还是负相关;方向问题留到标题里的「opposite ways」才展开。读时别提前代入结论。

b109We assess this expectation in Figure 2, which shows the scatter plots of the observed logarithmic propensities versus the zero-mean energy. We group together all proteins, and we distinguish pairs with direct, indirect and no native contacts and with short and long contact ranges.

承接上句的预测,这段说明用 Figure 2 的散点图来检验它:横轴是零均值能量,纵轴是观测到的对数倾向性。这是领域常识:倾向性指某对位置共变的偏好程度,零均值能量是能量参数去中心化后的值——上句刚下了预测,这段就给出检验方案。

分组有两套维度:按天然接触分直接、间接、无三类,按接触距离分长短。

值得留意:此段只交代了图怎么分组,还没给任何结论或数值,别把它读成"已经验证了预测"。

b110Comparing the different types of contacts, we see that the propensities of native contacts present strong negative correlation with the contact energies, as predicted. This means that stabilizing pairs of amino acids are more frequent. The corresponding slope is much stronger for long range native contacts (Fig. 2D) than for short range native contacts (Fig. 2A), as expected since short range contacts are subject to “negative design” that tends to make them less favorable.

这段在干什么

对比不同接触类型,验证「天然接触的倾向与接触能量负相关」,并解释长程比短程斜率更陡。

需要解释的地方

  • 倾向(propensity):某对氨基酸出现的偏好程度。
  • 接触能量(contact energy):衡量两残基相互作用是否稳定,越低越稳定。
  • 负相关:能量越低、倾向越高,即稳定配对更常出现。
  • 长程 / 短程接触:序列上距离远的 vs 近的残基对。
  • 负设计(negative design):领域常识,指进化主动让某些配对变得不利,以免错误折叠。

值得留意

长程斜率更强,是因为短程接触受「负设计」压制而变得不利——这是作者给出的解释,不是数据本身。

b111In agreement with negative design, short range pairs without native contacts (Fig. 2B) and with indirect contacts (Fig. 2C) are positively correlated with the contact energy, as predicted by our formula. This means that destabilizing pairs of amino acids are more frequent. However, this correlation disappears for long range pairs without contact, which are less subject to negative design (Fig. 2E) and it becomes negative for long range indirect contacts, which are more similar to native contacts than to no contacts (Fig. 2F).

承接上一段短程接触受负设计影响,这段把图2B–F的分类逐一说清楚。

这段在干什么:按“有无天然接触 × 短程/长程”四类,分别报告它们与接触能的相关性方向。

需要解释的地方:

  • 负设计(领域常识):进化倾向于让非天然接触变得不稳定,避免错误折叠。
  • 间接接触:两个位点不直接接触,但通过中间残基相连。

值得留意:

  • 短程无接触、间接接触都是正相关,说明去稳定配对更常见。
  • 长程无接触相关性消失,长程间接接触反而变负——越像天然接触,趋势越反转。

b112We can see that the amino acid pair Cysteine-Cysteine (CC) has one of the largest propensities in all plots. This pair, which can form disulfide bridges, has the most stabilizing contact energy. However, its nature of outlier is not due to its effect on stability but it is due to the fact that CC has the largest global propensity, which we attribute to the mutation process and to global selective pressures (see below).

这段在干什么

用 CC(半胱氨酸对)做例子,说明它倾向值大但不是因为稳定作用,而是突变过程和全局选择压。

需要解释的地方

  • propensity(倾向值):这里指氨基酸对在蛋白质里一起出现的频率偏好,图中越高的柱子代表越常被选中。
  • disulfide bridge(二硫键):两个半胱氨酸之间形成的化学键。领域常识:它能锁住蛋白结构、增加稳定性。
  • contact energy(接触能):这里指该残基对接触带来的稳定化能量。

值得留意

作者特意把"稳定作用最强"和"倾向值最大"拆开:CC 虽接触能最稳定,但它的异常高倾向被归因于突变与全局选择,原因留到下文。

b113Scatter plots of the observed logarithmic propensities of amino acid pairs versus the zero-mean energy of the same amino acids, distinguishing pairs in contact (A, D), pairs without contacts (B, E) and pairs with indirect contacts (C, F). The first line represents short range contacts (A, B, C), the second line represents long range contacts (D, E, F)

这段在干什么

这是图注,用散点图展示氨基酸对的观测对数倾向(propensity)与能量的关系,并按接触类型分组。

需要解释的地方

  • propensity(倾向):某氨基酸对出现的观测频率相对随机期望的比值,取对数后正负表示偏好或回避。
  • zero-mean energy(零均值能量):接触能量做了去均值处理,便于比较。
  • contact / non-contact / indirect contact:直接接触、不接触、间接接触三类,对应不同子图。
  • short/long range:第一行是短程接触,第二行是长程接触(领域常识:按序列距离远近区分)。

值得留意

正文段落讲"两类接触中倾向与能量的关系方向相反",但这条图注本身只描述图形,没有复述这个结论——别把正负方向归到图注里。

b114Figure 3 plots the scatter plots of observed versus predicted propensities. They represent the same type of pairs as Fig. 2.

这段是过渡句,把读者的注意力从上一段的图 2 转到图 3,交代图 3 画的是什么。

需要解释的地方

  • scatter plots(散点图):横轴是预测值,纵轴是观测值,每个点代表一对位置。
  • observED vs predictED propensities:观测到的倾向性 vs 模型预测的倾向性——散点越贴近对角线,说明预测越准。
  • Fig. 2:原文只说「同样类型的 pairs」,具体指哪些对,这段没展开。

值得留意

  • 图 3 与图 2 是同一类 pair,两图可对照看,别当成独立的两组数据。
  • 上一段结尾说的 A–F 分行(短程/长程),本段没提,但图 3 大概率沿用同样的分组方式。

b115We can see that the SCA approximation allows reasonable predictions of the logarithmic propensities. The quality of the prediction is higher for long range native contacts (Fig. 3D) than for short range native contacts (Fig. 3A) that are subject to the contrasting forces of negative and positive design.

这段在干什么

拿 SCA 预测的对数倾向性跟观测值对比,报告拟合质量:长程天然接触预测得比短程天然接触好。

需要解释的地方

  • SCA 近似:一种从序列协变推断残基间耦合的简化方法(领域常识),这里用来预测倾向性。
  • 对数倾向性:把倾向性取对数后的量,作者直接拿它和预测值比。
  • 长程/短程天然接触:指在天然结构中彼此接触的两个位置,相隔远近不同。

值得留意

短程天然接触预测差,作者归因于它同时受"负设计和正设计"两种相反力量拉扯——这句是解释原因的关键,别只看到"预测差"就跳过。

b116For sites without contact, the prediction is much better for long range pairs that are not subject to negative design (Fig. 3E). This is due to the fact that the prediction in this case is exclusively based on the global propensities \(q^\textrm{glob}\), which may suffer from overfitting despite the regularization that we adopted.

讲解

1. 这段在干什么

解释一个现象:对于没有接触的位点,长程配对的预测反而更好;并说明原因是此时预测只依赖\(q^\textrm{glob}\),可能过拟合。

2. 需要解释的地方

  • without contact / 无接触配对:指在天然结构中两个位点并不相互接触。
  • negative design(负设计):领域常识,指进化在选择时"避免"某些不利相互作用,不单是"促进"有利的。
  • \(q^\textrm{glob}\)(全局倾向):整体统计得到的倾向性,与具体配对无关。
  • overfitting(过拟合):模型把噪声也学进去了,泛化变差。作者说已用 regularization(正则化)抑制,但仍可能出现。

3. 值得留意

作者把"预测更好"归因于只用了全局倾向,暗示这种"好"可能是假象——因为没有局部细节可依赖,也就没有过拟合的来源被摊薄。这一段是在为前面的对比做补充说明,不是新论点。

b117Indirect contacts show lower agreement (Fig. 3C and F). In all cases, the amino acid pair CC presents one of the highest propensities.

这段在干什么

接着上句比较不同接触类型的倾向性:指出间接接触的吻合度更低,并强调 CC 这种氨基酸对在各类情形下倾向性都偏高。

需要解释的地方

  • Indirect contacts(间接接触):残基在空间上靠近但并非直接配对,属于领域常识中的一类接触。
  • propensity(倾向性):此处指某氨基酸对出现共变的倾向强弱。

值得留意

  • "in all cases"指哪种范围内的"所有情形",这段没说明,需回看上下文。
  • CC 倾向性高,但原文没说它是否对应高或低接触能,方向性留待后文。

b118Observed versus predicted logarithmic propensities of amino acid pairs for pairs of sites in contact (A, D), without contacts (B, E) and with indirect contacts (C, F). The first line represents short range contacts (A, B, C), the second line represents long range contacts (D, E, F)

这段在干什么

这是图3的图注,说明六个子图(A–F)分别对应哪类位点对和哪种作用距离,为正文比较“观测 vs 预测倾向性”提供坐标。

需要解释的地方

  • 倾向性(propensity):领域常识,指某氨基酸对出现的偏好程度,这里取对数。
  • contact / non-contact / indirect contact:直接接触、无接触、间接接触(隔着别的残基)。
  • short / long range:序列上近 vs 远的接触。

值得留意

本段只说图的布局(A–F 对应关系),不含任何结论或数值;结论在上文与后文。

Dependence on Contact Range and Buriedness

b120In Fig. 4, we show how the contact range \(|i-j|\) influences the weighted correlation coefficients between the logarithmic propensities of each amino acid pair, \(\log (1+Q^\textrm{obs}(a,b))\), and the zero-mean of the free energy matrix, Eq.(22), on which our theoretical prediction is based. The figure depicts pairs of sites in contact (top row), not in contact (middle row) and in indirect contact (bottom row). Each column represents one of the four classes of proteins (Short hydrophobic, short polar, Long hydrophobic and Long polar) and each symbol represent a type of pairs of sites (surface-surface, surface-core and core-core).

这段在干什么:介绍 Fig. 4 的坐标含义——看接触距离 \(|i-j|\) 如何影响氨基酸对的加权相关系数(观测倾向 vs. 理论自由能矩阵)。

需要解释的地方:

  • \(|i-j|\):序列上两个位点隔多远,即接触范围。
  • \(\log(1+Q^\textrm{obs}(a,b))\):观测到的氨基酸对共变倾向取对数。
  • 自由能矩阵零均值:理论预测所依据的能量矩阵(去均值)。
  • 四种蛋白分类、三类位点对(表面-表面、表面-核心、核心-核心):图的列与符号。

值得留意:图分三行——接触、不接触、间接接触;课文上一段提的"short/long range"对应这里的列分类。

b121The correlations are negative for pairs of sites in contact (top row, Fig. 4A, B, C, D), as expected based on positive design, but they are less negative for short range contacts as expected based on negative design, since short range contacts have large contact frequency \(\langle C_{ij}\rangle\), i.e. they are frequently formed in misfolded conformations. The strength of the correlations and their dependence on contact range is larger for pairs of sites both in the core (triangles) than pairs on the surface (circles). Correlations are relatively weaker for short polar proteins (Fig. 4B).

讲解

1. 这段在干什么

解读 Fig. 4A–D 的结果:位点间的相关性为负,并随接触距离、埋藏程度和蛋白类型而变化,用以支持“正设计/负设计”共变的解释。

2. 需要解释的地方

  • 相关性为负:两个位点氨基酸偏好此消彼长。
  • contact range(接触距离):序列上相隔多远。短程=序列上靠近。
  • \(\langle C_{ij}\rangle\) 接触频率:某对位点接触出现的频繁程度。
  • 正/负设计:领域常识——正设计求正确折叠稳定,负设计避免错误折叠。
  • triangles/circles:图里的符号,分别代表核心、表面位点对。

3. 值得留意

短程接触负相关“较弱”,作者归因于它们易在错误折叠中出现——这是一个解释,不是直接观测。相关性对距离的依赖,在核心内比表面更强。

b122For pairs of sites that do not form native contacts (middle row, Fig. 4E, F, G, H) the correlation between propensities and contact energies is positive for short range contacts, as expected based on negative design, and it becomes negative for long range contacts. However, these negative correlations at long ranges are not attributable to positive design, as for pairs of sites in contact, but to the negative correlation between the contact energies and the global propensities \(q^\textrm{glob}\), which present positive propensities between pairs of hydrophobic amino acids. In fact, \(q^\textrm{glob}\) determines almost completely the propensities of long range pairs without contacts (Fig. 3E). In long proteins (Fig. 4G,H), the strongest correlations for non-contact pairs happen for pairs of site in the core (triangles), as it is the case for contact pairs, while for short proteins (Fig. 4E,F) they happen for surface-surface sites (the core is not so well defined for short proteins, and many of these proteins are forming complexes, but quaternary contacts are not taken into account in our analysis).

这段在干什么

解释非接触位点对的相关性随距离的变化:短程为正、长程为负,且负相关另有原因。

需要解释的地方

  • native contacts(天然接触):蛋白折叠态里实际挨在一起的两个位点。
  • propensity(倾向性):某氨基酸出现在某类位置的偏好程度。
  • negative design(负设计):演化"刻意避免"某些不利组合,领域常识术语,不是本文新造。
  • \(q^\textrm{glob}\):全局倾向性。

值得留意

长程负相关不是正设计造成的,而是被 \(q^\textrm{glob}\) 主导——这是本段的关键转折。另外短蛋白的核(core)本身定义就模糊,且文中忽略了四级接触。

b123Similar to no-contact pairs, indirect contact pairs (Fig. 4I,J,K,L) show positive correlations at short range, as expected based on negative design, and negative at long range, where correlations are more negative than for no-contact pairs, i.e. long range pairs that form indirect contacts tend to interact more favourably than pairs that do not form contacts, which suggests that some kind of positive design is taking place for indirect contacts.

这段在干什么:报告间接接触对的相关性随距离变化的规律,并据此推测间接接触存在正设计。

需要解释的地方:

  • 间接接触对:两个位点不直接接触,但通过中间位点间接关联(领域常识)。
  • 负设计/正设计:负设计指避免不利相互作用,正设计指主动优化有利相互作用(领域常识)。
  • 短程正相关:与「无接触对」类似,作者说这符合负设计的预期。

值得留意:长程间接接触对的负相关比无接触对更强,作者据此推断它们相互作用更有利——这是相关性推断,非直接实验证据。

b124weighted Pearson correlation coefficients between the propensities of each amino acid pair, \(Q^\textrm{obs}(a,b)\), and the zero-mean of the free energy matrix, Eq.(22), on which our theoretical prediction is based, as a function of the contact range \(|i-j|\), for pairs in contact (first row, plots A,B,C,D), pairs not in contact (middle row, E,F,G,H) and pairs with indirect contact (bottom row,I,J,K,L). The three curves show three different pairs of sites. The four columns show the different classes of proteins

逐段带读

1. 这段在干什么

段首用一整句长句交代这张图的"画法":横轴是接触距离 \(|i-j|\),纵轴是观测倾向 \(Q^\textrm{obs}(a,b)\) 与自由能矩阵零均值之间的加权 Pearson 相关系数。承接上一段"间接接触存在正设计"的推断,这段把相关性按接触类型分层画出来。

2. 需要解释的地方

  • 加权 Pearson 相关:领域常识,就是相关系数的一种算法形式。
  • 零均值:文中只说"矩阵减去均值",没解释为什么减,这段没提。
  • 接触 / 不接触 / 间接接触:三种配对,分别对应三行图。

3. 值得留意

三行 × 四列共 12 张图,但第二句才交代曲线的身份——三条曲线代表三对位点;四个列对应不同蛋白类别。原文没给出任何趋势结论,只是"说明书",结论要看图本身。

b125We show in Fig. 5 the quality of the predictions, assessed through the weighted Pearson correlation coefficients between observed and predicted propensities, which are based on four global propensities matrices, one for each group of proteins. As in the previous figure, the three rows represent the three types of contacts (native contact, no contact and indirect contacts) and the four columns represent the four groups of proteins.

这段在干什么:交代图5要展示预测质量,并说明图5的行列布局。

需要解释的地方:

  • 加权皮尔逊相关系数:衡量观测倾向与预测倾向吻合程度的指标,越接近1越准(领域常识)。
  • 倾向矩阵(propensity matrix):每个蛋白组各一个的全局统计表,预测的基础。
  • 三类接触:native contact(天然接触)、no contact(无接触)、indirect contacts(间接接触)。

值得留意:图和上一段图共用同一套行列结构——三行=三类接触,四列=四组蛋白;上一段说的"三条曲线""不同位点对"是指另一张图,别混。

b126Comparing the contacts, one can see that the predictions are generally worse for indirect contacts (bottom row), which are not well represented in the SCA approximation. Comparing pairs of sites, we see that the predictions tend to be comparatively slightly worse for core-surface pairs (green curves). Comparing the different contact ranges, we see that the predictions are generally worse for short range contacts, which are more strongly affected by negative design, than for long range contacts, for essentially all types of contacts. Finally, comparing the four types of proteins, we can see that the worst predictions are achieved for short polar proteins, which are less affected by selection for negative design, which is an important component of our model. The peculiarity of short polar proteins could also be related with the fact that many of them occur in multichain complexes, so that the distinction between core and surface sites is not so clear-cut.

这段在干什么

逐条比较不同类别的预测效果,总结哪类接触/位点/蛋白预测更差,作为对上文图表的解读收尾。

需要解释的地方

  • 间接接触:两个位点在空间上挨得近,但不是直接相互作用,模型(SCA 近似)难以刻画。
  • negative design(负设计):领域常识,指演化不只"选中"某结构,还"排除"会错误折叠/聚集的序列。
  • short polar proteins:序列短、极性残基多的蛋白。
  • core-surface pairs:一个埋在内部、一个露在表面的位点对。

值得留意

  • 作者对 short polar 蛋白预测差给了两个并列解释(负设计影响弱;多为多链复合物,core/surface 界限模糊),并未定论哪个为主。
  • "indirect contacts 在 SCA 近似中表征不佳"——这是方法本身的局限,不只是数据问题。

b127Weighted Pearson correlation coefficients between the observed and predicted propensities of each amino acid pair, \(Q^\textrm{obs}(a,b)\), and \(Q^{\textrm{pred}}(a,b)\) as a function of the contact range \(|i-j|\), for pairs in contact (first row, A, B, C, D), pairs not in contact (middle row, E, F, G, H) and pairs with indirect contact (bottom row, I, J, K, L). The three curves show three different pairs of sites. The four columns show the different classes of proteins. The dashed line is a reference for the eye at \(r=0.5\)

这段是图注,交代一张按接触范围分层展示预测效果的大图怎么读。

关键概念:\(Q^\textrm{obs}\) 与 \(Q^\textrm{pred}\) 是某氨基酸对在真实数据与模型预测下的共变倾向(领域常识:打分/倾向性);接触范围 \(|i-j|\) 指两个位点在序列上隔多远;indirect contact 指不直接接触但通过第三方间接关联。横轴即 \(|i-j|\),纵轴是加权 Pearson 相关系数,衡量预测与观测的吻合度。

值得留意:图分三类(接触/不接触/间接接触)×四列蛋白类别,每条曲线是"三对位点",容易误读成三类各一条。虚线 \(r=0.5\) 只是眼睛的参考线,不是判定阈值。

Mutual Information Increases with the Contact Range for Pairs in Contact and Decreases for Pairs Not in Contact

b129As the result of the different selective pressures that act on native contacts, no-contacts and indirect contacts and on short and long contact ranges, the information content varies for different types of pairs.

这段话是过渡句:把前面「不同选择压力」和后面「信息量因配对类型而异」串起来,为下文分类讨论做铺垫。

需要解释的地方:

  • native contacts(天然接触):蛋白质在实际折叠结构里相互靠近的残基对(领域常识)。
  • no-contacts / indirect contacts:不接触、以及不直接接触但间接相关的配对。
  • 接触范围(contact range):序列上相隔多远,分短程和长程。
  • information content(信息量):这里指互信息,衡量两个位点是否协同变化。

值得留意:作者把「接触类型」和「范围长短」当作两个独立维度同时影响信息量,说明后面分析会交叉分类,而不是只看一种。

b130To illustrate this fact, we computed the observed mutual information for every class of pairs, removing the mutual information obtained from the global propensities \(q^\textrm{glob}\) in order to focus only on the site-specific contributions to mutual information.

这段在干什么

接上一段"信息量因配对类型而异",用具体计算来展示这一事实。

需要解释的地方

  • observed mutual information(观测互信息):领域常识,衡量两个位点取值是否相关。
  • global propensities q^glob(全局倾向):指只由氨基酸种类本身带来的背景关联,作者把它扣掉。
  • site-specific contributions(位点特异性贡献):扣掉背景后剩下的、真正来自具体位置的信号。

值得留意

扣掉 q^glob 是为了聚焦位点本身的贡献,不是否定全局倾向的存在——作者只是想让信号更干净。

b131We show the results in Fig. 6, lumping all proteins together and distinguishing contact types and pairs of sites (surface-surface in Fig. 6A, core-surface in Fig. 6B and core-core in Fig. 6C). As expected, the highest mutual information occurs for pairs in contact, followed by indirect contacts, while pairs that do not form contacts present almost vanishing values (i.e., most of their mutual information arises from the global propensities), except at short range where they are influenced by negative design.

这段把 Fig. 6 的结果按接触类型拆开讲:作者把上一段处理过的 MI(扣掉全局倾向、只留位点特异贡献)拿来看不同接触方式的差异。

这段在干什么:交代 Fig. 6 的核心观察——接触对的 MI 最高,间接接触次之,不接触的几乎为零。

需要解释的地方:

  • contact types:按位置分表面-表面(6A)、核心-表面(6B)、核心-核心(6C),这是领域里常见的分区。
  • global propensities:氨基酸本身出现频率带来的背景相关,不是位点特异的。
  • negative design:领域常识,指选择偏向"避免某些搭配"。

值得留意:不接触对"几乎消失"是扣掉全局倾向后的结论;但短程处它们仍受 negative design 影响,这个例外容易被略过。

b132As expected, the mutual information of sites in contact tends to increase for longer ranges. In fact, we expect that short range contacts present small mutual information because they are subject to the contrasting forces of positive and negative design.

这段在干什么

顺着上一段往下推进:给出"接触位点随距离增大互信息更高"这个观察,并预告短程接触为何互信息偏低。

需要解释的地方

  • 接触范围(range):此处指序列上两残基的距离,不是空间距离。
  • 互信息:衡量两个位点变异是否关联的指标(领域常识)。
  • 正/负设计:正设计维持功能、负设计避免错误互作,两者拉扯短程接触。

值得留意

"contrasting forces"是作者对短程偏低原因的预期解释,不是已证实的结论;原文只给方向,没给数字。

b133The contrary holds for sites without contacts and with indirect contacts, which present the strongest mutual information at short range, as expected based on negative design.

这段在干什么

用「相反的情况」对上一段做对照:无接触/间接接触位点的模式恰好相反,并给出解释。

需要解释的地方

  • contrary(相反):指与前文「接触位点」的趋势相反——不是长程最强,而是短程最强。
  • short range / 长程:这里指序列上相隔的距离(领域常识)。
  • negative design(负设计):领域常识,指进化主动避免某些不该发生的相互作用。

值得留意

  • 作者把「短程最强」归因于 negative design,是预期之内的,语气是印证而非新发现。
  • 「indirect contacts」也被归入这一类,说明关键不只是有没有直接接触。

b134Nevertheless, unexpectedly, pairs of surface sites in contact (Fig. 6A circles) present higher mutual information than pairs of core sites in contact (Fig. 6C circles). Selection for protein folding stability cannot explain this difference, since surface sites contribute less to protein stability. It is intriguing to speculate that this result may be due to selection for function and/or for proper protein interfaces.

讲解

1. 这段在干什么

抛出一个反常结果:表面接触对(Fig. 6A 圆圈)的互信息反而比核心接触对(Fig. 6C 圆圈)高,并给出推测性解释。

2. 需要解释的地方

  • 互信息:衡量两个位点"协同变化"强度的量,越高说明越相关(领域常识)。
  • in contact / 接触对:在三维结构里空间上挨着的一对残基。
  • 核心 vs 表面:核心位点埋在蛋白内部,表面位点露在外面。
  • folding stability(折叠稳定性):位点对稳定性的贡献。

3. 值得留意

作者明确说稳定性选择解释不了这个差异——因为表面位点对稳定性贡献本来就小。所以结尾只是"猜测(speculate)"可能源于功能选择或界面选择,并非结论。

b135Mutual information of each class of pairs of sites as a function of the contact range \(|i-j|\), for pairs in contact (cicles), pairs not in contact (squares) and pairs with indirect contact (triangles), distinguishing surface-surface (plot A), core-surface (plot B) and core-core sites (plot C)

这段在干什么:这是图注,说明用互信息(MI)衡量位点对随接触距离 \(|i-j|\) 的变化趋势,并把位点分为接触/不接触/间接接触三类,再按表层-表层、核心-表层、核心-核心分三张图。

需要解释的地方(领域常识):互信息是衡量两个位点氨基酸组成是否协同变化的量;接触距离 \(|i-j|\) 指序列上相隔多远;cicles/squares/triangles 是图上的圆形、方形、三角标记。

值得留意:标题说接触对 MI 随距离升高、不接触对降低,但图注本身只描述坐标轴与分组,未给趋势结论,别把标题当图注内容。

Long Proteins are Subject to Stronger Selection

b137Figure 4 suggests that long and hydrophobic proteins are subject to stronger selection for protein folding stability than short and polar proteins. To clarify this aspect, we plotted in Fig. 7A the selection parameter \(\Lambda\) fitted for the four different classes of proteins, considering one global propensity matrix \(q^\textrm{glob}\) for each protein class (Qglob4, black bars), only one \(q^\textrm{glob}\) for all protein class (Qglob1, green bars) and no \(q^\textrm{glob}\) at all (Qglob0, orange bars). We also compared models that consider core and surface sites (NC2, solid bars) with models that consider only one type of site, abolishing the distinction between core and surface sites (NC1, dotted bars).

这段在干什么

承接上文「长蛋白、疏水蛋白受到更强的折叠稳定性选择」的推测,用 Fig. 7A 做定量对比:比较不同蛋白分类下拟合出的选择参数 Λ。

需要解释的地方

  • Λ(选择参数):衡量选择压力强弱的拟合量,越大表示选择越强。
  • q^glob:全局倾向性矩阵,用于刻画氨基酸突变的偏好;Qglob4/1/0 是三种设定——每类蛋白一个、所有蛋白共用一个、完全不用。
  • NC2 / NC1:是否区分核心与表面位点。NC2 区分(实心条),NC1 不区分(虚线条)。

值得留意

图中用黑色、绿色、橙色、实心/虚线区分不同模型,读图时要先分清「条形颜色管 q^glob 设定,实虚管位点区分」,别混在一起。

(注:本段只描述图的设计,未给出任何结论或数值。)

b138As expected, the more detailed a model is, the lower the selection coefficient is, which indicates that it complies well with the mimimum selection principle. The lowest selection coefficients \(\Lambda\) corresponds to the most accurate global model NC2 Qglob4 (two sites and one global propensity matrix for each protein class).

这段在干什么

解释模型越细、选择系数越低,说明结果符合「最小选择原则」;并点名最低系数来自最精细的模型 NC2 Qglob4。

需要解释的地方

  • 选择系数 Λ:这里可理解为模型需要施加多大的「选择压力」才能解释观测到的氨基酸共变。数值越低,说明模型越省力、越符合奥卡姆剃刀(领域常识,非本文结论)。
  • 最小选择原则:选择压力越小越可取——模型不靠额外假设硬凑结果。
  • NC2 Qglob4:本文中的一种全局模型,每个蛋白类别用两个位点加一个全局倾向矩阵。

值得留意

「模型越细、系数越低」是作者说「符合预期」的,不是意外的发现;且最低值恰好给了最精细的模型,这个对应关系是这段的落点。是否唯一最低,原文未明说。

b139In almost all cases, we find larger selection coefficients for long proteins, which suggests that longer proteins are subject to stronger selection for protein folding stability. This is consistent with our previous conclusion that long proteins are more affected by misfolding problems, and they are subject to stronger negative design (Minning et al. 2013).

讲解

这段在干什么:顺着上一段的模型结果,报告一个发现——长蛋白的 selection coefficient 更大,并把它和之前"长蛋白更容易受错折叠困扰、受到更强的 negative design"的结论接上。

需要解释的地方:

  • selection coefficient(选择系数):这里衡量的是"选择压力有多强"的数值,越大说明该位点/蛋白越受进化选择约束(领域常识)。
  • negative design:指进化上"避免"某些不利结构或相互作用的设计压力(领域常识)。
  • Qglob4 / NC2:上一段提到的全局模型名,是本段的来源背景,本段未再展开。

值得留意:作者用"almost all cases"而非"全部",说明这一趋势是普遍的但非绝对;且结论是"支持/一致于"旧结论,而非新证据本身。

b140Nevertheless, when we compare polar and hydrophobic proteins, we find results contrary to our expectations. Misfolding problems are expected to be more serious for hydrophobic proteins, which are subject to stronger negative design (Minning et al. 2013), but the value of \(\Lambda\) is slightly larger for long polar proteins than for long hydrophobic proteins.

逐段讲解

1. 这段在干什么

报告一个与预期相反的结果:疏水蛋白本该问题更大,但长极性蛋白的 \(\Lambda\) 反而略大于长疏水蛋白。属于转折/反例段落。

2. 需要解释的地方

  • 极性 / 疏水蛋白:按氨基酸侧链亲水性划分的两类蛋白(领域常识)。
  • \(\Lambda\):本文用来衡量位置间共变强度的量,具体定义这段未给。
  • 负设计:进化上"避免错误折叠"的选择压力(领域常识)。

3. 值得留意

  • 作者只说"略大"(slightly larger),没给具体数值,也没下结论,只是陈述与预期冲突。
  • "contrary to our expectations"说明前文预期是疏水蛋白压力更强——这里出现了反证。

b141To further investigate this issue, we plot in Fig. 7C the maximum correlation between propensities and contact energies across no-contact pairs. This correlation allows quantifying negative design without any model. It tends to be larger for hydrophobic than for polar proteins and for long than for short proteins, as expected. Thus, we conclude that our qualitative expectation on negative design is confirmed, but the fitted value of \(\Lambda\) may not be accurate enough to show subtle differences.

这段在干什么:用一张新图(Fig. 7C)换个角度验证前面关于负设计的猜测,并承认拟合出的 \(\Lambda\) 可能不够灵敏。

需要解释的地方:

  • propensity vs. contact energy:前者是某氨基酸出现在某位置的偏好,后者是两残基接触的能量——这是领域常识。
  • no-contact pairs:不接触的位点对,用它们算出的最大相关可看作"负设计"的信号。
  • 负设计:避免错误互作的选择压力,属领域常识。

值得留意:作者明说定性预期被证实,但 \(\Lambda\) 的拟合值精度不足——这是主动承认局限;图里的相关只是"跨所有无接触对取最大值",不是平均值。

b142Similarly, we expect that short proteins experience stronger positive design for stabilizing the native state against unfolding, since the number of contacts per residues increases with chain length due to the surface to volume scaling, and they present fewer native interactions for compensating the loss of conformational entropy upon folding (Bastolla and Demetrius 2005). We quantify positive design for stabilizing the native state against unfolding through the maximum negative correlation between propensities and contact energies across all contact pairs (Fig. 7B). We expect that positive design is stronger for polar proteins, but we find that it is weaker both when assessed through the maximum negative correlation and when assessed through the selection coefficient \(\Lambda\) of short proteins (Fig. 7A), contrary to our expectation.

讲解

1. 这段在干什么

提出一个预期(短蛋白应受更强的正设计),然后用两种指标去检验,结果发现与预期相反,属于「预期—反证」的过渡段。

2. 需要解释的地方

  • 正设计(positive design):领域常识,指序列进化偏向于让天然态更稳定、抵抗去折叠。
  • 表面/体积比:链越长,单位残基的接触数越多,这是作者给出的推理依据。
  • 最大负相关:把倾向性与接触能放在所有接触对上,取负相关最强的值当作正设计强度的量化指标。

3. 值得留意

作者明确说结果「contrary to our expectation」——短蛋白正设计反而更弱。这与上一段「负设计已确认」形成对照,暗示两类选择信号强弱不同。

b143This discrepancy may be due to the fact that many proteins in our data set, in particularly short proteins, tend to occur in protein complexes, and they do not need to be stable on their own. To test this hypothesis, we restricted our data set only to monomeric proteins determined by Xray crystallography, and we tested two thresholds of length, 150 amino acids and 90 amino acids that is often considered the threshold for proteins that fold through two-states folding mechanisms without intermediate states (Barrick 2009). Indeed, for these short lengths the negative design score is very reduced (Fig. 7F orange bars), consistent with the absence of metastable intermediates, and the selection coefficient \(\Lambda\) is larger for short polar proteins than for short hydrophobic proteins (Fig. 7D orange bars), as we expected. However, the positive design score is again larger for short hydrophobic proteins than for short polar proteins (Fig. 7E).

这段在干什么

针对上段"短蛋白结果与预期相反"的疑问,作者提出假说:这可能是因短蛋白多以复合物形式存在、无需独立稳定,并只用单体蛋白重新检验来验证。

需要解释的地方

  • monomeric:单体蛋白,即单独一条链就能存在、不靠"搭伙"成复合物。
  • two-states folding:领域常识,指蛋白折叠中途没有稳定中间态,一步到位。
  • negative / positive design score:衡量序列对"避免错误结构 / 稳定正确结构"的贡献。

值得留意

作者假设"复合物中的短蛋白不需独立稳定",但结论并不完全干净利落:负设计分如预期降低,但正设计分反而在短疏水蛋白里更高,作者也照实写出,没掩饰。

b144Selective parameters \(\Lambda\) fitted for the four different classes of proteins (plots A and D) and maximum correlations between propensities and contact energies for native contacts (B and E) and no contacts (C and F) as a function of different global propensities \(q^\textrm{glob}\) and number of types of sites NC1 and NC2. Plots D,E,F refer to monomeric proteins determined by X ray

这段在干什么

这是图注,说明四类蛋白质拟合出的选择参数 Λ,以及倾向性与接触能之间的最大相关性随 \(q^\textrm{glob}\)、NC1、NC2 变化的情况;D–F 特指 X 射线测定的单体蛋白。

需要解释的地方

  • 选择参数 Λ:衡量选择压力强弱的量,越大表示选择越强(领域常识)。
  • 倾向性 / 接触能:前者是某残基偏好出现在某类位点的程度,后者描述残基间接触的有利程度(领域常识)。
  • NC1 / NC2:原文只提了名字,未定义,这段没说明。
  • 单体蛋白:只有一个亚基的蛋白(领域常识)。

值得留意

提到"最大相关性"和"native contacts / no contacts"两组对比——A–C 与 D–F 应是不同蛋白集合,但这段没说 A–C 指什么。

Global Amino Acid Correlations are Influenced by the Genetic Code and General Selective Pressures

b146By definition, the global propensities are not site-specific. Therefore, they are not influenced by selection for protein folding stability. They mainly reflect variation of the mutational process and of global (not site-specific) selective forces across the lineages where the proteins evolve. Even if here we do not consider multiple sequence alignments (MSA) or phylogenetic trees from which we can derive evolutionary information, we expect that the global propensities are related with the phylogenetic propensities that can be derived from the MSA and do not depend on site-specific structural or functional constraints.

讲解

1. 这段在干什么

过渡段:说明全局倾向(global propensities)为什么与折叠稳定性选择无关,并预告它与系统发育倾向有关。

2. 需要解释的地方

  • global propensities:全局的氨基酸出现倾向,即不区分位点、整个蛋白平均来看某氨基酸多常见。原文说它"不是位点特异的"(not site-specific),所以不受稳定性选择影响,主要反映突变过程和全局选择压力。注意:这是原文的定义。
  • MSA / phylogenetic trees:多序列比对和系统发育树,领域常识中用来推断进化信息的两类数据。原文说这里不采用它们。

3. 值得留意

作者用"Even if…we expect…"作了让步:即便不用 MSA/树,仍预期 global propensities 与可从 MSA 推出的 phylogenetic propensities 相关,且不依赖位点特异的约束——这是推测(expect),不是已证结论。

b147We show in Fig. 8 the ranked values of the global propensities \(q^\textrm{glob}_{a,b}\), focusing on the amino acid pairs that present the most positive (Fig. 8A) and most negative (Fig. 8B) propensities.

这段在干什么

接着上段引出 Fig. 8:把全局倾向 \(q^\textrm{glob}_{a,b}\) 按数值排名,只挑正、负倾向最强的氨基酸对来展示。

需要解释的地方

  • \(q^\textrm{glob}_{a,b}\):一对氨基酸(a、b)的"全局倾向",即整体上二者共同出现的倾向强度;上段已说它来自 MSA,不含位点特异的结构/功能约束。
  • ranked values(排名值):按大小排序后的结果,方便挑两端极值。
  • most positive / most negative:倾向最高(偏好共变)与最低(回避共变)的两类。

值得留意

图只呈现"最极端"的两端,而非全部氨基酸对;正负分开写成 Fig. 8A、8B。这段没交代具体数值或排行结果,那些在图中。

b148The first observation is that pairs of the same amino acids have the strongest positive propensities. This is expected since, if the mutational preference for amino acid a is higher in some lineages and lower in other lineages, this variation will create positive self-correlations of the amino acid frequency. More in general, the mutational preferences between amino acids a and b are similar if they are coded by similar codons, predicting positive global propensities in this case. Conversely, if the codons that code two amino acids overlap less than expected by chance, their mutational preferences will be anticorrelated, possibly leading to negative propensities.

这段在干什么

解释上段图中那些正/负倾向(propensity)从何而来——用突变偏好和密码子相似性说明全局氨基酸相关性的成因。

需要解释的地方

  • propensity(倾向):这里指某氨基酸对在数据中出现的关联强度,正=倾向于同时出现,负=倾向于互斥。
  • 密码子:三个碱基编码一个氨基酸的单元(领域常识)。密码子越相似,两个氨基酸的突变偏好越同步,所以倾向偏正;密码子重叠少于随机预期时,偏好反相关,倾向偏负。

值得留意

  • 同种氨基酸自身(self-correlation)倾向最强,作者归因于谱系间突变偏好的差异,而非直接选择。
  • 这段讲的是全局相关性,是可能干扰真实结构信号的背景因素——后文应会设法排除它。

b149A similar mechanism holds for variations of the selective preference for any given amino acid. These global preferences may include several selective pressures, which act globally at the level of the whole protein and modulate the mutational propensities:

这段在干什么:承接上文关于"两个位点偏好相反会导致负倾向"的说法,把机制从"成对"推广到"整体"——任一氨基酸的选择偏好变动,同样受全局性选择压力调节。

需要解释的地方:

  • 全局偏好 / 全局选择压力:指作用于整个蛋白层面、而非某个具体位点的偏好,它会整体调高或调低某种突变出现的倾向(mutational propensity)。

值得留意:这段是从"两个位点相互影响"过渡到"整条序列层面的全局因素",为后文讨论遗传密码和一般选择压力如何影响相关性做铺垫。真正的内容(具体是哪些压力、如何调节)这段并未展开。

b150The pressure towards disulfide bridges in short proteins that exist in an oxidative environment (Bastolla and Demetrius 2005). Consistently, Cysteine-Cysteine presents the strongest global propensities \(q^\textrm{glob}_{\textrm{CC}}=0.312\) (Fig. 8).

这段在干什么

紧接上句「全局选择压力调节突变倾向」,举二硫键为例:氧化环境中的短蛋白倾向于形成二硫键,而数据里 Cys-Cys 的全局倾向最高(0.312)。

需要解释的地方

  • 二硫键:两个半胱氨酸(Cys)之间形成的化学键,能加固蛋白质结构——这是领域常识。
  • 全局倾向 \(q^\textrm{glob}_{\textrm{CC}}\):衡量 Cys 和 Cys 在蛋白中一起出现的整体倾向有多强,数值越高越常见。

值得留意

作者用「Consistently」把机制预期(该成键)和统计结果(最高值)对上,这是旁证式的呼应,属相关而非因果证明;文献引用 (Bastolla and Demetrius 2005) 支持的是前半句机制说法,不是 0.312 这个数字。

b151The pressure to avoid biosynthetically costly amino acids, which is particularly relevant in bacterial community, which present hints of cross-feeding where only some community members synthesize some specific amino acids and leak them to the environment (Mustonen and Lässig 2010; Puente-Sánchez et al. 2024), and is lower in organisms that are less limited by resources. Consistently, tryptophane, considered the most biosynthetically costly amino acid, present the second-largest value of the global propensity \(q^\textrm{glob}_{\textrm{WW}}=0.152\), tyrosine presents the 7th highest propensity \(q^\textrm{glob}_{\textrm{YY}}=0.064\) and \(q^\textrm{glob}_{\textrm{WY}}=0.042\). Phenylalanine belongs to the same cluster: \(q^\textrm{glob}_{\textrm{FF}}=0.038\), \(q^\textrm{glob}_{\textrm{FW}}=0.018\) and \(q^\textrm{glob}_{\textrm{FY}}=0.027\).

承接上句「半胱氨酸-半胱氨酸 propensity 最高」,这段在论证全局氨基酸相关性还受合成成本和遗传密码影响,为后面解释共变来源做铺垫。

需解释

  • *biosynthetically costly*:合成耗能高的氨基酸,细菌常靠别的成员"漏"出来共享(领域常识)。
  • *propensity \(q^\textrm{glob}\)*:某氨基酸对在全局统计中被观察到的倾向强度,越大越常见。

值得留意

作者把高成本氨基酸(W、Y、F,同属芳香簇)与高 propensity 挂钩,但 W 只说"第二大",没直接说因果。

b152The pressure to avoid codons that frequently mutate to stop codons and can interrupt the translated protein, which mainly applies to the same amino acids discussed above (Tryptophane, Tyrosine and Cysteine), whose codons are close to stop codons.

这段紧接着上一段的 \(q^\textrm{glob}\) 数值,转而提出一个选择压力来源:密码子若靠近终止密码子,易突变成终止子,中断翻译。

需要解释的地方:

  • 终止密码子(领域常识):翻译的"句号",密码子突变成它,蛋白质合成会提前截断。
  • 括号里点名的三种氨基酸(色氨酸、酪氨酸、半胱氨酸),它们的密码子离终止密码子近,所以受这种压力影响最大。

值得留意:

作者用"mainly applies to the same amino acids discussed above"把它和上文的三种氨基酸绑定,说明这可能是解释它们协变的一个候选机制——但这段本身还没说它如何影响协变,别读成结论。

b153The pressure favoring positively charged amino acids in proteins that interact with nucleic acids that are negatively charged. Consistently, Lysine and Arginine have the third and the sixth-largest global propensity, \(q^\textrm{glob}_{\textrm{KK}}=0.100\) and \(q^\textrm{glob}_{\textrm{RR}}=0.064\). However, \(\textrm{RK}\) is disfavored, \(q^\textrm{glob}_{\textrm{RK}}=-0.058\), possibly because there are specific selective pressures that differentiate Arg and Lys besides their common charge.

这段在干什么

承接上段"密码子邻近终止子"的偏好话题,换角度说明:氨基酸的全局关联还同时受电荷等一般性选择压力影响。

需要解释的地方

  • \(q^\textrm{glob}\):全局倾向/关联强度的一个指标,正数表示偏好、负数表示排斥。(符号含义按原文数值合理推断)
  • Lys、Arg:赖氨酸、精氨酸,都是带正电的氨基酸(领域常识)。

值得留意

  • 原文只是假设("possibly")Arg/Lys 分化源于电荷之外的特定压力,并未给出证据。
  • 三个数字(0.100 / 0.064 / −0.058)是该指标数值:前两者偏正、RK 为负,说明"同为正电"并不等于彼此被偏好。

b154The pressure towards hydrophobic amino acids in membrane proteins. However, pairs of different hydrophobic residues are not favored, except Methionine and Isoleucine (\(q^\textrm{glob}_{\textrm{MI}}=0.015\)) that are coded by very similar codons.

逐段带读

1. 这段在干什么

承接上文对氨基酸配对偏好的讨论,指出膜蛋白里虽偏好疏水残基,但不同疏水残基之间并不互相偏好——是个例外说明。

2. 需要解释的地方

  • 疏水氨基酸:怕水、爱待在膜内部的氨基酸(领域常识)。
  • \(q^\textrm{glob}\):前文引入的全局配对相关性指标,值越大表示该氨基酸对越常一起出现。
  • 密码子相似的 MI:Met 和 Ile 的遗传密码很接近,可能因此共变。

3. 值得留意

作者用"码子相似"来解释 MI 的偏号,暗示这种相关性未必来自选择压力,而可能只是密码子结构的副产品——这正是本小节标题所谓"遗传密码影响相关性"的一个例证。

b155Interestingly, the four polar amino acids Serine, Threonine, Asparagine and Glutamine form a cluster, since they are positively correlated with each other (the highest value is \(q^\textrm{glob}_{\textrm{ST}}=0.030\) and the lowest is \(q^\textrm{glob}_{\textrm{TQ}}=0.003\)). These correlations may also be caused by the similarity of the codons that code for them.

这段紧接上句,继续补充「全局氨基酸相关性」的成因。

这段在干什么:指出 Ser、Thr、Asn、Gln 四种极性氨基酸彼此正相关,形成一个簇,并推测原因可能是密码子相似。

需要解释的地方:

  • \(q^\textrm{glob}\):衡量两个位点氨基酸全局相关强度的量,越大越相关(领域常识:\(q\) 类指标常这样解读)。
  • 密码子:mRNA 上三个核苷酸决定一个氨基酸(领域常识)。

值得留意:这段只是"may also be caused"的推测,并未给出证据;最低值 TQ=0.003 说明相关性其实很弱,别误读成强关联。

b156Global propensities \(q^\textrm{glob}_{ab }\). Plot A shows the 25 most positive propensities and plot B shows the 25 most negative ones

这段在干什么:这是图注的开头。上一段末尾提到相关系数可能来自密码子相似性,本段引入「全局倾向 \(q^\textrm{glob}_{ab}\)」,并说明两张子图分别展示最正/最负的 25 个倾向。

需要解释的地方:\(q^\textrm{glob}_{ab}\) 是某对氨基酸 \(a,b\) 在全局范围内一起出现的倾向(领域常识:下标 \(a,b\) 通常指两个氨基酸,glob 指全局统计,而非某蛋白内部)。

值得留意:原文只说 A 图给最正的 25 个、B 图给最负的 25 个,没有给具体数值或结论;上一段末句的因果猜测,这里也没正面回答。

Discussion

b158In this paper, we introduce the stability constrained model of protein evolution (SCPE) in the pairwise approximation, which generalizes our previous SCPE model that assumes that the protein sites are independent (Arenas et al. 2015). We model the site-specific amino acid distribution of a protein family based on minimal selection for folding stability, requiring that the probability distribution of the amino acids has minimal Kullback-Leibler (KL) divergence from the global (not site-specific) amino acid distribution, while constraining the protein folding stability \(\Delta G\) with respect to both the unfolded state and compact but uncorrectly folded misfolded conformations.

好的,我们来看这一段。它出现在讨论部分,是作者对自己工作的定位与概括。

1. 这段在干什么

交代本文提出了什么模型:把此前"位点相互独立"的 SCPE 推广到成对近似版本。

2. 需要解释的地方

  • SCPE:用"选择压力约束折叠稳定性"来建模蛋白序列演化的框架。
  • 成对近似:不再假设各位置独立,而是考虑两两位置间的关联。
  • KL 散度:衡量两个概率分布差多远,越小说明越接近。
  • ΔG:折叠自由能,越负越稳定,这里是说模型把稳定性作为约束条件。(这是领域常识)

3. 值得留意

模型不只约束"折叠态 vs 去折叠态",还额外约束了错误折叠构象——这点容易被一扫而过。

b159Minimum KL divergence coincides with the more familiar maximum entropy principle when amino acids are equiprobable, since maximum entropy is the same as minimum KL divergence from the uniform distribution, but it is more suitable for describing real situations in which the amino acids are not equally distributed. The main assumption of the model is that protein fitness can be approximated with the stability of the native state.

这一段先替模型“选尺子”,再交代核心假设。

1. 这段在干什么:说明为什么用最小 KL 散度来描述氨基酸分布,并点出模型的中心假设。

2. 需要解释的地方:

  • KL 散度:衡量两个分布的差异;散度越小,两个分布越像。
  • 最大熵:分布最“均匀、无偏”的状态(领域常识)。
  • 作者说:当各氨基酸等概率时,最小 KL 散度就等于最大熵原则;但真实中氨基酸并不等概率,所以前者更合用。

3. 值得留意:第二句“蛋白质适应度可用天然态稳定性近似”是全文的关键假设,不是已证结论。

b160Importantly, we do not fit stability constraints from the data, but we predict them from the protein sequence and the experimentally known protein structure with a simple contact model that also takes into account the stability of the misfolded state modelled with the Random Energy Model (Minning et al. 2013). In this way, we predict the parameters of a stability constrained Potts model that describes the amino acid distribution of the protein family in the pairwise approximation (Lapedes et al. 2002; Weigt et al. 2009; Marks et al. 2012; Morcos et al. 2011) only based on global amino acid propensities and site-specific stability constraints. Within the REM, the free energy of the misfolded ensemble depends on groups of four sites, while we approximate it here as a sum of pairwise terms. This approximation may explain in part the reduced accuracy of our predictions for short range pairs, which are dominated by selection on misfolding.

这段在干什么

承接上段「fitness≈native稳定性」的假设,说明预测 Potts 模型参数的来源:不拟合数据,而是从序列+实验结构出发预测。

需要解释的地方

  • 不 fit、而是 predict:约束不是从家族数据反推,而是由结构算出来——这是本文卖点。
  • contact model:用残基接触近似稳定性能量,领域常识。
  • Random Energy Model (REM):刻画错误折叠态稳定性的模型,即不只考虑正确折叠。
  • Potts model / pairwise:用成对耦合描述家族氨基酸分布的统计模型,领域常识。

值得留意

  • 作者只用了全局氨基酸倾向 + 位点特异性稳定性约束,没别的信息。
  • 把 REM 中「四残基一组」的错误折叠自由能近似成成对项之和——这是主动承认的简化,并据此解释短程残基对预测偏弱。

Selective Forces Considered by the SCPE Model

b162We adopt here the small-couplings approximation, i.e. we consider only first order terms in the Potts couplings, which is related with the mutual information approximation originally adopted for inferring the statistical couplings. We obtain an explicit formula that predicts the couplings between amino acid \(a_i\) at site i and amino acid \(a_j\) at site j, and depends on:

这段在干什么

引入小耦合近似,为下一句「预测任意两个位点上氨基酸间耦合的显式公式」做铺垫。

需要解释的地方

  • Potts 耦合:刻画两个位点氨基酸偏好是否相互影响的参数(领域常识)。
  • 一阶项:耦合太弱时,只保留最简单的线性贡献,高阶影响忽略。
  • 互信息近似:只从两两共变统计推断耦合,不看更高阶关联,与这里的思路一脉相承。
  • 显式公式:指能得到可写的解析表达式——但具体形式这段没给出。

值得留意

原文最后停在「depends on:」,冒号后内容被截断,公式依赖什么并未出现在本段。

b163The global (non-site-specific) couplings \(q^\textrm{glob}_{a_i,a_j}\) between the amino acids \(a_i\) and \(a_j\). In our interpretation, global couplings are the result of variations across organismic lineages of the mutation process and the selective forces that favor or disfavor \(a_i\) and \(a_j\). The global couplings are generally small, and they are strongest when \(a_i=a_j\), since in this case the mutational and selective forces that favor them are perfectly correlated. In particular, we find the strongest self-coupling for Cysteine, which is subject to global selective force for the formation of disulfide bridges. The second one is for Triptophane, which is subject to global selective force for sparing this metabolically costly amino acid. The same applies to Tyrosine and, to a limited extent, Phenilanine, and the three form a cluster of positively correlated amino acids. Another cluster is formed by the polar amino acids Serine, Threonine, Glutamine and Asparagine.

这段在干什么

这一段在介绍 SCPE 模型公式里的第一个输入项——氨基酸之间的全局耦合 \(q^\textrm{glob}\),也就是上一段公式冒号后面列出的第一类参数。

需要解释的地方

  • 全局耦合:不看具体位点,只看两个氨基酸种类之间的统计关联强度,是跨物种谱系平均下来的结果。
  • 自耦合(\(a_i=a_j\)):同种氨基酸配对,原文说这时最强,因为"偏好它的突变压力和选择压力完全相关"。

值得留意

  • 作者给每个强弱排序都配了生物学理由:半胱氨酸(二硫键)、色氨酸(代谢昂贵、节约使用),后者也适用于酪氨酸、部分适用于苯丙氨酸,三者成一簇;另一簇是极性氨基酸 Ser/Thr/Gln/Asn。
  • 注意作者说的是"我们的解释(in our interpretation)",即这些机制解读是模型假设,不是直接观测。

b164Site-specific selective forces due to positive design that stabilize the native state with respect to the unfolded state. They mainly act at pairs of sites that form native contacts. The predicted amino acid propensities of native contacts are approximately equal to minus the interaction free energy of the two amino acids, i.e. amino acid pairs that stabilize the native state are predicted to be more frequently observed at these pairs of sites, with contrasting influence from negative design on short range contacts.

讲解

1. 这段在干什么

介绍 SCPE 模型里的一类选择压力——来自 positive design、稳定天然态的位点特异性作用力。

2. 需要解释的地方

  • positive design(正向设计):领域常识,指进化选择那些让蛋白质折叠态更稳定的氨基酸,与"negative design"(避免错误聚集、错配等)相对。
  • 天然接触(native contacts):在天然结构里空间上相互靠近、发生作用的残基对。
  • 预测倾向 ≈ 负的相互作用自由能:相互作用越有利(能量越负),这对氨基酸就越常被观察到。

3. 值得留意

末尾一句提到 short range contacts 还受 negative design 的相反影响——这句话是补充,别漏读,它说明近距离接触的规律比天然接触更复杂。

b165Site-specific selective forces due to negative design that destabilize misfolded states. They mainly act at pairs of sites that are at short range along the sequence and do not form native contacts, and they predict amino acid propensities that are approximately equal to the interaction free energy of the two amino acids, i.e. amino acid pairs that destabilize misfolded conformations are more frequently observed at these pairs of sites.

这段在干什么

承接上文,具体说明 SCPE 模型里“负设计”这种选择压力:它怎么作用、作用在哪些位点对、对氨基酸偏好做出什么预测。

需要解释的地方

  • 负设计:不是稳住正确折叠,而是让错误折叠的状态变得不稳定(领域常识)。
  • 短程 / 非天然接触位点对:序列上离得近、但在天然结构里并不接触的两个位置。
  • 氨基酸偏好≈相互作用自由能:作者预测这类位点上出现的氨基酸组合,其偏好强度大致等于两氨基酸的相互作用自由能。

值得留意

负设计专门盯“短程、非天然接触”的位点对,而非天然接触对——这是它区别于正设计的关键。

b166We test these predictions on a representative subset of 37108 non-redundant protein chains from the PDB, grouping together 45 classes of pairs of sites with similar properties (native contacts versus no contacts and indirect contacts, 5 different distances along the sequence, core or surface sites). We also consider four types of proteins (long or short, hydrophobic or polar and their combinations).

这段在干什么:交代验证上一段预测所用的数据规模和分组方式——把 PDB 中 37108 条非冗余蛋白链按位点对性质分成 45 类,并另按蛋白类型分四类。

需要解释的地方:

  • non-redundant(非冗余):领域常识,指去掉了序列高度相似的蛋白,避免同一家族重复计数。
  • native contacts vs no contacts and indirect contacts:位点对在三维结构上直接接触、不接触、间接接触,属于分组维度之一。
  • 5 different distances along the sequence:沿序列的 5 种间距,即把位点对按序列距离分档。

值得留意:45 = 3(接触类型)× 5(序列距离)× 3(core/surface 之类),三种属性交叉分组;而"long or short, hydrophobic or polar 及其组合"是另一套独立的四类划分,两套分组别混。

Test of the SCPE Model Within the SCA Approximation

b168We examine the observed propensities of each class of pairs of sites p, \(Q_{p}(a,b)=P_{p}(a,b)/\left( P_{i}(a)P_{j}(b)\right) -1\), which in the small coupling approximation is proportional to the corresponding coupling.

这段在干什么

提出用观测倾向 \(Q_p(a,b)\) 来量化每类位点对,并说明它在弱耦合近似下正比于耦合强度,为后文检验 SCPE 模型提供可观测的代理量。

需要解释的地方

  • \(Q_p(a,b)\):位点对 \(p\) 上氨基酸 \(a,b\) 的观测频率,相对于两者独立出现时期望频率的偏离,再减 1。大于 0 表示共现比随机多,小于 0 表示少。
  • 小耦合近似:弱耦合时,这个偏离大致正比于真实耦合强度——所以能拿它当耦合的"影子"来用。

值得留意

这里说的是"正比",不是"等于",中间的比例系数未给出,因此 \(Q_p\) 只能做相对比较,不能直接读出耦合绝对值。

b169The qualitative dependence of the observed couplings on the contact free energy is for each type of pair in line with what our model predicts (see Fig. 4), and it is modulated by native contacts, the buriedness of the sites, their distance along the sequences, and the length of the proteins as the SCPE model predicts. In particular, the correlations between the couplings and the contact energy are negative for pairs that form native contacts, as expected based on positive design, and they are less negative for short range pairs, as expected based on negative design. They are more negative for core-core sites, as expected based on their importance for protein stability, and they are more negative for longer proteins. For pairs that do not form native contacts, the correlation is positive for short range pairs, as expected based on negative design, and they are stronger for longer proteins, which are expected to be subject to stronger selection against misfolding. This also holds for pairs that form indirect contacts through another residue. On the other hand, the couplings of long range pairs that do not form native contacts are essentially determined by the global couplings, with which they correlate very strongly.

这段在干什么

检验上文模型预测的耦合强度与接触自由能的关系是否在真实数据中成立,逐类配对逐一对照。

需要解释的地方

  • coupling(耦合):两个位点氨基酸偏好是否联动,前面已由小耦合近似正比于某个量。
  • positive/negative design:领域常识,前者指天然序列偏好形成天然接触,后者指避免非天然接触(防错折叠)。
  • buriedness:位点埋藏程度,越在核心越重要。
  • global couplings:全局耦合,长程非接触对主要由它决定。

值得留意

作者不说“完全吻合”,只说“in line with / as expected”——是定性趋势一致,不是定量证明。

b170The agreement with our predictions is reasonably good, with correlation coefficients r larger than 0.5 for most pairs of sites (see Fig. 3, where the dashed line represents \(r=0.5\)).

讲解

1. 这段在干什么

给出 SCPE 模型预测与实际数据吻合程度的定量结论:大部分位点对的相关系数 r > 0.5,算「相当好」。

2. 需要解释的地方

  • 相关系数 r:衡量两组数值同向变化程度的指标,取 −1 到 1,越大越正相关。这里 r 是预测值与观测协变之间的相关。
  • dashed line(虚线):图 3 里画 \(r=0.5\) 作为参照线,方便一眼看出哪些点在上面、哪些在下面。这是领域常识里常见的画法。

3. 值得留意

  • 说的是「most pairs」和「reasonably good」,不是全部、也不完美——作者用词留了余地。
  • r = 0.5 只是文章自设的宽松门槛,本身并非统计显著性的标准。

(约 150 字)

b171However, one can see that there are some regimes in which our predictions do not perform very well. In particular, predictions are poor for pairs that form indirect contacts (bottom line in Fig. 3), which are neglected in the small coupling approximation and can be computed only a posteriori. Predictions are also poor for short polar proteins (second column in Fig. 3). This is in part due to the fact that most of these chains (59%) appear in multimeric proteins, so that they may be unstable as monomers. Also, 49% of these chains have been determined by NMR and may be different from average crystal structures that dominate for long proteins. Finally, predictions are poor for short range contacts. There are two likely explanations of this fact: one is the breakdown of the small couplings approximation for short range pairs that do not form native contacts, whose couplings are large (see Fig. 6), and the other one is the approximation of the free energy of the misfolded ensemble predicted through the REM as a sum of pairwise terms, Eq.(), which is necessary for applying the pairwise model but may be not sufficiently accurate.

讲解

1. 这段在干什么

紧接上一段的乐观结论(多数位点对 \(r>0.5\)),这段转折,列出预测效果差的几类情形,并逐一交代可能原因。

2. 需要解释的地方

  • 间接接触:两个位点不直接挨着,通过第三个位点发生关联。
  • 小耦合近似:把弱相互作用当成可忽略来简化计算;间接接触恰在此近似中被丢掉,只能事后补算。
  • 短极性蛋白:链短、带极性残基多的蛋白。
  • 多聚体 vs 单体:多聚体是几条链组装在一起;单拎一条出来可能不稳定。
  • NMR 结构:溶液中的结构,与晶体结构可能有差异(领域常识)。
  • REM / 成对项求和:把错折叠态自由能拆成两两相加来估算,是成对模型成立的前提。

3. 值得留意

作者只是罗列"可能原因"(likely explanations, may be not accurate),并未给出确证;数据(59%、49%)只针对短极性蛋白这一类。

Unexpected Observations

b173These results suggest that a big part of the couplings in protein evolution originate from the variation of global mutational and selective forces, which act on most pairs of sites, plus couplings due to positive selection for enhancing the stability of the native state, which mainly act at pairs of sites that form native contacts, plus couplings due to negative design for destabilizing pairs that are close along the sequence and do not form native contacts.

这段在干什么:论文观察到蛋白质演化中的耦合(couplings)不只一种来源,这里把它们归纳成三部分叠加。

需要解释的地方:

  • 耦合(coupling):两个位点协同变化、彼此不独立的现象(领域常识)。
  • 全局突变与选择压力:作用于大多数位点对的整体背景力量。
  • 正选择稳定天然态:偏好维持稳定的选择,主要作用于形成天然接触的位点对。
  • 负设计:避免序列上邻近、却不接触的位点对互相破坏(领域常识)。

值得留意:作者说耦合是「三部分之和」,正好呼应上一段「pairwise 之和」的说法——但这里说的是来源,不是模型本身。

b174Perhaps unexpectedly, a larger number of couplings are influenced by negative design against misfolding (positive correlations between observed propensities and contact energies) than those influenced by positive design for stabilizing the native state against unfolding.

这段在干什么

抛出一个"意料之外"的观察结论:受负设计影响的耦合,比受正设计影响的更多。

需要解释的地方

  • 正设计(positive design):为稳定天然态、对抗去折叠而做的选择——领域常识,即"选稳定的"。
  • 负设计(negative design):为对抗错误折叠而做的选择,针对那些序列上相近却不成天然接触的配对——即"防选错的"。
  • 耦合(couplings):两个位点间协同变化的关系。
  • 倾向 vs 接触能的正相关:括号里指观测到的倾向与接触能同向变化。

值得留意

作者用 "Perhaps unexpectedly" 点明这结果与直觉相悖——通常以为稳定天然态才是主导。且这是观察量多寡的比较,不代表负设计作用更强。

b175Another interesting and perhaps unexpected result is that pairs of surface-surface sites in contact possess higher mutual information than pairs of core-core suites in contact (compare circles in Fig. 6A and C) despite the fact that the latters contribute more to protein folding stability. This suggests that surface-surface contacts are subject to other selective pressures besides protein folding stability. Strong candidates are functional movements and interface properties.

逐段带读

1. 这段在干什么

提出一个反直觉的观察:表面-表面接触位点对的互信息,反而比核心-核心接触位点对更高。

2. 需要解释的地方

  • 互信息(mutual information):领域常识,用来量化两个位点是否"协同变化"——一个变,另一个也跟着变,值就高。
  • 表面-表面 / 核心-核心:按位点在蛋白里的位置分。核心埋在里面,表面朝外。
  • 接触:两个位点在三维结构上挨着。

3. 值得留意

  • 作者自己点破反差:核心位点对其实对折叠稳定性贡献更大,按理该受更强选择、互信息更高,结果却相反。
  • 所以作者给的解释是:表面接触还受别的选择压力(功能运动、界面性质)影响,不只是折叠稳定性。这正是"Unexpected Observations"想引出的问题。

(原文只给"强候选"是功能运动和界面性质,没说已证实。)