主题
数据验收
F0 · 可选步骤 · 图 5 张 · 发表版表格 4 个 (流程见 v5 定稿)
流程 v5 的可选步骤,不参与候选筛选。检查输入数据本身是否可用,并用 LDSC 描述四个结局之间的遗传相关性。
图
图 F0a 六个结局之间的遗传相关(LDSC)

结论. 四种糖尿病并发症在遗传层面高度重叠,宜作为一组强相关的结局共同解读;而 1 型与 2 型糖尿病之间几乎不存在共享的遗传基础,本研究全程将二者视为彼此独立的糖尿病参照性状,不作合并。
结果. LD 得分回归显示,四种并发症两两之间的遗传相关普遍很强,以视网膜病变与肾病之间最高(rg = 0.97 ± 0.09),其次依次为视网膜病变与黄斑病变(0.95 ± 0.10)与黄斑病变与肾病(0.94 ± 0.10);神经病变与其余三者的相关中等偏高(0.60–0.75)。四种并发症与 2 型糖尿病的遗传相关(0.59–0.74)普遍高于与 1 型糖尿病(0.36–0.45)。所检验的 15 对性状中有 14 对在 Bonferroni 校正后仍然显著,唯一的例外是 1 型糖尿病与 2 型糖尿病(rg = 0.09 ± 0.03,校正后 P = 0.16)。
完整图注(英文,与投稿版一致)
Figure F0a. Genetic correlation (rg) between diabetic complication outcomes and the two diabetes control arms, estimated by LD score regression.
Cells show the rg point estimate. An asterisk marks the 14 of the 15 trait pairs tested that pass Bonferroni correction (Bonferroni-adjusted P < 0.05); the single pair without an asterisk is named below. Only one triangle is shown; the matrix is symmetric. Diagonal cells are rg of a trait with itself and are therefore 1.00 by definition; they are not tests and carry no asterisk. Cell fill is the same rg on a blue-white-red scale fixed to -1..1, so colour and printed value carry the same quantity; the printed value switches from black to white where |rg| > 0.6 purely for contrast against the darker fill, and does not denote a different category. 14 of the 15 pairs pass Bonferroni correction. Type 1 diabetes versus Type 2 diabetes reaches nominal but not corrected significance (rg = 0.09, unadjusted P = 0.011, adjusted P = 0.16); the two diabetes arms are therefore treated as separate control arms throughout and are never pooled. Standard errors, z statistics and P values for every pair are given in the supplementary table ldsc_genetic_correlation_pairwise.csv; they are omitted from the cells here to keep the 6 x 6 panel legible at print size. SNP-based heritability of each outcome is reported separately (Table: ldsc_heritability.csv) and is not shown in this matrix. Data: FinnGen R9 endpoints (diabetic retinopathy, maculopathy, nephropathy, neuropathy); type 1 diabetes GWAS Catalog GCST90824163; type 2 diabetes DIAGRAM Mahajan et al. 2018 (European, no UKBB). Sample sizes and accessions are listed in Table 1 (data_sources.csv). LD scores: European reference. Method: LD score regression (LDSC) on munged summary statistics; effective sample size Neff = 4/(1/ncase + 1/ncontrol) was supplied for each binary trait.
Role in the analysis pipeline: data acceptance is an OPTIONAL, non-standard step in this study's workflow. It characterises the outcomes; it does NOT filter, rank, or veto any protein candidate. No result downstream depends on these estimates.
Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
图 F0b 遗传相关网络,按数据来源

结论. 上一图中并发症之间的高遗传相关,有相当一部分来自样本重叠,因此不宜直接解读为共同的生物学机制。
结果. LDSC 在估计两个性状的遗传协方差时会同时给出一个交叉性状截距,它反映两份 GWAS 之间的样本重叠:当两项研究不含共同个体时其期望为 0。本研究中,四种并发症两两之间的交叉性状截距为 0.24–0.63,而它们与两个糖尿病参照性状之间仅为 0.017–0.069,相差约一个数量级。这与数据结构一致 —— 四个并发症端点同出 FinnGen R9 的同一批参与者,而两项糖尿病 GWAS 为外部独立研究。图中结点面积反映各结局的有效样本量,以 2 型糖尿病最大(约 193,400),神经病变最小(约 11,300)。
完整图注(英文,与投稿版一致)
Figure F0b. Genetic correlations between the 6 study outcomes as a network.
Panel. Every one of the 15 pairs is drawn. With 6 nodes the graph is complete, so its topology carries no information and the layout is fixed on a circle rather than produced by a force-directed algorithm, which for a complete graph would give an arbitrary and non-reproducible arrangement. Edge width and colour are the genetic correlation; the exact values and their significance are given in the matrix figure (F0a) and are deliberately NOT repeated here. What this figure adds is the structure: node colour is the data source and node area increases with the effective sample size of that outcome but is NOT proportional to it: the scale is compressed so that the smallest node stays legible, so node sizes should be read against the legend keys rather than compared as ratios.
★ What the node colours show. The four complications come from a single FinnGen release and share their control set; the two diabetes arms are external studies. The tightly interconnected block is therefore partly a property of the data sources, not only of the biology, and is quantified in the companion figure.
No significance threshold applies to this figure: every pair is drawn whatever its P value, so no edge is selected in or out. For reference, 14 of the 15 pairs pass Bonferroni correction; the exception is marked in the matrix figure (F0a) and is the thinnest, palest edge here.
Method: LD score regression on munged summary statistics, European reference LD scores. Correlations and their standard errors are read from the analysis output table; nothing is recomputed here.
Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
图 F0c 遗传相关与配对有效样本量

结论. 在本研究的数据结构下,数据来源与样本量彼此混杂、无法分离,因此既不能用「样本量不足」解释掉这些遗传相关,也不能用这张图去证明它们不是假象;恰当的做法是在正文中如实交代共享对照与相互嵌套的表型定义。
结果. 遗传相关的估计精度依赖样本量,样本量偏小的性状对其估计更易出现偏差。因此文献中常把遗传相关对配对有效样本量作回归,用以说明所报告的相关并非功效不足所致的假象。有效样本量定义为 Neff = 4 /(1/病例数 + 1/对照数),在病例与对照数量悬殊时会远小于总人数;每一对性状取两者的几何平均。
本研究中该回归是显著的(R²adj = 0.27,P = 0.027,共 15 对),且斜率为负 —— 照字面读,正是上述假象的表现。但按数据来源分组后可以看到,这一关系被来源本身完全解释:两端都来自 FinnGen 的 6 对,遗传相关中位 0.85、有效样本量中位 18,217;一端来自 FinnGen 的 8 对为 0.52 / 49,465;两端都不是的 1 对为 0.09 / 123,927。同一生物库内部的端点共享对照、表型定义相互嵌套,遗传相关天然偏高;而它们又都是病例数较少的端点,有效样本量天然偏小,外部糖尿病 GWAS 则恰好相反。三组在两条坐标轴上完全同序,来源与样本量在设计上共线,回归系数无法归因到其中任何一方,故图面上不报斜率。
完整图注(英文,与投稿版一致)
Figure F0c. Genetic correlation of each of the 15 outcome pairs against the effective sample size of the pair, coloured by whether the two outcomes come from the same cohort.
Why this is not the usual sample-size check. A regression of correlation on effective sample size is often reported to show that correlations are not an artefact of power. Here that regression is statistically significant and negative (adjusted R-squared 0.270, P = 0.027, 15 pairs), which at face value would suggest exactly such an artefact. It does not. The three groups are ordered identically on both axes: Both from FinnGen, median rg 0.85 at median effective sample size 18,217; One from FinnGen, median rg 0.52 at median effective sample size 49,465; Neither from FinnGen, median rg 0.09 at median effective sample size 123,927. Pairs drawn from the same FinnGen release share around 308,543 controls and have partly nested phenotype definitions, which raises their correlation, and they are also the smaller endpoints, which lowers their effective sample size. Source and sample size are therefore collinear by design and cannot be separated in these data.
No slope is printed on the figure. With 15 pairs, one of the three groups containing a single pair, and the two explanatory quantities collinear, no coefficient from this model is stable enough to report as a result. The figure is shown to make the confounding visible, not to estimate it.
Consequence for the study. None of the downstream results depends on these estimates: LD score regression is an optional acceptance step in this workflow and is not used to filter, rank or veto any protein candidate. What it does affect is wording: the high correlations among the four complications must be described together with their shared controls and nested definitions, and not as a purely biological result.
No significance threshold applies to this figure: all pairs are plotted whatever their P value, and the dashed line is shown only to make the naive analysis visible. The correlation P values themselves are reported in the matrix figure.
Method: effective sample size Neff = 4/(1/ncase + 1/ncontrol) for each outcome, computed from the outcome manifest; the pair value is the geometric mean of the two. Genetic correlations from LD score regression as in the companion figures.
Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
图 F0d 各结局的 SNP 遗传度(观测量表)

结论. 六项 GWAS 的 LDSC 质控均可接受,足以支撑上述遗传相关分析;但观测量表下的遗传度本身受各数据集病例比例影响,⛔ 不可在结局之间直接比较大小。
结果. SNP 遗传度指全部常见变异合起来所能解释的性状变异比例。对二分类疾病,它可以报在观测量表或易感性量表上:前者取决于样本中病例所占的比例,后者需给定人群患病率方能换算。本研究未指定各 FinnGen 端点的人群患病率,故只报观测量表。六个性状的观测量表遗传度为 0.12–0.31;遗传度 Z 值(点估计与其标准误之比)为 5.8–17.7,均超过通常要求的 4,说明各性状的遗传信号足以支撑遗传相关的估计。
LDSC 的做法是把每个位点的检验统计量 χ² 对其连锁不平衡得分作回归:斜率反映多基因遗传度,截距则估计与多基因性无关的那部分膨胀(人群分层、隐性亲缘等混杂),无混杂时期望为 1。本研究六个性状的截距为 0.98–1.08,接近 1。LDSC 另给出 ratio =(截距−1)/(平均 χ²−1),表示统计量膨胀中可归因于混杂的比例,越接近 0 越好。⚠️ 本研究有 5 个性状的 ratio 为 0.22–0.36,只有 2 型糖尿病为 -0.04,表面上并不接近 0。但这几个性状的平均 χ² 仅 1.08–1.34,分母(平均 χ²−1)只有 0.08–0.34,截距上很小的偏离即被放大,其标准误也达到 0.09;功效最高的 2 型糖尿病(平均 χ² = 1.44)ratio 即为 -0.04。因此此处的 ratio 主要反映功效不足,不宜据以判定存在两三成的混杂。
完整图注(英文,与投稿版一致)
Figure F0d. SNP-based heritability of each outcome on the observed scale, estimated by LD score regression. Bars are ordered by point estimate; whiskers are the 95% confidence interval as reported by LDSC.
Estimates are on the OBSERVED scale. Conversion to the liability scale was not performed because population prevalences for these FinnGen endpoints were not specified in this study; LDSC documentation notes that conversion to the liability scale affects only the heritability estimate and not the regression intercept. Observed-scale heritability depends on the case ascertainment of each dataset, so values should not be compared directly between outcomes with very different case-control ratios. Genetic correlation is unaffected by this: LDSC documentation states that there is no notion of observed or liability scale genetic correlation, so Figure F0a (genetic correlation) requires no scale conversion.
Data: FinnGen R9 endpoints (diabetic retinopathy, maculopathy, nephropathy, neuropathy); type 1 diabetes GWAS Catalog GCST90824163; type 2 diabetes DIAGRAM Mahajan et al. 2018 (European, no UKBB). Case and control counts are listed in Table 1 (data_sources.csv). Effective sample size Neff = 4/(1/ncase + 1/ncontrol) was supplied to LDSC for each binary trait.
No significance threshold applies to this figure: heritability is reported as a point estimate with its 95% confidence interval, not tested against a cut-off. Role in the analysis pipeline: data acceptance is an OPTIONAL, non-standard step. It does not filter, rank, or veto any protein candidate.
Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
Figure F0e. Input acceptance checks on the outcome GWAS files

完整图注(英文,与投稿版一致)
Figure F0e. Input acceptance checks on the outcome GWAS files.
(a) Direction of the reported effect-allele frequency (EAF), adjudicated against an external reference (1000 Genomes EUR). For each outcome, non-palindromic variants with unambiguous allele identity were compared; the bar shows the correlation between the frequency reported in the source file and the 1000G EUR frequency of the same allele. Type 1 diabetes (GCST90824163) is inverted in the raw source file (r = -0.99, 93.7% of comparable variants flipped). This orientation is declared in the study manifest and the pipeline applies 1 - EAF when reading the file, so the analysis uses the corrected values. The check is shown on the RAW file deliberately: it is the check that detected the inversion. Type 2 diabetes (Mahajan et al. 2018) cannot be adjudicated by this method and is therefore absent from panel (a): its variants are identified as chromosome:position rather than rsID, so they cannot be matched to the rsID-keyed 1000G reference. This is 'not assessable', NOT 'tested and found correct'.
(b) Effective sample size per outcome, Neff = 4/(1/ncase + 1/ncontrol), the quantity supplied to LD score regression for binary traits. Case and control counts are listed in Table 1 (data_sources.csv).
No significance threshold applies to this figure: panel (a) reports a correlation coefficient used to detect an inverted allele-frequency column, and panel (b) reports effective sample size; neither is tested against a cut-off. Genome build harmonisation is asserted inside the pipeline at read time rather than as a standalone product, and is therefore not shown here.
Role in the analysis pipeline: data acceptance is an OPTIONAL, non-standard step. It does not filter, rank, or veto any protein candidate.
Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
发表版表格
| 文件 | 下载 |
|---|---|
input_check_allele_frequency.csv | 下载 |
ldsc_genetic_correlation.csv | 下载 |
ldsc_genetic_correlation_pairwise.csv | 下载 |
ldsc_heritability.csv | 下载 |
本页图为网页版 PNG(最宽 1600 px)。投稿用的矢量 PDF 与 600 dpi TIFF 体积较大,留在仓库
results/ldsc_R9_dm/figures/,不随文档站分发。