主题
工具选择与 MR
步骤 1–2 · 图 4 张(另补充 3 张) · 发表版表格 8 个 (流程见 v5 定稿)
cis-pQTL 工具的筛选漏斗,以及全蛋白组 MR 加多重检验校正的结果。★ 本流程不设 MHC 排除步骤,27 个 MHC 蛋白全部保留。
图
Figure F1. Selection of cis-pQTL instruments, in units of proteins

完整图注(英文,与投稿版一致)
Figure F1. Selection of cis-pQTL instruments, in units of proteins.
(a) Cascade from all cis-pQTL proteins in UKB-PPP to the set actually tested. Numbers to the right of each bar are proteins remaining; orange numbers to the left are proteins lost at that step. Counts are for the four FinnGen diabetic-complication endpoints, which share an identical instrument set (n = 1,587 each). The two diabetes control arms differ because their variant coverage differs: type 1 diabetes (GCST90824163) n = 1,722 and type 2 diabetes (Mahajan et al. 2018) n = 1,614.
(b) Itemised composition of the 77 proteins lost at allele harmonisation. Palindromic variants with intermediate allele frequency were removed because strand cannot be resolved; indels were lost because the two datasets represent insertion/deletion alleles differently; five further variants had alleles that could not be reconciled (a sixth, rs1006298, is already counted among the palindromic losses).
Mapping to the pre-specified workflow: (i) genome build and variant-identifier harmonisation, (ii) the cis window (UKB-PPP defines cis as within 1 Mb of the gene body), (iii) instrument strength. Steps (i) and (ii) account for all attrition. Step (iii) removed nothing: the weakest instrument had F = 43.8, so no instrument fell below F >= 10. Clumping is not applicable because a single sentinel variant per protein was used. MR-Steiger is not part of instrument selection in this workflow: it is computed after the association test, on the associations that pass FDR < 0.05, and is reported with reverse Mendelian randomization as the two layers of the directionality step. No variant is removed from the association test on directionality grounds.
NO exclusion of the MHC region was applied at this step. This is a deliberate departure from the common practice of removing MHC-encoded proteins during instrument selection; the 27 MHC-flagged proteins were carried forward and are handled at the colocalisation step instead.
Chromosome coverage was not restricted to autosomes. Of the 39 proteins whose cis sentinel lies on a sex chromosome, none matched any of the FinnGen complication endpoints, whereas 38 are present in the type 1 diabetes control arm.
All counts in this figure were measured directly from the instrument table, the outcome files and the pipeline logs on 2026-08-09.
Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
图 F2c 31 个显著蛋白在四种并发症与两型糖尿病中的效应估计

结论. 四种并发症的显著蛋白高度集中在 MHC 区,「同一蛋白横跨多个并发症」与「同时与 1 型糖尿病相关」这两个现象也主要由该区域驱动;因此本图上部的广谱蛋白不能直接当作并发症共有的生物学机制来读。
结果. 在至少一个并发症中通过 FDR < 0.05 的 31 个蛋白与六个结局构成 186 个格,其中 185 格有效应估计;唯一的空缺是 LTB 与 2 型糖尿病,该蛋白的 cis 哨兵变异不在后者的汇总统计中,图上以灰格标出,属未测而非阴性。四种并发症合计 57 对关联通过校正(视网膜病变 24 对、黄斑病变 19 对、肾病 9 对、神经病变 5 对)。行自上而下按图上的星号数(六列合计)分组,2 个蛋白 5 颗、7 个蛋白 4 颗、4 个蛋白 3 颗、8 个蛋白 2 颗、10 个蛋白 1 颗,组内按最小 FDR 排列,没有蛋白在六个结局上全部显著;关联到两种及以上并发症的 15 个蛋白里有 10 个编码在 MHC 区,而全部 31 个蛋白中为 11 个。就两型糖尿病而言,这 31 个蛋白里有 15 个在 1 型糖尿病、4 个在 2 型糖尿病同样显著;与 1 型的重叠明显更多,尽管 2 型的有效样本量(193,440)是 1 型(79,393)的 2.4 倍。这一不对称几乎全部来自 MHC:与 1 型重叠的 15 个里有 9 个位于 MHC,而与 2 型重叠的 4 个无一位于 MHC,去除 MHC 区后两者分别为 6 个与 4 个。
完整图注(英文,与投稿版一致)
Figure F2c. Effect estimates for every protein that passed FDR < 0.05 in at least one diabetic complication, shown across the four complications and across type 1 and type 2 diabetes.
Rows are the 31 proteins significant in at least one complication. They are ordered from the top down by the number of outcomes in which the protein is significant, counting all six columns shown -- that is, by the number of asterisks in the row (2 with 5, 7 with 4, 4 with 3, 8 with 2, 10 with 1) -- and within each of those groups by smallest FDR. Columns are the four complications, then type 1 and type 2 diabetes; the two blocks are separated so that the diabetes columns are not read as complications. Fill is log2(odds ratio) per one standard deviation increase in genetically predicted plasma protein level. An asterisk marks the 76 protein-outcome pairs passing FDR < 0.05; unmarked cells are estimates that did not pass correction, not missing data. No protein is significant for all six outcomes.
185 of the 186 cells carry an estimate. The single grey cell, LTB x Type 2 diabetes, was never tested: the cis sentinel variant for that protein is absent from the type 2 diabetes summary statistics. The outcome datasets do not cover identical protein sets (1,587 proteins for each complication, 1,722 for type 1 diabetes, 1,614 for type 2 diabetes), and the sets are not nested. A grey cell is therefore an absence of data, not an estimate of no effect.
The diabetes columns are shown so that each complication association can be read against the association with diabetes liability itself. Of these 31 proteins, 15 also reach FDR < 0.05 for type 1 diabetes and 4 for type 2 diabetes. The larger overlap is with type 1 diabetes even though its effective sample size is the smaller of the two: 9 of those 15 proteins are encoded in the MHC region, against 0 of 4 for type 2 diabetes, and outside the MHC the two counts are 6 and 4.
A black triangle after a protein name marks proteins encoded in the MHC region (11 of the 31). 10 of the 15 proteins associated with two or more complications are MHC-encoded: long-range linkage disequilibrium across the MHC can drive apparent associations for several proteins and several outcomes at once, so a row with many asterisks is not by itself evidence of shared biology.
The colour scale is capped at the 98th percentile of |log2(OR)| across all estimated cells (+/- 2.791); the 4 more extreme estimates are drawn at the scale limit. Capping is applied because a small number of very large point estimates would otherwise compress all remaining cells to near-white.
The stratified comparison of complication effects against diabetes effects is presented separately in Figure F-DM; this figure only places the estimates side by side.
Method: Wald ratio from a single cis sentinel variant per protein. Multiple-testing correction: Benjamini-Hochberg FDR applied within each outcome.
Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
图 F2d 95 个显著蛋白在六个结局之间的重叠

结论. 筛查得到的蛋白绝大多数只与单一结局相关,数量上最大的两组是只与糖尿病本身相关的蛋白;而在与并发症相关的蛋白中又有近六成同时与糖尿病相关 —— 并发症特异的信号只占全部阳性结果的一小部分。
结果. 六个结局合计 142 对关联通过 FDR < 0.05,涉及 95 个不同蛋白,构成 18 种结局组合。72 个蛋白(95 个中的 76%)仅在一个结局中显著;最大的两个单结局组合是 1 型糖尿病(40 个蛋白)与 2 型糖尿病(22 个),合计 64 个蛋白只与糖尿病相关而与四种并发症均无关联。反过来看,至少与一种并发症相关的 31 个蛋白中有 18 个(58%)同时与至少一种糖尿病相关(1 型 15 个、2 型 4 个,其中 1 个两者兼有),这正是图 F-DM 逐对量化的那一部分重叠。同时在两型糖尿病中显著的蛋白全局仅 3 个,与两者之间遗传相关近于零(图 F0a)的结果方向一致。
完整图注(英文,与投稿版一致)
Figure F2d. Overlap between outcomes of the 95 proteins that reached FDR < 0.05 in the step-2 Mendelian randomization screen.
What is plotted. Vertical bars count proteins per outcome combination; the dot matrix below identifies the combination; horizontal bars give the number of significant proteins per outcome. The two diabetes traits are included deliberately: the question this figure answers is how much of the complication signal is shared with diabetes liability itself.
Reading it. 72 of the 95 proteins are significant for exactly one outcome, and the two largest single-outcome blocks are the two diabetes traits (type 1 diabetes 40 proteins, type 2 diabetes 22). In total 64 proteins are significant for a diabetes trait but for none of the four complications.
Read in the other direction, 18 of the 31 proteins significant for at least one complication are also significant for at least one diabetes trait (58%): 15 for type 1 diabetes, 4 for type 2 diabetes, 1 for both. This is the overlap quantified association by association in Figure F-DM, and it is the reason the two diabetes traits are shown here rather than left out.
This figure is at the association level only. No colocalisation, replication or directionality evidence enters it; sharing here can reflect linkage disequilibrium as easily as shared biology.
Method: Wald ratio from a single cis sentinel variant per protein; Benjamini-Hochberg FDR applied within each outcome, threshold 0.05.
Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
Figure F2e. Protein-complication network of the 57 associations reaching FDR < 0.05 in the step-2 Mendelian randomization screen, involving 31 proteins and the four diabetic complications

完整图注(英文,与投稿版一致)
Figure F2e. Protein-complication network of the 57 associations reaching FDR < 0.05 in the step-2 Mendelian randomization screen, involving 31 proteins and the four diabetic complications.
Edges. One edge per protein-complication association, coloured by the direction of the estimated effect: an edge is drawn when a genetically predicted increase in the plasma protein is associated with higher or with lower odds of that complication. Edge colour is direction only; edge width carries no information.
Nodes. The four complications are the large nodes. Proteins are coloured by whether they are encoded in the MHC region: 11 of the 31 proteins here are. ⛔ No downstream information is used -- colocalisation, external replication and the epitope investigation all come later in the workflow and are deliberately not shown. ⛔ No candidate grading is shown either, because this study does not grade candidates.
Labels. The four complications and the 15 proteins associated with two or more complications are labelled, of which 10 are MHC-encoded; the remaining protein names are omitted to keep the figure legible at print size and are listed in the supplementary association table.
What this figure is not. It is an association-level map, drawn before reverse Mendelian randomization, external replication, the epitope investigation and colocalisation. Sharing of a protein between two complications here can reflect linkage disequilibrium or shared liability to diabetes as easily as a shared causal mechanism; the proteins shared across several complications are mostly MHC-region proteins. The two diabetes control arms are deliberately not drawn: including them would make most edges non-complication edges while the figure is titled a complication network. Their overlap is shown in Figures F2d and F5d instead.
Method: Wald ratio from a single cis sentinel variant per protein; Benjamini-Hochberg FDR within each outcome, threshold 0.05. Layout: Fruchterman-Reingold force-directed, with a fixed random seed so the figure is reproducible.
Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls.
补充图
以下 3 张不进正文,作为补充材料。图与图注均为完整版,只是版面位置不同。
Figure F2a. Proteome-wide cis-pQTL Mendelian randomization screen, one panel per outcome

完整图注(英文,与投稿版一致)
Figure F2a. Proteome-wide cis-pQTL Mendelian randomization screen, one panel per outcome.
Each point is one protein-outcome association estimated by the Wald ratio from a single cis sentinel variant. The x axis is log2(odds ratio) per one standard deviation increase in genetically predicted plasma protein level; the y axis is -log10(P). The dotted horizontal line in each panel marks the P value corresponding to FDR < 0.05 within that outcome. Up to eight proteins per outcome are labelled, chosen by smallest P.
Triangles mark proteins encoded in the MHC region as flagged in the UKB-PPP sentinel table. The MHC region was NOT excluded during instrument selection in this study, which is a deliberate departure from common practice; these symbols show what that decision contributes to the screen.
Number of proteins tested and passing FDR < 0.05 in each outcome: diabetic retinopathy 1,587 tested / 24 passing (9 in MHC); diabetic maculopathy 1,587 / 19 (9); diabetic nephropathy 1,587 / 9 (8); diabetic neuropathy 1,587 / 5 (4); type 1 diabetes 1,722 / 57 (18); type 2 diabetes 1,614 / 28 (1). The three groups differ in the number of proteins tested because the outcome files differ in variant coverage, not because different selection rules were applied.
The two diabetes arms (type 1, type 2) are shown here because they were part of the same screen. They serve as control arms; the stratified comparison of complication effects against diabetes effects is presented separately (Figure F-DM), not here.
Multiple-testing correction: Benjamini-Hochberg FDR applied within each outcome.
Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls. Diabetes control arms: Type 1 diabetes 20,355 cases / 797,363 controls; Type 2 diabetes 55,005 cases / 400,308 controls.
Figure F2b. All 57 protein-complication associations passing FDR < 0.05, grouped by outcome

完整图注(英文,与投稿版一致)
Figure F2b. All 57 protein-complication associations passing FDR < 0.05, grouped by outcome.
Points are odds ratios per one standard deviation increase in genetically predicted plasma protein level, estimated by the Wald ratio from a single cis sentinel variant; whiskers are 95% confidence intervals. The x axis is on a log scale and the dashed line marks the null (OR = 1). Within each outcome, associations are ordered by odds ratio. The 57 associations involve 31 distinct proteins. Odds ratios range from 0.071 to 3.080; 28 associations are in the risk direction (OR > 1) and 29 in the protective direction (OR < 1).
Triangles mark proteins encoded in the MHC region. Thirty of the 57 associations (53%), involving 11 of the 31 proteins, are MHC-region proteins. The MHC region was not excluded during instrument selection; colocalisation at these loci is reported separately and is interpreted with the caveats described there.
Multiple-testing correction: Benjamini-Hochberg FDR applied within each outcome. The two diabetes control arms are not shown here; see Figure F2a for the full screen and Figure F-DM for the stratified comparison.
Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls.
Figure F2f. The 15 plasma proteins whose genetically predicted level is associated with 2 or more of the four diabetic complications, showing all 41 associations

完整图注(英文,与投稿版一致)
Figure F2f. The 15 plasma proteins whose genetically predicted level is associated with 2 or more of the four diabetic complications, showing all 41 associations.
Panel. One row per protein, one estimate per complication in which that protein reached the significance threshold, offset vertically so overlapping estimates stay visible. Points are odds ratios with 95% confidence intervals on a logarithmic scale, taken directly from the screening table. Horizontal rules separate proteins by how many complications they are associated with: 3 proteins in 4 complications; 5 proteins in 3 complications; 7 proteins in 2 complications.
† marks the 10 of 15 proteins encoded in the major histocompatibility complex, identified from the MHC column of the instrument table rather than from coordinates applied here. This is the main message of the figure and not an aside: long-range linkage disequilibrium across the MHC means that one haplotype can drive apparent associations for many different proteins at once, so sharing across complications in this dataset is largely a linkage-disequilibrium phenomenon rather than evidence of shared biology. By group: 3 of 3 in the 4-complication group; 3 of 5 in the 3-complication group; 4 of 7 in the 2-complication group.
Counting rule. Only the four complications are counted. The two diabetes control arms are deliberately excluded from the count: an earlier version of this figure in a previous version of the pipeline was titled as showing proteins shared across three or more complications while in fact counting the control arms as well, so that most of the proteins shown did not meet the stated criterion. Here the criterion and the count are the same thing, and 41 of the 57 significant protein-complication associations in the screen involve these 15 proteins.
What this figure is not. These are step-2 screening associations, drawn before reverse Mendelian randomization, external replication, the epitope investigation and colocalisation. An association appearing for several complications is not evidence that the protein acts on each of them independently: the complications are themselves strongly genetically correlated, and the shared liability to diabetes is an alternative explanation examined separately.
Method: Wald ratio from a single cis sentinel variant per protein; Benjamini-Hochberg false discovery rate within each outcome, threshold 0.05. Confidence intervals are those reported in the screening table and were not recomputed here.
Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls.
发表版表格
| 文件 | 下载 |
|---|---|
mr_all_maculopathy.csv | 下载 |
mr_all_nephropathy.csv | 下载 |
mr_all_neuropathy.csv | 下载 |
mr_all_retinopathy.csv | 下载 |
mr_all_type1_diabetes.csv | 下载 |
mr_all_type2_diabetes.csv | 下载 |
mr_overlap_matrix.csv | 下载 |
mr_significant_summary.csv | 下载 |
本页图为网页版 PNG(最宽 1600 px)。投稿用的矢量 PDF 与 600 dpi TIFF 体积较大,留在仓库
results/screen_R9_dm/figures/,不随文档站分发。