Skip to content

跨组学定向验证

步骤 6-3 · 只标注 · 图 5 张 · 发表版表格 0 个 (流程见 v5 定稿

视网膜 / 全血 / 组织 eQTL 与剪接 sQTL 四层,以及 WARS 的跨层位点对齐。★ 只能加分,不能减分 —— 跨层顶信号常常并非同一个变异。


Figure F6-0b. Which of the eleven prioritised proteins could actually be tested in each of the eight downstream molecular layers

layer_coverage
完整图注(英文,与投稿版一致)

Figure F6-0b. Which of the eleven prioritised proteins could actually be tested in each of the eight downstream molecular layers. The plasma protein layer is the discovery layer and is covered for every protein by construction, so it is not shown.

Why this panel exists. A reader looking at Figure F6-0 cannot tell whether a cell that is not positive means the layer was tested and returned nothing, or that the layer could not test that gene at all. Of the 88 gene-by-layer combinations here, 45 were analysed and 43 were not, for the reasons colour-coded above.

Two examples that must not be conflated. IFNAR1 was tested for splicing in all three tissues and returned a clear negative in each, so its splicing result is an informative negative and not missing data. SIGLEC5 has no annotated intron cluster in those tissues, so no splicing test exists for it at all. Both would appear as a non-positive cell in Figure F6-0.

Kidney cortex requires particular care. The GTEx kidney cortex expression sample is small (n = 73) and is flagged as underpowered throughout. Four genes additionally have no detectable cis-eQTL in a larger kidney reference (n = 686); for those the correct statement is that no cis-eQTL was detected, not that none exists.

Method and threshold. This panel is a coverage audit, not a test: for each gene and layer it records whether an interpretable cis signal exists at all (an eGene call, an annotated intron cluster, a significant splicing cluster, or sufficient sample size), so that a non-positive cell in Figure F6-0 can be read correctly. No significance threshold applies to this figure; the colocalisation thresholds used in Figure F6-0 are strong PP.H4 >= 0.8 and moderate 0.5 to 0.8. Sources. Retina: Advani et al. 2024. Whole blood: eQTLGen phase 1 (n = 31,684) and GTEx v8. Tibial nerve and kidney cortex expression and splicing: GTEx v8. Per-gene counts of tested variants, intron clusters and gene-level p values are given in the supplementary coverage tables.

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL.

Figure F6-0. Colocalisation of each of the eleven prioritised proteins with each diabetic complication, across nine molecular layers

crosslayer_colocalisation
完整图注(英文,与投稿版一致)

Figure F6-0. Colocalisation of each of the eleven prioritised proteins with each diabetic complication, across nine molecular layers. Rows are the 44 protein-by-outcome combinations, grouped by protein; columns are the layers.

Cell states. Of the 396 cells, 228 were analysed and 50 were not measured in that layer. Among the analysed cells, 20 show strong evidence of a shared causal variant (PP.H4 >= 0.8), 9 moderate (0.5 to 0.8), 155 were tested and did not colocalise, and 162 were tested but are not assessable. Numbers are printed only where PP.H4 reaches the moderate threshold, to keep the panel legible; the full matrix is given in the supplementary table.

Why not assessable is kept separate from not colocalised. These are different claims and conflating them is the main way a figure like this misleads. A cell is marked not assessable when the layer could not deliver an interpretable test for that gene: the kidney cortex expression data are underpowered at n = 73; four genes have no detectable cis-eQTL in kidney cortex even in the larger n = 686 reference; three genes are not expressed as eGenes in retina; APOE has no significant cis-eQTL in whole blood; and several genes have either no annotated intron cluster or no significant splicing cluster in a given tissue. In each of those cells a posterior probability can still be computed, but it does not carry evidence, so it is not shown as a negative result.

Thresholds are inherited from the main colocalisation analysis (step 5) and are not redefined here: strong PP.H4 >= 0.8, moderate 0.5 to 0.8. Applying a lower cut would change which cells look positive, so none was applied.

Reading the figure. WARS and diabetic retinopathy is supported in 4 layers at the strong threshold and a further 3 at the moderate threshold, spanning plasma protein, retina, whole blood expression and splicing; the two layers that do not support it are tibial nerve and kidney cortex expression. IFNAR1 illustrates why every outcome is shown for every protein: its plasma signal is with maculopathy, but its cross-layer expression support sits with neuropathy, in two independent whole-blood datasets. Restricting the rows to protein-outcome pairs that colocalised in plasma would have removed that observation entirely.

Method. coloc.abf on a 1 Mb cis window in every layer, priors p1 = p2 = 1e-4 and p12 = 1e-5. Expression layers: retina (Advani et al. 2024), whole blood (eQTLGen phase 1, n = 31,684; and GTEx v8), tibial nerve and kidney cortex (GTEx v8). Splicing layers: GTEx v8 sQTL in the same three tissues. Layer-level coverage and the reason for each non-assessable cell are given in Figure F6-0b.

Role in the analysis pipeline: step 6 is descriptive annotation. It does not filter, rank, veto or promote any protein. Nothing elsewhere in this study is conditional on a cell in this figure.

Complication endpoints from FinnGen R9 (general-population controls); type 1 diabetes GCST90824163; type 2 diabetes Mahajan et al. 2018 (non-UK Biobank).

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls.

Figure F6-0c. Posterior probability of a shared causal variant (PP.H4) for each of the 11 candidates across 9 molecular layers and the four diabetic complications

crosslayer_ph4_by_candidate
完整图注(英文,与投稿版一致)

Figure F6-0c. Posterior probability of a shared causal variant (PP.H4) for each of the 11 candidates across 9 molecular layers and the four diabetic complications.

Panels and points. One panel per candidate; within a panel, one row per molecular layer and one point per complication, offset vertically so that overlapping estimates remain visible. The horizontal position is the PP.H4 from coloc.abf; the dashed and dotted vertical lines are the 0.8 and 0.5 thresholds used at step 5 and not redefined here. Candidates are ordered by the number of interpretable estimates reaching the higher threshold, then by their highest PP.H4.

Three states are kept visually distinct, and this matters more here than the estimates themselves. Of the 396 gene-layer-outcome cells, 228 carry a PP.H4. Of those, 184 are interpretable and are drawn as filled points; the remaining 44 carry a number that should not be read as evidence and are drawn as open points. The 168 cells with no estimate at all are not drawn, and where a layer yielded nothing at all for a gene the whole row is shaded. An empty position therefore never means 'no colocalisation'; it means nothing was measured.

Why the kidney cortex row is entirely open. Every cell in that layer has a PP.H4, yet none is interpretable: the GTEx kidney cortex sample is no cis-eQTL detected even at n = 686. Drawn as filled points these would read as a whole layer of negative results, which is the single most likely misreading of this figure.

Interpretable estimates reaching the higher threshold, by candidate: WARS 4 of 8 layers; IFNAR1 3 of 5 layers; APOE 2 of 3 layers; NUDT5 2 of 5 layers; ERMAP 2 of 7 layers; NOTCH2 2 of 4 layers; PAM 1 of 6 layers; GALNT3 1 of 5 layers; SIGLEC5 1 of 4 layers; ACRBP 1 of 5 layers; LACTB2 1 of 4 layers.

Method and role. Colocalisation by coloc.abf (Giambartolomei et al. 2014) in a 1 Mb cis window, with the default priors used throughout this study. This is step-6 descriptive annotation: no candidate was promoted, demoted or vetoed on the basis of this figure, and the layers are not independent replications of one another. Molecular layer data: retina eQTL from Advani et al.; whole blood from eQTLGen and GTEx; tibial nerve, kidney cortex and all splicing quantitative trait loci from GTEx via the eQTL Catalogue.

This figure shows the same estimates as the colocalisation state matrix in the supplement, on a quantitative axis rather than as colour bands.

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls.

Figure F6-0d. Mendelian randomization estimates for the 11 candidates in each of the 8 molecular layers, for the four diabetic complications

layer_mr_effects
完整图注(英文,与投稿版一致)

Figure F6-0d. Mendelian randomization estimates for the 11 candidates in each of the 8 molecular layers, for the four diabetic complications.

Panels and points. One panel per molecular layer, one row per candidate and one estimate per complication, offset vertically so that overlapping estimates stay visible. Points are odds ratios with 95% confidence intervals on a logarithmic scale; filled points are significant at FDR < 0.05 within that layer, open points are not. 196 of the 352 gene-layer-outcome cells in these layers carry both an effect estimate and a standard error and are drawn; where a candidate yielded nothing in a layer the row is lightly shaded. An empty position means nothing was estimated, not a null result.

⚠ Do not compare the size of the odds ratios between panels. The horizontal scale is free between panels because the exposure is a different quantity in each layer: normalised gene expression for the expression panels and the intron excision ratio for the splicing panels, each normalised as in its source dataset. Only the direction of effect and the significance are comparable across panels. Within a panel the odds ratios among interpretable estimates run from Whole blood (GTEx) 0.78 to 1.52; Tibial nerve (GTEx) 0.81 to 1.22; Retina 0.98 to 1.10; Whole blood (eQTLGen) 0.63 to 1.61; Splicing, whole blood 0.88 to 1.16; Splicing, tibial nerve 0.88 to 1.13; Splicing, kidney cortex 1.08 to 1.14.

Why one panel is shaded. Kidney cortex (GTEx) is shaded because every estimate in it comes from a sample that is no cis-eQTL detected even at n = 686; underpowered (n = 73), so the estimates are shown for completeness but should not be read as evidence either way. The layer is kept rather than dropped: dropping it would turn 'tested but not interpretable' into 'not tested'.

Estimates drawn and significant, by layer: Retina 12 of 44 cells, 5 at FDR < 0.05; Whole blood (eQTLGen) 40 of 44 cells, 18 at FDR < 0.05; Whole blood (GTEx) 36 of 44 cells, 2 at FDR < 0.05; Tibial nerve (GTEx) 44 of 44 cells, 1 at FDR < 0.05; Kidney cortex (GTEx) 32 of 44 cells, 1 at FDR < 0.05; Splicing, whole blood 16 of 44 cells, 4 at FDR < 0.05; Splicing, tibial nerve 12 of 44 cells, 6 at FDR < 0.05; Splicing, kidney cortex 4 of 44 cells, 4 at FDR < 0.05.

Method and role. Inverse-variance weighted Mendelian randomization where more than one instrument was available, Wald ratio otherwise; the method used for each cell is recorded in the accompanying table. Benjamini-Hochberg false discovery rate applied within each layer. The candidate ordering is the same as in the colocalisation figure so the two can be read row by row. This is step-6 descriptive annotation: no candidate was promoted, demoted or vetoed on the basis of this figure, and the layers are not independent replications of one another. Molecular layer data: retina eQTL from Advani et al.; whole blood from eQTLGen and GTEx; tibial nerve, kidney cortex and all splicing quantitative trait loci from GTEx via the eQTL Catalogue.

The plasma protein layer is not shown here because it is the subject of two earlier figures; this figure covers only the eight molecular layers downstream of it.

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls; Diabetic maculopathy 3,572 cases / 308,547 controls; Diabetic nephropathy 4,111 cases / 308,539 controls; Diabetic neuropathy 2,843 cases / 271,817 controls.

Figure F6-0e. The WARS locus in 8 molecular layers and in retinopathy, on a common genomic axis

crosslayer_locus_WARS_Retinopathy
完整图注(英文,与投稿版一致)

Figure F6-0e. The WARS locus in 8 molecular layers and in retinopathy, on a common genomic axis.

Panels and points. One track per layer, all on the same 500 kb window centred on the anchor variant. Each point is one variant; the vertical axis is the two-sided association P value on a negative log scale, with a free scale per track because the layers differ in power by many orders of magnitude. Points are coloured by linkage disequilibrium with the anchor, computed from a European reference panel; variants absent from the panel are shown in grey as unknown rather than as low. The open diamond and the dashed line mark the anchor.

The anchor is 14:100376300, the variant with the highest per-variant posterior of being the shared causal variant in the plasma-protein-to-outcome colocalisation (SNP.PP.H4 = 0.996). It is the same anchor used in the single-layer locus figures, so the two can be compared directly. Linkage disequilibrium could be resolved for 62% of the variants shown.

Track order and labels. Tracks are ordered by the region-level colocalisation posterior from the frozen colocalisation table, which is printed in each track label so this figure can be read row by row against the cross-layer posterior figure. Strongest peak per track: Kidney cortex (GTEx) 3.1; Retinopathy 5.8; Plasma protein 198.7; Retina 11.9; Splicing, kidney cortex 7.7; Splicing, tibial nerve 105.9; Splicing, whole blood 16.4; Tibial nerve (GTEx) 3.2; Whole blood (GTEx) 37.3.

⚠ One track is sparse and its peak shape is not comparable with the others. Retina. The source for that layer reports conditionally independent signals from a stepwise forward-backward analysis rather than full nominal statistics, so it contributes 92 variants in this window against 10,734 for the densest track. It is kept rather than dropped because it is one of the layers that reached a high colocalisation posterior; it is labelled so that its thin appearance is not read as absence of signal.

What this figure does and does not establish. Alignment of the peaks across tracks is the visual counterpart of the colocalisation posteriors; it is not an independent test, and the layers are not independent of one another. The expression-quantitative and splicing-quantitative trait loci come from the same donors within a tissue, so agreement between them is expected. Genome build is GRCh38 for every track shown; the eQTLGen whole-blood layer is reported on GRCh37 and is therefore not shown here, and it did not reach the higher colocalisation threshold for this pair in any case.

Method: colocalisation by coloc.abf (Giambartolomei et al. 2014) in a 1 Mb cis window with the default priors used throughout; per-variant posteriors from the same fit. Linkage disequilibrium from PLINK on a European reference panel. Molecular layer data: retina eQTL from Advani et al.; whole blood, tibial nerve, kidney cortex expression and splicing from GTEx via the eQTL Catalogue.

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL. Outcomes: FinnGen R9 -- Diabetic retinopathy 10,413 cases / 308,633 controls.


发表版表格

本页的表格与 [靶点 pheWAS 与成药性](/dm-complications/29-靶点 pheWAS 与成药性) 共用(同属 09_descriptive_annotation)。

本页图为网页版 PNG(最宽 1600 px)。投稿用的矢量 PDF 与 600 dpi TIFF 体积较大,留在仓库 results/crossomics_eqtl/figures/,不随文档站分发。

个人科研与运维文档 · 内容持续修订