Skip to content

靶点 pheWAS 与成药性

步骤 6-1 / 6-2 · 只标注 · 图 4 张 · 发表版表格 13 个 (流程见 v5 定稿

6-1 on-target pheWAS(在靶获益与风险、脱靶表型域、具名性状)与 6-2 成药性。⛔ 这一步不剔除任何候选。


Figure F6-1b. All phenome-wide associations of the instruments for the prioritised proteins, grouped by GWAS Catalog EFO parent term

offtarget_phenotype_domains
完整图注(英文,与投稿版一致)

Figure F6-1b. All phenome-wide associations of the instruments for the prioritised proteins, grouped by GWAS Catalog EFO parent term. Each point is one association; 978 associations across 890 datasets and 9 proteins are shown.

Coverage. 59.0% of all associations and 55.7% of Bonferroni-significant associations map to an EFO parent term and are plotted. The remaining 681 associations come from sources with no ontology mapping, predominantly Neale-lab UK Biobank scans, and are reported in the supplementary table rather than assigned to a category we cannot verify.

The dashed line is the Bonferroni threshold, 3e-05, computed over all 1,659 associations returned. P values are capped at 1e-300 for display. Point shape and colour give the direction of the trait association for the allele that raises the protein; direction is shown for completeness and is not interpretable as benefit or harm for non-disease traits.

One protein dominates. APOE accounts for the large majority of significant associations here, consistent with it being a highly pleiotropic locus. The labelled point in each category is the single most significant association in that category.

Category assignment. EFO parent terms are resolved in three steps, in order: by GWAS Catalog study accession for ebi-a datasets, then by publication identifier, then by exact match of the trait name against the official EFO term label. Ambiguous EFO labels that map to more than one parent are discarded rather than guessed. Only the official EFO term column is used for name matching; the reported trait string is not, because its parent term reflects the study as a whole and produces incorrect assignments.

This mapping is written to dataset_domains_efo.csv and is separate from the three-way domain classification used in Figure F6-1, which comes from OpenGWAS metadata. The two answer different questions and are deliberately not merged.

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL.

Figure F6-1. Phenome-wide associations of the cis variant used to instrument each prioritised protein, classified by phenotype domain and, for disease phenotypes, by direction relative to the modelled target effect

ontarget_benefit_risk
完整图注(英文,与投稿版一致)

Figure F6-1. Phenome-wide associations of the cis variant used to instrument each prioritised protein, classified by phenotype domain and, for disease phenotypes, by direction relative to the modelled target effect.

Panel a. Every association returned for the instrument is counted, split into disease phenotypes, quantitative traits and molecular measures. The burden is extremely uneven: APOE returns 926 associations, of which 131 are disease phenotypes, while the next largest is NUDT5 with 29 and the smallest is GALNT3 with 2. APOE is a well-known highly pleiotropic locus and the figure is not scaled to hide that.

Panel b. Disease associations are aligned to the direction in which the target would be perturbed to reduce the index complication, and are then labelled potential benefit or potential risk. Only three proteins have any disease association at all: APOE (25 benefit, 106 risk), NOTCH2 (0 benefit, 3 risk) and NUDT5 (1 benefit, 0 risk). For the remaining proteins the disease count is zero, which here means that no disease phenotype passed the significance threshold, not that the variant was not tested.

LACTB2 is marked not assessable rather than zero. Its instrument is too rare for a phenome-wide scan to return interpretable associations, so the analysis cannot distinguish an absence of pleiotropy from an absence of power. Plotting it as zero would assert the former. This is the same distinction applied throughout step 6.

Direction is only meaningful for disease phenotypes. A higher or lower value of a quantitative trait or a molecular measure is not intrinsically beneficial or harmful, so those two classes are counted in panel a but are deliberately not assigned a direction in panel b.

Method: phenome-wide association scan of each instrument across OpenGWAS, with significance assessed at a Bonferroni threshold over all associations returned. Phenotype domains are assigned from OpenGWAS metadata by analysis/48_trait_domains.R. Full per-association results are in the supplementary tables.

Role in the analysis pipeline: this is descriptive annotation. The phenome-wide results were never used as a filter at any step of this study.

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL.

Figure F6-1c. The 5 strongest phenome-wide associations of the cis instrument for each prioritised protein, named

phewas_top_traits
完整图注(英文,与投稿版一致)

Figure F6-1c. The 5 strongest phenome-wide associations of the cis instrument for each prioritised protein, named.

Panel. One row per association, grouped by protein and ordered within a protein by P value; proteins are ordered by their total number of significant off-target associations. Colour is the direction of the association for the instrument allele that raises the protein. 38 rows are shown, drawn from 1,151 significant off-target associations across 790 distinct traits.

Why this figure exists alongside the other two phenome-wide figures. The burden figure counts associations by phenotype class and the domain figure groups them by ontology term; neither names a single trait. With 790 distinct traits behind those summaries, naming the strongest ones is the only way a reader can judge whether the pleiotropy is plausibly mechanism-related or simply reflects a gene-dense locus.

Exclusions and caveats, all computed from the source table. 6 associations in which the trait is the protein's own gene were removed: an instrument affecting the expression of its own gene is on-target, not pleiotropy. 42 associations report a P value at or below the smallest positive double-precision number and are drawn at that limit and marked; their true P values are smaller than can be represented and are not comparable with one another. 0 of the rows shown are labelled by Ensembl gene identifier rather than by a trait name, because the source dataset is a molecular measurement labelled that way; they are shown as the source labels them rather than renamed here.

LACTB2. Its instrument returned no phenome-wide association at all, so it contributes no row here. That is an absence of retrievable data, not evidence that the instrument is free of off-target effects.

Role in the analysis. This is step-6 descriptive annotation. No candidate was filtered, ranked or vetoed on the basis of these associations. The instruments are single cis variants, so an association here can arise from the protein, from another gene in linkage disequilibrium with the same variant, or from a neighbouring gene in the same cis window; this figure does not distinguish between those.

Data: OpenGWAS phenome-wide scan of each cis sentinel variant. Threshold: Bonferroni correction across all datasets queried, as recorded in the source table. Ontology terms from the GWAS Catalog EFO mapping where available.

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL.

Figure F6-2. Existing pharmacological tools for each of the eleven prioritised proteins: whether a tractable modality exists, whether the target is recorded in ChEMBL, how many approved drugs act on it, and whether those drugs act in the direction that the Mendelian randomisation estimate implies would be beneficial

druggability_landscape
完整图注(英文,与投稿版一致)

Figure F6-2. Existing pharmacological tools for each of the eleven prioritised proteins: whether a tractable modality exists, whether the target is recorded in ChEMBL, how many approved drugs act on it, and whether those drugs act in the direction that the Mendelian randomisation estimate implies would be beneficial.

Approved drug counts are taken from ChEMBL, not from Open Targets. For IFNAR1 the two disagree: Open Targets counts 12 approved drugs while ChEMBL records 11 at phase 4, because albinterferon alfa-2b is phase 3 in ChEMBL. Development of that agent was discontinued in 2010. ChEMBL is upstream of Open Targets and carries the finer phase annotation, so it is used here; the discrepancy is recorded in the source table as OT12/ChEMBL11.

A mechanistic constraint that the direction column does not convey on its own. IFNAR1 is the only protein here with approved drugs: 10 of the 11 act in the same direction as the protective effect estimated by Mendelian randomisation, and 1 acts in the opposite direction. However, the 10 same-direction agents are all interferons, and interferons signal through the IFNAR1 and IFNAR2 heterodimer rather than through IFNAR1 alone. The only approved agent that targets the IFNAR1 subunit specifically is anifrolumab, and it is a receptor antagonist, that is, the one drug in the opposite direction. Any repurposing argument therefore has to address the fact that selectivity for the subunit and the required direction of effect currently point to different molecules.

Two different kinds of absence, which the figure keeps apart. For 4 proteins ChEMBL lists the target but holds no mechanism-of-action record, meaning the database was searched and returned nothing. For 5 proteins the target does not appear in ChEMBL at all, so no search was possible. Neither of these is evidence that no drug exists: both Open Targets and ChEMBL only cover compounds with recorded bioactivity. The correct statement for these nine proteins is that no known drug is recorded in either resource, which is a stronger statement than the previous wording that referred to Open Targets alone.

Tractability is taken from Open Targets and reflects whether a small-molecule pocket or an antibody-accessible epitope has been identified; it is a prediction about feasibility, not evidence that a compound exists.

No significance threshold applies to this figure: it reports database annotation (tractability, approved-drug counts and mechanism direction), not a statistical test. Role in the analysis pipeline: this is descriptive annotation. Nothing in the candidate selection depended on it, and a protein with no pharmacological tools is not thereby downgraded.

Data sources and sample sizes. Exposure: UKB-PPP plasma proteome (Olink Explore 3072), 34,557 European participants, 1,954 proteins with a cis-pQTL.


发表版表格

文件下载
druggability_chembl_crosscheck.csv下载
druggability_open_targets.csv下载
eqtl_retina.csv下载
eqtl_retina_coverage.csv下载
eqtl_tissue.csv下载
eqtl_tissue_coverage.csv下载
eqtl_whole_blood.csv下载
eqtl_whole_blood_coverage.csv下载
ontarget_phewas_benefit_risk.csv下载
ontarget_phewas_summary.csv下载
singlecell_expression.csv下载
sqtl.csv下载
sqtl_coverage.csv下载

本页图为网页版 PNG(最宽 1600 px)。投稿用的矢量 PDF 与 600 dpi TIFF 体积较大,留在仓库 results/ontarget_phewas/figures/,不随文档站分发。

个人科研与运维文档 · 内容持续修订