Description

Genes that arose from a common ancestor by duplication, called paralogs, often keep a similar protein sequence and three-dimensional structure long after they have taken on different roles. Because equivalent positions in a family of paralogs tend to play equivalent structural or functional roles, a variant known to cause disease at one position in one gene points to a candidate disease-associated position at the aligned residue in each of its paralogs. This idea, known as paralogue annotation, was introduced for the cardiac ion channel genes and has been used to help classify variants of uncertain significance.

This track container takes every ClinVar variant that changes a protein residue and projects it onto the equivalent residue of each paralog of the gene it occurs in. It has two parts:

The mapping is purely positional: it indicates that a residue is the alignment equivalent of a position that is variant in a paralog. It does not itself assign pathogenicity, and a projection is only as reliable as the underlying alignment, so the percent-identity and residue-conservation fields should be used to judge confidence.

Display Conventions and Configuration

In the Paralog Variants track, each feature marks the codon in this gene that aligns to a protein-changing ClinVar variant in a paralog. Features are colored by the ClinVar clinical significance of the source variant:

  Pathogenic or likely pathogenic
  Benign or likely benign
  Uncertain significance
  Conflicting classifications
  Other or risk factor

These are the same colors used by the ClinVar track.

The details page for each mapped variant links to the original ClinVar record and to the position of the source variant on the genome. The track can be filtered by clinical significance, ClinVar review status (0 to 4 stars), molecular consequence, residue conservation (identical, similar or different between the two genes), Ensembl paralog percent identity, and the percent identity of the pairwise alignment. Because most protein-changing ClinVar variants are classified as uncertain, filtering to pathogenic and likely pathogenic classifications and to higher alignment identities is recommended for variant-interpretation use.

The Paralog Alignments track shows each pairwise protein alignment as a blocked feature; the label is the partner gene, and the aligned blocks mark the regions that could be compared between the two proteins.

Methods

Paralog relationships and their percent identities were taken from Ensembl BioMart (release 116) as within-species paralog pairs between Ensembl genes. For each protein-coding gene, the single representative transcript and protein were taken from the MANE Select set, so that residue numbering matches the transcripts used in clinical variant reporting. Each pair of paralogs whose proteins share at least 20% identity was aligned with a global Needleman-Wunsch alignment (BLOSUM62), and the aligned residue positions were recorded. Pairs below 20% identity were not aligned, because alignments in that range are not reliable enough to transfer positions.

Every ClinVar variant with a protein-residue consequence (missense, stop gain or loss, start loss and in-frame insertions or deletions) was assigned to the codon of the overlapping MANE Select transcript, and its reference residue was taken from the translated protein. Each variant was then projected across the alignment to the equivalent residue of every paralog, and that residue was mapped back to its genomic codon to place the feature. Variants that align to a gap, or that occur in a gene with no paralog, were not projected. The pairwise alignments themselves were converted to PSL to form the Paralog Alignments track.

The ClinVar data were taken from the UCSC ClinVar track for this assembly. The paralog list was downloaded from the Ensembl BioMart service at https://www.ensembl.org/biomart/martservice, and the MANE Select set from the UCSC MANE track. The commands used to build the track are documented in the makeDoc, and the processing scripts are in the kent source tree.

Data Access

The data can be explored interactively in table format with the Table Browser or the Data Integrator and exported from there to spreadsheet or tab-sep tables. From scripts, the data can be accessed through our API, track=clinvarMappedParalog.

For automated download and analysis, the annotation is stored in bigBed files that can be downloaded from our download server. The files for this track are called clinvarMappedParalog.bb and clinvarMappedParalogAln.bb. Individual regions or the whole genome annotation can be obtained using our tool bigBedToBed, which can be compiled from the source code or downloaded as a precompiled binary for your system. Instructions for downloading source code and binaries can be found here. The tool can also be used to obtain features within a given range, e.g. bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/clinvarMapped/clinvarMappedParalog.bb -chrom=chr2 -start=165000000 -end=166000000 stdout.

Credits

Thanks to Ensembl for the paralog annotations, to the MANE collaboration between NCBI and EMBL-EBI for the reference transcript set, and to ClinVar for the variant data. The paralogue annotation concept was developed by Ware, Walsh, Cook and colleagues.

References

Ware JS, Walsh R, Cunningham F, Birney E, Cook SA. Paralogous annotation of disease-causing variants in long QT syndrome genes. Hum Mutat. 2012 Aug;33(8):1188-1191. PMID: 22581653; PMC: PMC4640174

Walsh R, Peters NS, Cook SA, Ware JS. Paralogue annotation identifies novel pathogenic variants in patients with Brugada syndrome and catecholaminergic polymorphic ventricular tachycardia. J Med Genet. 2014 Jan;51(1):35-44. PMID: 24136861; PMC: PMC3888601

Morales J, Pujar S, Loveland JE, Astashyn A, Bennett R, Berry A, Cox E, Davidson C, Ermolaeva O, Farrell CM et al. A joint NCBI and EMBL-EBI transcript set for clinical genomics and research. Nature. 2022 Apr;604(7905):310-315. PMID: 35388217; PMC: PMC9007741