Description

The Human Pangenome Reference Consortium (HPRC) is building a collection of high-quality, nearly complete genome assemblies from many people, so that human genetic variation is represented by many reference genomes rather than a single one. Each person contributes two assemblies, one for each copy of their chromosomes (the maternal and paternal haplotypes). Comparing these assemblies to the standard reference genome GRCh38 reveals where individual genomes differ, including places where sequence has been gained, lost, inverted or duplicated.

This track collection shows the HPRC Release 2 assemblies compared against GRCh38. Release 2 comprises 462 haplotype assemblies from 231 individuals. Rather than displaying each assembly separately, the tracks summarize, at every position of GRCh38, how the whole set of assemblies aligns: how many of them cover a region, where their alignments break, and where they carry insertions, deletions, inversions or duplications relative to GRCh38.

Display Conventions and Configuration

The collection contains several tracks:

By default the rearrangement tracks show only events of 50 bp or larger to keep the display readable; this threshold can be changed with the Event size filter on the track configuration page. The number of assemblies sharing an event can also be used as a filter.

Colors in the rearrangement tracks match the structural-variant type colors used by the Long-read SVs track collection:

  Insertion — sequence present in the assemblies but absent from GRCh38
  Deletion — sequence present in GRCh38 but absent from the assemblies
  Inversion — segment aligned in the opposite orientation
  Duplication — segment aligned more than once
  Complex — sequence unalignable in both GRCh38 and the assemblies

Methods

HPRC Release 2 assemblies were built from long-read sequencing and aligned to the references using the Minigraph-Cactus pangenome pipeline; the consortium released per-assembly alignment chains of each haplotype against GRCh38. Starting from those chains, the tracks in this collection were derived at UCSC. For each of the 462 haplotypes, the chain was oriented so that GRCh38 is the target, then a net was computed. Coverage was measured by taking the single-cover projection of each assembly's aligning blocks onto GRCh38 and counting, per base, how many assemblies cover it, normalized by the number of assemblies. Insertions, deletions and complex events were extracted from the top-level chains of each net; inversions and duplications were extracted directly from the chains. Events at identical GRCh38 positions were merged across assemblies, recording how many assemblies share each event. Alignment breaks were collected from the chains and colored by how many assemblies break at the same position.

The HPRC Release 2 chain files were downloaded from the human-pangenomics annotation data tables (files *_vs_GRCh38.chain.gz in the free s3://human-pangenomics/ bucket). Processing steps are documented in the makeDoc, and the scripts used are in the kent source tree.

Data Access

The data can be explored interactively with the Table Browser or the Data Integrator. For automated analysis, the annotations are stored in bigWig and bigBed files that can be downloaded from our download server. bigBed files can be queried with the bigBedToBed tool, and bigWig files with bigWigToBedGraph; both can be downloaded from the utilities page. The original HPRC assemblies and alignments are available from the Human Pangenome Project.

Credits

Thanks to the Human Pangenome Reference Consortium for producing and releasing the Release 2 assemblies and alignments.

References

Liao WW, Asri M, Ebler J, Doerr D, Haukness M, Hickey G, Lu S, Lucas JK, Monlong J, Abel HJ et al. A draft human pangenome reference. Nature. 2023 May;617(7960):312-324. PMID: 37165242; PMC: PMC10172123