Researchers have developed a nearly complete diploid genome benchmark that could help laboratories more accurately assess sequencing and genomic analysis methods, particularly in regions that have been difficult to evaluate using existing reference-based approaches.
The study focused on HG002, a widely used reference sample developed by the Genome in a Bottle (GIAB) Consortium. The researchers assembled both the maternally and paternally inherited copies of the genome from telomere to telomere, covering 99.4 percent of the approximately 6-billion-base diploid genome without detectable errors.
Most genomic sequencing workflows identify variants by comparing a patient's sequencing data with a standard reference genome. This approach works well across much of the genome but can be less reliable in repetitive, duplicated, and structurally complex regions. As a result, some regions are excluded from existing benchmarks used to assess the accuracy of sequencing tests and bioinformatics pipelines.
The new T2T-HG002 version 1.1 benchmark added 701.4 Mb of high-confidence autosomal sequence that was not included in the GIAB version 4.2.1 variant benchmark. It also included both sex chromosomes. The researchers reported that 99.35 percent of the diploid genome was free of detectable errors. Most remaining low-confidence sequence involved ribosomal DNA arrays, with other uncertain regions including satellites and segmental duplications.
The researchers also annotated the maternal and paternal genomes separately, allowing them to examine differences between the two chromosome sets. These included copy-number differences affecting genes such as DUSP22, CFHR1, CFHR3, GSTT1, and GSTM1. Such regions can be challenging to characterize when sequencing data are analyzed solely against a standard reference genome.
To support laboratory assessment of genomic methods, the researchers developed Genome Quality Checker software. The software can compare sequencing reads, phased variant calls, and genome assemblies directly with the diploid benchmark, evaluating factors such as base-level accuracy, coverage, structural accuracy, and whether variants are assigned to the correct parental haplotype.
The study also illustrated the difference between conventional variant calling and de novo genome assembly. Personalized genomes constructed from variants called against GRCh38 covered approximately 93 percent of HG002. Using the more complete T2T-CHM13 reference increased coverage to approximately 96 to 98 percent. By comparison, the best-performing de novo assembly assessed in the study covered 99.94 percent of the genome.
The benchmark provides a way to evaluate genomic assays in regions that have often fallen outside standard performance assessments. However, the researchers noted that the resource is not error-free, some repetitive regions remain difficult to validate, and genome-benchmarking metrics have not yet been standardized for clinical use.
