• Home page
  • Resources
  • BLOG
  • Unlock the Full Potential of Genomics with PacBio HiFi Long‑Read Sequencing

Unlock the Full Potential of Genomics with PacBio HiFi Long‑Read Sequencing


Release time:2026-07-10 18:53:11


1. What Are You Missing with Short Read Sequencing?

Short-read sequencing has delivered a lot over the years, and it's still the go‑to for SNPs (single nucleotide polymorphisms) and small InDels (insertions and deletions). But those are the easy ones. The real challenge is that genomes are full of repetitive sequences, segmental duplications, and complex structural rearrangements that short reads simply cannot resolve. And that matters, because those hard‑to‑reach regions often hold the key to understanding some of the most important biology.

So what exactly are you missing?

  • Structural variants (SVs): Large insertions, deletions, inversions, and translocations that affect gene dosage and regulation. Short reads are too short to detect them reliably, and most SVs lie in regions that are difficult to map.

  • Copy number variations (CNVs): Gains or losses of genomic segments that are known drivers of cancer, neurodevelopmental disorders, and adaptive traits. Short-read algorithms struggle to quantify CNVs accurately, especially in repetitive contexts.

  • Repeat expansions: The hallmark of dozens of neurological and neuromuscular diseases. Short reads cannot span these regions, let alone size them accurately.

  • DNA methylation: The key layer of epigenetic regulation that responds to development and environment. Traditional bisulfite sequencing suffers from uneven coverage and fails in repetitive elements, leaving large gaps in your methylome.

  • Chromatin accessibility: The "on/off" switches of gene expression. Standard methods like ATAC-seq require separate library preparation, fragmented data, and offer no singlemolecule resolution.

When you rely on short-read data alone, you're working with an incomplete picture. You end up with assemblies that have gaps. You miss pathogenic SVs that could be the key to your study. You cannot connect methylation changes to specific haplotypes because the reads just aren't long enough to phase them. And on top of all that, you burn time and budget on separate assays—WGBS here, ATAC-seq there—that really could have been combined from the start.

So what does that add up to? You get missed discoveries, assemblies that fall apart when you look closer, and conclusions that might not survive a second look through a more complete lens.

2.The Solution: Fiber-seq—One Run, Four Layers of Data

PacBio HiFi (High-Fidelity) sequencing was designed to solve exactly these problems. It combines two capabilities that no other single platform offers: long read length and exceptional base-level accuracy.

HiFi reads typically come in at 15 to 25 kilobases, which is long enough to span most SVs, repeat expansions, and other complex regions. But length alone would not be enough if the reads were error‑prone. Older long‑read technologies had to trade accuracy for length; HiFi avoids that trade‑off by using Circular Consensus Sequencing (CCS). The same DNA molecule is read multiple times in a circle, and the consensus is built from those passes. The result is a read that is both long and highly accurate—≥99.9% (Q30+)—something that was considered impossible not long ago. With HiFi, you can assemble genomes with fewer gaps, call SVs with confidence, and phase haplotypes directly from the reads—without needing trio data or statistical imputation.

The real breakthrough comes with Fiber-seq.

Fiber-seq is a single-molecule chromatin fiber sequencing assay that simultaneously profiles chromatin accessibility, DNA methylation, and genetic variation on individual DNA molecules. Here is how it works:

A non-sequence-specific adenine methyltransferase (m6A-MTase) is used to mark accessible regions of chromatin. Wherever chromatin is open, which allows transcription factors and other regulatory proteins to bind—the enzyme adds a methyl group to adenine bases in the DNA. These methylated bases are then detected by the PacBio sequencing instrument, alongside the canonical A, C, G, and T bases, in the same read.

From a single sequencing run—combining a Fiber-seq gDNA library and a MAS-Seq cDNA library in one SMRT Cell—you obtain four layers of data:

  • Whole-genome sequence: For assembly and variant calling (SNPs, InDels, SVs, CNVs)

  • 5mC methylation: The full DNA methylome, including previously inaccessible repetitive regions

  • Chromatin accessibility: Open chromatin regions (FIRE peaks), nucleosome positioning, and transcription factor footprints

  • Haplotype-resolved regulatory architecture: Because long reads carry phase information, you can link regulatory elements to specific parental haplotypes

To get comparable data the traditional way, you would need three separate experiments: WGS for the genome, WGBS for methylation, and ATAC‑seq for chromatin. Each has its own library prep, its own sequencing run, and its own bioinformatics pipeline, which means batch effects can creep in at multiple stages. You also end up using three times as much sample, and the data layers cannot be directly linked at the single‑molecule level because they come from different DNA aliquots.

Fiber‑seq eliminates all of that. You get everything from a single library and a single sequencing run—four integrated data layers, no batch effects, no redundant sample prep. Weeks of work get compressed into a single workflow.

This is not just an incremental improvement. It is a fundamentally different way of seeing the genome.

3.Data Quality: Validated Against the Gold Standard

Before you trust any new technology with your project, you want to know it actually works. The question is not whether long reads can span complex regions - we already know they can. What really matters is whether the data is accurate enough to support confident biological conclusions.

We have validated our pipeline against the GIAB (Genome in a Bottle) benchmark datasets, which are widely regarded as the industry gold standard for human genome accuracy. GIAB provides highly curated reference genomes for well‑characterised cell lines, so you can directly compare your variant calls against a truth set.

Using these benchmarks, HiFi genome sequencing consistently achieves:

Metric

Performance

SNV F1-score

>99.8%—matching the highest published standards

InDel accuracy

>99%—outperforming shortread methods across all sequencing depths

SV detection

High sensitivity and precision in repetitive regions where short reads fail

Haplotype phasing

Direct phasing from long reads without trio data

Short‑read sequencing, by contrast, struggles to call InDels accurately in homopolymer runs and repetitive sequences, and it simply cannot detect most SVs above 50 bp at all. As genomes get more complex, that performance gap only widens.

But accuracy is not just about getting variants right—coverage completeness matters just as much. In head‑to‑head comparisons with whole‑genome bisulfite sequencing, or WGBS, which is still the gold standard for methylation analysis, HiFi picked up roughly 5.6 million more CpG sites, particularly in repetitive elements and regions where WGBS coverage tends to be low. Coverage is also significantly more uniform: over 90% of CpGs achieve ≥10× coverage with HiFi, compared to only about 65% with WGBS.

The reason is actually quite straightforward. Short reads cannot map uniquely to repetitive regions, so they simply get discarded during alignment. HiFi reads, on the other hand, are long enough to anchor uniquely in flanking unique sequence, carrying methylation information from previously inaccessible loci straight into the analysis.

The bottom line is this: you get clinical‑grade variant accuracy, complete coverage across the hardest‑to‑map regions, and integrated methylation profiling—all from a single HiFi run. No compromises, no separate assays, no data gaps.

4.Case Study 1: Unravelling Plant Chromatin Dynamics with Fiber-seq

Title: Single-molecule views of chromatin accessibility and structure during photomorphogenesis
Journal: Proceedings of the National Academy of Sciences (PNAS), 2025
Authors: Lei I, Guanyu Chen, Guangquan Zhu, et al.

 

Plants are constantly sensing and responding to their environment, and light is one of the most important signals they receive. It drives photosynthesis, sure—but it also triggers profound developmental changes. The transition from dark-grown to light-grown seedlings involves reprogramming thousands of genes. Until recently, however, the chromatin-level mechanisms behind this reprogramming were largely a black box.

The question was simple but technically brutal: how does light reshape chromatin architecture at singlemolecule resolution?

Previous methods could not give a complete answer. ATAC-seq and DNase-seq could map open chromatin, but they required fragmentation, lost phasing information, and could not capture DNA methylation on the same molecules. And in complex genomes like maize—where repeats make up over 80% of the genome—short-read methods simply fail to map in the most interesting regions.

The team applied Fiber-seq to Arabidopsis thaliana across five light-exposure time points and multiple mutants, as well as to maize (B73 and Jing724). The data gave them near-nucleotide resolution maps of chromatin accessibility, nucleosome positioning, and cytosine methylation—all from single DNA molecules.

What they found:

Light exposure drove significant, locusspecific changes in chromatin accessibility—both increases and decreases—particularly in genes involved in photosynthesis, hormone signalling, and development. The same loci showed coordinated changes in DNA methylation, pointing to an integrated regulatory response.

Mutant analysis (cop1-6, pifq, and hy5hyh) confirmed that classical light signalling pathways directly regulate chromatin accessibility. In one dataset, they could trace the path from genotype to chromatin state to phenotype.

Importantly, because Fiber-seq uses long reads, the team could profile DNA methylation in previously inaccessible repetitive regions—including 5S rRNA gene clusters and CEN180 satellite repeats. These regions had been essentially invisible to shortread methylation assays (Figure 1).

 

db92c400-6e1a-4e73-9fc4-0d27690e8a8f_看图王.jpg

 

Figure 1. Fiber‑seq enables high‑accuracy genome assembly and structural variant detection in complex, repeat‑rich genomes. Open chromatin regions and fine‑scale SVs identified in maize (B73) using Fiber‑seq. (Source: Li et al., PNAS 2025)

In maize, the results were even more striking. Fiber-seq picked up a broader range of biologically relevant open chromatin regions than short-read methods, enabling both high-accuracy de novo genome assembly and the detection of fine-scale structural variants that earlier approaches had missed.

What this study proved:

Fiber-seq is not just for model organisms. It works in complex, repeat-rich plant genomes. It delivers assembly-grade sequence data alongside chromatin and methylation information. And it reveals regulatory dynamics that short-read methods simply cannot see.

For plant researchers, this means a single assay can now replace the traditional workflow of separate genome assembly, methylation profiling, and chromatin mapping projects—compressing months of work into weeks, and eliminating the technical inconsistencies that come from comparing data across separate experiments.

5.Case Study 2: Solving a Rare Mendelian Disease with Synchronized Multi-omics

Title: Synchronized long-read genome, methylome, epigenome, and transcriptome profiling resolve a Mendelian condition
Journal: Nature Genetics, 2025
Authors: Vollger MR, Stergachis AB, et al.

Some patients remain undiagnosed despite years of testing. Their conditions are rare, their genomes are complex, and standard diagnostic pipelines—which rely heavily on shortread exome or genome sequencing—return negative results.

One such patient was a 9monthold infant with bilateral retinoblastomas, developmental delays, and other symptoms. Previous shortread sequencing had failed to identify a pathogenic variant. The clinical team suspected a complex genomic rearrangement, but standard methods could not resolve it.

The challenge: The patient's condition likely involved multiple molecular mechanisms—not just a single mutation—but no single assay could interrogate all of them simultaneously. Testing sequentially would have taken months and consumed precious sample.

The approach: Researchers applied synchronized Fiber-seq and Kinnex (MAS-seq) on the PacBio Revio system. This single workflow integrated longread genomic, methylomic, epigenomic, and transcriptomic data—all from a single patient sample.

A single X;13 balanced translocation was identified (Figure 2). But here is where the integrated multi-omics data proved transformative: the translocation disrupted four key genes, each through a different molecular mechanism (Table 1).

 

770b17fb-0a4c-46aa-93d2-33763f4ec8ae_看图王.jpg

 

Figure 2. Long‑read genome assembly precisely localises the X;13 translocation breakpoint. The breakpoint falls within intron 41 of NBEA, disrupting the gene and causing haploinsufficiency. (Source: Vollger et al., Nat Genet 2025)

 

Table 1. Summary of multi‑omic findings

Data Layer

Finding

Genome

Translocation breakpoint in NBEA intron 41, causing haploinsufficiency

Transcriptome

PDK3MAB21L1 fusion transcript with a 66aminoacid extension

Methylome

CpG hypermethylation silencing the fusion partner promoter

Epigenome

Enhancer adoption driving ectopic PDK3 expression

One patient. One sequencing run. Four distinct disease mechanisms (Figure 3).

 

67d2a665-2761-4958-a47d-270cbe00206b_看图王.jpg

 

Figure 3. Four genes, four distinct disease mechanisms — all resolved from a single Fiber‑seq + Kinnex run. The X;13 translocation disrupts NBEA (genome), generates a PDK3‑MAB21L1 fusion (transcriptome), silences the fusion partner promoter via CpG hypermethylation (methylome), and drives ectopic PDK3 expression through enhancer adoption (epigenome). (Source: Vollger et al., Nat Genet 2025)

Without the integrated approach, the fusion transcript would have been missed by DNA sequencing alone. The methylation silencing would have required a separate WGBS experiment. The enhancer adoption would have required ATAC-seq or ChIPseq. Each assay would have consumed additional sample, introduced batch effects, and taken weeks to complete. And critically, none of them would have revealed the complete picture because the findings are causally linked—they are not independent observations.

What this study proved:

Complex genetic diseases are often driven by multiple interacting mechanisms. Short-read sequencing alone cannot resolve them. But with synchronized long-read multi-omics, you can interrogate the genome, epigenome, and transcriptome in a single experiment—and identify the causal chain linking mutation to methylation to transcription to phenotype.

For researchers studying rare diseases, cancer, or any condition where genetic and epigenetic layers interact, this marks a paradigm shift. The era of sequential single-layer assays is over.

References:

1.Hammond N, et al. Analytical validation of germline small variant detection using longread HiFi genome sequencing. Genome Research. 2025.

2.Powered by PacBio: Selected publications from August 2025. PacBio Blog. 2025.

3.How pairing EpiCypher’s Fiberseq with HiFi sequencing delivers an allinone multiomic view. PacBio Blog. 2025.

4.Stergachis AB, et al. Synchronized longread genome, methylome, epigenome, and transcriptome profiling resolve a Mendelian condition. Nature Genetics. 2025.

5.Li I, Chen G, Zhu G, et al. Single-molecule views of chromatin accessibility and structure during photomorphogenesis. PNAS. 2025;122(48):e2516708122.

Contact Us

If you are interested in our long-read sequencing services or potential collaboration, please contact us. Our team is ready to support your research with tailored solutions. We also welcome feedback from users to help us improve our services.

Contact Us
%{tishi_zhanwei}%

Contact Us

E-mail:service@sailgene.com

WhatsApp:1-(617)-223-7544

Tel:16172237544

Email:service@sailgene.com

中企跨境-全域组件 制作前进入CSS配置样式

在线客服添加返回顶部

右侧在线客服样式 1,2,3 1

图片alt标题设置: SAILGENE TECHNOLOGY INC.

表单验证提示文本: Content cannot be empty!

循环体没有内容时: Sorry,no matching items were found.

CSS / JS 文件放置地

Welcome to leave an online message, we will contact you promptly

%{tishi_zhanwei}%