Sailgene Technology
From 20 kb to 200 kb: How ONT Read Length Shapes Genome Resolution
Release time:2026-09-20 11:07:17
Long-read sequencing has transformed genome research by enabling direct observation of genomic structures that were previously difficult to resolve. However, as sequencing technologies continue to improve, researchers are facing a more practical question:
How long do sequencing reads need to be for a specific biological question?
Unlike sequencing coverage, which increases confidence through repeated observations, read length determines whether a single sequencing molecule can physically connect genomic regions that are separated by repeats, structural rearrangements, or haplotype differences.
For genome assembly, structural variant analysis, haplotype phasing, and repetitive region characterization, the appropriate read-length strategy is therefore not determined by maximizing sequencing output alone. Instead, it depends on whether individual reads are long enough to bridge the genomic structures researchers aim to resolve.
Understanding Different Read-Length Strategies: From 20 kb to 200 kb
Long-read sequencing does not have a single optimal read length for all research applications. The required read-length distribution depends on the genomic complexity, the size of the structures being resolved, and the level of biological resolution required.
A read-length strategy around 20 kb can already provide substantial advantages over short-read sequencing for many genome analysis projects. Reads at this scale can improve assembly continuity, support genome annotation, and enable detection of many small- to medium-scale structural variants.
Increasing read lengths toward the 50 kb range provides additional genomic context, improving the ability to bridge repetitive regions, connect heterozygous variants, and generate more contiguous assemblies, particularly for genomes with moderate complexity.
When researchers aim to resolve larger repetitive regions, complex structural variants, or chromosome-scale genome organization, read lengths approaching 100 kb provide greater connectivity. These longer molecules can span genomic regions that remain fragmented in conventional long-read assemblies and are particularly useful for haplotype phasing and challenging genome reconstruction.
For the most complex genome projects, read lengths around 150 kb can further increase the probability of spanning difficult genomic regions. Under appropriate sample conditions with high-quality high-molecular-weight DNA, ultra-long molecules can improve continuity across large repeats, centromeric regions, segmental duplications, and other challenging genomic structures.
At the extreme end, reads approaching 200 kb represent highly optimized ultra-long sequencing strategies that are typically considered for projects requiring maximum molecule-level connectivity. These reads are not required for every project, but they can provide unique advantages when the biological question depends on connecting genomic information across exceptionally long distances, such as resolving large structural rearrangements, complex haplotypes, or highly repetitive chromosome regions.
Importantly, these values should not be interpreted as strict thresholds. A genome does not become “solvable” at 100 kb but impossible at 50 kb. Instead, increasing read length gradually increases the probability that individual sequencing molecules will span the genomic structures researchers aim to resolve.
From Long Reads to Ultra-Long Reads: What Changes?
The value of ultra-long reads was demonstrated in human genome assembly studies. Jain et al. generated ultra-long ONT reads with a read N50 of approximately 99.7 kb, and adding approximately 5× ultra-long sequencing coverage on top of existing long-read datasets substantially improved assembly continuity. Importantly, the ultra-long reads complemented conventional long-read sequencing rather than replacing it as an independent low-coverage strategy (Jain et al., 2018).
More recent benchmarking studies have further shown that read-length distribution can influence structural variant detection. A recent study using ONT datasets with different read-length distributions demonstrated that datasets enriched for ultra-long reads improved recovery of complex structural variants compared with standard long-read distributions, highlighting that not only sequencing depth but also molecule length affects variant resolution (Hemker et al., 2026).

Figure 1: Ultra-long data finds the most comprehensive set of structural variants.
- Read-length histograms. b) The total structural variant counts. c) The relative proportions of variant types
(Source: https://doi.org/10.1093/g3journal/jkag043)
A recent study using Oxford Nanopore ultra-long sequencing demonstrated the value of extended read lengths for plant genome completion. By optimizing DNA extraction and sequencing strategies, researchers generated ultra-long reads with average read-length N50 above 80 kb and maximum reads exceeding several hundred kilobases. These long molecules helped resolve remaining assembly gaps and improved the feasibility of telomere-to-telomere genome reconstruction in complex plant genomes (Zhang et al., 2024).
This example highlights that ultra-long sequencing is not simply about generating longer reads. The additional genomic span provides direct biological information by connecting regions that are otherwise separated by repetitive sequences, making it particularly valuable for chromosome-level assembly, haplotype reconstruction, and difficult genomic regions.

Figure 2: Advancements in ONT sequencing improve read lengths.
(Source: https://doi.org/10.1016/j.molp.2024.10.008)
Read Length Is a Distribution, Not a Single Number
Although read-length targets such as 20 kb, 50 kb, 100 kb, 150 kb, and 200 kb are commonly used to describe long-read sequencing strategies, the final dataset is not defined by a single read length.
Read N50, a commonly used metric summarizing the read-length distribution, represents the read length at which 50% of the total sequenced bases are contained in reads of that length or longer. Read N50 provides a summary of sequencing read distribution, but it does not describe how many reads reach extremely long lengths. Two datasets may have similar read N50 values but different proportions of reads above 100 kb or 200 kb, which may lead to different assembly and variant-resolution capabilities.
Therefore, selecting a sequencing strategy should consider not only the average performance metric, but also the complete read-length distribution and the biological structures that need to be resolved.
Sample Quality Defines the Upper Limit of Read Length
Long-read sequencing performance ultimately depends on the physical integrity of DNA molecules.
Generating ultra-long reads requires high-quality, high-molecular-weight DNA because sequencing cannot recover information from molecules that have already been fragmented during extraction or handling. Even with advanced sequencing platforms, poor DNA integrity will limit the achievable read length.
For this reason, sample preparation is not a separate step from sequencing strategy. It is a critical factor determining whether a project can benefit from longer read approaches.
Supporting Advanced Long-Read Sequencing with Sailgene
Sailgene provides ONT long-read sequencing solutions designed around different genome research objectives. This includes optimization of DNA extraction, sequencing design, and bioinformatics analysis to balance read-length performance, sequencing coverage, and project objectives.
For genome assembly, structural variant analysis, and haplotype-resolved projects, Sailgene supports long-read sequencing strategies optimized for generating the required genomic continuity. For challenging genomes requiring additional long-range information, ultra-long sequencing approaches can be considered when high-quality DNA is available.
By integrating sequencing optimization with bioinformatics analysis, Sailgene helps researchers select an appropriate long-read strategy to address increasingly complex genomic questions.

Figure 3: Representative ultra-long sequencing results across multiple species. The datasets show estimated read N50 values above 100 kb, demonstrating Sailgene's ultra-long sequencing capability across diverse biological samples.
FAQ
- Does higher sequencing coverage replace longer reads?
No. Sequencing coverage and read length solve different problems. Coverage improves confidence in observed sequences, while longer reads provide additional genomic connections.
- Is the highest read N50 always the best choice?
Not necessarily. The optimal read-length distribution depends on genome complexity and the biological objective. Extremely long reads are most valuable when researchers need to resolve large genomic structures.
- When should researchers consider ultra-long sequencing?
Ultra-long sequencing is particularly useful for complex genome assembly, repetitive regions, large structural variants, and haplotype-resolved reconstruction where longer physical connections are required.
- Does achieving longer reads always require more sequencing output?
No. Longer reads depend primarily on DNA molecule integrity rather than sequencing volume alone. Generating ultra-long reads requires high-quality HMW DNA extraction and optimized sample handling.
- Can ultra-long DNA extracted for ONT sequencing be directly used for standard long-read sequencing?
No. Ultra-long DNA extraction aims to preserve extremely long DNA molecules, while standard ONT sequencing requires a suitable fragment-size distribution and sequencing yield. Therefore, ultra-long DNA typically requires additional size selection or purification steps before being used for standard long-read sequencing.
- How are ONT sequencing data filtered for quality control?
ONT sequencing data quality filtering is typically performed based on the Phred quality score (Q score), which depends on the sequencing chemistry and basecalling model.
Common thresholds include:
R9 chemistry: HAC ≥ Q7; SUP ≥ Q9
R10 chemistry: HAC ≥ Q9; SUP ≥ Q10
The final dataset usually contains both pass and fail reads, with pass reads used for downstream analysis.
- If a T2T reference genome is available, can standard ONT resequencing determine the true telomere length?
No. Telomeres consist of highly repetitive tandem sequences, and their actual repeat copy number can vary between cells. Mapping ONT reads to an existing T2T reference can confirm sequence similarity, but it cannot determine whether the telomere is completely resolved or measure the exact telomere length without dedicated analysis strategies.
- Does increasing read length always improve genome assembly quality?
Not necessarily. Longer reads increase the probability of spanning complex genomic regions, such as repeats, segmental duplications, and structural variants. However, assembly quality also depends on genome complexity, sequencing coverage, DNA quality, and complementary data such as Hi-C. The optimal read-length strategy should match the biological question rather than simply maximize read length.
- Can ultra-long ONT reads replace Hi-C sequencing?
No. Ultra-long reads improve contig continuity by resolving difficult genomic regions, but they do not provide chromosome-scale ordering information. Hi-C captures three-dimensional chromatin interactions and is still commonly required to scaffold assembled contigs into chromosome-level assemblies.
- Does a longer ONT read mean higher sequencing accuracy?
No. Read length and base accuracy are independent parameters. Longer reads provide more genomic connectivity, while sequencing accuracy depends on nanopore chemistry, basecalling algorithms, and downstream polishing strategies.
- Should researchers always choose the longest possible ONT reads?
Not necessarily. Different biological questions require different read-length strategies. For example, ~20 kb reads may be sufficient for many genome surveys and structural variant studies, while 100–200 kb ultra-long reads are more valuable for highly repetitive genomes, complex structural variants, h2aplotype resolution, and T2T-level assembly.
References
Jain, M. et al. (2018).Nanopore sequencing and assembly of a human genome with ultra-long reads.
Nature Biotechnology, 36, 338–345.
https://doi.org/10.1038/nbt.4060
Logsdon, G.A., Vollger, M.R. and Eichler, E.E. (2020).Long-read human genome sequencing and its applications.Nature Reviews Genetics, 21, 597–614.
https://doi.org/10.1038/s41576-020-0236-x
Hemker, J.A., Gellert, H.R., Smiley-Rhodes, J.A., Kim, B.Y. and Petrov, D.A. (2026)Manual validation finds ultra-long-read sequencing best enables faithful, population-level structural variant calling in Drosophila melanogaster euchromatin with nanopore. G3 Genes|Genomes|Genetics, Volume 16, Issue 5, May
https://doi.org/10.1093/g3journal/jkag043
Zhang, X. et al. (2024).Nanopore ultra-long sequencing and adaptive sampling spur plant complete telomere-to-telomere genome assembly.Molecular Plant, 17(11), 1773–1786.DOI: https://doi.org/10.1016/j.molp.2024.10.008
Previous Page
Previous Page
Contact Us
If you are interested in our long-read sequencing services or potential collaboration, please contact us. Our team is ready to support your research with tailored solutions. We also welcome feedback from users to help us improve our services.
Contact Us
Add: One Innovation Drive, Suite B3-406, Worcester, MA 01605, USA
在线客服添加返回顶部
右侧在线客服样式 1,2,3 1
图片alt标题设置: SAILGENE TECHNOLOGY INC.
表单验证提示文本: Content cannot be empty!
循环体没有内容时: Sorry,no matching items were found.
CSS / JS 文件放置地