Sailgene Technology
Why Plant Pan‑Genomics Is Essential for Breeding and Evolutionary Research
Release time:2026-08-15 18:32:03
1.Are You Really Ready for a Plant Pangenome Project?
Whether you are breeding a better watermelon or tracing the evolutionary history of a wild snapdragon lineage, a single reference genome is a biased lens.
Have you ever run into any of these situations?
-
You spent years building a high-quality reference genome for your crop, but when you ran GWAS to find the QTL for disease resistance, nothing showed up. Then you mapped the same data to a different cultivar’s reference — and the QTL suddenly appeared.
-
You generated a beautiful T2T genome, but when you used it to guide breeding, you found that many trait-relevant genes from elite varieties weren’t even assembled — because your reference came from a single genotype that doesn’t represent your breeding pool.
-
Your crop is an allopolyploid(like wheat, cotton, or oilseed rape). You ran your data through a standard diploid assembler, and the subgenomes collapsed into a mess — months of bioinformatics work down the drain.
If you keep relying on a single reference genome to represent an entire species, you are:
-
Wasting genetic resources— disease-resistance and stress-tolerance alleles in wild relatives get filtered out because they aren't in the reference.
-
Building markers that fail— molecular markers developed against one reference often don't work in other ecotypes or varieties.
-
Missing the real drivers of phenotypes— structural variants (SVs) are more likely than SNPs to influence domestication traits, but most SVs are invisible to short-read GWAS.
A simple fix: Before you commit to a large-scale project, do a stratified sampling design — map out your phylogeny, geography, and phenotype first, then decide which accessions to sequence, at what quality, and with what assembly strategy.
2.How Plant Pan‑Genomics Addresses These Challenges
A single reference genome is, by definition, a single individual. It cannot capture the genetic diversity of an entire species — especially in plants, where structural variation, polyploidy, and large repetitive regions are common.
A pan‑genome — the union of all genomic sequences from multiple individuals — solves this problem by:
Recovering missing genes and regulatory elements that are absent from the reference but present in wild relatives or other cultivars
Enabling comprehensive SV discovery across the entire species, not just within one genotype
Providing a graph‑based reference that allows any new sample to be mapped against the full diversity of the species, rather than against a single biased genome
The case studies below demonstrate how this approach has already transformed our understanding of crop genetics and breeding — and why the same principles apply to evolutionary research in any plant species.
3.A Four-Dimensional Strategic Framework for Plant Pan-Genomics
Dimension 1: Sample Size — How Many Accessions Are Enough?
There is no universal number. But you can estimate it using saturation curve analysis: randomly subsample your genomes in increasing numbers, plot pan-genome size against sample count, and see where the curve plateaus.
Methodological reference: The chicken pangenome study used 20 genomes and explicitly performed this analysis — their Figure 2 shows the core genome stabilising while the dispensable genome continues to grow slowly. For most outcrossing plant species, 15–30 well-chosen accessions capture >95% of the variable gene content.
Dimension 2: Stratified Sampling — Phylogeny + Geography + Phenotype
Random sampling is inefficient. Use a three-layer strategy:
-
Phylogenetic layer— select accessions that represent the major evolutionary lineages of your species or genus.
-
Geographic layer— cover the full native or cultivated distribution range.
-
Phenotypic layer— include extremes for your traits of interest (disease-resistant vs. susceptible, high-yield vs. low-yield).
The watermelon super-pangenome is the textbook example: the authors strategically selected 27 representative accessions based on phylogenetic relationships and the geographical distributions of 429 accessions, covering all seven species of the Citrullus genus (Figure 1). This wasn't random — it was deliberate, systematic, and reproducible.

Figure 1. 27 accessions: phylogeny + geography + phenotype
(Source: https://www.nature.com/articles/s41588-024-01823-6/figures/1)
Dimension 3: Wild Relatives — The Untapped Reservoir of Adaptive Alleles
Crop domestication creates genetic bottlenecks. Cultivated watermelons, even when collected from different geographical regions, typically exhibit low genetic diversity. If you only sequence cultivated varieties, you miss the adaptive alleles present in wild populations — especially for disease resistance and stress tolerance.
The watermelon study explicitly states that there has been a “shifted focus toward characterizing and using the genetic variations within watermelon's crop wild relatives (CWRs)”. They included six wild or semi-wild species and successfully identified multidisease-resistant loci from Citrullus amarus and Citrullus mucosospermus that were introduced into cultivated Citrullus lanatus.
The Brassica oleracea pangenome took a similar approach: 27 high-quality genomes representing all morphotypes and their wild relatives.
Dimension 4: Technology Choice — T2T vs. Haplotype-Resolved vs. Draft
|
Assembly Type |
Best for |
Example |
|
T2T (telomere-to-telomere) |
Small- to medium-sized genomes where complete structural resolution is critical |
Grapevine (Nat Genet, 2024); Watermelon (Nat Genet, 2024) |
|
Haplotype-resolved (phased) |
Highly heterozygous diploids where allele-specific expression matters |
Grapevine (phased T2T assemblies) |
|
High-quality draft (HiFi/ONT) |
Large, complex, or polyploid genomes where T2T is currently impractical |
Many crop genomes >3 Gb |
Key insight: You don't always need T2T. But when your traits are driven by complex SVs in repetitive regions — as in grapevine — T2T is a game-changer.
4.Case Study Gallery: Three Plant Pan-Genomes in Action
Case 1: Grapevine — T2T + SV-GWAS + Machine Learning Breeding
Paper: Liu et al., Nature Genetics, November 2024
DOI: 10.1038/s41588-024-01967-5
|
Sample Design |
Assembly Strategy |
Key Impact |
|
18 newly generated phased T2T assemblies + 11 published assemblies = 29 haplotype-resolved genomes |
Phased telomere-to-telomere (T2T) |
Built Grapepan v.1.0, a graph-based pangenome reference |
What they found:
-
A variation map with 9,105,787 short variants and 236,449 structural variations (SVs) from resequencing data of 466 grapevine cultivars (Figure 2)
-
148 QTLs for 29 agronomic traits, of which 50.7% were newly identified
-
12 traits significantly contributed by SVs
-
Estimated heritability improved by 22.78% on average when SVs were included
-
The MC-based pangenome (Grapepan v.1.0) reached 1.43 Gb, which is 2.88 times that of the PNT2T genome

Figure 2. PCA of 466 grape accessions
(Source: https://www.nature.com/articles/s41588-024-01967-5)
Why this matters for your project: This study proves that SVs are not rare exceptions — they are major contributors to complex traits. If your GWAS only looks at SNPs, you are leaving >20% of heritability on the table.
Case 2: Watermelon Super-Pangenome — All 7 Species of the Genus
Paper: Zhang et al., Nature Genetics, July 2024
DOI: 10.1038/s41588-024-01823-6
|
Sample Design |
Assembly Strategy |
Key Impact |
|
27 distinct genotypes, encompassing all seven Citrullusspecies |
Telomere-to-telomere (T2T) assemblies |
Expanded the previous reference genome by 399.2 Mb and 11,225 genes |
What they found:
-
Cultivated watermelons exhibit low genetic diversitydue to domestication bottlenecks
-
Multidisease-resistant loci from amarus and C. mucosospermus were successfully introduced into cultivated C. lanatus
-
SVs in lanatuswere inherited not only from cordophanus but also from C. mucosospermus, suggesting additional ancestors beyond the previously recognised single origin
-
The super-pangenome covers 768.5 Mb and 32,513 gene families— 1.5 times the size of a single watermelon genome
Why this matters for your project: If your crop has wild relatives, you cannot afford to ignore them. The disease-resistance alleles that matter most for breeding are often locked in wild species, not in your cultivated reference.
Case 3: Brassica oleracea — SVs as Bidirectional Regulators of Gene Expression
Paper: Li et al., Nature Genetics, February 2024
DOI: 10.1038/s41588-024-01655-4
|
Sample Design |
Assembly Strategy |
Key Impact |
|
27 high-quality genomes representing all morphotypes (cabbage, broccoli, cauliflower, kale, Brussels sprouts, kohlrabi, Chinese kale) and their wild relatives (Figure 3) |
PacBio / ONT + Illumina |
First evidence that SVs act as bidirectional dosage regulators of gene expression |

Figure 3. Phylogenetic tree of 704 accessions showing all morphotypes
(Source: https://www.nature.com/articles/s41588-024-01655-4/figures/1)
What they found:
-
SVs exert bidirectional effectson gene expression — suppressing through DNA methylation or promoting by harbouring transcription factor-binding elements
-
Specific examples:
-
-
SVs promoting BoPNY and suppressing BoCKX3 in cauliflower/broccoli
-
SVs suppressing BoKAN1and BoACS4 in cabbage
-
SVs promoting BoMYBtfin ornamental kale
-
-
Phylogenetic analysis using SNPs classified the 704 accessions into three main groups (Figure 1a)
-
Retrotransposons (Copia and Gypsy) have been continuously expanding in all genomes since four million years ago (Figure 1d)
Why this matters for your project: SVs are not just “neutral” structural changes — they are active drivers of gene regulation. If your reference genome doesn't capture them, you are missing the regulatory logic behind your crop's most important traits.
5.Decision Matrix: Which Strategy Fits Your Crop?
|
Your Crop Type |
Recommended Sample Size |
Must-Include Groups |
Recommended Assembly |
Key Risk to Avoid |
|
Inbreeding crop with narrow germplasm (e.g., rice, wheat) |
15–20 accessions |
Wild relatives + landraces |
T2T if ≤3 Gb; otherwise HiFi draft |
Ignoring wild gene pool |
|
Outcrossing crop with high diversity (e.g., maize, grape) |
20–30 accessions |
Diverse ecotypes + wild species |
Phased T2T (for SV resolution) |
Under-sampling phenotypic extremes |
|
Polyploid crop (e.g., wheat, cotton, oilseed rape) |
25–40 accessions |
Representatives of each subgenome origin |
HiFi/ONT + ploidy-aware assemblers |
Using diploid assemblers on polyploid data |
|
Orphan / understudied crop |
5–10 (pilot) + expand |
Distinct ecotypes from major growing regions |
Long-read draft + short-read validation |
Scaling up before saturation analysis |
6.FAQ: Plant Pan-Genomics
Q1: My crop is an allopolyploid. Do I need to sample differently?
A: Yes. You need accessions that represent different subgenome origins. Allotetraploids (like cotton) and autotetraploids (like potato) have very different k-mer distribution patterns. Use smudgeplot or GenomeScope2 to estimate ploidy before committing to assembly.
Q2: Can I build a plant pan-genome with short reads alone?
A: No. Short reads cannot resolve complex SVs and repetitive regions. All three case studies above used long-read (HiFi/ONT) technologies as the primary assembly platform. Short reads are useful for validation and population-scale SV-GWAS — but only after you have a high-quality pan-genome reference.
Q3: My budget is limited. How many accessions should I start with?
A: Start with a pilot of 5–10 accessions at 10–15× coverage to run a saturation analysis. The duck pan-genome (5 genomes) and the chicken pan-genome (20 genomes) both show that even modest sample sizes can deliver high-impact results when accessions are strategically chosen.
Q4: How do I know if my crop needs T2T assembly?
A: Ask yourself: Are your traits of interest controlled by SVs in repetitive or subtelomeric regions? If yes (as in grapevine), T2T is worth the investment. If your genome is >3 Gb or highly polyploid, a high-quality haplotype-resolved draft may be more cost-effective.
Q5: I'm not a breeder — I study evolution. Does this guide apply to me?
A: Yes. The strategic framework (sample size estimation, stratified sampling, wild relative inclusion, and technology choice) is species-agnostic. Whether you are studying domestication in Brassicaor speciation in a wild orchid, the same principles apply — a single genome cannot capture the evolutionary history of a species.
7.Sailgene's Plant PanGenomics Service Package
At Sailgene, we offer a fully integrated Plant PanGenomics Service that covers every strategic decision discussed above:
|
Phase |
Deliverables |
|
Preproject consultation |
Literature review, diversity assessment, ploidy test (flow cytometry), saturation analysis, customised sampling plan |
|
Sequencing |
HiFi / ONT long reads + Illumina short reads (as needed) for optimal SV discovery and genome assembly |
|
Assembly & annotation |
T2T or haplotyperesolved assemblies, gene annotation, repeat masking |
|
Pangenome graph construction |
Multiple alignment, graphbased reference, structural variant (SV) discovery |
|
Downstream analysis |
SVGWAS, selection scans, comparative genomics, transcriptomics integration |
|
Report & training |
Comprehensive PDF report + raw data + optional bioinformatics training |
Our capabilities include:
-
Longread sequencing (HiFi / ONT) for complex plant genomes
-
Telomeretotelomere (T2T) genome assembly
-
Phased / haplotyperesolved assemblies for highly heterozygous species
-
Structural variant (SV) analysis and integration with GWAS
-
Pangenome graph construction and comparative genomics
Turnaround: 4–8 weeks (depending on genome size and ploidy)
Get Your Plant PanGenome Roadmap Today
Contact us for a free consultation and a customized quote.
Reference:
-
Liu, Z., Wang, N., Su, Y. et al. Grapevine pangenome facilitates trait genetics and genomic breeding. Nat Genet56, 2804–2814 (2024). https://doi.org/10.1038/s41588-024-01967-5.
-
Zhang, Y., Zhao, M., Tan, J. et al. Telomere-to-telomere Citrullus super-pangenome provides direction for watermelon breeding. Nat Genet56, 1750–1761 (2024). https://doi.org/10.1038/s41588-024-01823-6
-
Li, X., Wang, Y., Cai, C. et al.Large-scale gene expression alterations introduced by structural variation drive morphotype diversification in Brassica oleracea. Nat Genet 56, 517–529 (2024). https://doi.org/10.1038/s41588-024-01655-4
Contact Us
If you are interested in our long-read sequencing services or potential collaboration, please contact us. Our team is ready to support your research with tailored solutions. We also welcome feedback from users to help us improve our services.
Contact Us
Add: One Innovation Drive, Suite B3-406, Worcester, MA 01605, USA
在线客服添加返回顶部
右侧在线客服样式 1,2,3 1
图片alt标题设置: SAILGENE TECHNOLOGY INC.
表单验证提示文本: Content cannot be empty!
循环体没有内容时: Sorry,no matching items were found.
CSS / JS 文件放置地