• Home page
  • Resources
  • BLOG
  • Why Plant Pan‑Genomics Is Essential for Breeding and Evolutionary Research

Why Plant Pan‑Genomics Is Essential for Breeding and Evolutionary Research


Release time:2026-08-15 18:32:03


1.Are You Really Ready for a Plant Pangenome Project?

Whether you are breeding a better watermelon or tracing the evolutionary history of a wild snapdragon lineage, a single reference genome is a biased lens.

Have you ever run into any of these situations?

  • You spent years building a high-quality reference genome for your crop, but when you ran GWAS to find the QTL for disease resistance, nothing showed up. Then you mapped the same data to a different cultivar’s reference — and the QTL suddenly appeared.

  • You generated a beautiful T2T genome, but when you used it to guide breeding, you found that many trait-relevant genes from elite varieties weren’t even assembled — because your reference came from a single genotype that doesn’t represent your breeding pool.

  • Your crop is an allopolyploid(like wheat, cotton, or oilseed rape). You ran your data through a standard diploid assembler, and the subgenomes collapsed into a mess — months of bioinformatics work down the drain.

 

If you keep relying on a single reference genome to represent an entire species, you are:

  • Wasting genetic resources— disease-resistance and stress-tolerance alleles in wild relatives get filtered out because they aren't in the reference.

  • Building markers that fail— molecular markers developed against one reference often don't work in other ecotypes or varieties.

  • Missing the real drivers of phenotypes— structural variants (SVs) are more likely than SNPs to influence domestication traits, but most SVs are invisible to short-read GWAS.

A simple fix: Before you commit to a large-scale project, do a stratified sampling design — map out your phylogeny, geography, and phenotype first, then decide which accessions to sequence, at what quality, and with what assembly strategy.

 

2.How Plant Pan‑Genomics Addresses These Challenges

A single reference genome is, by definition, a single individual. It cannot capture the genetic diversity of an entire species — especially in plants, where structural variation, polyploidy, and large repetitive regions are common.

A pan‑genome — the union of all genomic sequences from multiple individuals — solves this problem by:

Recovering missing genes and regulatory elements that are absent from the reference but present in wild relatives or other cultivars

Enabling comprehensive SV discovery across the entire species, not just within one genotype

Providing a graph‑based reference that allows any new sample to be mapped against the full diversity of the species, rather than against a single biased genome

The case studies below demonstrate how this approach has already transformed our understanding of crop genetics and breeding — and why the same principles apply to evolutionary research in any plant species.

 

3.A Four-Dimensional Strategic Framework for Plant Pan-Genomics

Dimension 1: Sample Size — How Many Accessions Are Enough?

There is no universal number. But you can estimate it using saturation curve analysis: randomly subsample your genomes in increasing numbers, plot pan-genome size against sample count, and see where the curve plateaus.

Methodological reference: The chicken pangenome study used 20 genomes and explicitly performed this analysis — their Figure 2 shows the core genome stabilising while the dispensable genome continues to grow slowly. For most outcrossing plant species, 15–30 well-chosen accessions capture >95% of the variable gene content.

 

Dimension 2: Stratified Sampling — Phylogeny + Geography + Phenotype

Random sampling is inefficient. Use a three-layer strategy:

  • Phylogenetic layer— select accessions that represent the major evolutionary lineages of your species or genus.

  • Geographic layer— cover the full native or cultivated distribution range.

  • Phenotypic layer— include extremes for your traits of interest (disease-resistant vs. susceptible, high-yield vs. low-yield).

The watermelon super-pangenome is the textbook example: the authors strategically selected 27 representative accessions based on phylogenetic relationships and the geographical distributions of 429 accessions, covering all seven species of the Citrullus genus (Figure 1). This wasn't random — it was deliberate, systematic, and reproducible.

3.jpg

Figure 1. 27 accessions: phylogeny + geography + phenotype

(Source: https://www.nature.com/articles/s41588-024-01823-6/figures/1)

 

Dimension 3: Wild Relatives — The Untapped Reservoir of Adaptive Alleles

Crop domestication creates genetic bottlenecks. Cultivated watermelons, even when collected from different geographical regions, typically exhibit low genetic diversity. If you only sequence cultivated varieties, you miss the adaptive alleles present in wild populations — especially for disease resistance and stress tolerance.

The watermelon study explicitly states that there has been a “shifted focus toward characterizing and using the genetic variations within watermelon's crop wild relatives (CWRs)”. They included six wild or semi-wild species and successfully identified multidisease-resistant loci from Citrullus amarus and Citrullus mucosospermus that were introduced into cultivated Citrullus lanatus.

The Brassica oleracea pangenome took a similar approach: 27 high-quality genomes representing all morphotypes and their wild relatives.

Dimension 4: Technology Choice — T2T vs. Haplotype-Resolved vs. Draft

Assembly Type

Best for

Example

T2T (telomere-to-telomere)

Small- to medium-sized genomes where complete structural resolution is critical

Grapevine (Nat Genet, 2024); Watermelon (Nat Genet, 2024)

Haplotype-resolved (phased)

Highly heterozygous diploids where allele-specific expression matters

Grapevine (phased T2T assemblies)

High-quality draft (HiFi/ONT)

Large, complex, or polyploid genomes where T2T is currently impractical

Many crop genomes >3 Gb

Key insight: You don't always need T2T. But when your traits are driven by complex SVs in repetitive regions — as in grapevine — T2T is a game-changer.

 

4.Case Study Gallery: Three Plant Pan-Genomes in Action

Case 1: Grapevine — T2T + SV-GWAS + Machine Learning Breeding

Paper: Liu et al., Nature Genetics, November 2024
DOI: 10.1038/s41588-024-01967-5

Sample Design

Assembly Strategy

Key Impact

18 newly generated phased T2T assemblies + 11 published assemblies = 29 haplotype-resolved genomes

Phased telomere-to-telomere (T2T)

Built Grapepan v.1.0, a graph-based pangenome reference

 

What they found:

  • A variation map with 9,105,787 short variants and 236,449 structural variations (SVs) from resequencing data of 466 grapevine cultivars (Figure 2)

  • 148 QTLs for 29 agronomic traits, of which 50.7% were newly identified

  • 12 traits significantly contributed by SVs

  • Estimated heritability improved by 22.78% on average when SVs were included

  • The MC-based pangenome (Grapepan v.1.0) reached 1.43 Gb, which is 2.88 times that of the PNT2T genome

 

22.jpg

Figure 2. PCA of 466 grape accessions

(Source: https://www.nature.com/articles/s41588-024-01967-5)

 

Why this matters for your project: This study proves that SVs are not rare exceptions — they are major contributors to complex traits. If your GWAS only looks at SNPs, you are leaving >20% of heritability on the table.

Case 2: Watermelon Super-Pangenome — All 7 Species of the Genus

Paper: Zhang et al., Nature Genetics, July 2024
DOI: 10.1038/s41588-024-01823-6

Sample Design

Assembly Strategy

Key Impact

27 distinct genotypes, encompassing all seven Citrullusspecies

Telomere-to-telomere (T2T) assemblies

Expanded the previous reference genome by 399.2 Mb and 11,225 genes

 

What they found:

  • Cultivated watermelons exhibit low genetic diversitydue to domestication bottlenecks

  • Multidisease-resistant loci from  amarus and C. mucosospermus were successfully introduced into cultivated C. lanatus

  • SVs in  lanatuswere inherited not only from cordophanus but also from C. mucosospermus, suggesting additional ancestors beyond the previously recognised single origin

  • The super-pangenome covers 768.5 Mb and 32,513 gene families— 1.5 times the size of a single watermelon genome

Why this matters for your project: If your crop has wild relatives, you cannot afford to ignore them. The disease-resistance alleles that matter most for breeding are often locked in wild species, not in your cultivated reference.

Case 3: Brassica oleracea — SVs as Bidirectional Regulators of Gene Expression

Paper: Li et al., Nature Genetics, February 2024
DOI: 10.1038/s41588-024-01655-4

Sample Design

Assembly Strategy

Key Impact

27 high-quality genomes representing all morphotypes (cabbage, broccoli, cauliflower, kale, Brussels sprouts, kohlrabi, Chinese kale) and their wild relatives (Figure 3)

PacBio / ONT + Illumina

First evidence that SVs act as bidirectional dosage regulators of gene expression

 

 c773c728-d39c-498a-87a0-e37dc53bf391_看图王.jpg

Figure 3. Phylogenetic tree of 704 accessions showing all morphotypes

(Source: https://www.nature.com/articles/s41588-024-01655-4/figures/1)

 

What they found:

  • SVs exert bidirectional effectson gene expression — suppressing through DNA methylation or promoting by harbouring transcription factor-binding elements

  • Specific examples:

    • SVs promoting BoPNY and suppressing BoCKX3 in cauliflower/broccoli

    • SVs suppressing BoKAN1and BoACS4 in cabbage

    • SVs promoting BoMYBtfin ornamental kale

  • Phylogenetic analysis using SNPs classified the 704 accessions into three main groups (Figure 1a)

  • Retrotransposons (Copia and Gypsy) have been continuously expanding in all genomes since four million years ago (Figure 1d)

Why this matters for your project: SVs are not just “neutral” structural changes — they are active drivers of gene regulation. If your reference genome doesn't capture them, you are missing the regulatory logic behind your crop's most important traits.

 

5.Decision Matrix: Which Strategy Fits Your Crop?

Your Crop Type

Recommended Sample Size

Must-Include Groups

Recommended Assembly

Key Risk to Avoid

Inbreeding crop with narrow germplasm (e.g., rice, wheat)

15–20 accessions

Wild relatives + landraces

T2T if ≤3 Gb; otherwise HiFi draft

Ignoring wild gene pool

Outcrossing crop with high diversity (e.g., maize, grape)

20–30 accessions

Diverse ecotypes + wild species

Phased T2T (for SV resolution)

Under-sampling phenotypic extremes

Polyploid crop (e.g., wheat, cotton, oilseed rape)

25–40 accessions

Representatives of each subgenome origin

HiFi/ONT + ploidy-aware assemblers

Using diploid assemblers on polyploid data

Orphan / understudied crop

5–10 (pilot) + expand

Distinct ecotypes from major growing regions

Long-read draft + short-read validation

Scaling up before saturation analysis

 

6.FAQ: Plant Pan-Genomics

Q1: My crop is an allopolyploid. Do I need to sample differently?
A: Yes. You need accessions that represent different subgenome origins. Allotetraploids (like cotton) and autotetraploids (like potato) have very different k-mer distribution patterns. Use smudgeplot or GenomeScope2 to estimate ploidy before committing to assembly.

Q2: Can I build a plant pan-genome with short reads alone?
A: No. Short reads cannot resolve complex SVs and repetitive regions. All three case studies above used long-read (HiFi/ONT) technologies as the primary assembly platform. Short reads are useful for validation and population-scale SV-GWAS — but only after you have a high-quality pan-genome reference.

Q3: My budget is limited. How many accessions should I start with?
A: Start with a pilot of 5–10 accessions at 10–15× coverage to run a saturation analysis. The duck pan-genome (5 genomes) and the chicken pan-genome (20 genomes) both show that even modest sample sizes can deliver high-impact results when accessions are strategically chosen.

Q4: How do I know if my crop needs T2T assembly?
A: Ask yourself: Are your traits of interest controlled by SVs in repetitive or subtelomeric regions? If yes (as in grapevine), T2T is worth the investment. If your genome is >3 Gb or highly polyploid, a high-quality haplotype-resolved draft may be more cost-effective.

Q5: I'm not a breeder — I study evolution. Does this guide apply to me?
A: Yes. The strategic framework (sample size estimation, stratified sampling, wild relative inclusion, and technology choice) is species-agnostic. Whether you are studying domestication in Brassicaor speciation in a wild orchid, the same principles apply — a single genome cannot capture the evolutionary history of a species.

 

7.Sailgene's Plant PanGenomics Service Package

At Sailgene, we offer a fully integrated Plant PanGenomics Service that covers every strategic decision discussed above:

Phase

Deliverables

Preproject consultation

Literature review, diversity assessment, ploidy test (flow cytometry), saturation analysis, customised sampling plan

Sequencing

HiFi / ONT long reads + Illumina short reads (as needed) for optimal SV discovery and genome assembly

Assembly & annotation

T2T or haplotyperesolved assemblies, gene annotation, repeat masking

Pangenome graph construction

Multiple alignment, graphbased reference, structural variant (SV) discovery

Downstream analysis

SVGWAS, selection scans, comparative genomics, transcriptomics integration

Report & training

Comprehensive PDF report + raw data + optional bioinformatics training

 

Our capabilities include:

  • Longread sequencing (HiFi / ONT) for complex plant genomes

  • Telomeretotelomere (T2T) genome assembly

  • Phased / haplotyperesolved assemblies for highly heterozygous species

  • Structural variant (SV) analysis and integration with GWAS

  • Pangenome graph construction and comparative genomics

Turnaround: 4–8 weeks (depending on genome size and ploidy)

 

Get Your Plant PanGenome Roadmap Today
Contact us for a free consultation and a customized quote.

Reference:

  1. Liu, Z., Wang, N., Su, Y. et al. Grapevine pangenome facilitates trait genetics and genomic breeding. Nat Genet56, 2804–2814 (2024). https://doi.org/10.1038/s41588-024-01967-5.

  2. Zhang, Y., Zhao, M., Tan, J. et al. Telomere-to-telomere Citrullus super-pangenome provides direction for watermelon breeding. Nat Genet56, 1750–1761 (2024). https://doi.org/10.1038/s41588-024-01823-6

  3. Li, X., Wang, Y., Cai, C. et al.Large-scale gene expression alterations introduced by structural variation drive morphotype diversification in Brassica oleracea. Nat Genet 56, 517–529 (2024). https://doi.org/10.1038/s41588-024-01655-4 

 

 

Contact Us

If you are interested in our long-read sequencing services or potential collaboration, please contact us. Our team is ready to support your research with tailored solutions. We also welcome feedback from users to help us improve our services.

Contact Us
%{tishi_zhanwei}%

Contact Us

E-mail:service@sailgene.com

Inquiry

Tel:16172237544

Email:service@sailgene.com

中企跨境-全域组件 制作前进入CSS配置样式

在线客服添加返回顶部

右侧在线客服样式 1,2,3 1

图片alt标题设置: SAILGENE TECHNOLOGY INC.

表单验证提示文本: Content cannot be empty!

循环体没有内容时: Sorry,no matching items were found.

CSS / JS 文件放置地

Welcome to leave an online message, we will contact you promptly

%{tishi_zhanwei}%