1 Project Overview

1.1 Project Information

1.2 Project Background

1)Species name: XXX;

2)Estimated gene count in contract: None;

3)Analysis plan: genome structural annotation and functional annotation.


2 Analysis Summary

The genome is the total set of genetic material of an organism. Genes are indispensable to all life—they store information for processes such as growth, development and apoptosis. After genome assembly, the sequences must be interpreted, including coding and noncoding annotations as well as functional annotation. These analyses allow us to further elucidate genome structure and characteristics.


3 Technical Route

Genome annotation uses bioinformatics methods and tools—combining de novo prediction, homology-based prediction, and structural definition—to search for and define sequence elements, and to determine their positions, sequences, structures and functions. The resulting functional element information is written into genome annotation files (e.g., GFF) for downstream analyses such as functional‑gene analysis, comparative genomics, and target‑gene analyses.


4 Analysis Pipeline

Genome annotation mainly consists of three parts: (1) identification/annotation of repetitive sequences; (2) prediction of noncoding RNAs; (3) prediction of protein‑coding gene structures and their functional annotation.

survey pipeline

Figure 1 Analysis pipeline


5 Results

5.1 Preview of Conclusions

1)Repetitive sequence proportion: 52.42%;

2)Number of coding genes: 57,690;

3)BUSCO used the embryophyta_odb10 database; assessment result: 99.4%.

5.2 Repeat Annotation

We combined two strategies: (i) homology‑based annotation against public databases using RepeatMasker; and (ii) de novo annotation using LTR_retriever and RepeatModeler. Results from both were integrated.

Table 1 Basic statistics of interspersed repeats

TE protiens
De novo + repbase
Combined TEs
Type Length (Bp) % in genome Length (Bp) % in genome Length (Bp) % in genome
DNA 5,722,587 1.44 94,743,409 23.78 94,781,995 23.79
LINE 1,730,281 0.43 11,701,364 2.94 11,750,839 2.95
SINE 0 0.00 1,148,810 0.29 1,148,810 0.29
LTR 22,930,977 5.76 96,459,928 24.21 97,966,614 24.59
LTR-Gypsy 16,905,393 4.24 60,098,481 15.09 61,701,406 15.49
LTR-Copia 5,723,048 1.44 12,108,003 3.04 12,820,990 3.22
Other 0 0.00 820 0.00 820 0.00
Unknown 0 0.00 5,867,447 1.47 5,867,447 1.47
Total 30,383,520 7.63 207,266,692 52.03 208,831,911 52.42

Note: (1) DNA: DNA transposons, including MITEs, Helitrons, etc.; (2) LINE (long interspersed nuclear element): non‑LTR retrotransposons such as L1, R2, Jockey; (3) SINE (short interspersed element): non‑LTR retrotransposons such as V‑SINE, AmnSINE, CORE‑SINE; (4) LTR (long terminal repeat): retrotransposons (mainly Gypsy and Copia). They do not encode proteins but contain regulatory elements such as promoters and enhancers; (5) Other: unclassified DNA transposons (e.g., Mirage, P‑element, Transib); (6) Unknown: marked but unclassified repeats.

Tandem repeats were predicted with TRF and MISA; results are as follows:

Table 2 Tandem repeat statistics

Note: Tandem repeats are categorized by the length of the repeat unit: microsatellite (unit 1–9 bp), minisatellite (unit 10–99 bp), and satellite (unit ≥100 bp).

5.3 Noncoding RNA Annotation

tRNA, rRNA, miRNA and snRNA were predicted with tRNAscan‑SE, RNAmmer and INFERNAL, respectively. Final results are summarized in GFF format.

Table 3 Summary of noncoding RNA annotation

Note: miRNA (microRNA) regulates gene expression; tRNA (transfer RNA) carries amino acids to the translation complex; rRNA (ribosomal RNA) forms the ribosome; snRNA (small nuclear RNA) participates in splicing, transcription, and processing in the nucleus.

5.4 Protein‑coding Gene Structural Annotation

We performed gene prediction using transcriptome evidence, protein homology from closely related species, and de novo models in parallel. The initial integration used EVidenceModeler/Maker; then an in‑house integrator produced the final set of protein‑coding genes in GFF format. CDS and protein sequences were extracted for downstream evaluation and statistics.

Table 4 Summary of all evidence sets

Note: “Method” indicates evidence type; “Software” lists tools used; “Species” shows Latin names of close relatives used for homology; “Gene number” is the number of genes supported by that evidence on this genome (for “Homology”, this is the number mapped from homologous proteins, not the donor species’ native gene count); “Average gene length” is the average mRNA length; “Average CDS length” is average CDS length; “Average exon per gene” is mean exons per gene; “Average exon length” is mean exon length; “Average intron length” is mean intron length.

Table 5 Statistics of coding‑gene prediction

Note: “the total number of gene”: total gene count; “the average of gene_length”: average gene length; “the average of CDS_length”: average CDS length; “the average of exon_number”: mean exons per gene; “the average of exon_length”: mean exon length; “the average of intron_length”: mean intron length; “the total number of exon”: total exon count; “the total number of intron”: total intron count; “the total intron length”: total intron length.

Because closely related species typically share similar distributions of gene lengths, CDS lengths, intron lengths, and exon lengths, comparing these distributions between the focal species and its relatives provides an intuitive, global check on annotation quality. The distributions are plotted below:

Figure 2 Comparison across related species

Note: The four plots’ x‑axes represent gene length, CDS length, intron length, and exon length, respectively; the y‑axes show the percentage of genes in each length bin. Different colors correspond to different species; the red line indicates the focal species.

5.5 Functional Annotation of Protein‑coding Genes

Protein sequences from the annotated coding genes were mapped to functional databases using sequence similarity (diamond) and domain/motif searches (hmmscan), yielding functional annotations.

Table 6 Summary of functional annotation

Note: “All”: all genes; “Annotation”: genes with at least one annotation; “KEGG”: genes annotated to KEGG orthologs; “Pathway”: genes mapped to KEGG pathways via KEGG orthologs; “Nr”: genes annotated to the NCBI NR database; “Uniprot”: genes annotated to the UniProt database; “GO”: genes annotated to the GO database; “KOG/COG”: genes annotated to the Cluster of Orthologous Groups of proteins (KOG for eukaryotes, COG for prokaryotes); “Pfam”: genes annotated to the Pfam database; “Interpro”: genes annotated to the InterPro database.

Figure 3 Upset plot of functional‑annotation categories

5.6 Evaluation of Genome Annotation

We used BUSCO to assess the completeness of the gene set. As a rule of thumb, a BUSCO score above 90% indicates a good annotation.

Table 7 BUSCO assessment of the gene set (busco_lineage: embryophyta_odb10)

Assembly
Annotation
Proteins Percentage (%) Proteins Percentage (%)
Complete BUSCOs 1,608 99.6 1,604 99.4
Complete Single-Copy BUSCOs 1,566 97.0 1,565 97.0
Complete Duplicated BUSCOs 42 2.6 39 2.4
Fragmented BUSCOs 6 0.4 6 0.4
Missing BUSCOs 0 0.0 4 0.2
Total BUSCO groups searched 1,614 100.0 1,614 100.0

Note: “Complete BUSCOs”: complete BUSCOs, defined as genes with lengths within the 95% confidence interval of the average length in the BUSCO orthogroup; “Complete and single BUSCOs”: complete and single‑copy; “Complete and duplicated BUSCOs”: complete but duplicated; “Fragmented BUSCOs”: fragmented (incomplete); “Missing BUSCOs”: missing (not detected); “Total BUSCO group searched”: total number of BUSCOs in the selected database.

5.7 Result File Structure and Notes


6 Materials and Methods

6.1 Bioinformatics Methods

6.1.1 Repeat Annotation

Repeats are divided into interspersed and tandem repeats. Interspersed repeats are transposable elements (TEs), including four major types: LTR, LINE, SINE and DNA transposons. First, de novo prediction was performed with RepeatModeler to obtain a genome‑specific library; LTR_FINDER and LTRharvest were used to predict LTRs, and LTR_retriever was applied to de‑duplicate the two LTR sets to obtain a non‑redundant LTR library. The two de novo libraries were then merged; unknown entries were reclassified by teclass; public RepBase and the de novo library were merged and RepeatMasker was run to produce the “De novo + repbase” set. RepeatProteinMask (a subtool of RepeatMasker) predicted TE_protein‑type repeats (the “TE protiens” set). Finally, all predictions were merged and deduplicated to yield the final interspersed‑repeat set (“Combined TEs”).

6.1.2 Noncoding RNA Annotation

Noncoding RNAs are RNAs that are not translated into proteins yet have important biological functions. tRNA and rRNA directly participate in protein synthesis. tRNAs are identified by tRNAscan‑SE based on structural features; rRNAs are predicted by RNAmmer; other ncRNAs (e.g., snRNA, miRNA) are detected with INFERNAL using the Rfam database.

6.1.3 Structural Annotation of Protein‑coding Genes

Gene structure prediction is essential for detailed gene distribution/structure information and is the foundation of functional annotation and evolutionary analyses. This process includes predicting gene loci, ORFs, translation start/stop, introns/exons, promoters, alternative splicing, and coding sequences. Here we combined transcriptome‑based, homology‑based, and de novo predictions.

1) Transcriptome‑based prediction

Transcriptome‑based prediction includes short‑read and long‑read RNA‑seq:

  1. Short‑read RNA‑seq: fastp filters raw reads; hisat2 maps reads to the genome; stringtie reconstructs transcripts from BAM files. If only short‑read data are available, TransDecoder predicts ORFs to yield coding genes.

  1. Long‑read RNA‑seq:

    1. ONT full‑length reads: NanoFilt filters data; pychopper identifies full‑length reads. minimap2 maps to the genome; stringtie reconstructs transcripts from BAM files. If only ONT full‑length data are available, TransDecoder predicts ORFs.

    1. PacBio full‑length reads: two input modes:

      1. With subreads: ccs is called in smrtlink to generate CCS reads; isoseq3 performs FL detection, polishing and clustering; polished non‑redundant FL reads are mapped with pbmm2 and reconstructed with isoseq3.

      1. With clustered fasta: minimap2 mapping; cDNA_Cupcake reconstructs transcripts from BAM. If only PacBio FL data are available, TransDecoder predicts ORFs.

  1. If both short‑ and long‑read datasets are available, stringtie merges predicted transcripts and TransDecoder performs ORF prediction.

2) Homology‑based prediction

Protein sequences from closely related species are aligned with miniprot for homology‑based predictions. Protein sets from the BUSCO database can also serve as indirect homology evidence.

3) De novo prediction

Analyses are performed on the repeat‑masked genome:

  1. augustus is trained with transcriptome and homology evidence, then used for prediction.

  1. genscan’s built‑in models predict genes in animal genomes.

  1. glimmerhmm is trained on homology evidence for predictions in plants and fungi.

Finally, maker/EVidenceModeler integrates all evidence into a unified gene set. Transcriptome annotations serve as EST evidence; homology predictions as protein evidence; de novo predictions are combined as inputs to gene prediction. The final set is validated by ORF correctness, start/stop accuracy, and gene‑length filters. Based on maker/EVidenceModeler results, we prioritize genes supported by multiple evidence types and supplement with single‑evidence genes to meet the expected gene count, giving priority to transcriptome, homology, and de novo evidence. This integrated strategy ensures accuracy and completeness of gene structure predictions.

6.1.4 Functional Annotation of Protein‑coding Genes

Functional annotation assigns gene functions and metabolic pathways using existing databases (motifs, domains, protein functions, and pathways). Two main approaches are used:

1) Sequence‑similarity search:

Protein sequences are compared to UniProt and NR. UniProt links protein families to Gene Ontology and to orthologous groups (COG, KOG), enabling function inference for proteins; this part uses diamond.

2) Motif‑similarity search:

InterProScan searches InterPro component databases (CDD, Gene3D, Hamap, Panther, Phobius, Pirsf, Pirsr, Prints, Prosite, Sfld, Smart, Superfamily, Tigrfam, Tmhmm) to identify conserved sequences, motifs, and domains. Additionally, hmmscan is used for domain prediction. KEGG pathway annotation uses the internal KOfam HMM database; Pfam is a large collection of protein families based on alignments and HMMs.

6.2 Software, Databases, and References

Table 8 Summary of software, databases, and references

Workflow step Software/Database Version Parameter 参考文献或网址
Repeat sequence annotation RepeatModeler 2.0.6 -database mydb -threads 16 Flynn, Jullien M et al. “RepeatModeler2 for automated genomic discovery of transposable element families.” Proceedings of the National Academy of Sciences of the United States of Americavol. 117,17 (2020): 9451-9457. doi:10.1073/pnas.1921046117 https://www.repeatmasker.org/RepeatModeler
Repeat sequence annotation teclass 2.1.3 default Abrusán, György et al. “TEclass–a tool for automated classification of unknown eukaryotic transposable elements.” Bioinformatics (Oxford, England) vol. 25,10 (2009): 1329-30. doi:10.1093/bioinformatics/btp084 https://www.compgen.uni-muenster.de/tools/teclass/index.hbi
Repeat sequence annotation LTR_FINDER Official release of LTR_FINDER_parallel -threads 16 -harvest_out -size 1000000 -time 300 Xu, Zhao, and Hao Wang. “LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons.” Nucleic acids research vol. 35,Web Server issue (2007): W265-8. doi:10.1093/nar/gkm286 https://github.com/oushujun/LTR_FINDER_parallel
Repeat sequence annotation LTRharvest 1.6.5 -minlenltr 100 -maxlenltr 7000 -mintsd 4 -maxtsd 6 -motif TGCA -motifmis 1 -similar 85 -vic 10 -seed 20 -seqids yes Ellinghaus, David et al. “LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons.” BMC bioinformatics vol. 9 18. 14 Jan. 2008, doi:10.1186/1471-2105-9-18 http://genometools.org/pub/binary_distributions
Repeat sequence annotation LTR_retriever 3.0.1 -threads 16 -noanno Ou, Shujun, and Ning Jiang. “LTR_retriever: A Highly Accurate and Sensitive Program for Identification of Long Terminal Repeat Retrotransposons.” Plant physiology vol. 176,2 (2018): 1410-1422. doi:10.1104/pp.17.01310 https://github.com/oushujun/LTR_retriever
Repeat sequence annotation RepeatMasker 4.1.7 -noLowSimple -pvalue 0.0001 Smit A, Hubley R, Green P. “RepeatMasker Open-4.0.” 2013. https://www.repeatmasker.org/RepeatMasker/
Repeat sequence annotation RepeatProteinMask 4.1.7 -noLowSimple -pvalue 0.0001 Smit A, Hubley R, Green P. “RepeatMasker Open-4.0.” 2013. https://www.repeatmasker.org/RepeatProteinMask.html
Repeat sequence annotation TRF 4.09 2 7 7 80 10 50 2000 -d Benson, G. “Tandem repeats finder: a program to analyze DNA sequences.” Nucleic acids research vol. 27,2 (1999): 573-80. doi:10.1093/nar/27.2.573 https://tandem.bu.edu/trf/trf.html
Repeat sequence annotation MISA v2.1 default Beier, Sebastian et al. “MISA-web: a web server for microsatellite prediction.” Bioinformatics (Oxford, England) vol. 33,16 (2017): 2583-2585. doi:10.1093/bioinformatics/btx198 https://webblast.ipk-gatersleben.de/misa/
Non-coding RNA annotation tRNAscan-SE v.2.0.12 -E -j tRNA.gff -o tRNA.result -f tRNA.struct –thread 16 Lowe T M, Eddy S R. “tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence.” Nucleic acids research, 1997, 25(5): 955-964. https://github.com/UCSC-LoweLab/tRNAscan-SE
Non-coding RNA annotation INFERNAL 1.1.4 –cut_ga –rfam –nohmmonly –fmt 2 Nawrocki, Eric P, and Sean R Eddy. “Infernal 1.1: 100-fold faster RNA homology searches.” Bioinformatics (Oxford, England) vol. 29,22 (2013): 2933-5. doi:10.1093/bioinformatics/btt509 http://eddylab.org/infernal/
Non-coding RNA annotation RNAmmer 1.2 -S euk -m tsu,lsu,ssu Lagesen, Karin et al. “RNAmmer: consistent and rapid annotation of ribosomal RNA genes.” Nucleic acids research vol. 35,9 (2007): 3100-8. doi:10.1093/nar/gkm160 https://services.healthtech.dtu.dk/services/RNAmmer-1.2/
Protein-coding gene expression evidence annotation fastp 0.23.2 -L -w 16 Chen, Shifu et al. “fastp: an ultra-fast all-in-one FASTQ preprocessor.” Bioinformatics (Oxford, England) vol. 34,17 (2018): i884-i890. doi:10.1093/bioinformatics/bty560 https://github.com/OpenGene/fastp
Protein-coding gene expression evidence annotation fastqc v0.11.9 default Andrews S. “FastQC: a quality control tool for high throughput sequence data.” 2010. https://www.bioinformatics.babraham.ac.uk/projects/fastqc/
Protein-coding gene expression evidence annotation hisat2 2.2.1 –dta -p 16 Kim, Daehwan et al. “Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype.” Nature biotechnology vol. 37,8 (2019): 907-915. doi:10.1038/s41587-019-0201-4 http://daehwankimlab.github.io/hisat2/
Protein-coding gene expression evidence annotation samtools 1.17 sort Danecek, Petr et al. “Twelve years of SAMtools and BCFtools.” GigaScience vol. 10,2 (2021): giab008. doi:10.1093/gigascience/giab008 https://github.com/samtools/samtools
Protein-coding gene expression evidence annotation NanoFilt 2.8.0 -l 50 -q 7 De Coster, Wouter et al. “NanoPack: visualizing and processing long-read sequencing data.” Bioinformatics (Oxford, England) vol. 34,15 (2018): 2666-2669. doi:10.1093/bioinformatics/bty149 https://github.com/wdecoster/nanofilt
Protein-coding gene expression evidence annotation pychopper v2.7.5 -t 16 Schuster, Jakob et al. “Restrander: rapid orientation and artefact removal for long-read cDNA data.” NAR genomics and bioinformatics vol. 5,4 lqad108. 23 Dec. 2023, doi:10.1093/nargab/lqad108 https://github.com/epi2me-labs/pychopper
Protein-coding gene expression evidence annotation seqkit v2.4.0 stat Shen, Wei et al. “SeqKit2: A Swiss army knife for sequence and alignment processing.” iMeta vol. 3,3 e191. 5 Apr. 2024, doi:10.1002/imt2.191 https://bioinf.shenwei.me/seqkit/
Protein-coding gene expression evidence annotation minimap2 2.26-r1175 -t 16 -ax splice -uf –secondary=no Li, Heng. “Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences.” Bioinformatics (Oxford, England) vol. 32,14 (2016): 2103-10. doi:10.1093/bioinformatics/btw152 https://github.com/lh3/minimap2
Protein-coding gene expression evidence annotation smrtlink v12.0 default
Protein-coding gene expression evidence annotation ccs 7.0.0 default
Protein-coding gene expression evidence annotation isoseq3 3.9.0 default
Protein-coding gene expression evidence annotation pbmm2 1.10.0 align -j 32 –preset ISOSEQ –sort
Protein-coding gene expression evidence annotation cDNA_Cupcake 1.0 default
Protein-coding gene expression evidence annotation stringtie 2.2.1 -p 16 -R -L Pertea, Mihaela et al. “StringTie enables improved reconstruction of a transcriptome from RNA-seq reads.” Nature biotechnology vol. 33,3 (2015): 290-5. doi:10.1038/nbt.3122 https://ccb.jhu.edu/software/stringtie/
Protein-coding gene expression evidence annotation TransDecoder v5.7.0 default
Protein-coding gene homology annotation exonerate 2.4.0 default Slater, G. S. C., and Birney, E. (2005), “Automated generation of heuristics for biological sequence comparison,” BMC bioinformatics, Springer, 6, 1–11. https://github.com/nathanweeks/exonerate
Protein-coding gene homology annotation miniprot 0.13 –gff -Iut50 Li, Heng. “Protein-to-genome alignment with miniprot.” Bioinformatics (Oxford, England) vol. 39,1 (2023): btad014. doi:10.1093/bioinformatics/btad014 https://github.com/lh3/miniprot
Protein-coding gene ab initial prediction augustus 3.5.0 –uniqueGeneId=true –noInFrameStop=true –gff3=on –strand=both Stanke, Mario et al. “Using native and syntenically mapped cDNA alignments to improve de novo gene finding.” Bioinformatics (Oxford, England) vol. 24,5 (2008): 637-44. doi:10.1093/bioinformatics/btn013 https://github.com/Gaius-Augustus/Augustus
Protein-coding gene ab initial prediction genscan 1.0 default Burge C, Karlin S. “Prediction of complete gene structures in human genomic DNA.” Journal of molecular biology, 1997, 268(1): 78-94. http://hollywood.mit.edu/GENSCAN.html
Protein-coding gene ab initial prediction glimmerhmm 3.0.4 -f -g Delcher, Arthur L et al. “Identifying bacterial genes and endosymbiont DNA with Glimmer.” Bioinformatics (Oxford, England) vol. 23,6 (2007): 673-9. doi:10.1093/bioinformatics/btm009 http://ccb.jhu.edu/software/glimmerhmm/
Integrated gene annotation maker 3.01.03 default Holt, Carson, and Mark Yandell. “MAKER2: an annotation pipeline and genome-database management tool for second-generation genome projects.” BMC bioinformatics vol. 12 491. 22 Dec. 2011, doi:10.1186/1471-2105-12-491 https://www.yandell-lab.org/software/maker.html
Integrated gene annotation EVidenceModeler v2.1.0 default Haas, Brian J et al. “Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments.” Genome biology vol. 9,1 R7. 11 Jan. 2008, doi:10.1186/gb-2008-9-1-r7 https://github.com/EVidenceModeler/EVidenceModeler
Annotation assessment BUSCO 5.7.0 -m prot -c 16 Manni, Mosè et al. “BUSCO Update: Novel and Streamlined Workflows along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes.” Molecular biology and evolution vol. 38,10 (2021): 4647-4654. doi:10.1093/molbev/msab199 https://busco.ezlab.org/
Functional annotation diamond 2.1.8 –evalue 1e-05 Buchfink, Benjamin et al. “Fast and sensitive protein alignment using DIAMOND.” Nature methods vol. 12,1 (2015): 59-60. doi:10.1038/nmeth.3176 https://github.com/bbuchfink/diamond
Functional annotation hmmscan 3.3.2 –cpu 16 -E 1e-5 Eddy, Sean R. “A new generation of homology search tools based on probabilistic inference.” Genome informatics. International Conference on Genome Informatics vol. 23,1 (2009): 205-11. http://hmmer.org/download.html
Functional annotation InterProScan 5.55-88.0 default Mulder, Nicola, and Rolf Apweiler. “InterPro and InterProScan: tools for protein sequence classification and comparison.” Methods in molecular biology (Clifton, N.J.) vol. 396 (2007): 59-70. doi:10.1007/978-1-59745-515-2_5 https://github.com/ebi-pf-team/interproscan
Functional annotation Rfam 14.10 - Kalvari, Ioanna et al. “Rfam 14: expanded coverage of metagenomic, viral and microRNA families.” Nucleic acids research vol. 49,D1 (2021): D192-D200. doi:10.1093/nar/gkaa1047 https://rfam.org/
Functional annotation InterPro 5.55-88.0 - Blum, Matthias et al. “The InterPro protein families and domains database: 20 years on.” Nucleic acids research vol. 49,D1 (2021): D344-D354. doi:10.1093/nar/gkaa977 https://www.ebi.ac.uk/interpro/
Functional annotation Uniprot 2024-05-29 - Apweiler, Rolf et al. “UniProt: the Universal Protein knowledgebase.” Nucleic acids research vol. 32,Database issue (2004): D115-9. doi:10.1093/nar/gkh131 https://www.uniprot.org/
Functional annotation GO 2024-05-29 - Ashburner, M et al. “Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.” Nature genetics vol. 25,1 (2000): 25-9. doi:10.1038/75556 https://geneontology.org/
Functional annotation kofam 2024-05-27 - Aramaki, Takuya et al. “KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold.” Bioinformatics (Oxford, England)vol. 36,7 (2020): 2251-2252. doi:10.1093/bioinformatics/btz859 https://www.genome.jp/tools/kofamkoala/
Functional annotation Pfam 37.0 - Finn R D, Bateman A, Clements J, et al. “Pfam: the protein families database.” Nucleic acids research, 2014, 42(D1): D222-D230. https://pfam.xfam.org/
Functional annotation NR 2024-02-07 - Deng Y Y, Li J Q, Wu S F, et al. “Integrated nr database in protein annotation system and its localization.” Comput. Eng, 2006, 32(5): 71-72. https://ftp.ncbi.nlm.nih.gov
Homology alignment ncbi-blast+ 2.11.0 -outfmt 6 -evalue 1e-5 Altschul, S F et al. “Basic local alignment search tool.” Journal of molecular biology vol. 215,3 (1990): 403-10. doi:10.1016/S0022-2836(05)80360-2 https://blast.ncbi.nlm.nih.gov/Blast.cgi
Syntenic gene pairs MCScanX 0.8 default Wang, Yupeng et al. “MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity.” Nucleic acids research vol. 40,7 (2012): e49. doi:10.1093/nar/gkr1293 https://github.com/wyp1125/MCScanX
Circos plot circlize 0.4.16 default Gu, Zuguang et al. “circlize Implements and enhances circular visualization in R.” Bioinformatics (Oxford, England) vol. 30,19 (2014): 2811-2. doi:10.1093/bioinformatics/btu393 https://jokergoo.github.io/circlize/

7 Frequently Asked Questions

Q: Which software opens .xls files?

Microsoft Excel is recommended; text editors such as Sublime Text can also open them.

Q: Notes on GFF/GFF3 format

GFF/GFF3 files are tab‑delimited. Lines starting with “#” contain summary annotations. Below are the column definitions for non‑comment rows.

Q: Why is the number of homologous protein sequences from related species much larger than the number of genes annotated in this species?

Homology sets may include multiple isoforms from the same gene; these are not the true gene counts. Applying filters such as “longest transcript per gene” yields the actual gene numbers for the related species.

Q: What’s the difference between “Assembly” and “Annotation” values in busco_assessment.xls?

“Assembly” refers to BUSCO assessment on the genomic DNA assembly, while “Annotation” assesses the translated protein set from the annotation. BUSCO provides protein orthologs; assembly assessments use protein‑vs‑nucleotide tools (tblastn, metaeuk, miniprot), whereas annotation assessments use HMM domain matches. Therefore, differences between assembly and annotation BUSCO scores are expected.

Q: DIAMOND BLAST output format

Q: weGO.xls format

This is a reshaped version of go.xls: the first column is the gene/transcript ID, followed by all GO terms assigned to that ID—formatted for WEGO input.

Q: iprscan.tsv.xls format

Q: hmmscan result format

See https://www.omicsclass.com/article/499.

Q: What’s the difference between pfam.xls and pfam_stringent_hmm.xls?

pfam.xls is a subset of pfam_stringent_hmm.xls, with alignment scores and coverage columns removed—leaving only sequence IDs and corresponding Pfam entries. Use pfam_stringent_hmm.xls for downstream filtering under custom thresholds.


8 References

See this report’s section: Software, Databases, and References.


9 Contact Us

Customer Service Telephone: 1-(617)-223-7544

Customer Service Email:

Company Website: https://www.sailgene.us/

Company Address: 82 Wendell Avenue, Pittsfield, Massachusetts, USA