Sb bicolor BTx623 v3 (Sorghum_bicolor_NCBIv3) ▼

Sb bicolor BTx623 v3 Assembly and Gene Annotation

About Sorghum bicolor BTx623

Sorghum bicolor (L.) Moench subsp. bicolor, is a widely grown cereal crop, particularly in Africa, ranking 5th in global cereal production (FAOSTAT 2008; http://www.fao.org/in-action/inpho/crop-compendium/cereals-grains/). It is a C4 grass also used for sugar production, brewing, feedstock, and as a biofuel crop. Its diploid genome (~730 Mbp) has a haploid chromosome number of 10. The inbred variety ‘BTx623’ is the current reference genome for sorghum. It has short stature and an early maturing genotype used primarily to produce grain sorghum hybrids. It is a line susceptible to sugarcane aphid and sensitive to low nitrogen, and therefore often used in functional comparative studies.

Germplasm

U.S. National Plant Germplasm System (GRIN - Global) identifier for BTx623: PI 564163.

This germplasm is part of the following population panels:

Assembly

The genome assembly of Sorghum bicolor cv. Moench was published in 2009 (Paterson et al, 2009). The present assembly corresponds to v3.1.1 at the US Department of Energy Joint Genome Institute (JGI) described in (McCormick et al, 2018), and is also known as the NCBIv3 assembly. Sequencing by the JGI's Community Sequencing Program in collaboration with the Plant Genome Mapping Laboratory at the University of Georgia, followed a whole-genome shotgun strategy reaching 8X coverage with scaffolds -where possible- being assigned to the genetic map. JGI did two additional rounds of improvements. The most recent update of release v3.0 included ~351 Mb of finished sorghum sequence. A total of 349 clones were manually inspected, then finished and validated using a variety of technologies including Sanger, 454 and Illumina. They were integrated into chromosomes by aligning to v1.0 assembly. As a result, 4,426 gaps were closed, and a total of 4.96 Mb of sequence was added to the assembly. Overall contiguity (contig N50) increased by a factor of 5.8X from 204.5 Kb to 1.2 Mb. For more details, see Phytozome.

NCBI accession: GCA_000003195.3.

Annotation

Gene predictions resulted from combining homology-based and ab initio methods with expressed sequences from sorghum, maize and sugarcane, using the JGI annotation pipeline (Goodstein et al, 2012). The SorghumBase browser presents data from the current JGI v3.1.1 release, which comprises the v3.0.1 assembly and v3.1.1 gene set (February 2017). Read more at Phytozome.

This is a modern annotation using resources used in the original v1.0 release (Sbi1 assembly and Sbi1.4 gene set) and geneAtlas RNA-seq data. The main genome is in 10 chromosomes with small unmapped pieces, some of which contain annotated genes. The NCBIv3 release (Phytozome v3.1.1) is essentially the same as Phytozome v3.1 except for 82 genes/loci that were inactivated due to 4 scaffolds entirely present in chromosome(s) that were removed.

Assembly information
Assembly name Sorghum_bicolor_NCBIv3
Assembly date June 2017
Assembly accession GCA_000003195.3
WGS accession ABXC00000000
Assembly provider
Sequencing description Sequencing technologies: Sanger; Illumina
Sequencing method
Genome coverage: 8x
Assembly description Assembly methods: ARACHNE_modified v. 200721016
Construction of pseudomolecules
Finishing strategy
NCBI submission
Publication: Paterson *et al* (2009); McCormick *et al* (2018)
Assembly statistics
Number of contigs 2,688
Total assembly length (Mb) 732
Contig N50 (Mb) 1
Annotations stats
Total number of genes 34,118
Total number of transcripts 47,110
Average gene length 3,714
Exons per transcript 5

Source: NCBI, April 2021.

Repeats

Repeats were annotated with the Ensembl Genomes repeat feature pipeline (Aken et al, 2016), which uses six classes of repeats loaded from ENA.

Repeat feature Frequency Coverage (Mb) % of the genome covered
Low complexity (Dust) features 685,783 29 4
RepeatMasker (with RepBase library) 455,749 451 62.1
RepeatMasker (with REdat library) 392,778 409 56.2
Tandem repeats (TRF) features 245,654 41 5

Repeat feature Frequency Coverage (Mb) % of the genome covered Low complexity (Dust) features 685,783 29 4 RepeatMasker (with RepBase library) 455,749 451 62.1 RepeatMasker (with REdat library) 392,778 409 56.2 Tandem repeats (TRF) features 245,654 41 5.7

Variation

Variation in SorghumBase is available for short variants (genetic variation, which in turn may be naturally occurring or chemically induced) and QTL variants associated with physical traits.

Genetic Variation

Genetic variation data for a sorghum gene is available graphically and in tabular form, and for each variant, a Variant page provides more detailed information. Below are examples of each of these data representations.

Single Nucleotide Polymorphisms (SNPs). Currently in SorghumBase, there are 46 million SNPs (of which 41 million have standard rsIDs assigend by the European Variation Archive, EVA) from two SNP data sets mapped to sorghum BTx623:

  • The Lozano et al (2021) data set includes early 13 million naturally occurring SNPs in 499 sorghum accessions (including 14 duplicate samples). Included are accessions from the TERRA-MEPP and TERRA-REF population panels, and lines previously genotyped by Emma Mace and collaborators in 2013.
  • The Boatwright et al (2022) includes nearly 44 million genetic variants including about 38 million SNPs and 5 million indels determined in the 400 Sorghum Association Panel (SAP) accessions via whole-genome sequencing (WGS).
    Chemically induced variation

Ethyl methanesulfonate (EMS)-induced mutations. Currently in SorghumBase, there are three collections of EMS-induced mutant lines. EMS is a chemical commonly used to cause point mutations, that is, to change single nucleotides in the DNA of a plant seed. Variants in the Jiao dataset were recalled and added back in release 5, and a new dataset by Dr. Xin is being added in release 6.

  • The Addo-Quaye et al (2018) data set includes over 2.6 million point mutations identified in 486 sorghum accessions corresponding to the M3 generation of an EMS-mutagenized sorghum population.
  • The Xin EMS dataset (Jiao et al, 2016) features over 1.7 million variants recalled from the original EMS-induced G/C to A/T transition mutations data set annotated from 252 M3 families selected from the 6,400 sorghum mutant library in BTx623 background described by Xin and colleagues (Xin et al, 2008). Genomic DNA used for sequencing was pooled from 20 M3 plants per M2 family (Jiao et al, 2016).
  • About 13.8 million new EMS-induced mutations were recently released by Dr. Zhanguo Xin and collaborators (manuscript in preparation).
    Loss of function mutations

Loss-of-function (LOF) mutations are those in which the altered gene product lacks the molecular function of the wild type gene. In SorghumBase, the following seven functional consequences in a genetic variant are predicted to result in LOF of a protein coding gene: splice acceptor variant, splice donor variant, stop gained, frameshift variant, stop gained, start lost, and missense variant. In the present release, for each of the above mentioned BTx623 genetic variation studies (Boatwright et al (2022), Lozano et al (2021); Xin et al, manuscript in preparation; Addo-Quaye et al (2018), and Jiao et al (2016)), we are making available a list of LOF mutations including +/- 250 nt flanking sequences for each variant. This data may be downloaded in bulk from SorghumBase's FTP.

Phenotypic Variation

Genome-wide association studies (GWAS). GWAs hits from a meta-analysis of 234 phenotypic trait datasets of 40 studies (25 qualified for inclusion) on 406 SAP accessions performed by Mural et al (2021).

Example of SNP associated with nutritional traits (e.g, protein/starch/fat content).

Quantitative Trait Locus (QTLs). Data corresponding to 5,843 QTL features for 220 sorghum traits were imported from Sorghum QTL Atlas and are provided with predicted syntenic locations in maize and rice.

Example region with QTLs associated with multiple traits including greenbug resistance, fresh biomass, and flag leaf height. Hint: For additional regions with QTL data in the current sorghum assembly (v.3), use the physical or genetic (cM) coordinates kindly provided by the Sorghum QTL Atlas team.

References

Addo-Quaye C, Tuinstra M, Carraro N, Weil C, Dilkes BP. Whole-Genome Sequence Accuracy Is Improved by Replication in a Population of Mutagenized Sorghum. G3 . 2018;8: 1079–1094. doi: 10.1534/g3.117.300301.

Aken, Bronwen L., Sarah Ayling, Daniel Barrell, Laura Clarke, Valery Curwen, Susan Fairley, Julio Fernandez Banet, et al. 2016. “The Ensembl Gene Annotation System.” Database: The Journal of Biological Databases and Curation. PMID: 27337980. doi: 10.1093/database/baw093.

Boatwright JL, Sapkota S, Jin H, Schnable JC, Brenton Z, Boyles R, Kresovich S. 2022. "Sorghum Association Panel whole-genome sequencing establishes cornerstone resource for dissecting genomic diversity." Plant J. PMID: 35653240. doi: 10.1111/tpj.15853.

Brenton, Zachary W., Elizabeth A. Cooper, Mathew T. Myers, Richard E. Boyles, Nadia Shakoor, Kelsey J. Zielinski, Bradley L. Rauh, William C. Bridges, Geoffrey P. Morris, and Stephen Kresovich. 2016. “A Genomic Resource for the Development, Improvement, and Exploitation of Sorghum for Bioenergy.” Genetics 204 (1): 21–33. PMID: 27356613. doi: 10.1534/genetics.115.183947.

Casa, Alexandra M., Gael Pressoir, Patrick J. Brown, Sharon E. Mitchell, William L. Rooney, Mitchell R. Tuinstra, Cleve D. Franks, and Stephen Kresovich. 2008. “Community Resources and Strategies for Association Mapping in Sorghum.” Crop Science 48 (1): 30–40. doi: 10.2135/cropsci2007.02.0080.

Goodstein, David M., Shengqiang Shu, Russell Howson, Rochak Neupane, Richard D. Hayes, Joni Fazo, Therese Mitros, et al. 2012. “Phytozome: A Comparative Platform for Green Plant Genomics.” Nucleic Acids Research 40 (Database issue): D1178–86. PMID: 22110026. doi: 10.1093/nar/gkr944.

Jiao, Yinping, John J. Burke, Ratan Chopra, Gloria Burow, Junping Chen, Bo Wang, Chad Hayes, Yves Emendack, Doreen Ware, and Zhanguo Xin. 2016. “A Sorghum Mutant Resource as an Efficient Platform for Gene Discovery in Grasses.” The Plant Cell. PMID: 27354556. doi: 10.1105/tpc.16.00373.

Lozano R, Gazave E, Dos Santos JPR, Stetter MG, Valluru R, Bandillo N, et al. Comparative evolutionary genetics of deleterious load in sorghum and maize. Nat Plants. 2021;7: 17–24. PMID: 33452486. doi: 10.1038/s41477-020-00834-5.

McCormick, Ryan F., Sandra K. Truong, Avinash Sreedasyam, Jerry Jenkins, Shengqiang Shu, David Sims, Megan Kennedy, et al. 2018. “The Sorghum Bicolor Reference Genome: Improved Assembly, Gene Annotations, a Transcriptome Atlas, and Signatures of Genome Organization.” The Plant Journal: For Cell and Molecular Biology 93 (2): 338–54. PMID: 29161754. doi: 10.1111/tpj.13781.

Paterson, A. H., J. E. Bowers, R. Bruggmann, I. Dubchak, J. Grimwood, H. Gundlach, G. Haberer, et al. 2009. “The Sorghum Bicolor Genome and the Diversification of Grasses.” Nature 457 (7229): 551–56. PMID: 19189423. doi: 10.1038/nature07723.

Xin, Zhanguo, Ming Li Wang, Noelle A. Barkley, Gloria Burow, Cleve Franks, Gary Pederson, and John Burke. 2008. “Applying Genotyping (TILLING) and Phenotyping Analyses to Elucidate Gene Function in a Chemically Induced Sorghum Mutant Population.” BMC Plant Biology. PMID: 18854043. doi: 10.1186/1471-2229-8-103

Gene Expression

Baseline Gene Expression (Atlas)

Baseline gene expression data from seven sorghum BTx623 datasets curated and processed by the EMBL-EBI Expression Atlas Team.

Click here for an example of baseline gene expression for the msd2 gene.

More information

General information about this species can be found in Wikipedia.

More information

General information about this species can be found in Wikipedia.

Statistics

Summary

AssemblySorghum_bicolor_NCBIv3, INSDC Assembly GCA_000003195.3, Jun 2017
Database version108.30
Golden Path Length708,735,318
Genebuild by
Genebuild methodGenerated from ENA annotation

Gene counts

Coding genes34,118
Gene transcripts48,471