Sb bicolor BTx623 v5 Assembly and Gene Annotation
About Sorghum bicolor BTx623 v5.1 (JGI)
Sorghum bicolor (L.) Moench subsp. bicolor, is a widely grown cereal crop, particularly in Africa, ranking 5th in global cereal production (FAOSTAT 2008; http://www.fao.org/in-action/inpho/crop-compendium/cereals-grains/). It is a C4 grass also used for sugar production, brewing, feedstock, and as a biofuel crop. Its diploid genome (~730 Mbp) has a haploid chromosome number of 10. The inbred variety ‘BTx623’ is the current reference genome for sorghum. It has short stature and an early maturing genotype used primarily to produce grain sorghum hybrids. It is a line susceptible to sugarcane aphid and sensitive to low nitrogen, and therefore often used in functional comparative studies. It is used as a biofuel crop and potential cellulosic feedstock.
The Department of Energy Joint Genome Institute (JGI)'s Sorghum bicolor BTx623 assembly version 5.1 is a wholly resequenced genome applying PacBio long-read data. The annotation is an update that uses all the v3.1 resources, but with additional RNA-seq from JGI projects and full-length transcripts included to further improve the completeness of the gene set. Additional details in JGI Phytozome.
The v5.1 assembly was released before scientific publication according to the Fort Lauderdale Accord. The accord restricts publication of articles containing analyses of genes or genomic data on this chromosome-scale assembly prior to publication of a comprehensive genome analysis by JGI and/or its collaborators. Therefore, we are only providing a basic genome browser for the new v5.1 assembly.
Germplasm
U.S. National Plant Germplasm System (GRIN - Global) identifier for BTx623: PI 564163.
This germplasm is part of the following population panels:
- Sorghum Association Panel (SAP) - 407 accessions (Casa et al, 2008)
- Sorghum Bioenergy Association Panel (BAP) - 386 accessions (Brenton et al, 2016)
Assembly
The main assembly consisted of 122.94x of PACBIO coverage (11,008 bp average read size), and was assembled using MECAT. The resulting sequence was polished using RACON. A combination of syntenic markers from BTx642 and primary annotated genes from BTx623 were used to identify misjoins in the assembly. Syntenic markers consisted of a total of 32,400 unique, non-repetitive, non-overlapping 1 KB sequences that were generated using the version 1.0 S. bicolor BTx642 genome release. The gene set consisted of 12,641 uniquely aligned BTx642 annotated primary genes. Both the syntenic markers and the primary gene set were aligned to the polished assembly. Misjoins were identified as an abrupt change in linkage group. A total of 9 breaks were made. The broken contigs were then ordered, oriented, and assembled into 10 chromosomes using combined syntenic markers and genes from version 1.0 S. bicolor BTx642. A total of 32 joins were made during this process. Adjacent alternative haplotypes were identified on the joined contig set. Althap regions were collapsed using the longest common substring between the two haplotypes. A total of 8 adjacent altHaps were collapsed. Care was taken to ensure that contigs terminating in telomere were properly oriented in the chromosomes, and the resulting sequence was screened for retained vector and/or contaminants. Finally, Homozygous SNPs and INDELs were corrected in the release sequence using ~53.3x of Illumina reads (2x150, 400bp insert). For more details, see Phytozome.
Annotation
Illumina RNA-seq reads were used to construct transcript assemblies using PERTRAN, which conducts genome-guided transcriptome short read assembly via GSNAP, and builds splice alignment graphs after alignment validation, realignment and correction. To obtain 748K putative full-length transcripts, 12M PacBio Iso-Seq CCSs were corrected and collapsed by a genome guided correction pipeline, which aligns CCS reads to genome with GMAP, and clusters alignments when all introns are the same or 95% overlap for single exon. Subsequently, 581,006 transcript assemblies were constructed using PASA from ESTs, corrected CCS, and RNA-seq transcript assemblies above. Version 3.1 gene models on assembly v3.0 were lifted over to assembly v5.0 and improved with JGI's in-house gene model improvement (GMI) pipeline. The final gene model proteins were assigned PFAM and PANTHER domains, and gene models were further filtered for transposable element domains. Locus model name was assigned by mapping forwad v3.1 locus models if using JGI's locus name mapping pipeline, otherwise, a new name was given using the locus naming convention used in JGI v3.1 locus model naming. In summary, the model names of approximately 96% of non-TE associated v3.1 models were mapped forward to v5.1. For more details, see Phytozome.
Genetic and Phenotypic Variation, Gene Expression & Pathways
Due to JGI's restriction for using the v5.1 assembly in comprehensive genome analyses, genetic and phenotypic variation, gene expression and pathways data for BTx623 genes is only available on assembly version 3.1.
References
Brenton, Zachary W., Elizabeth A. Cooper, Mathew T. Myers, Richard E. Boyles, Nadia Shakoor, Kelsey J. Zielinski, Bradley L. Rauh, William C. Bridges, Geoffrey P. Morris, and Stephen Kresovich. 2016. “A Genomic Resource for the Development, Improvement, and Exploitation of Sorghum for Bioenergy.” Genetics 204 (1): 21–33. PMID: 27356613. doi: 10.1534/genetics.115.183947.
Casa, Alexandra M., Gael Pressoir, Patrick J. Brown, Sharon E. Mitchell, William L. Rooney, Mitchell R. Tuinstra, Cleve D. Franks, and Stephen Kresovich. 2008. “Community Resources and Strategies for Association Mapping in Sorghum.” Crop Science 48 (1): 30–40. doi: 10.2135/cropsci2007.02.0080.
Statistics
Summary
| Assembly | Sb-BTX623-REFERENCE-JGI-5.1, May 2020 |
| Database version | 108.5 |
| Golden Path Length | 719,894,357 |
| Genebuild by | |
| Genebuild method | Import |
Gene counts
| Coding genes | 32,160 |
| Gene transcripts | 44,210 |

