Improved prediction of bacterial transcription start sites

Gordon, John J., Towsey, Michael W., Hogan, James M., Mathews, Sarah A., & Timms, Peter (2006) Improved prediction of bacterial transcription start sites. Bioinformatics, 22(2), pp. 142-148.

View at publisher


Motivation: Identifying bacterial promoters is an important step toward understanding gene regulation. In this paper, we address the problem of predicting the location of promoters and their transcription start sites (TSSs) in Escherichia coli. The accepted method for this problem is to use position weight matrices (PWMs), which define conserved motifs at the sigma-factor binding site. However this method is known to result in a large numbers of false positive predictions.

Results: Our approaches to TSS prediction are based upon an ensemble of support vector machines (SVMs) employing a variant of the mismatch string kernel. This classifier is sub-sequently combined with a PWM and a model based on distribution of distances from TSS to gene start. We investi-gate the effect of different scoring techniques and quantify performance using area under a detection-error tradeoff curve. When tested on a biologically realistic task, our method provides performance comparable or superior to the best reported for this task. False positives are significantly reduced, an improvement of great significance to biologists.

Impact and interest:

40 citations in Scopus
Search Google Scholar™
37 citations in Web of Science®

Citation counts are sourced monthly from Scopus and Web of Science® citation databases.

These databases contain citations from different subsets of available publications and different time periods and thus the citation count from each is usually different. Some works are not in either database and no count is displayed. Scopus includes citations from articles published in 1996 onwards, and Web of Science® generally from 1980 onwards.

Citations counts from the Google Scholar™ indexing service can be viewed at the linked Google Scholar™ search.

Full-text downloads:

175 since deposited on 14 May 2007
53 in the past twelve months

Full-text downloads displays the total number of times this work’s files (e.g., a PDF) have been downloaded from QUT ePrints as well as the number of downloads in the previous 365 days. The count includes downloads for all files if a work has more than one.

ID Code: 7549
Item Type: Journal Article
Refereed: Yes
Keywords: bacterial promoters, support vector machines
DOI: 10.1093/bioinformatics/bti771
ISSN: 1460-2059
Subjects: Australian and New Zealand Standard Research Classification > MATHEMATICAL SCIENCES (010000) > APPLIED MATHEMATICS (010200) > Biological Mathematics (010202)
Australian and New Zealand Standard Research Classification > INFORMATION AND COMPUTING SCIENCES (080000) > ARTIFICIAL INTELLIGENCE AND IMAGE PROCESSING (080100) > Pattern Recognition and Data Mining (080109)
Australian and New Zealand Standard Research Classification > INFORMATION AND COMPUTING SCIENCES (080000) > ARTIFICIAL INTELLIGENCE AND IMAGE PROCESSING (080100) > Artificial Intelligence and Image Processing not elsewhere classified (080199)
Divisions: Past > QUT Faculties & Divisions > Faculty of Science and Technology
Current > Institutes > Institute of Health and Biomedical Innovation
Copyright Owner: Copyright 2006 (The authors): Licensed to Oxford University Press
Deposited On: 14 May 2007 00:00
Last Modified: 29 Feb 2012 13:18

Export: EndNote | Dublin Core | BibTeX

Repository Staff Only: item control page