Conference proceeding
GO for gene documents
Proceedings of the 1st international workshop on text mining in bioinformatics, pp.43-51
TMBIO '06
11/10/2006
DOI: 10.1145/1183535.1183546
Abstract
Annotating genes and their products with Gene Ontology codes is an important area of research. One approach for doing this is to use the information available about these genes in the biomedical literature. Our goal, based on this approach, is to develop automatic methods for annotation that could supplement the expensive manual annotation processes currently in place. Using a set of Support Vector Machines (SVM) classifiers we were able to achieve Fscores of 0.48, 0.4 and 0.32 for codes of the molecular function, cellular component and biological process GO hierarchies respectively. We explore thresholding of SVM scores, the relationship of performance to hierarchy level and to the number of positives in the training sets. We find that hierarchy level is important especially for the molecular function and biological process hierarchies. We find that the cellular component hierarchy stands apart from the other two in many respects. This may be due to fundamental differences in link semantics. This research also exploits the hierarchical structures by defining and testing a relaxed criteria for classification correctness.
Details
- Title: Subtitle
- GO for gene documents
- Creators
- Xin Ying QiuPadmini Srinivasan - University of Iowa, Computer Science
- Resource Type
- Conference proceeding
- Publication Details
- Proceedings of the 1st international workshop on text mining in bioinformatics, pp.43-51
- Series
- TMBIO '06
- DOI
- 10.1145/1183535.1183546
- Publisher
- ACM
- Language
- English
- Date published
- 11/10/2006
- Academic Unit
- Nursing; Computer Science; Business Analytics
- Record Identifier
- 9984003796202771
Metrics
18 Record Views