HS3D, A DATASET OF HOMO SAPIENS SPLICE REGIONS, AND ITS EXTRACTION PROCEDURE FROM A MAJOR PUBLIC DATABASE
Pasquale Pollastro () and
Salvatore Rampone ()
Additional contact information
Pasquale Pollastro: Facoltà di Scienze, Università del Sannio, Via Port'arsa 11, Benevento, I-82100, Italy
Salvatore Rampone: Department of Geological and Environmental Studies, Università del Sannio, Via Port'arsa 11, Benevento, I-82100, Italy
International Journal of Modern Physics C (IJMPC), 2002, vol. 13, issue 08, 1105-1117
Abstract:
The aim of this work is to describe a cleaning procedure of GenBank data, producing material to train and to assess the prediction accuracy of computational approaches for gene characterization. A procedure (GenBank2HS3D) has been defined, producing a dataset (HS3D — Homo Sapiens Splice Sites Dataset) of Homo Sapiens Splice regions extracted from GenBank (Rel.123 at this time). It selects, from the complete GenBank Primate Division, entries of Human Nuclear DNA according with several assessed criteria; then it extracts exons and introns from these entries (actually 4523 + 3802). Donor and acceptor sites are then extracted as windows of 140 nucleotides around each splice site (3799 + 3799). After discarding windows not including canonical GT–AG junctions (65 + 74), including insufficient data (not enough material for a 140 nucleotide window) (686 + 589), including not AGCT bases (29 + 30), and redundant (218 + 226), the remaining windows (2796 + 2880) are reported in the dataset. Finally, windows of false splice sites are selected by searching canonical GT–AG pairs in not splicing positions (271 937 + 332 296). The false sites in a range +/- 60 from a true splice site are marked as proximal. HS3D, release 1.2 at this time, is available at the Web server of the University of Sannio:.
Keywords: Extraction algorithm; GenBank; splice sites; dataset (search for similar items in EconPapers)
Date: 2002
References: Add references at CitEc
Citations:
Downloads: (external link)
http://www.worldscientific.com/doi/abs/10.1142/S0129183102003796
Access to full text is restricted to subscribers
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:wsi:ijmpcx:v:13:y:2002:i:08:n:s0129183102003796
Ordering information: This journal article can be ordered from
DOI: 10.1142/S0129183102003796
Access Statistics for this article
International Journal of Modern Physics C (IJMPC) is currently edited by H. J. Herrmann
More articles in International Journal of Modern Physics C (IJMPC) from World Scientific Publishing Co. Pte. Ltd.
Bibliographic data for series maintained by Tai Tone Lim ().