EconPapers    
Economics at your fingertips  
 

HS3D, A DATASET OF HOMO SAPIENS SPLICE REGIONS, AND ITS EXTRACTION PROCEDURE FROM A MAJOR PUBLIC DATABASE

Pasquale Pollastro () and Salvatore Rampone ()
Additional contact information
Pasquale Pollastro: Facoltà di Scienze, Università del Sannio, Via Port'arsa 11, Benevento, I-82100, Italy
Salvatore Rampone: Department of Geological and Environmental Studies, Università del Sannio, Via Port'arsa 11, Benevento, I-82100, Italy

International Journal of Modern Physics C (IJMPC), 2002, vol. 13, issue 08, 1105-1117

Abstract: The aim of this work is to describe a cleaning procedure of GenBank data, producing material to train and to assess the prediction accuracy of computational approaches for gene characterization. A procedure (GenBank2HS3D) has been defined, producing a dataset (HS3D — Homo Sapiens Splice Sites Dataset) of Homo Sapiens Splice regions extracted from GenBank (Rel.123 at this time). It selects, from the complete GenBank Primate Division, entries of Human Nuclear DNA according with several assessed criteria; then it extracts exons and introns from these entries (actually 4523 + 3802). Donor and acceptor sites are then extracted as windows of 140 nucleotides around each splice site (3799 + 3799). After discarding windows not including canonical GT–AG junctions (65 + 74), including insufficient data (not enough material for a 140 nucleotide window) (686 + 589), including not AGCT bases (29 + 30), and redundant (218 + 226), the remaining windows (2796 + 2880) are reported in the dataset. Finally, windows of false splice sites are selected by searching canonical GT–AG pairs in not splicing positions (271 937 + 332 296). The false sites in a range +/- 60 from a true splice site are marked as proximal. HS3D, release 1.2 at this time, is available at the Web server of the University of Sannio:.

Keywords: Extraction algorithm; GenBank; splice sites; dataset (search for similar items in EconPapers)
Date: 2002
References: Add references at CitEc
Citations:

Downloads: (external link)
http://www.worldscientific.com/doi/abs/10.1142/S0129183102003796
Access to full text is restricted to subscribers

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:wsi:ijmpcx:v:13:y:2002:i:08:n:s0129183102003796

Ordering information: This journal article can be ordered from

DOI: 10.1142/S0129183102003796

Access Statistics for this article

International Journal of Modern Physics C (IJMPC) is currently edited by H. J. Herrmann

More articles in International Journal of Modern Physics C (IJMPC) from World Scientific Publishing Co. Pte. Ltd.
Bibliographic data for series maintained by Tai Tone Lim ().

 
Page updated 2025-03-20
Handle: RePEc:wsi:ijmpcx:v:13:y:2002:i:08:n:s0129183102003796