Cuisine: Classification using stylistic feature sets and/or name‐based feature sets
Yaakov HaCohen‐Kerner,
Hananya Beck,
Elchai Yehudai,
Mordechay Rosenstein and
Dror Mughaz
Journal of the American Society for Information Science and Technology, 2010, vol. 61, issue 8, 1644-1657
Abstract:
Document classification presents challenges due to the large number of features, their dependencies, and the large number of training documents. In this research, we investigated the use of six stylistic feature sets (including 42 features) and/or six name‐based feature sets (including 234 features) for various combinations of the following classification tasks: ethnic groups of the authors and/or periods of time when the documents were written and/or places where the documents were written. The investigated corpus contains Jewish Law articles written in Hebrew–Aramaic, which present interesting problems for classification. Our system CUISINE (Classification UsIng Stylistic feature sets and/or NamE‐based feature sets) achieves accuracy results between 90.71 to 98.99% for the seven classification experiments (ethnicity, time, place, ethnicity&time, ethnicity&place, time&place, ethnicity&time&place). For the first six tasks, the stylistic feature sets in general and the quantitative feature set in particular are enough for excellent classification results. In contrast, the name‐based feature sets are rather poor for these tasks. However, for the most complex task (ethnicity&time&place), a hill‐climbing model using all feature sets succeeds in significantly improving the classification results. Most of the stylistic features (34 of 42) are language‐independent and domain‐independent. These features might be useful to the community at large, at least for rather simple tasks.
Date: 2010
References: Add references at CitEc
Citations: View citations in EconPapers (1)
Downloads: (external link)
https://doi.org/10.1002/asi.21350
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:bla:jamist:v:61:y:2010:i:8:p:1644-1657
Ordering information: This journal article can be ordered from
https://doi.org/10.1002/(ISSN)1532-2890
Access Statistics for this article
More articles in Journal of the American Society for Information Science and Technology from Association for Information Science & Technology
Bibliographic data for series maintained by Wiley Content Delivery ().