EconPapers    
Economics at your fingertips  
 

Cuisine: Classification using stylistic feature sets and/or name‐based feature sets

Yaakov HaCohen‐Kerner, Hananya Beck, Elchai Yehudai, Mordechay Rosenstein and Dror Mughaz

Journal of the American Society for Information Science and Technology, 2010, vol. 61, issue 8, 1644-1657

Abstract: Document classification presents challenges due to the large number of features, their dependencies, and the large number of training documents. In this research, we investigated the use of six stylistic feature sets (including 42 features) and/or six name‐based feature sets (including 234 features) for various combinations of the following classification tasks: ethnic groups of the authors and/or periods of time when the documents were written and/or places where the documents were written. The investigated corpus contains Jewish Law articles written in Hebrew–Aramaic, which present interesting problems for classification. Our system CUISINE (Classification UsIng Stylistic feature sets and/or NamE‐based feature sets) achieves accuracy results between 90.71 to 98.99% for the seven classification experiments (ethnicity, time, place, ethnicity&time, ethnicity&place, time&place, ethnicity&time&place). For the first six tasks, the stylistic feature sets in general and the quantitative feature set in particular are enough for excellent classification results. In contrast, the name‐based feature sets are rather poor for these tasks. However, for the most complex task (ethnicity&time&place), a hill‐climbing model using all feature sets succeeds in significantly improving the classification results. Most of the stylistic features (34 of 42) are language‐independent and domain‐independent. These features might be useful to the community at large, at least for rather simple tasks.

Date: 2010
References: Add references at CitEc
Citations: View citations in EconPapers (1)

Downloads: (external link)
https://doi.org/10.1002/asi.21350

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:bla:jamist:v:61:y:2010:i:8:p:1644-1657

Ordering information: This journal article can be ordered from
https://doi.org/10.1002/(ISSN)1532-2890

Access Statistics for this article

More articles in Journal of the American Society for Information Science and Technology from Association for Information Science & Technology
Bibliographic data for series maintained by Wiley Content Delivery ().

 
Page updated 2025-03-19
Handle: RePEc:bla:jamist:v:61:y:2010:i:8:p:1644-1657