EconPapers    
Economics at your fingertips  
 

Theoretical measures of relative performance of classifiers for high dimensional data with small sample sizes

Peter Hall, Yvonne Pittelkow and Malay Ghosh

Journal Of The Royal Statistical Society Series B, 2008, vol. 70, issue 1, pages 159-173

Abstract: We suggest a technique, related to the concept of 'detection boundary' that was developed by Ingster and by Donoho and Jin, for comparing the theoretical performance of classifiers constructed from small training samples of very large vectors. The resulting 'classification boundaries' are obtained for a variety of distance-based methods, including the support vector machine, distance-weighted discrimination and "k"th-nearest-neighbour classifiers, for thresholded forms of those methods, and for techniques based on Donoho and Jin's higher criticism approach to signal detection. Assessed in these terms, standard distance-based methods are shown to be capable only of detecting differences between populations when those differences can be estimated consistently. However, the thresholded forms of distance-based classifiers can do better, and in particular can correctly classify data even when differences between distributions are only detectable, not estimable. Other methods, including higher criticism classifiers, can on occasion perform better still, but they tend to be more limited in scope, requiring substantially more information about the marginal distributions. Moreover, as tail weight becomes heavier the classification boundaries of methods designed for particular distribution types can converge to, and achieve, the boundary for thresholded nearest neighbour approaches. For example, although higher criticism has a lower classification boundary, and in this sense performs better, in the case of normal data, the boundaries are identical for exponentially distributed data when both sample sizes equal 1. Copyright 2008 Royal Statistical Society.

Date: 2008

Downloads: (external link)
http://www.blackwell-synergy.com/doi/abs/10.1111/j.1467-9868.2007.00631.x link to full text (text/html)
Access to full text is restricted to subscribers.

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: http://EconPapers.repec.org/RePEc:bla:jorssb:v:70:y:2008:i:1:p:159-173

Ordering information: This journal article can be ordered from
http://www.blackwell ... bs.asp?ref=1369-7412

Access Statistics for this article

Journal Of The Royal Statistical Society Series B is edited by C. Robert and A. T. A. Wood

More articles in Journal Of The Royal Statistical Society Series B from Royal Statistical Society
Series data maintained by Christopher F. Baum ().

 
Page updated 2009-11-23
Handle: RePEc:bla:jorssb:v:70:y:2008:i:1:p:159-173