EFFICIENT UNSUPERVISED MINING FROM NOISY CO-OCCURRENCE DATA
Hiroshi Mamitsuka ()
Additional contact information
Hiroshi Mamitsuka: Institute for Chemical Research, Kyoto University, Gokasho, Uji 611-0011, Japan
New Mathematics and Natural Computation (NMNC), 2005, vol. 01, issue 01, 173-193
Abstract:
We consider the problem of mining from noisy unsupervised data sets. The data point we call noise is an outlier in the current context of data mining, and it has been generally defined as the one locates in low probability regions of an input space. The purpose of the approach for this problem is to detect outliers and to perform efficient mining from noisy unsupervised data. We propose a new iterative sampling approach for this problem, using both model-based clustering and the likelihood given to each example by a trained probabilistic model for finding data points of such low probability regions in an input space. Our method uses an arbitrary probabilistic model as a component model and repeats two steps of sampling non-outliers with high likelihoods (computed by previously obtained models) and training the model with the selected examples alternately. In our experiments, we focused on two-mode and co-occurrence data and empirically evaluated the effectiveness of our proposed method, comparing with two other methods, by using both synthetic and real data sets. From the experiments using the synthetic data sets, we found that the significance level of the performance advantage of our method over the two other methods had more pronounced for higher noise ratios, for both medium- and large-sized data sets. From the experiments using a real noisy data set of protein–protein interactions, a typical co-occurrence data set, we further confirmed the performance of our method for detecting outliers from a given data set. Extended abstracts of parts of the work presented in this paper have appeared in Refs. 1 and 2.
Keywords: Unsupervised learning; selective sampling; model-based clustering; co-occurrence data; protein–protein interactions (search for similar items in EconPapers)
Date: 2005
References: View complete reference list from CitEc
Citations:
Downloads: (external link)
http://www.worldscientific.com/doi/abs/10.1142/S1793005705000093
Access to full text is restricted to subscribers
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:wsi:nmncxx:v:01:y:2005:i:01:n:s1793005705000093
Ordering information: This journal article can be ordered from
DOI: 10.1142/S1793005705000093
Access Statistics for this article
New Mathematics and Natural Computation (NMNC) is currently edited by Paul P Wang
More articles in New Mathematics and Natural Computation (NMNC) from World Scientific Publishing Co. Pte. Ltd.
Bibliographic data for series maintained by Tai Tone Lim ().