Parallel hierarchical clustering using weighted confidence affinity
Baoying Wang, 
Imad Rahal and 
Aijuan Dong
International Journal of Data Mining, Modelling and Management, 2011, vol. 3, issue 2, 110-129
Abstract:
There have been many attempts for clustering categorical data such as market basket dataset. However, most of categorical clustering approaches belong to partitional clustering which requires at least one input parameter (e.g., the minimum intra-cluster similarity or the desired number of clusters). In this paper, we propose a parallelised hierarchical clustering approach for categorical data (PH-clustering) using vertical data structures. In order to minimise the impact of low support items, we devise a weighted confidence (WC) affinity function to compute the similarity between clusters. Based on our analysis of the major clustering steps, we adopt a partial local and partial global approach to reduce computation time as well as to keep network communication at minimum. Load balance issues are addressed especially during the data partitioning phase. Our experimental results on standardised market basket data show that the proposed weighted confidence affinity measure is more accurate than other contemporary affinity measures in the literature and that our parallel clustering approach provides magnitudes of time improvements over sequential clustering especially over larger data sizes. Our results also indicate that the number of items/attributes in the dataset has a more drastic impact on performance than the number of transactions/tuples.
Keywords: parallel clustering; hierarchical clustering; weighted confidence affinity; message passing interface; market basket data; data mining; data partitioning; parallel computing; vertical data structures; categorical clustering; load balancing. (search for similar items in EconPapers)
Date: 2011
References: Add references at CitEc 
Citations: 
Downloads: (external link)
http://www.inderscience.com/link.php?id=41491 (text/html)
Access to full text is restricted to subscribers.
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX 
RIS (EndNote, ProCite, RefMan) 
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:ids:ijdmmm:v:3:y:2011:i:2:p:110-129
Access Statistics for this article
More articles in International Journal of Data Mining, Modelling and Management  from  Inderscience Enterprises Ltd
Bibliographic data for series maintained by Sarah Parker ().