EconPapers    
Economics at your fingertips  
 

Clustering multivariate count data via Dirichlet-multinomial network fusion

Xin Zhao, Jingru Zhang and Wei Lin

Computational Statistics & Data Analysis, 2023, vol. 179, issue C

Abstract: Clustering of multivariate count data has widespread applications in areas such as text analysis and microbiome studies. The need to account for overdispersion generally results in a nonconvex loss function, which does not fit into the existing convex clustering framework. Moreover, prior knowledge of a network over the samples, often available from citation or similarity relationships, is not taken into account. We introduce Dirichlet-multinomial network fusion (DMNet) for clustering multivariate count data, which models the samples via Dirichlet-multinomial distributions with individual parameters and employs a weighted group L1 fusion penalty to pursue homogeneity over a prespecified network. To circumvent the nonconvexity issue, we present two exponential family approximations to the Dirichlet-multinomial distribution, which are amenable to efficient optimization and theoretical analysis. We derive an ADMM algorithm and establish nonasymptotic error bounds for the proposed methods. Our bounds involve a trade-off between the connectivity of the network and its fidelity to the true parameter. The usefulness of our methods is illustrated through simulation studies and two text clustering applications.

Keywords: Convex clustering; Exponential family approximation; Group L1 fusion; Nonasymptotic error bound; Overdispersion; Text analysis (search for similar items in EconPapers)
Date: 2023
References: View references in EconPapers View complete reference list from CitEc
Citations:

Downloads: (external link)
http://www.sciencedirect.com/science/article/pii/S0167947322002146
Full text for ScienceDirect subscribers only.

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:eee:csdana:v:179:y:2023:i:c:s0167947322002146

DOI: 10.1016/j.csda.2022.107634

Access Statistics for this article

Computational Statistics & Data Analysis is currently edited by S.P. Azen

More articles in Computational Statistics & Data Analysis from Elsevier
Bibliographic data for series maintained by Catherine Liu ().

 
Page updated 2025-03-19
Handle: RePEc:eee:csdana:v:179:y:2023:i:c:s0167947322002146