Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points
Yunlai Yang and
Zhenzhu Wan
Journal of Applied Mathematics, 2026, vol. 2026, 1-10
Abstract:
In biomedical sciences, agricultural sciences, and geosciences, similarities are assessed among samples by using their compositional data for various applications. Pearson correlation coefficient and cosine similarity are commonly applied for these assessments. In this paper, we demonstrate that, for one type of samples with a data range over multiple orders of magnitude and clustered data points, it is not proper to use Pearson correlation coefficient or cosine similarity to measure the similarity. This is because the effect of individual data points in the cluster is suppressed, i.e., not equally treated when implementing Pearson correlation analysis; and the effect of low value data points is reduced in cosine similarity analysis. To properly assess the similarity for this special type of samples, based on the meaning of similarity, we propose a new similarity coefficient. Similarity is actually about the compositional proportion of two samples, the closer the compositional proportion among the data points, the higher the similarity between the two samples. Therefore, the similarity of the ratios (gradients), which are the measure of compositional proportion of the data points between two samples, can be used to measure their similarity. Because the gradients are independent of actual values of individual data points, the effect of each data point on the similarity coefficient is treated equally. This new similarity coefficient is thus scale-independent and not affected by data clustering. Therefore, the new similarity coefficient can represent similarity more accurately than Pearson correlation coefficient or cosine similarity for this type of samples. The limitation of Pearson correlation coefficient and cosine similarity and advantage of the new similarity coefficient are demonstrated here by analyzing three sets of natural samples.
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
http://downloads.hindawi.com/journals/jam/2026/6492494.pdf (application/pdf)
http://downloads.hindawi.com/journals/jam/2026/6492494.xml (application/xml)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:hin:jnljam:6492494
DOI: 10.1155/jama/6492494
Access Statistics for this article
More articles in Journal of Applied Mathematics from Hindawi
Bibliographic data for series maintained by Mohamed Abdelhakeem ().