Computation of Large Spatial Datasets with the M function
Eric Marcon () and
Florence Puech ()
Additional contact information
Eric Marcon: UMR AMAP - Botanique et Modélisation de l'Architecture des Plantes et des Végétations - Cirad - Centre de Coopération Internationale en Recherche Agronomique pour le Développement - CNRS - Centre National de la Recherche Scientifique - IRD [Occitanie] - Institut de Recherche pour le Développement - délégation Occitanie - IRD - Institut de Recherche pour le Développement - INRAE - Institut National de Recherche pour l’Agriculture, l’Alimentation et l’Environnement - UM - Université de Montpellier, AgroParisTech
Florence Puech: UMR PSAE - Paris-Saclay Applied Economics - AgroParisTech - Université Paris-Saclay - INRAE - Institut National de Recherche pour l’Agriculture, l’Alimentation et l’Environnement
Working Papers from HAL
Abstract:
Since agglomeration is a core question in regional science, spatial concentration measures are widely employed in that field to evaluate the spatial distribution of activities. Increasing access to large geo-referenced datasets, coupled with the development of computing power, has encouraged the search for suitable spatial statistical tools. Distance-based methods have been extensively developed to detect spatial concentration, dispersion or independence of entities at any distance and without any bias. Recently, Tidu et al. (2024) highlighted the qualities of Marcon and Puech's M function, a relative distance-based measure, and also expressed reservations about the computation time required. Herein, we explore two possible ways to reduce the computation burden of large geo-located datasets: approximating the position of points and thinning the point pattern. In both cases, the deterioration extent of the M results is estimated and discussed as the gains it provides in computation time, using the R software. We also discuss implications of these findings in the field of regional science. We notably provide evidence that the individual location approximation generates information loss at substantially small distances, implying a trade-off between the smallest distance at which spatial interactions could be detected and computing performance. We also give support that thinning is an efficient method to analyze large datasets with very good accuracy. The R code used in the article is given for the reproducibility of our results.
Keywords: Distance-based method; M-function; Performance test; R Package dbmss; MAUP (search for similar items in EconPapers)
Date: 2026-07-17
Note: View the original document on HAL open archive server: https://hal.science/hal-05512154v2
References: Add references at CitEc
Citations:
Downloads: (external link)
https://hal.science/hal-05512154v2/document (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:hal:wpaper:hal-05512154
Access Statistics for this paper
More papers in Working Papers from HAL
Bibliographic data for series maintained by CCSD ().