How Small is Big Enough? Open Labeled Datasets and the Development of Deep Learning

de Souza, Daniel; Geuna, Aldo; Rodr\'iguez, Jeff

How Small is Big Enough? Open Labeled Datasets and the Development of Deep Learning

Daniel de Souza, Aldo Geuna and Jeff Rodr\'iguez

Abstract: We investigate the emergence of Deep Learning as a technoscientific field, emphasizing the role of open labeled datasets. Through qualitative and quantitative analyses, we evaluate the role of datasets like CIFAR-10 in advancing computer vision and object recognition, which are central to the Deep Learning revolution. Our findings highlight CIFAR-10's crucial role and enduring influence on the field, as well as its importance in teaching ML techniques. Results also indicate that dataset characteristics such as size, number of instances, and number of categories, were key factors. Econometric analysis confirms that CIFAR-10, a small-but-sufficiently-large open dataset, played a significant and lasting role in technological advancements and had a major function in the development of the early scientific literature as shown by citation metrics.

Date: 2024-08
New Economics Papers: this item is included in nep-big
References: View references in EconPapers View complete reference list from CitEc
Citations:

Downloads: (external link)
https://arxiv.org/pdf/2408.10359 Latest version (application/pdf)

Related works:
Journal Article: How small is big enough? Open labeled datasets and the development of deep learning (2025)
Working Paper: How Small is Big Enough? Open Labeled Datasets and the Development of Deep Learning (2025)
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:arx:papers:2408.10359

Access Statistics for this paper

More papers in Papers from arXiv.org
Bibliographic data for series maintained by arXiv administrators ().