EconPapers    
Economics at your fingertips  
 

Automated Class Numbers Prediction for Books: an AI/ML Based Approach Using Annif

Soumik Kerketta () and Parthasarathi Mukhopadhyay
Additional contact information
Soumik Kerketta: University of Kalyani, Junior Research Fellow (JRF), Department of Library and Information Science
Parthasarathi Mukhopadhyay: University of Kalyani, Professor, Department of Library and Information Science

A chapter in Proceedings of the International Conference on Marching Beyond the Libraries (ICMBL): Leadership, Creativity, and Innovation (ICMBL 2024), 2025, pp 140-147 from Springer

Abstract: Abstract In this research study as reported here, we endeavor to explore the possibilities of an AI/ML-based automated indexing system for the vast collections in a library. Library classification systems are considered pre-coordinated indexing approaches a while ago. Various machine learning techniques are applying to synthesizing classification numbers. A recently popular technique involves using a supervised learning algorithm to train a model on a set of documents that have been manually indexed/classified by their corresponding annotations using standardized terminology by trained library professionals’ experts using controlled vocabularies. The trained model learns patterns from the reference data and then predict the subject and class number for new documents. In the preliminary phase, we gathered a substantial collected around 2 lacks MARC-21 formatted bibliographic records where Tag 082 (DDC Call Number), Tag 245 (Title of Document), Tag 520 (Summary Note), and Tag 650 (Subject Descriptors) are contained in the datasets. After that processed this data using the data wrangling software named OpenRefine. Then dataset was subsequently divided into three sections: (i) a training dataset, (ii) a validation dataset and (ii) a test dataset. Here We usedAnnif, an open-source AI environment to analyze the dataset using the Dewey Decimal Classification (DDC) Scheme. Training Annif involved utilizing a substantial set of bibliographic records, based on the MARC-21 tags mentioned previously. In the next stage, the framework was trained using a various of backend algorithms, such asOmikuji, fastText, SVC (associative group), and simple and neural network (ensemble)based on neural network model. In order to assess the effectiveness of these models, all of these machine learning backends were finally compared using two crucial retrieval metrics: F1@5 and NDCG. When it comes to automated class number building, we have discovered that the neural network model outperforms rather than all other backends. This overall framework based on open-source software, an open dataset, and open standards.

Keywords: Annif; Automated indexing; DDC; F1@5; Library classification; Neural network; NDCG (search for similar items in EconPapers)
Date: 2025
References: Add references at CitEc
Citations:

There are no downloads for this item, see the EconPapers FAQ for hints about obtaining it.

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:spr:advbcp:978-94-6463-712-0_12

Ordering information: This item can be ordered from
http://www.springer.com/9789464637120

DOI: 10.2991/978-94-6463-712-0_12

Access Statistics for this chapter

More chapters in Advances in Economics, Business and Management Research from Springer
Bibliographic data for series maintained by Sonal Shukla () and Springer Nature Abstracting and Indexing ().

 
Page updated 2026-07-24
Handle: RePEc:spr:advbcp:978-94-6463-712-0_12