EconPapers    
Economics at your fingertips  
 

Performance Analysis of Transformer-Enabled Semantic Crawlers for Scalable Text Retrieval

Anil Kumar Sinha, Khushboo Mishra, Md Alimul Haque and B. K. Mishra
Additional contact information
Anil Kumar Sinha: Department of Computer Science, V.K.S. University Ara, India
Khushboo Mishra: P.G. Department of Physics, V.K.S. University, Ara , India
Md Alimul Haque: Department of Computer Science, V.K.S. University Ara, India
B. K. Mishra: Department of Computer Science, V.K.S. University Ara, India

Diginomics, 2026, vol. 5, 309

Abstract: With the exponential growth of web-based content, efficient retrieval of contextually relevant textual information starting from seed URLs has become a critical challenge in web content mining and information retrieval. Traditional crawling and search methods—such as breadth-first search (BFS), depth-first search (DFS), best-first (focused crawling), topic-sensitive PageRank, and context-graph models—typically suffer from limitations such as parameter tuning overhead, lack of contextual understanding, requirement of large training datasets, high computational cost, and the need for specialised infrastructure. This research presents a comprehensive comparative study of multiple search and crawling models applied to textual retrieval from seed URLs, with a particular focus on their performance in diverse web‐structures (static vs dynamic) and content types. Employing a unified experimental framework implemented in Python with MySQL backend, we evaluate each algorithm using standard performance metrics (precision, recall, F1-score) alongside newer metrics such as coverage, relevance score, search time, memory usage, throughput and harvest rate. Machine-learning enabled variants (for example semantic-BFS and semantic-DFS using transformer-based embeddings) are also incorporated to assess their value over purely structural methods. Our results demonstrate that while semantic-enhanced BFS (Semantic-BFS) yields higher coverage, better relevance and faster response time in many scenarios, it shows limitations in classical metrics like precision/recall/F1 when ground-truth labels are inadequate for semantic relevance. The study provides insights into algorithmic trade-offs, suitability for different web architectures, and proposes hybrid strategies for next-generation crawlers and retrieval systems. The findings contribute toward the design of more adaptive, semantic-aware, and scalable web content mining frameworks.

Keywords: Web Content Mining; Information Retrieval; Seed URL; Text Search Models; Link Analysis; Context Graph; BFS; DFS; Semantic Search; Algorithm Comparison; Machine Learning (search for similar items in EconPapers)
Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://diginomics.ar/index.php/digi/article/view/309 Abstract page (text/html)
https://diginomics.ar/index.php/digi/article/download/309/249 Full text (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:cwg:digino:v:5:y:2026:id:309

DOI: 10.56294/digi2026309

Access Statistics for this article

More articles in Diginomics from Centro de Estudios de Economía Digital
Bibliographic data for series maintained by Prof. Carlos Alberto Gómez Cano ().

 
Page updated 2026-07-19
Handle: RePEc:cwg:digino:v:5:y:2026:id:309