Performance Analysis of Transformer-Enabled Semantic Crawlers for Scalable Text Retrieval
Anil Kumar Sinha,
Khushboo Mishra,
Md Alimul Haque and
B. K. Mishra
Additional contact information
Anil Kumar Sinha: Department of Computer Science, V.K.S. University Ara, India
Khushboo Mishra: P.G. Department of Physics, V.K.S. University, Ara , India
Md Alimul Haque: Department of Computer Science, V.K.S. University Ara, India
B. K. Mishra: Department of Computer Science, V.K.S. University Ara, India
Diginomics, 2026, vol. 5, 309
Abstract:
With the exponential growth of web-based content, efficient retrieval of contextually relevant textual information starting from seed URLs has become a critical challenge in web content mining and information retrieval. Traditional crawling and search methods—such as breadth-first search (BFS), depth-first search (DFS), best-first (focused crawling), topic-sensitive PageRank, and context-graph models—typically suffer from limitations such as parameter tuning overhead, lack of contextual understanding, requirement of large training datasets, high computational cost, and the need for specialised infrastructure. This research presents a comprehensive comparative study of multiple search and crawling models applied to textual retrieval from seed URLs, with a particular focus on their performance in diverse web‐structures (static vs dynamic) and content types. Employing a unified experimental framework implemented in Python with MySQL backend, we evaluate each algorithm using standard performance metrics (precision, recall, F1-score) alongside newer metrics such as coverage, relevance score, search time, memory usage, throughput and harvest rate. Machine-learning enabled variants (for example semantic-BFS and semantic-DFS using transformer-based embeddings) are also incorporated to assess their value over purely structural methods. Our results demonstrate that while semantic-enhanced BFS (Semantic-BFS) yields higher coverage, better relevance and faster response time in many scenarios, it shows limitations in classical metrics like precision/recall/F1 when ground-truth labels are inadequate for semantic relevance. The study provides insights into algorithmic trade-offs, suitability for different web architectures, and proposes hybrid strategies for next-generation crawlers and retrieval systems. The findings contribute toward the design of more adaptive, semantic-aware, and scalable web content mining frameworks.
Keywords: Web Content Mining; Information Retrieval; Seed URL; Text Search Models; Link Analysis; Context Graph; BFS; DFS; Semantic Search; Algorithm Comparison; Machine Learning (search for similar items in EconPapers)
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
https://diginomics.ar/index.php/digi/article/view/309 Abstract page (text/html)
https://diginomics.ar/index.php/digi/article/download/309/249 Full text (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:cwg:digino:v:5:y:2026:id:309
DOI: 10.56294/digi2026309
Access Statistics for this article
More articles in Diginomics from Centro de Estudios de Economía Digital
Bibliographic data for series maintained by Prof. Carlos Alberto Gómez Cano ().