Fraudulent Patterns: Unveiling Deception Through Text Analysis and Topic Modeling
Abdelrahim Al Aqqad ()
Chapter Chapter 19 in Fraud Analytics in Action, 2026, pp 459-490 from Springer
Abstract:
Abstract This chapter introduces text mining and topic modeling as practical tools for detecting fraud in unstructured data such as corporate emails. Using the Enron email dataset as a case study, the chapter presents a complete natural language processing pipeline covering tokenization, stopword removal, lemmatization, and stemming to convert raw text into clean, structured data. Readers learn to perform keyword searches on DataFrames using pandas string operations, build fraud dictionaries with multiple search terms, and create binary flag variables to identify suspicious content. The chapter then covers Latent Dirichlet Allocation (LDA), an unsupervised topic modeling technique, demonstrating how to construct a bag-of-words corpus using Gensim, train an LDA model, and interpret its topic distributions. Visualization with pyLDAvis is introduced for interactive exploration of topic clusters. The final section shows how to assign dominant topics to individual documents and flag emails strongly associated with suspect themes. Lab exercises use Python with NLTK, Gensim, pandas, and NumPy, accompanied by a Jupyter notebook on the companion GitHub repository. By the end, readers can apply text analytics to real-world fraud detection workflows.
Date: 2026
References: Add references at CitEc
Citations:
There are no downloads for this item, see the EconPapers FAQ for hints about obtaining it.
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:spr:sprchp:978-3-032-16023-2_19
Ordering information: This item can be ordered from
http://www.springer.com/9783032160232
DOI: 10.1007/978-3-032-16023-2_19
Access Statistics for this chapter
More chapters in Springer Books from Springer
Bibliographic data for series maintained by Sonal Shukla () and Springer Nature Abstracting and Indexing ().