Modeling of Hyperparameter Tuned Deep Learning Model for Automated Image Captioning

Omri, Mohamed; Abdel-Khalek, Sayed; Khalil, Eied M.; Bouslimi, Jamel; Joshi, Gyanendra Prasad

Modeling of Hyperparameter Tuned Deep Learning Model for Automated Image Captioning

Mohamed Omri, Sayed Abdel-Khalek, Eied M. Khalil, Jamel Bouslimi and Gyanendra Prasad Joshi
Additional contact information
Mohamed Omri: Deanship of Scientific Research, King Abdulaziz University, Jeddah 21589, Saudi Arabia
Sayed Abdel-Khalek: Mathematics Department, Faculty of Science, Taif University, Taif 21944, Saudi Arabia
Eied M. Khalil: Department of Mathematics, Faculty of Science, Taif University, Taif 21944, Saudi Arabia
Jamel Bouslimi: Physics Department, Faculty of Science, Taif University, Taif 21944, Saudi Arabia
Gyanendra Prasad Joshi: Department of Computer Science and Engineering, Sejong University, Seoul 05006, Korea

Mathematics, 2022, vol. 10, issue 3, 1-20

Abstract: Image processing remains a hot research topic among research communities due to its applicability in several areas. An important application of image processing is the automatic image captioning technique, which intends to generate a proper description of an image in a natural language automated. Image captioning is a recently developed hot research topic, and it started to receive significant attention in the field of computer vision and natural language processing (NLP). Since image captioning is considered a challenging task, the recently developed deep learning (DL) models have attained significant performance with increased complexity and computational cost. Keeping these issues in mind, in this paper, a novel hyperparameter tuned DL for automated image captioning (HPTDL-AIC) technique is proposed. The HPTDL-AIC technique encompasses two major parts, namely encoder and decoder. The encoder part utilizes Faster SqueezNet with the RMSProp model to generate an effective depiction of the input image via insertion into a predefined length vector. At the same time, the decoder unit employs a bird swarm algorithm (BSA) with long short-term memory (LSTM) model to concentrate on the generation of description sentences. The design of RMSProp and BSA for the hyperparameter tuning process of the Faster SqueezeNet and LSTM models for image captioning shows the novelty of the work, which helps to accomplish enhanced image captioning performance. The experimental validation of the HPTDL-AIC technique is carried out against two benchmark datasets, and the extensive comparative study pointed out the improved performance of the HPTDL-AIC technique over recent approaches.

Keywords: image captioning; deep learning; machine learning; encoder; decoder; hyperparameter tuning (search for similar items in EconPapers)
JEL-codes: C (search for similar items in EconPapers)
Date: 2022
References: View complete reference list from CitEc
Citations: View citations in EconPapers (4)

Downloads: (external link)
https://www.mdpi.com/2227-7390/10/3/288/pdf (application/pdf)
https://www.mdpi.com/2227-7390/10/3/288/ (text/html)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:gam:jmathe:v:10:y:2022:i:3:p:288-:d:726992

Access Statistics for this article

Mathematics is currently edited by Ms. Emma He

More articles in Mathematics from MDPI
Bibliographic data for series maintained by MDPI Indexing Manager ().