A Survey of Available Techniques for Naive Artificial Intelligence based System for Conversing with a Human
Rayyan Hashmi and
Ayan Rajput
International Journal of Scientific Research in Science and Technology, 2024, vol. 11, issue 3, 865-876
Abstract:
Visual questions selectively target different areas of an image, including background details and underlying context. As a result, a system that succeeds at VQA typically needs a more detailed understanding of the image and complex reasoning than a system producing generic image captions. Moreover, VQA is amenable to automatic evaluation, since many open-ended answers contain only a few words or a closed set of answers that can be provided in a multiple-choice format. We provide a dataset containing ∼0.25M images, ∼0.76M questions, and ∼10M answers (www.visualqa.org), and discuss the information it provides. Numerous baselines and methods for VQA are provided and compared with human performance. ”
Keywords: Visual Question Answering; visualqa; AI; Natural Language Processing (search for similar items in EconPapers)
Date: 2024
References: Add references at CitEc
Citations:
Downloads: (external link)
https://ijsrst.com/home/article/view/IJSRST2411362 Abstract page (text/html)
https://ijsrst.com/home/article/download/IJSRST2411362/IJSRST2411362 Full text (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:etm:ijsrst:v11:y2024:i3:id:272
DOI: 10.32628/IJSRST2411362
Access Statistics for this article
More articles in International Journal of Scientific Research in Science and Technology from Technoscience Academy
Bibliographic data for series maintained by Pankaj Sharma ().