Region Collaborative Network for Detection-Based Vision-Language Understanding
Linyan Li,
Kaile Du,
Minming Gu,
Fuyuan Hu () and
Fan Lyu
Additional contact information
Linyan Li: Suzhou Institute of Trade & Commerce, Suzhou 215009, China
Kaile Du: Electronic & Information Engineering, Suzhou University of Science and Technology, Suzhou 215009, China
Minming Gu: Electronic & Information Engineering, Suzhou University of Science and Technology, Suzhou 215009, China
Fuyuan Hu: Electronic & Information Engineering, Suzhou University of Science and Technology, Suzhou 215009, China
Fan Lyu: College of Intelligence and Computing, Tianjin University, Tianjin 300000, China
Mathematics, 2022, vol. 10, issue 17, 1-12
Abstract:
Given a query language, a Detection-based Vision-Language Understanding (DVLU) system needs to respond based on the detected regions (i.e.,bounding boxes). With the significant advancement in object detection, DVLU has witnessed great improvements in recent years, such as Visual Question Answering (VQA) and Visual Grounding (VG). However, existing DVLU methods always process each detected image region separately but ignore that they were an integral whole. Without the full consideration of each region’s context, the image’s understanding may contain more bias. In this paper, to solve the problem, a simple yet effective Region Collaborative Network (RCN) block is proposed to bridge the gap between independent regions and the integrative DVLU task. Specifically, the Intra-Region Relations (IntraRR) inside each detected region are computed by a position-wise and channel-wise joint non-local model. Then, the Inter-Region Relations (InterRR) across all the detected regions are computed by pooling and sharing parameters with IntraRR. The proposed RCN can enhance the features of each region by using information from all other regions and guarantees the dimension consistency between input and output. The RCN is evaluated on VQA and VG, and the experimental results show that our method can significantly improve the performance of existing DVLU models.
Keywords: detection-based vision-language understanding; region collaborative network; non-local network (search for similar items in EconPapers)
JEL-codes: C (search for similar items in EconPapers)
Date: 2022
References: View complete reference list from CitEc
Citations:
Downloads: (external link)
https://www.mdpi.com/2227-7390/10/17/3110/pdf (application/pdf)
https://www.mdpi.com/2227-7390/10/17/3110/ (text/html)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:gam:jmathe:v:10:y:2022:i:17:p:3110-:d:901356
Access Statistics for this article
Mathematics is currently edited by Ms. Emma He
More articles in Mathematics from MDPI
Bibliographic data for series maintained by MDPI Indexing Manager ().