arrow
返回

Toward Visual Syntactical Understanding

delete2024-01-01
delete0
delete
OA
AI
S
Sayeed Shafayet Chowdhury *
S
Soumyadeep Chandra
K
Kaushik Roy
DOI:10.1109/ACCESS.2024.3397061delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Syntax is usually studied in the realm of linguistics and refers to the arrangement of words in a sentence. Similarly, an image can be considered as a visual 'sentence', with the semantic parts of the image acting as 'words'. While visual syntactic understanding occurs naturally to humans, it is interesting to explore whether deep neural networks (DNNs) are equipped with such reasoning. To that end, we alter the syntax of natural images (e.g. swapping the eye and nose of a face), referred to as 'incorrect' images, to investigate the sensitivity of DNNs to such syntactic anomaly. Through our experiments, we discover an intriguing property of DNNs where we observe that state-of-the-art convolutional neural networks, as well as vision transformers, fail to discriminate between syntactically correct and incorrect images, when trained on only correct ones. To counter this issue and enable visual syntactic understanding with DNNs, we propose a three-stage framework- (i) the 'words' (or the sub-features) in the image are detected, (ii) the detected words are sequentially masked and reconstructed using an autoencoder, (iii) the original and reconstructed parts are compared at each location to determine syntactic correctness. The reconstruction module is trained with BERT-like masked autoencoding for images, with the motivation to leverage language model inspired training to better capture the syntax. Note, our proposed approach is unsupervised in the sense that the incorrect images are only used during testing and the correct versus incorrect labels are never used for training. We perform experiments on CelebA, and AFHQ datasets and obtain classification accuracy of 92.10%, and 90.89%, respectively. Notably, the approach generalizes well to ImageNet samples which share common classes with CelebA and AFHQ without explicitly training on them.
Keyword:
Syntactics
Image reconstruction
Visualization
Face recognition
Semantics
Training
Linguistics
Artificial neural networks
Image syntax
interpretability of DNNs
novel problem of DNNs
syntactic understanding
vision transformers

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

Purdue University System 封面图
Purdue University System
学者数:
3.9W
论文数: 3.6W
被引数: 66
引用论文

引用论文

The serine/threonine protein phosphatase SIT4 modulates yeast‐to‐hypha morphogenesis and virulence in Candida albicans
err2004-01-20
err0
errOAAI
errChang‐Muk Lee; André Nantel; Linghuo Jiang; Malcolm Whiteway; Shi‐Hsiang Shen
err分享
err收藏
err分享
err收藏
YAC transgene-mediated olfactory receptor gene choiceYAC转基因介导的嗅觉受体基因选择
err2000-02-01
err0
errOAAI
errFarah A.W. Ebrahimi; James Edmondson; Rodney Rothstein; Andrew Chess
err分享
err收藏
err分享
err收藏
Structural Neural Substrates of Reading the Mind in the Eyes在眼睛中阅读心灵的结构神经基质
err2016-04-11
err0
errOAAI
errWataru Sato; Takanori Kochiyama; Shota Uono; Reiko Sawada; Yasutaka Kubota; Sayaka Yoshimura; Motomi Toichi
err分享
err收藏
学者 查看更多内容