返回
Multi-Task Deep Relative Attribute Learning for Visual Urban Perception
DOI:10.1109/TIP.2019.2932502.png)
摘要
En 中文
Visual urban perception aims to quantify perceptual attributes (e.g., safe and depressing attributes) of physical urban environment from crowd-sourced street-view images and their pairwise comparisons. It has been receiving more and more attention in computer vision for various applications, such as perceptive attribute learning and urban scene understanding. Most existing methods adopt either 1) a regression model trained using image features and ranked scores converted from pairwise comparisons for perceptual attribute prediction or 2) a pairwise ranking algorithm to independently learn each perceptual attribute. However, the former fails to directly exploit pairwise comparisons while the latter ignores the relationship among different attributes. To address them, we propose a multi-task deep relative attribute learning network (MTDRALN) to learn all the relative attributes simultaneously via multi-task Siamese networks, where each Siamese network will predict one relative attribute. Combined with deep relative attribute learning, we utilize the structured sparsity to exploit the prior from natural attribute grouping, where all the attributes are divided into different groups based on semantic relatedness in advance. As a result, MTDRALN is capable of learning all the perceptual attributes simultaneously via multi-task learning. Besides the ranking sub-network, MTDRALN further introduces the classification sub-network, and these two types of losses from two sub-networks jointly constrain parameters of the deep network to make the network learn more discriminative visual features for relative attribute learning. In addition, our network can be trained in an end-to-end way to make deep feature learning and multi-task relative attribute learning reinforces each other. Extensive experiments on the large-scale Place Pulse 2.0 dataset validate the advantage of our proposed network. Our qualitative results along with visualization of saliency maps also show that the proposed network is able to learn effective features for perceptual attributes.
Keyword:
Visualization
Task analysis
Deep learning
Urban areas
Correlation
Computer vision
Predictive models
Visual urban perception
relative attribute
multi-task learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
13.7
论文数:
1.0W
被引数:
8.4W
机构
引用论文
Social sensing from street-level imagery: A case study in learning spatio-temporal urban mobility patterns街道层面图像的社会感知: 学习时空城市流动模式的案例研究
Using deep learning and Google Street View to estimate the demographic makeup of neighborhoods across the United States使用深度学习和Google街景视图估算美国社区的人口构成

