arrow
Return

Grounded Sequence to Sequence Transduction

delete2020-03-01
delete0
PRE
AI
L
Lucia Specia
L
Loïc Barrault
O
Ozan Çağlayan
A
Amanda Duarte
D
Desmond Elliott
S
Spandana Gella
N
Nils Holzenberger
S
Sun Jae Lee
J
Jindřich Libovický
P
Pranava Madhyastha
F
Florian Metze
K
Karl Mulligan
S
Shruti Palaskar *
J
Josiah Wang
R
Raman Arora
DOI:10.1109/JSTSP.2020.2998415delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Speech recognition and machine translation have made major progress over the past decades, providing practical systems to map one language sequence to another. Although multiple modalities such as sound and video are becoming increasingly available, the state-of-the-art systems are inherently unimodal, in the sense that they take a single modality - either speech or text - as input. Evidence from human learning suggests that additional modalities can provide disambiguating signals crucial for many language tasks. In this article, we describe the How2 dataset , a large, open-domain collection of videos with transcriptions and their translations. We then show how this single dataset can be used to develop systems for a variety of language tasks and present a number of models meant as starting points. Across tasks, we find that building multimodal architectures that perform better than their unimodal counterpart remains a challenge. This leaves plenty of room for the exploration of more advanced solutions that fully exploit the multimodal nature of the How2 dataset , and the general direction of multimodal learning with other datasets as well.
Keywords:
Visualization
Feature extraction
Speech recognition
Task analysis
Training
Adaptation models
Grounding
multimodal machine learning
speech recognition
machine translation
representation learning
summarization
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Journal of Selected Topics in Signal Processing cover
IEEE Journal of Selected Topics in Signal Processing
IF:
13.7
Papers:
1.9K
Citations:
1.1W

Organization

U
University of Sheffield
Scholars:
3.0W
Papers: 2.9W
Citations: 3.9W
U
University of Copenhagen
Scholars:
7.6W
Papers: 6.6W
Citations: 86
C
Carnegie Mellon University
Scholars:
1.4W
Papers: 1.4W
Citations: 2.7W
U
university of pennsylvania
Scholars:
9.2W
Papers: 7.8W
Citations: 153
U
University of Munich
Scholars:
5.7W
Papers: 4.2W
Citations: 68
W
Worcester Polytechnic Institute
Scholars:
3.6K
Papers: 3.0K
Citations: 28
J
Johns Hopkins University
Scholars:
10.2W
Papers: 8.8W
Citations: 13.0W
U
universitat politecnica de catalunya
Scholars:
1.9W
Papers: 1.6W
Citations: 17
I
Imperial College London
Scholars:
8.3W
Papers: 7.3W
Citations: 11.1W
U
University of Edinburgh
Scholars:
5.1W
Papers: 4.6W
Citations: 71
researcher View more organizations