arrow
Return

Semantic hashing

delete2009-07-01
delete939
delete
OA
AI
R
Ruslan Salakhutdinov *
G
Geoffrey E. Hinton
DOI:10.1016/j.ijar.2008.11.006delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We show how to learn a deep graphical model of the word-count vectors obtained from a large set of documents. The values of the latent variables in the deepest layer are easy to infer and give a much better representation of each document than Latent Semantic Analysis. When the deepest layer is forced to use a small number of binary variables (e.g. 32), the graphical model performs semantic hashing: Documents are mapped to memory addresses in such a way that semantically similar documents are located at nearby addresses. Documents similar to a query document can then be found by simply accessing all the addresses that differ by only a few bits from the address of the query document. This way of extending the efficiency of hash-coding to approximate matching is Much faster than locality sensitive hashing, which is the fastest current method. By using semantic hashing to filter the documents given to TF-IDF, we achieve higher accuracy than applying TF-IDF to the entire document set. (C) 2008 Elsevier Inc. All rights reserved.
Keywords:
Information retrieval
Graphical models
Unsupervised learning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

International Journal of Approximate Reasoning cover
International Journal of Approximate Reasoning
IF:
3
Papers:
2.9K
Citations:
5.1K

Organization

U
university of toronto
Scholars:
14.7W
Papers: 12.0W
Citations: 165