arrow
返回

Text mining with n-gram variables

delete2018-11-19
delete28
delete
OA
AI
M
Matthias Schonlau *
N
Nick Guenther
I
Ilia Sucholutsky
DOI:10.1177/1536867X1701700406delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Text mining is the process of turning free text into numerical variables and then analyzing them with statistical techniques. We introduce the command ngram, which implements the most common approach to text mining, the bag of words. An n-gram is a contiguous sequence of words in a text. Broadly speaking, ngram creates hundreds or thousands of variables, each recording how often the corresponding n-gram occurs in a given text. This is more useful than it sounds. We illustrate ngram with the categorization of text answers from two open-ended questions.
Keyword:
st0502
ngram
bag of words
sets of words
unigram
gram
statistical learning
machine learning

期刊

S
Stata Journal
IF:
2.4
论文数:
1.2K
被引数:
8.4K

机构

U
University of Waterloo
学者数:
2.2W
论文数: 2.3W
被引数: 3.3W
引用论文

引用论文

err分享
err收藏
Support vector machines支持向量机
err2016-12-01
err141
errOAAI
errGuenther, Nick; Schonlau, Matthias
err分享
err收藏
Streptococcus agalactiae invasion of human brain microvascular endothelial cells is promoted by the laminin-binding protein Lmb
err2007-05-01
err0
PREAI
errTobias Tenenbaum; Barbara Spellerberg; Rüdiger Adam; Markus Vogel; Kwang Sik Kim; Horst Schroten
err分享
err收藏
Adaptabilidade e estabilidade da produção de borracha e seleção em progênies de seringueira
err2009-10-01
err0
errOAAI
errCecília Khusala Verardi; Marcos Deon Vilela de Resende; Reginaldo Brito da Costa; Paulo de Souza Gonçalves
err分享
err收藏
学者 查看更多内容