arrow
Return

Short Text Classification Using Contextual Analysis

delete2021-01-01
delete8
delete
OA
AI
S
Sami Al Sulaimani *
A
Andrew Starkey
DOI:10.1109/ACCESS.2021.3125768delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Micro blogging tools provide a real time service for the public to express opinions, to broadcast news and information and offer an opportunity to comment and respond to such output. Word usage in social media is continually evolving. Micro bloggers may use different sets of words to describe a specific event and they may use new words (i.e. neither exist in the training dataset nor in informal or formal dictionaries) or use words in new contexts. Dynamically capturing new words and their potential meaning from their context can help to reflect the words relationship in social media, which then can be useful for solving various problems, like the event classification task. Different approaches have been proposed in this regard, one of them is Contextual Analysis. This paper focuses on examining the potential of this approach for grouping short texts (tweets) talking about the same event into the same category. A new transparent method for text multi-class categorization is presented. It uses the Contextual Analysis approach to capture the most important words in the context of an event and to detect the usage of similar words in different contexts. In order to test the efficacy in these areas, this study evaluates the performance of the proposed method and other well known methods, such as Naive Bayes, Support Vector Machines, K-Nearest Neighbors and Convolutional Neural Networks. On average, the experiments' results show that the proposed multi-class classification method can effectively categorize tweets into various event groups, with a high f1-measure score f1>97.09% and f1>95.27%, in the imbalanced classes and high number of classes experiments, respectively. However, similar to the baseline methods, the performance is negatively influenced by the imbalanced dataset. The Convolutional Neural Networks method produces the best performance among the other algorithms with f1>97.74% in all experiments, which is 1.73% and 2.72% higher than the lowest performance of Naive Bayes and K-Nearest Neighbors, respectively, but does not meet the requirements of transparency of results.
Keywords:
Blogs
Support vector machines
Social networking (online)
Convolutional neural networks
Machine learning algorithms
Task analysis
Context modeling
Text analysis
event classification
contextual analysis
supervised machine learning

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

U
University of Aberdeen
Scholars:
1.3W
Papers: 1.3W
Citations: 2.0W
Cited Papers

Cited Papers

Diseases and Molecular Diagnostics: A Step Closer to Precision Medicine
err2017-08-22
err0
errOAAI
errShailendra Dwivedi; Purvi Purohit; Radhieka Misra; Puneet Pareek; Apul Goel; Sanjay Khattri; Kamlesh Kumar Pant; Sanjeev Misra; Praveen Sharma
errShare
errSave
Directed acyclic graph kernels for structural RNA analysis
err2008-07-22
err0
errOAAI
errKengo Sato; Toutai Mituyama; Kiyoshi Asai; Yasubumi Sakakibara
errShare
errSave
A naive Bayes strategy for classifying customer satisfaction: A study based on online reviews of hospitality services
err2019-08-01
err60
errOAAI
errSanchez-Franco, Manuel J.; Navarro-Garcia, Antonio; Javier Rondan-Cataluna, Francisco
errShare
errSave
errShare
errSave
errShare
errSave
Two-dimensional bricklayer arrangements of tolans using halogen bonding interactions
err2015-01-01
err0
PREAI
errFanny Frausto; Zachary C. Smith; Terry E. Haas; Samuel W. Thomas III
errShare
errSave
researcher View more