arrow
返回

A Deep-Learned Embedding Technique for Categorical Features Encoding

delete2021-01-01
delete102
delete
OA
AI
M
Mwamba Kasongo Dahouda
I
Inwhee Joe *
DOI:10.1109/ACCESS.2021.3104357delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Many machine learning algorithms and almost all deep learning architectures are incapable of processing plain texts in their raw form. This means that their input to the algorithms must be numerical in order to solve classification or regression problems. Hence, it is necessary to encode these categorical variables into numerical values using encoding techniques. Categorical features are common and often of high cardinality. One-hot encoding in such circumstances leads to very high dimensional vector representations, raising memory and computability concerns for machine learning models. This paper proposes a deep-learned embedding technique for categorical features encoding on categorical datasets. Our technique is a distributed representation for categorical features where each category is mapped to a distinct vector, and the properties of the vector are learned while training a neural network. First, we create a data vocabulary that includes only categorical data, and then we use word tokenization to make each categorical data a single word. After that, feature learning is introduced to map all of the categorical data from the vocabulary to word vectors. Three different datasets provided by the University of California Irvine (UCI) are used for training. The experimental results show that the proposed deep-learned embedding technique for categorical data provides a higher F1 score of 89% than 71% of one-hot encoding, in the case of the Long short-term memory (LSTM) model. Moreover, the deep-learned embedding technique uses less memory and generates fewer features than one-hot encoding.
Keyword:
Encoding
Numerical models
Machine learning
Data models
Training
Biological neural networks
Computational modeling
Data preprocessing
categorical variables
natural language processing
machine learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

H
hanyang university
学者数:
2.9W
论文数: 2.7W
被引数: 36
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
Arabic Word Segmentation With Long Short-Term Memory Neural Networks and Word Embedding
err2019-01-01
err15
errOAAI
errAlmuhareb, Abdulrahman; Alsanie, Waleed; Al-Thubaity, Abdulmohsen
err分享
err收藏
Acetylcholine stimulates the release of prostacyclin by rabbit aorta endothelium
err1983-04-01
err0
PREAI
errJohan R Beetens; Cor van Hove; Marc Rampart; Arnold G Herman
err分享
err收藏
学者 查看更多内容