arrow
Return

stringi: Fast and Portable Character String Processing in R

delete2022-01-01
delete43
delete
OA
AI
M
Marek Gągolewski *
DOI:10.18637/jss.v103.i02delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Effective processing of character strings is required at various stages of data analysis pipelines: from data cleansing and preparation, through information extraction, to report generation. Pattern searching, string collation and sorting, normalization, transliteration, and formatting are ubiquitous in text mining, natural language processing, and bioinfor-matics. This paper discusses and demonstrates how and why stringi, a mature R package for fast and portable handling of string data based on ICU (International Components for Unicode), should be included in each statistician's or data scientist's repertoire to complement their numerical computing and data wrangling skills.
Keywords:
stringi
character strings
text
ICU
Unicode
regular expressions
data cleansing
natural language processing
R

Journal

Journal of Statistical Software cover
Journal of Statistical Software
IF:
8.1
Papers:
622
Citations:
4.6W

Organization

D
Deakin University
Scholars:
2.0W
Papers: 2.1W
Citations: 2.8W