arrow
Return

ydata-profiling: Accelerating data-centric AI with high-quality data

delete2023-10-01
delete6
PRE
AI
F
Fabiana Martins Clemente
G
Gonçalo Martins Ribeiro
A
Alexandre Quemy
M
Miriam Seoane Santos *
R
Ricardo Cardoso Pereira
A
Alex Barros
DOI:10.1016/j.neucom.2023.126585delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
ydata-profiling is an open-source Python package for advanced exploratory data analysis that enables users to generate data profiling reports in a simple, fast, and efficient manner, fostering a standardized and visual understanding of the data. Beyond traditional descriptive properties and statistics, ydata-profiling follows a Data-Centric AI approach to exploratory analysis, as it focuses on the automatic detection and highlighting of complex data characteristics often associated with potential data quality issues, such as high ratios of missing or imbalanced data, infinite, unique, or constant values, skewness, high correlation, high cardinality, non-stationarity, seasonality, duplicate records, and other inconsistencies. The source code, documentation, and examples are available in the GitHub repository: https://github.com/ydataai/ydataprofiling.
Keywords:
Exploratory data analysis
Data profiling
Data quality
Data-centric AI
Data Intrinsic Characteristics
Data Complexity

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

No organization information available