Return
ydata-profiling: Accelerating data-centric AI with high-quality data
DOI:10.1016/j.neucom.2023.126585.png)
Abstract
En 中文
ydata-profiling is an open-source Python package for advanced exploratory data analysis that enables users to generate data profiling reports in a simple, fast, and efficient manner, fostering a standardized and visual understanding of the data. Beyond traditional descriptive properties and statistics, ydata-profiling follows a Data-Centric AI approach to exploratory analysis, as it focuses on the automatic detection and highlighting of complex data characteristics often associated with potential data quality issues, such as high ratios of missing or imbalanced data, infinite, unique, or constant values, skewness, high correlation, high cardinality, non-stationarity, seasonality, duplicate records, and other inconsistencies. The source code, documentation, and examples are available in the GitHub repository: https://github.com/ydataai/ydataprofiling.
Keywords:
Exploratory data analysis
Data profiling
Data quality
Data-centric AI
Data Intrinsic Characteristics
Data Complexity
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
Organization
No organization information available

