arrow
返回

Enhancing Ovarian Tumor Dataset Analysis Through Data Mining Preprocessing Techniques

delete2024-01-01
delete0
delete
OA
AI
R
Roopashri Shetty
M
M. Geetha *
U
U. Dinesh Acharya
G
G. Shyamala
DOI:10.1109/ACCESS.2024.3450520delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The early detection and treatment of ovarian cancer face considerable hurdles due to its complexity and lethal nature. Because of its high death rates and heterogeneity, ovarian cancer poses a significant challenge to oncology. In-depth study of ovarian tumor datasets is crucial to improve the knowledge on this complicated illness and to develop new diagnostic and treatment approaches. The accuracy of the information utilized for training and analysis has a substantial impact on how well computer models predict and comprehend ovarian cancer. Data mining methods mostly rely on the quality of data. Hence, in order to improve the accuracy and dependability of ensuing studies, this work is carried out to examine the critical preprocessing methods that are used on ovarian tumor dataset. A novel ovarian tumor dataset is collected and this raw dataset has missing values, incomplete data, noisy data, redundant data and outliers and these anomalies degrade the performance of mining results. In this study, we explore the application of data mining preprocessing methods to enhance the analysis of ovarian tumor datasets. Through the use of methods like feature selection, data cleaning, normalization, and dimensionality reduction, we aim to improve the quality of the data, and make it easier to find significant patterns and biomarkers linked to ovarian cancer. The work emphasizes the importance of preprocessing in maximizing the potential of ovarian tumor datasets and expanding the field's understanding of this debilitating illness in order to improve detection and treatment process. Preprocessing performance indicators namely accuracy, sensitivity, and specificity are used to assess the efficiency. It is found that, after preprocessing of the dataset, an accuracy of 88% is achieved when classified as benign or malignant using Logistic Regression. Upon applying every feature selection technique on the dataset, it is evident that features obtained through Recursive Feature Elimination technique and feature importance yield greater accuracy of 92% when classified with respect to Logistic Regression and Support Vector Machine. It is expected that the knowledge gathered from these preprocessing techniques result in more precise and trustworthy computer models, which could enhance patient outcomes in the field of ovarian cancer.
Keyword:
Data mining
Tumors
Ovarian cancer
Imputation
Feature extraction
Cleaning
Accuracy
Classification algorithms
Supervised learning
Medical diagnosis
classification
data mining
preprocessing
supervised learning technique

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

暂无机构信息
引用论文

引用论文

Progress in Outlier Detection Techniques: A Survey
err2019-01-01
err320
errOAAI
errWang, Hongzhi; Bah, Mohamed Jaward; Hammad, Mohamed
err分享
err收藏
Thin Film Nanofibrous Composite Membrane for Dead-End Seawater Desalination
err2016-01-01
err0
errOAAI
errBaturalp Yalcinkaya; Fatma Yalcinkaya; Jiri Chaloupek
err分享
err收藏
err分享
err收藏
New data preprocessing trends based on ensemble of multiple preprocessing techniques基于多种预处理技术集成的数据预处理新趋势
err2020-11-01
err247
errOAAI
errMishra, Puneet; Biancolillo, Alessandra; Roger, Jean Michel; Marini, Federico; Rutledge, Douglas N.
err分享
err收藏
err分享
err收藏
Medical data quality assessment: On the development of an automated framework for medical data curation医疗数据质量评估: 关于医疗数据管理自动化框架的开发
err2019-04-01
err73
PREAI
errPezoulas, Vasileios C.; Kourou, Konstantina D.; Kalatzis, Fanis; Exarchos, Themis P.; Venetsanopoulou, Aliki; Zampeli, Evi; Gandolfo, Saviana; Skopouli, Fotini; De Vita, Salvatore; Tzioufas, Athanasios G.; Fotiadis, Dimitrios I.
err分享
err收藏
学者 查看更多内容