arrow
Return

A study on data imputation and prediction modelling using maximum margin matrix factorization

delete2026-04-27
delete0
PRE
AI
A
Akshar Chintalapally
S
Sowmini Devi Veeramachaneni *
Y
Yaswanth Gavini
A
Arun K. Pujari
DOI:10.1007/s10115-026-02764-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Missing data is a pervasive challenge in real-world datasets, often distorting feature relationships and degrading predictive reliability. Traditional imputation techniques such as mean substitution, median filling or k-nearest neighbours frequently fail to capture complex inter-feature dependencies, particularly in structured or high-dimensional data. In this work, we investigate Maximum Margin Matrix Factorization (MMMF), originally developed for collaborative filtering, as a general-purpose framework for missing data imputation in supervised learning. To enable its use on continuous feature matrices, we introduce a discretization-based transformation pipeline that allows MMMF to operate beyond ordinal recommender systems on general tabular datasets. We evaluate the method using a two-task framework: (i) measuring imputation accuracy across varying sparsity levels and (ii) assessing the downstream impact of imputed data on regression models. Experiments on diverse datasets show that MMMF achieves strong reconstruction accuracy while preserving predictive structure, such that regression models trained on complete data produce comparable predictions when evaluated on corresponding imputed test data. These findings suggest that MMMF can serve as a robust imputation strategy for practical machine learning pipelines.
Keywords:
Maximum margin matrix factorization
Missing data
Data imputation
Prediction

Journal

Knowledge and Information Systems cover
Knowledge and Information Systems
IF:
3.1
Papers:
517
Citations:
5.2K

Organization

C
cse
Scholars:
18
Papers: 12
Citations: 0
A
AI
Scholars:
4
Papers: 2
Citations: 0