arrow
Return

Basis-driven learnable operator for MLP-mixers

delete2026-08-27
delete0
delete
OA
AI
A
AE Ahmed Elsheikh
M
Mohammed E. Fouda *
A
Ahmed M. Eltawil
DOI:10.3389/frai.2026.1769436delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
IntroductionExisting multi-layer perceptron (MLP)-mixer architectures either rely heavily on extensive training data or employ rigid, hand-engineered mixing operations. This study introduces a Basis-Driven Learnable Operator (BDLO) for MLP-mixers, targeting a balance between high-capacity flexible learning and computationally efficient, yet constrained, handcrafted solutions.MethodsBDLO is a drop-in replacement for the shifting block of shifting-based MLP-mixers. It approximates shifting-based mixing operations by learning the coefficients of a real, complete, discrete basis, yielding an input-dependent transformation matrix that is applied to both rows and columns of the token table, while the patch embedding, channel-mixing MLPs, skip connections and classification head are left unchanged. The operator was integrated into CycleMLP, HireMLP and AS-MLP at three model sizes each, and all models were trained from scratch under an identical configuration on CIFAR10, CIFAR100, a reduced ImageNet (32 × 32, 500 classes) and full ImageNet1K. Standard and discrete cosine transform bases were compared, hyperparameter sensitivity was assessed with Optuna, and differences were tested using the Wilcoxon signed-rank test.ResultsBDLO reduced the cost of the mixing layer to 1.68 GFLOPs, against 3.33–5.08 GFLOPs for the original layers, and reduced whole-model parameter counts by 21.8%–56.7% (mean reduction 12.84M), with a mean accuracy difference of 0.38% in favor of BDLO. The Wilcoxon test confirmed a significantly lower parameter distribution (p = 0.0039) and no statistically significant accuracy difference (p = 0.496 on CIFAR100, p = 0.0625 on CIFAR10, p = 0.5 on the reduced ImageNet). Results were invariant to the choice of basis (mean cosine distance 0.0049 between the learned coefficient vectors) and insensitive to the BDLO-specific hyperparameters.DiscussionComparability holds in aggregate and in the parameter-constrained regime, with model-specific exceptions for baselines that include channel mixing (HireMLP) and for high-capacity baselines on larger datasets (AS-MLP). BDLO behaves as an input-adaptive spectral modulator whose bounded coefficients provide implicit regularization. This confirms the effectiveness of BDLO as an efficient operator replacement for shifting-based MLP-mixers, most suitable for parameter-constrained, small-to-medium-scale models in image recognition tasks.
Keywords:
deep learning,neural networks,learnable operators,learnable spectral transformation bases,MLP-mixers

Journal

F
Frontiers in Artificial Intelligence
IF:
4.7
Papers:
2.4K
Citations:
4.4K

Organization

C
compumacy for artificial intelligence solutions
Scholars:
3
Papers: 3
Citations: 0
M
mathematics and engineering physics department
Scholars:
2
Papers: 1
Citations: 0
C
cemse division
Scholars:
2
Papers: 1
Citations: 0
researcher View more organizations