A reproducible, data-centric benchmarking framework evaluating the robustness of tabular machine learning models under systematic feature shift using OpenML-CC18 datasets and automated feature engineering.
benchmarkingautomated-feature-engineeringwilcoxon-testdata-centric-aidistribution-shiftmodel-reliabilitytabular-machine-learningfeature-shiftopenml-cc18wilcoxon-signed-rank-test
-
Updated
Aug 3, 2026 - Python