Learning to Validate the Predictions of Black Box Machine Learning Models on Unseen Data
4th Workshop on Human-In-the-Loop Data Analytics (HILDA) at SIGMOD, 2019 Workshop
Abstract
When end users apply a machine learning (ML) model on new unlabeled data, it is difficult for them to decide whether they can trust its predictions. Errors or shifts in the target data can lead to hard-to-detect drops in the predictive quality of the model. We therefore propose an approach to assist non-ML experts working with pretrained ML models. Our approach estimates the change in prediction performance of a model on unseen target data. It does not require explicit distributional assumptions on the dataset shift between the training and target data. Instead, a domain expert can declaratively specify typical cases of dataset shift that she expects to observe in real-world data. Based on this information, we learn a performance predictor for pretrained black box models, which can be combined with the model, and automatically warns end users in case of unexpected performance drops. We demonstrate the effectiveness of our approach on two models — logistic regression and a neural network, applied to several real-world datasets.
Citations
Showing 20 citing works with retrievable metadata.
- Injury Prediction in Sports using Artificial Intelligence Applications: A Brief ReviewG. Kumar, M. D. Kumar, S. Venkata, et al. · 2023 · Journal of Robotics and Control (JRC)
- Learning Prediction Intervals for Model PerformanceElder, Benjamin, Arnold, Matthew, Murthi, Anupama, et al. · 2021 · Proceedings of the AAAI Conference on Artificial Intelligence
- Use of Bi-Temporal ALS Point Clouds for Tree Removal Detection on Private Property in Racibórz, PolandPatrycja Przewoźna, Paweł Hawryło, Karolina Zięba-Kulawik, et al. · 2021 · Remote Sensing
- QuAcc: Using Quantification to Predict Classifier Accuracy Under Prior Probability ShiftLorenzo Volpi, Alejandro Moreo, Fabrizio Sebastiani · 2025 · Intelligenza Artificiale
- Learning to Validate the Predictions of Black Box Classifiers on Unseen DataSebastian Schelter, Tammo Rukat, Felix Biessmann · 2020 · SIGMOD Conference
- Performance Prediction Under Dataset ShiftSimona Maggio, Victor Bouvier, Leo Dreyfus-Schmidt · 2022 · 2022 26th International Conference on Pattern Recognition (ICPR)
- Automatic Data Extraction Utilizing Structural Similarity From A Set of Portable Document Format (PDF) FilesHadipurnawan Satria, Anggina Primanita · 2023 · Sriwijaya Journal of Informatics and Applications
- A hybrid Bi-GRU–bi-LSTM framework for multivariate climate finance forecasting: Application to the most vulnerable countriesAmna Farooqui Arsalan, Falak Khan · 2026 · Array
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, et al. · 2026
- TOLEBI: Learning Fault-Tolerant Bipedal Locomotion via Online Status Estimation and Fallibility RewardsHokyun Lee, Woo-Jeong Baek, Junhyeok Cha, et al. · 2026 · arXiv.org
- Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment SettingsAngéline Pouget, Mohammad Yaghini, Stephan Rabanser, et al. · 2025 · International Conference on Machine Learning
- Fin-Fed-OD: Federated Outlier Detection on Financial Tabular DataDayananda Herurkar, Sebastián M. Palacio, Ahmed Anwar, et al. · 2024 · arXiv.org
- Explaining Anomalies using Denoising Autoencoders for Financial Tabular DataTimur Sattarov, Dayananda Herurkar, Jörn Hees · 2022 · arXiv.org
- Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Y. Gal, et al. · 2022 · Neural Information Processing Systems
- Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte CarloI. Peis, Chao Ma, José Miguel Hernández-Lobato · 2022 · Neural Information Processing Systems
- Robust Variational Autoencoders for Outlier Detection and Repair of Mixed-Type DataSimao Eduardo, A. Nazábal, Christopher K. I. Williams, et al. · 2019 · International Conference on Artificial Intelligence and Statistics
- Robust Variational Autoencoders for Outlier Detection in Mixed-Type DataSimao Eduardo, A. Nazábal, Christopher K. I. Williams, et al. · 2019 · arXiv.org
- Automating Data Quality Validation for Dynamic Data IngestionS. Redyuk, Zoi Kaoudi, V. Markl, et al. · 2021 · International Conference on Extending Database Technology
- From Cleaning before ML to Cleaning for MLFelix Neutatz, Binger Chen, Ziawasch Abedjan, et al. · 2021 · IEEE Data Engineering Bulletin
- Ethical Considerations of Artificial Intelligence via Neural Ethical Considerations of Artificial Intelligence via Neural Networks Applied to Medical Applications Networks Applied to Medical ApplicationsJames Ternent, M. Thompson · 2020