Healthcare-Centric Confidence Assessment of Machine Learning Models via Average Precision and Feature-Subset Stability
DOI:
https://doi.org/10.57159/jcmm.5.3.26714Keywords:
Threshold-Agnostic Ranking, Model Calibration, Permutation Feature Importance, Convolutional Neural Network, Breast Cancer Diagnosis, SDG 3: Good Health and Well-BeingAbstract
The paper presents a confidence-based assessment system of machine learning models which uses the Average Precision (AP) which is the result of the precision-recall curve as a threshold-agnostic ranking metric and indicator of operational reliability. This method is not based on the accuracy of the class-labels of the model like the traditional metrics, but the consistency of the model performance at the entire range of decision thresholds. To act as an experimental test bed, a one-dimensional convolutional neural network (1D CNN) was trained using the Breast Cancer Wisconsin (Diagnostic) dataset. Permutation based feature importance was done after full-feature baseline AP to measure the AP degradation on a feature-by-feature basis and an even higher-order subset, such as pairs, triplets and quadruplets. Empirical analysis revealed worst concave points, worst concavity and mean texture as the main factors of individual AP drops. Within higher-order subsets, worst concavity, worst radius and mean texture always prevailed in the AP-based reliability profile. It is interesting to note that several four-feature combinations produced an AP score of the same size as or greater than the baseline of 30 features, which indicates a high level of informational redundancy. AP functions as a stable threshold-agnostic ranking and operational-performance measure, not as direct probabilistic confidence, as evidenced by its consistently high AP (>0.95) across operating thresholds. The paper finds that high AP-based operational reliability does not necessitate a complete feature set, but instead the consistency of AP-based reliability in small feature sets increases the interpretability, trust and deployability of machine learning systems in the studied medical diagnostic setting and in related safety-critical applications pending external validation.
References
[1] T. Dawood et al., "Uncertainty aware training to improve deep learning model calibration for classification of cardiac MR images," Medical Image Analysis, vol. 88, p. 102861, 2023.
[2] M. Chidambaram, H. Lee, C. McSwiggen, and S. Rezchikov, "How Flawed Is ECE? An Analysis via Logit Smoothing," in Proceedings of the 41st International Conference on Machine Learning (ICML), PMLR, vol. 235, pp. 8417–8435, 2024.
[3] R. Caruana and A. Niculescu-Mizil, "Data mining in metric space: An empirical analysis of supervised learning performance criteria," in Proceedings of the 10th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 69–78, 2004.
[4] D. Desai et al., "Model Quality in AI-based Bruise Detection: Rethinking IoU and Confidence Thresholds," AMIA Annual Symposium Proceedings, vol. 2024, pp. 293–302, 2024.
[5] U. Johansson, T. Löfström, and H. Boström, "Calibrating probability estimation trees using Venn-Abers predictors," in Proceedings of the SIAM International Conference on Data Mining (SDM), pp. 28–36, 2019.
[6] J. Hernández-Orallo, P. Flach, and C. Ferri, "A unified view of performance metrics: Translating threshold choice into expected classification loss," Journal of Machine Learning Research, vol. 13, pp. 2813–2869, 2012.
[7] N. Mehdiyev, M. Majlatow, and P. Fettke, "Integrating permutation feature importance with conformal prediction for robust Explainable Artificial Intelligence in predictive process monitoring," Engineering Applications of Artificial Intelligence, vol. 149, p. 110363, 2025.
[8] Y. Gao, D. Zhao, B. Liang, X. Yang, and X. Xue, "Remote Sensing Monitoring of Soil Salinization Based on Bootstrap-Boruta Feature Stability Assessment: A Case Study in Minqin Lake Region," Remote Sensing, vol. 18, no. 2, p. 245, 2026.
[9] A. Spooner, G. Mohammadi, P. S. Sachdev, H. Brodaty, and A. Sowmya, "Ensemble feature selection with data-driven thresholding for Alzheimer's disease biomarker discovery," BMC Bioinformatics, vol. 24, p. 9, 2023.
[10] Y. Liu et al., "Comparative analysis of convolutional neural networks and traditional machine learning models for IVF live birth prediction: A retrospective analysis of 48514 IVF cycles and an evaluation of deployment feasibility in resource-constrained settings," Frontiers in Endocrinology, vol. 16, p. 1556681, 2025.
[11] S. Wang, D. Upadhyay, M. Zaman, K. Naik, and S. N. Rai, "Diabetes risk modeling through tabular-to-image transformations and ensemble learning," Informatics in Medicine Unlocked, vol. 61, 2026.
[12] A. Chamma, B. Thirion, and D. Engemann, "Variable Importance in High-Dimensional Settings Requires Grouping," in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 11195–11203, 2024.
[13] J. P. B. Pereira, E. S. G. Stroes, A. H. Zwinderman, and E. Levin, "Covered Information Disentanglement: Model Transparency via Unbiased Permutation Importance," in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 7984–7992, 2022.
[14] L. Brocki and N. C. Chung, "Feature perturbation augmentation for reliable evaluation of importance estimators in neural networks," Pattern Recognition Letters, vol. 176, pp. 131–139, 2023.
[15] A. Bhadra, J. Datta, N. G. Polson, V. Sokolov, and J. Xu, "Merging two cultures: Deep and statistical learning," Wiley Interdisciplinary Reviews: Computational Statistics, vol. 16, no. 2, p. e1647, 2024.
[16] W. N. Street, W. H. Wolberg, and O. L. Mangasarian, "Nuclear feature extraction for breast tumor diagnosis," in Biomedical Image Processing and Biomedical Visualization, IS&T/SPIE International Symposium on Electronic Imaging: Science and Technology, vol. 1905, pp. 861–870, 1993.
Downloads
Published
How to Cite
Issue
Section
Categories
License
Copyright (c) 2026 The Author(s)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Articles published in the Journal of Computers, Mechanical and Management are licensed under a CC BY-NC 4.0 license. Authors retain copyright of their work and grant the journal a non-exclusive license to publish, distribute, and archive the article. Full terms are on the Copyright and Licensing page.