Healthcare-Centric Confidence Assessment of Machine Learning Models via Average Precision and Feature-Subset Stability

Authors

  • Kiran Veernapu Applied AI & Data Science, Brown University, Providence, RI, USA; College of Business & Management, Colorado Technical University, Colorado Springs, CO, USA https://orcid.org/0009-0005-6567-4684
  • Surya Teja Meesala Department of Computer Engineering, University of Houston-Clear Lake, Houston, TX, USA https://orcid.org/0009-0001-6247-6878
  • Nithesh Gudipuri Technology Modernization, Raymond James & Associates, St. Petersburg, FL, USA
  • Saffat Kamal Niloy Department of MBA in International Business, Pacific States University, Los Angeles, CA, USA
  • Debabrata Biswas Department of Information Systems, Pacific States University, Los Angeles, CA, USA
  • Sumaia Benta Arif Department of Information Systems, Pacific States University, Los Angeles, CA, USA
  • R. Vidya Department of CSE, Malla Reddy (MR) Deemed to be University, Hyderabad, Telangana, India https://orcid.org/0000-0002-5142-5557

DOI:

https://doi.org/10.57159/jcmm.5.3.26714

Keywords:

Threshold-Agnostic Ranking, Model Calibration, Permutation Feature Importance, Convolutional Neural Network, Breast Cancer Diagnosis, SDG 3: Good Health and Well-Being

Abstract

The paper presents a confidence-based assessment system of machine learning models which uses the Average Precision (AP) which is the result of the precision-recall curve as a threshold-agnostic ranking metric and indicator of operational reliability. This method is not based on the accuracy of the class-labels of the model like the traditional metrics, but the consistency of the model performance at the entire range of decision thresholds. To act as an experimental test bed, a one-dimensional convolutional neural network (1D CNN) was trained using the Breast Cancer Wisconsin (Diagnostic) dataset. Permutation based feature importance was done after full-feature baseline AP to measure the AP degradation on a feature-by-feature basis and an even higher-order subset, such as pairs, triplets and quadruplets. Empirical analysis revealed worst concave points, worst concavity and mean texture as the main factors of individual AP drops. Within higher-order subsets, worst concavity, worst radius and mean texture always prevailed in the AP-based reliability profile. It is interesting to note that several four-feature combinations produced an AP score of the same size as or greater than the baseline of 30 features, which indicates a high level of informational redundancy. AP functions as a stable threshold-agnostic ranking and operational-performance measure, not as direct probabilistic confidence, as evidenced by its consistently high AP (>0.95) across operating thresholds. The paper finds that high AP-based operational reliability does not necessitate a complete feature set, but instead the consistency of AP-based reliability in small feature sets increases the interpretability, trust and deployability of machine learning systems in the studied medical diagnostic setting and in related safety-critical applications pending external validation.

References

[1] T. Dawood et al., "Uncertainty aware training to improve deep learning model calibration for classification of cardiac MR images," Medical Image Analysis, vol. 88, p. 102861, 2023.

[2] M. Chidambaram, H. Lee, C. McSwiggen, and S. Rezchikov, "How Flawed Is ECE? An Analysis via Logit Smoothing," in Proceedings of the 41st International Conference on Machine Learning (ICML), PMLR, vol. 235, pp. 8417–8435, 2024.

[3] R. Caruana and A. Niculescu-Mizil, "Data mining in metric space: An empirical analysis of supervised learning performance criteria," in Proceedings of the 10th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 69–78, 2004.

[4] D. Desai et al., "Model Quality in AI-based Bruise Detection: Rethinking IoU and Confidence Thresholds," AMIA Annual Symposium Proceedings, vol. 2024, pp. 293–302, 2024.

[5] U. Johansson, T. Löfström, and H. Boström, "Calibrating probability estimation trees using Venn-Abers predictors," in Proceedings of the SIAM International Conference on Data Mining (SDM), pp. 28–36, 2019.

[6] J. Hernández-Orallo, P. Flach, and C. Ferri, "A unified view of performance metrics: Translating threshold choice into expected classification loss," Journal of Machine Learning Research, vol. 13, pp. 2813–2869, 2012.

[7] N. Mehdiyev, M. Majlatow, and P. Fettke, "Integrating permutation feature importance with conformal prediction for robust Explainable Artificial Intelligence in predictive process monitoring," Engineering Applications of Artificial Intelligence, vol. 149, p. 110363, 2025.

[8] Y. Gao, D. Zhao, B. Liang, X. Yang, and X. Xue, "Remote Sensing Monitoring of Soil Salinization Based on Bootstrap-Boruta Feature Stability Assessment: A Case Study in Minqin Lake Region," Remote Sensing, vol. 18, no. 2, p. 245, 2026.

[9] A. Spooner, G. Mohammadi, P. S. Sachdev, H. Brodaty, and A. Sowmya, "Ensemble feature selection with data-driven thresholding for Alzheimer's disease biomarker discovery," BMC Bioinformatics, vol. 24, p. 9, 2023.

[10] Y. Liu et al., "Comparative analysis of convolutional neural networks and traditional machine learning models for IVF live birth prediction: A retrospective analysis of 48514 IVF cycles and an evaluation of deployment feasibility in resource-constrained settings," Frontiers in Endocrinology, vol. 16, p. 1556681, 2025.

[11] S. Wang, D. Upadhyay, M. Zaman, K. Naik, and S. N. Rai, "Diabetes risk modeling through tabular-to-image transformations and ensemble learning," Informatics in Medicine Unlocked, vol. 61, 2026.

[12] A. Chamma, B. Thirion, and D. Engemann, "Variable Importance in High-Dimensional Settings Requires Grouping," in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 11195–11203, 2024.

[13] J. P. B. Pereira, E. S. G. Stroes, A. H. Zwinderman, and E. Levin, "Covered Information Disentanglement: Model Transparency via Unbiased Permutation Importance," in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 7984–7992, 2022.

[14] L. Brocki and N. C. Chung, "Feature perturbation augmentation for reliable evaluation of importance estimators in neural networks," Pattern Recognition Letters, vol. 176, pp. 131–139, 2023.

[15] A. Bhadra, J. Datta, N. G. Polson, V. Sokolov, and J. Xu, "Merging two cultures: Deep and statistical learning," Wiley Interdisciplinary Reviews: Computational Statistics, vol. 16, no. 2, p. e1647, 2024.

[16] W. N. Street, W. H. Wolberg, and O. L. Mangasarian, "Nuclear feature extraction for breast tumor diagnosis," in Biomedical Image Processing and Biomedical Visualization, IS&T/SPIE International Symposium on Electronic Imaging: Science and Technology, vol. 1905, pp. 861–870, 1993.

Downloads

Published

2026-06-30

How to Cite

Veernapu, K., Meesala, S. T., Gudipuri, N., Niloy, S. K., Biswas, D., Arif, S. B., & Vidya, R. (2026). Healthcare-Centric Confidence Assessment of Machine Learning Models via Average Precision and Feature-Subset Stability. Journal of Computers, Mechanical and Management, 5(3), 129–140. https://doi.org/10.57159/jcmm.5.3.26714