Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
S28: Philosophy of Cognition & AI 3 Location: 23.03 01.22 Session Chair: Gottfried Vosgerau | |
| Presentation 3 | |
3:00pm - 3:45pm
The benchmarking epistemology: What inferences can scientists draw from competitive comparisons of prediction models? University of Tübingen, Germany Benchmarking, the evaluation of machine learning (ML) models based on predictive performance and competitive ranking, is a cornerstone of ML research and an increasingly prominent tool in scientific arguments. This paper argues that benchmarking constitutes a scientific epistemology, offering a powerful framework for scientific inference. We identify four core types of inferences drawn from benchmarks: those about the best (1) model, (2) learning algorithm, (3) deployment decision, and (4) prediction. We demonstrate that the validity of each of these inference relies on additional assumptions, analogous to ensuring construct validity in psychological tests. Through case studies in image recognition, life outcomes prediction, and weather forecasting, we examine these assumptions and their implications for inference validity. Finally, we discuss the social roles of benchmarks in organizing scientific communities and their potential threats to validity, offering strategies to mitigate these challenges and improve benchmark design and interpretation. | |

