Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Moderators: Prof. Vaiva Hendrixson and Doc. Lina Zabulienė
Presentation 8
Generative AI and Supervised Written Examinations: Do Performance Trends Support Faculty Concerns?
Vaiva Hendrixson1, Viktor Riklefs2, Danila Silischev2
1: Vilnius University, Faculty of Medicine, Lithuania; 2: Karaganda Medical University, Kazakhstan
Background and Aim Generative artificial intelligence (AI) has raised concerns about score inflation and the validity of written assessments. This study examined whether performance in supervised medical examinations changed across three academic years of widespread AI availability.
Materials and Methods A retrospective repeated cross-sectional study analyzed 4,874 examination responses from Karaganda Medical University in 2023/2024–2025/2026: 2,850 in second-year basic sciences and 2,024 in fourth-year internal medicine. Examination format, supervision, and scoring remained unchanged; devices, reference materials, and AI were prohibited. Scores were compared using Kruskal–Wallis tests and epsilon-squared effect sizes. For structured expert review, five questions were randomly selected from each discipline, and five responses scoring above 90 were sampled per question and year (n=150). Five faculty experts independently assessed responses, blinded to year, using a standardized rubric covering organization, completeness, knowledge integration, clinical reasoning, and overall quality.
Results Clinical scores differed across years (H=11.78, p=0.003), but the effect was negligible (ε²=0.005); means were 80.95, 81.61, and 83.20. Basic-science scores varied more (H=94.98, p<0.001; ε²=0.033), with means of 69.76, 65.79, and 75.18. This non-linear pattern did not indicate consistent year-on-year score inflation. Expert review identified modest improvements in organization and readability, while completeness, knowledge integration, and clinical reasoning remained comparable.
Conclusions Across three generative-AI-era cohorts, supervised examination performance showed no uniform score inflation. The findings do not support abandoning written examinations solely because of perceived AI-related effects. Assessment should instead evolve towards authentic case-based tasks requiring justification of clinical decisions. As individual AI use was not measured and no pre-AI cohort was available, the results indicate temporal patterns rather than causal effects.