Quantifying delay: modeling the impact of timeliness on narrative feedback for Entrustable Professional Activity assessments in internal medicine training using artificial intelligence
DOI:
https://doi.org/10.36834/v1anvj69Abstract
Introduction: In competency-based medical education (CBME), specifically entrustable professional activities (EPAs), timely feedback is essential because delays can distort assessor recall and performance documentation. Prior work links longer timelag to reduced feedback quality, but its effects across thresholds and outcomes remain unclear. This study examines these relationships using statistical analyses, incorporating artificial intelligence (AI) tools to standardize extraction of narrative feedback metrics.
Methods: We performed a retrospective cross-sectional study of internal medicine resident EPA assessments, each including two narrative sections (‘General Feedback,’ ‘Next Steps’), entrustment scores (0-4), and completion dates. Pretrained transformer-based models estimated sentiment positivity (0-1) and Quality of Assessment for Learning (QuAL) scores (0-5). We manually refined outputs, subsequently fitting datapoints for ordinal logistic, binary logistic, and linear regression. Exploratory analyses identified timelag thresholds linked to changes in feedback metrics.
Results: 301 assessors evaluated 1,147 EPAs for 111 residents, with a 7-day median timelag (IQR: 21). Longer timelag was associated with lower QuAL (OR=0.997/day; p < 0.05), fewer words (-0.07words/day; p<0.05), and higher odds of empty comments (OR=1.003/day; p < 0.01). Effects emerged sequentially: transient increases in entrustment around seven days, followed by dose-dependent declines in quality from 14 days and reductions in feedback length thereafter. Beyond 100 days, comments became uniformly brief with higher positivity and entrustment, suggesting reduced specificity and recall.
Conclusion: Delayed completion meaningfully affects feedback quality, completeness, and tone. Findings support 7-14 day completion targets and flagging EPA assessments beyond 100 days as unreliable. AI-assisted extraction of narrative metrics may support standardized, scalable EPA assessment analysis in CBME.
References
1. Leclair R, Ho JSS, Braund H, et al. Exploring the quality of narrative feedback provided to residents during ambulatory patient care in medicine and surgery. J Med Educ Curric Dev. 2023 May 16;10:23821205231175734. https://doi.org/10.1177/23821205231175734.
2. Brown A, Currie D, Mercia M, et al. Does the implementation of competency-based medical education impact the quality of narrative feedback? a retrospective analysis of assessment data in a Canadian internal medicine residency program. Can J Gen Intern Med. 2022 Nov;17(4):67–85. https://doi.org/10.22374/cjgim.v17i4.640.
3. Liu L, Jiang Z, Qi X, et al. An update on current EPAs in graduate medical education: a scoping review. Med Educ Online. 26(1):1981198. https://doi.org/10.1080/10872981.2021.1981198.
4. Madrazo L, DCruz J, Correa N, Puka K, Kane SL. Evaluating the quality of written feedback within entrustable professional activities in an internal medicine cohort. J Grad Med Educ. 2023 Feb;15(1):74–80. https://doi.org/10.4300/JGME-D-22-00222.1.
5. ten Cate O. Entrustability of professional activities and competency-based training. Med Educ. 2005 Dec;39(12):1176–7. https://doi.org/10.1111/j.1365-2929.2005.02341.x.
6. Klein R, Snyder ED, Koch J, et al. Analysis of narrative assessments of internal medicine resident performance: are there differences associated with gender or race and ethnicity? BMC Med Educ. 2024 Jan 17;24(1):72. https://doi.org/10.1186/s12909-023-04970-2.
7. Branfield Day L, Miles A, Ginsburg S, Melvin L. Resident perceptions of assessment and feedback in competency-based medical education: a focus group study of one internal medicine residency program. Acad Med. 2020 Nov;95(11):1712–7. https://doi.org/10.1097/ACM.0000000000003315
8. Fernandes RD, de Vries I, McEwen L, et al. Evaluating the quality of narrative feedback for entrustable professional activities in a surgery residency program. Ann Surg. 2024 Apr 25. https://doi.org/10.1097/SLA.0000000000006308.
9. Marcotte L, Egan R, Soleas E, et al. Assessing the quality of feedback to general internal medicine residents in a competency-based environment. Can Med Educ J. 2019 Nov 28;10(4):e32–47. https://doi.org/10.36834/cmej.57323
10. Gundle KR, Mickelson DT, Hanel DP. Reflections in a time of transition: orthopaedic faculty and resident understanding of accreditation schemes and opinions on surgical skills feedback. Med Educ Online. 2016 Jan 1. https://doi.org/10.3402/meo.v21.30584.
11. Gordon M, Daniel M, Ajiboye A, et al. A scoping review of artificial intelligence in medical education: BEME guide no. 84. Med Teach. 2024 Apr 2;46(4):446–70. https://doi.org/10.1080/0142159X.2024.2314198.
12. Bello RJ, Sarmiento S, Rosson GD, et al. Understanding the role for operative performance rating tools in meeting surgical trainee feedback needs: a qualitative study. Plast Reconstr Surg Glob Open. 2016 Jun 29;4(6):e780. https://doi.org/10.1097/GOX.0000000000000777.
13. Tuma F, Nassar AK. Feedback in medical education. In: StatPearls. Treasure Island (FL): StatPearls Publishing; 2025.
14. Ahn E, LaDonna KA, Landreville JM, Mcheimech R, Cheung WJ. Only as strong as the weakest link: resident perspectives on entrustable professional activities and their impact on learning. J Grad Med Educ. 2023 Dec;15(6):676–84. https://doi.org/10.4300/JGME-D-23-00204.1.
15. Nathwani JN, Glarner CE, Law KE, et al. Integrating post-operative feedback into workflow: perceived practices and barriers. J Surg Educ. 2017;74(3):406–14. https://doi.org/10.1016/j.jsurg.2016.11.001.
16. Abeid M, Mbekenga C, Jusabani AM, et al. “It’s the time they take to give feedback; it has become a problem” exploring barriers and facilitators for research supervision in postgraduate medical education in Tanzania: a qualitative study. Adv Med Educ Pract. 2025 Apr 28;16:685–98. https://doi.org/10.2147/AMEP.S512498.
17. Hanmore T, Moon C, Curtis R, Hopman W, Baxter S. Is time really of the essence? timeliness of narrative feedback in ophthalmology CBME assessments. Med Teach. 2024;46(5):705–10. https://doi.org/10.1080/0142159X.2023.2274286.
18. Williams RG, Chen X (Phoenix), Sanfey H, et al. The measured effect of delay in completing operative performance ratings on clarity and detail of ratings assigned. J Surg Educ. 2014 Nov 1;71(6):e132–8. https://doi.org/10.1016/j.jsurg.2014.06.015.
19. Chen JX, Kozin E, Bohnen J, et al. Tracking operative autonomy and performance in otolaryngology training using smartphone technology: a single institution pilot study. Laryngoscope Investig Otolaryngol. 2019;4(6):578–86. https://doi.org/10.1002/lio2.323.
20. Williams RG, George BC, Bohnen JD, et al. A proposed blueprint for operative performance training, assessment, and certification. Ann Surg. 2021 Apr;273(4):701. https://doi.org/10.1097/SLA.0000000000004467.
21. Koduri S, Altshuler DB, Khalsa SSS, et al. Using a mobile application for evaluation of procedural learning in neurosurgery. World Neurosurg. 2020 Jun 1;138:e124–50. https://doi.org/10.1016/j.wneu.2020.02.049.
22. Abbott KL, Krumm AE, Kelley JK, et al. Surgical trainee performance and alignment with surgical program director expectations. Ann Surg. 2022 Dec;276(6):e1095. https://doi.org/10.1097/SLA.0000000000004990.
23. Abdelsattar JM, AlJamal YN, Ruparel RK, et al. Correlation of objective assessment data with general surgery resident in-training evaluation reports and operative volumes. J Surg Educ. 2018 Nov 1;75(6):1430–6. https://doi.org/10.1016/j.jsurg.2018.04.016.
24. Bello RJ, Meyer ML, Cooney DS, et al. The reliability of operative rating tool evaluations: how late is too late to provide operative performance feedback? Am J Surg. 2018 Dec 1;216(6):1052–5. https://doi.org/10.1016/j.amjsurg.2018.04.005.
25. Hiller KM, Waterbrook A, Waters K. Timing of emergency medicine student evaluation does not affect scoring. J Emerg Med. 2016 Feb 1;50(2):302–7. https://doi.org/10.1016/j.jemermed.2015.09.010.
26. Traser CJ. “Do I really have to complete another evaluation?” exploring relationships among physicians’ evaluative load, evaluative strain, and the quality of clinical clerkship evaluations. Indiana University–Purdue University Indianapolis; 2017. http://dx.doi.org/10.7912/C2/2110
27. Abdou H, Kidd-Romero S, Brown RF, Kavic SM, Kubicki NS. Keep it SIMPL: improved feedback after implementation of an app-based feedback tool. Am Surg. 2022 Jul 1;88(7):1475–8. https://doi.org/10.1177/00031348221082279.
28. Faiella W, Navjot S, Ramer S. Competency-based cardiology training: a simple approach to improve supervisor completion of entrustable professional activities. CJC Open. 2024 Oct 1;6(10):1248–53. https://doi.org/10.1016/j.cjco.2024.07.007.
29. Elentra https://elentra.com
30. Chen HC, van den Broek WES, ten Cate O. The case for use of entrustable professional activities in undergraduate medical education. Acad Med. 2015 Apr;90(4):431–6. https://doi.org/10.1097/ACM.0000000000000586
31. Ryan MS, Khan AR, Park YS, et al. Workplace-based entrustment scales for the core EPAs: a multisite comparison of validity evidence for two proposed instruments using structured vignettes and trained raters. Acad Med. 2022 Apr 1;97(4):544–51. https://doi.org/10.1097/ACM.0000000000004222
32. ten Cate O. A primer on entrustable professional activities. Korean J Med Educ. 2018 Mar;30(1):1–10. https://doi.org/10.3946/kjme.2018.76.
33. Sanh V, Debut L, Chaumond J, Wolf T. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv; 2020. https://doi.org/10.48550/arXiv.1910.01108.
34. Pang B, Lee L. Seeing stars: exploiting class relationships for sentiment categorization with respect to rating scales. ACL. 2005:115–24. https://doi.org/10.3115/1219840.1219855.
35. DistilBERT fine-tuned SST2. Hugging Face; 2024.
36. Chan TM, Sebok-Syer SS, Sampson C, Monteiro S. The quality of assessment of learning (Qual) score: validity evidence for a scoring system aimed at rating short, workplace-based comments on trainee performance. Teach Learn Med. 2020;32(3). https://doi.org/10.1080/10401334.2019.1708365.
37. Choo EK, Woods R, Walker ME, O’Brien JM, Chan TM. The quality of assessment for learning score for evaluating written feedback in anesthesiology postgraduate medical education: a generalizability and decision study. Can Med Educ J. 2023;14(6):78–85. https://doi.org/10.36834/cmej.75876.
38. Johnson AEW, Pollard T, Shen L et al. MIMIC-III, a freely accessible critical care database. Sci Data. 2016;3:160035. https://doi.org/10.1038/sdata.2016.35.
39. Alsentzer E, Murphy JR, Boag W, et al. Publicly available clinical BERT embeddings. arXiv; 2019. https://doi.org/10.48550/arXiv.1904.03323.
40. Spadafore M, Yilmaz Y, Rally V, et al. Using natural language processing to evaluate the quality of supervisor narrative comments in competency-based medical education. Acad Med. 2024;99(5):534. https://doi.org/10.1097/ACM.0000000000005634
41. Harris CR, Millman KJ, van der Walt SJ, et al. Array programming with NumPy. Nature. 2020;585:357–62. https://doi.org/10.1038/s41586-020-2649-2.
42. Pandas Development Team. Pandas. Zenodo; 2025. https://doi.org/10.5281/zenodo.17229934.
43. Hunter JD. Matplotlib: a 2D graphics environment. Comput Sci Eng. 2007;9(3):90–5.
44. Waskom ML. Seaborn: statistical data visualization. J Open Source Softw. 2021;6(60):3021. https://doi.org/10.21105/joss.03021.
45. Virtanen P, Gommers R, Oliphant TE et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17:261–72. https://doi.org/10.1038/s41592-019-0686-2.
46. Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825–30.
47. Seabold S, Perktold J. Statsmodels: econometric and statistical modeling with Python. 2010.
48. Massie J, Ali JM. Workplace-based assessment: a review of user perceptions and strategies to address the identified shortcomings. Adv Health Sci Educ Theory Pract. 2016 May;21(2). https://doi.org/10.1007/s10459-015-9614-0.
49. Poeppelman RS, Cho J, Nachbor K, et al. “But why?”: explanatory feedback is a reliable marker of high-quality narrative assessment of surgical performance. Acad Med. 2025 May;100(5):614–20. https://doi.org/10.1097/ACM.0000000000005985
50. Dudek NL, Marks MB, Regehr G. Failure to fail: the perspectives of clinical supervisors. Acad Med. 2005 Oct;80(10 Suppl):S84–7. https://doi.org/10.1097/00001888-200510001-00023.
51. Dexter F, Vasilopoulos T, Fahy BG. Adjusting for resident rater leniency or severity improves the reliability of routine resident evaluations of faculty anesthesiologists. Cureus. 2025 Jun 19;17(6). https://doi.org/10.7759/cureus.86366.
52. Curtis R, Moon CC, Hanmore T, Hopman WM, Baxter S. Use the right words: evaluating the effect of word choice and word count on quality of narrative feedback in ophthalmology competency-based medical education assessments. Can Med Educ J. 2024 Apr 29. https://doi.org/10.36834/cmej.76671.
53. Page M, Gardner J, Booth J. Validating written feedback in clinical formative assessment. Assess Eval High Educ. 2020 Jul 3;45(5):697–713. https://doi.org/10.1080/02602938.2019.1691974.
54. Marty AP, Linsenmeyer M, George B, et al. Mobile technologies to support workplace-based assessment for entrustment decisions: guidelines for programs and educators: AMEE guide no. 154. Med Teach. 2023 Nov 2;45(11):1203–13. https://doi.org/10.1080/0142159X.2023.2168527.
55. Kit Y, Mokji MM. Sentiment analysis using pre-trained language model with no fine-tuning and less resource. IEEE Access. 2022;10:107056–65. https://doi.org/10.1109/ACCESS.2022.3212367.
56. Gin BC, ten Cate O, O’Sullivan PS, Boscardin C. Assessing supervisor versus trainee viewpoints of entrustment through cognitive and affective lenses: an artificial intelligence investigation of bias in feedback. Adv Health Sci Educ. 2024;29(5):1571–92. https://doi.org/10.1007/s10459-024-10311-9
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Enoch Yu, Haozhong Tian, Farida Mohamad, Karen Schultz, Laura McEwen, Stephen Gauthier, Heather Braund, Nicholas Cofie, Nancy Dalgarno, Adam Szulewski, Farhana Zulkernine, Benjamin YM Kwan

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Submission of an original manuscript to the Canadian Medical Education Journal will be taken to mean that it represents original work not previously published, that it is not being considered elsewhere for publication. If accepted for publication, it will be published online and it will not be published elsewhere in the same form, for commercial purposes, in any language, without the consent of the publisher.
Authors who publish in the Canadian Medical Education Journal agree to release their articles under the Creative Commons Attribution-Noncommercial-No Derivative Works 4.0 Canada Licence. This licence allows anyone to copy and distribute the article for non-commercial purposes provided that appropriate attribution is given. For details of the rights an author grants users of their work, please see the licence summary and the full licence.


