Quantification du délai : modélisation de l’impact de la rapidité de la rétroaction narrative dans les évaluations des activités professionnelles confiables (EPA) au cours de la formation en médecine interne à l’aide de l’intelligence artificielle

Auteurs-es

  • Enoch Yu Queen’s University
  • Haozhong Tian Queen’s University
  • Farida Mohamad Queen’s University
  • Karen Schultz Queen’s University
  • Laura McEwen Queen’s University
  • Stephen Gauthier Queen’s University
  • Heather Braund Queen’s University
  • Nicholas Cofie Queen’s University
  • Nancy Dalgarno Queen’s University
  • Adam Szulewski Queen’s University
  • Farhana Zulkernine Queen’s University
  • Benjamin YM Kwan Queen’s University

DOI :

https://doi.org/10.36834/v1anvj69

Résumé

Introduction : Dans la formation médicale axée sur les compétences, en particulier les activités professionnelles confiables (EPA), une rétroaction rapide est essentielle, car les retards peuvent fausser la mémoire des évaluateurs et la documentation de la performance. Des travaux antérieurs associent un délai plus long à une rétroaction de moindre qualité, mais ses effets selon différents seuils et résultats demeurent incertains. Cette étude examine ces relations au moyen d’analyses statistiques, en intégrant des outils d’intelligence artificielle (IA) pour standardiser l’extraction de mesures issues de la rétroaction narrative.

Méthodes : Nous avons mené une étude rétrospective transversale d’évaluations EPA de résidents en médecine interne, comprenant chacune deux sections narratives (« Commentaires généraux », « Prochaines étapes »), des scores de confiance (0-4) et des dates de réalisation. Des modèles préentraînés fondés sur des transformeurs ont estimé la positivité du sentiment (0-1) et les scores de qualité de l’évaluation pour l’apprentissage (QuAL) (0-5). Nous avons affiné manuellement les résultats, puis ajusté les données au moyen de régressions logistique ordinale, logistique binaire et linéaire. Des analyses exploratoires ont permis d’identifier les seuils de délai associés à des changements dans les mesures de rétroaction.

Résultats : 301 évaluateurs ont évalué 1 147 EPA pour 111 résidents, avec un délai médian de 7 jours (IQR : 21). Un délai plus long était associé à une QuAL plus faible (OR=0,997/jour ; p<0,05), à moins de mots (-0,07 mot/jour ; p<0,05) et à une probabilité plus élevée de commentaires vides (OR=1,003/jour ; p<0,01). Les effets sont apparus de façon séquentielle : augmentation transitoire de la confiance autour de sept jours, suivie d’une baisse dose-dépendante de la qualité à partir de 14 jours, puis d’une réduction de la longueur de la rétroaction. Au-delà de 100 jours, les commentaires devenaient uniformément brefs, avec une positivité et une confiance plus élevées, suggérant une diminution de la spécificité et du rappel.

Conclusion : Les retards de réalisation affectent de façon significative la qualité, l’exhaustivité et le ton de la rétroaction. Les résultats appuient des cibles de réalisation de 7 à 14 jours et le signalement des évaluations EPA dépassant 100 jours comme non fiables. L’extraction assistée par IA de mesures narratives pourrait soutenir une analyse standardisée et évolutive des évaluations EPA dans la formation médicale axée sur les compétences.

Références

1. Leclair R, Ho JSS, Braund H, et al. Exploring the quality of narrative feedback provided to residents during ambulatory patient care in medicine and surgery. J Med Educ Curric Dev. 2023 May 16;10:23821205231175734. https://doi.org/10.1177/23821205231175734.

2. Brown A, Currie D, Mercia M, et al. Does the implementation of competency-based medical education impact the quality of narrative feedback? a retrospective analysis of assessment data in a Canadian internal medicine residency program. Can J Gen Intern Med. 2022 Nov;17(4):67–85. https://doi.org/10.22374/cjgim.v17i4.640.

3. Liu L, Jiang Z, Qi X, et al. An update on current EPAs in graduate medical education: a scoping review. Med Educ Online. 26(1):1981198. https://doi.org/10.1080/10872981.2021.1981198.

4. Madrazo L, DCruz J, Correa N, Puka K, Kane SL. Evaluating the quality of written feedback within entrustable professional activities in an internal medicine cohort. J Grad Med Educ. 2023 Feb;15(1):74–80. https://doi.org/10.4300/JGME-D-22-00222.1.

5. ten Cate O. Entrustability of professional activities and competency-based training. Med Educ. 2005 Dec;39(12):1176–7. https://doi.org/10.1111/j.1365-2929.2005.02341.x.

6. Klein R, Snyder ED, Koch J, et al. Analysis of narrative assessments of internal medicine resident performance: are there differences associated with gender or race and ethnicity? BMC Med Educ. 2024 Jan 17;24(1):72. https://doi.org/10.1186/s12909-023-04970-2.

7. Branfield Day L, Miles A, Ginsburg S, Melvin L. Resident perceptions of assessment and feedback in competency-based medical education: a focus group study of one internal medicine residency program. Acad Med. 2020 Nov;95(11):1712–7. https://doi.org/10.1097/ACM.0000000000003315

8. Fernandes RD, de Vries I, McEwen L, et al. Evaluating the quality of narrative feedback for entrustable professional activities in a surgery residency program. Ann Surg. 2024 Apr 25. https://doi.org/10.1097/SLA.0000000000006308.

9. Marcotte L, Egan R, Soleas E, et al. Assessing the quality of feedback to general internal medicine residents in a competency-based environment. Can Med Educ J. 2019 Nov 28;10(4):e32–47. https://doi.org/10.36834/cmej.57323

10. Gundle KR, Mickelson DT, Hanel DP. Reflections in a time of transition: orthopaedic faculty and resident understanding of accreditation schemes and opinions on surgical skills feedback. Med Educ Online. 2016 Jan 1. https://doi.org/10.3402/meo.v21.30584.

11. Gordon M, Daniel M, Ajiboye A, et al. A scoping review of artificial intelligence in medical education: BEME guide no. 84. Med Teach. 2024 Apr 2;46(4):446–70. https://doi.org/10.1080/0142159X.2024.2314198.

12. Bello RJ, Sarmiento S, Rosson GD, et al. Understanding the role for operative performance rating tools in meeting surgical trainee feedback needs: a qualitative study. Plast Reconstr Surg Glob Open. 2016 Jun 29;4(6):e780. https://doi.org/10.1097/GOX.0000000000000777.

13. Tuma F, Nassar AK. Feedback in medical education. In: StatPearls. Treasure Island (FL): StatPearls Publishing; 2025.

14. Ahn E, LaDonna KA, Landreville JM, Mcheimech R, Cheung WJ. Only as strong as the weakest link: resident perspectives on entrustable professional activities and their impact on learning. J Grad Med Educ. 2023 Dec;15(6):676–84. https://doi.org/10.4300/JGME-D-23-00204.1.

15. Nathwani JN, Glarner CE, Law KE, et al. Integrating post-operative feedback into workflow: perceived practices and barriers. J Surg Educ. 2017;74(3):406–14. https://doi.org/10.1016/j.jsurg.2016.11.001.

16. Abeid M, Mbekenga C, Jusabani AM, et al. “It’s the time they take to give feedback; it has become a problem” exploring barriers and facilitators for research supervision in postgraduate medical education in Tanzania: a qualitative study. Adv Med Educ Pract. 2025 Apr 28;16:685–98. https://doi.org/10.2147/AMEP.S512498.

17. Hanmore T, Moon C, Curtis R, Hopman W, Baxter S. Is time really of the essence? timeliness of narrative feedback in ophthalmology CBME assessments. Med Teach. 2024;46(5):705–10. https://doi.org/10.1080/0142159X.2023.2274286.

18. Williams RG, Chen X (Phoenix), Sanfey H, et al. The measured effect of delay in completing operative performance ratings on clarity and detail of ratings assigned. J Surg Educ. 2014 Nov 1;71(6):e132–8. https://doi.org/10.1016/j.jsurg.2014.06.015.

19. Chen JX, Kozin E, Bohnen J, et al. Tracking operative autonomy and performance in otolaryngology training using smartphone technology: a single institution pilot study. Laryngoscope Investig Otolaryngol. 2019;4(6):578–86. https://doi.org/10.1002/lio2.323.

20. Williams RG, George BC, Bohnen JD, et al. A proposed blueprint for operative performance training, assessment, and certification. Ann Surg. 2021 Apr;273(4):701. https://doi.org/10.1097/SLA.0000000000004467.

21. Koduri S, Altshuler DB, Khalsa SSS, et al. Using a mobile application for evaluation of procedural learning in neurosurgery. World Neurosurg. 2020 Jun 1;138:e124–50. https://doi.org/10.1016/j.wneu.2020.02.049.

22. Abbott KL, Krumm AE, Kelley JK, et al. Surgical trainee performance and alignment with surgical program director expectations. Ann Surg. 2022 Dec;276(6):e1095. https://doi.org/10.1097/SLA.0000000000004990.

23. Abdelsattar JM, AlJamal YN, Ruparel RK, et al. Correlation of objective assessment data with general surgery resident in-training evaluation reports and operative volumes. J Surg Educ. 2018 Nov 1;75(6):1430–6. https://doi.org/10.1016/j.jsurg.2018.04.016.

24. Bello RJ, Meyer ML, Cooney DS, et al. The reliability of operative rating tool evaluations: how late is too late to provide operative performance feedback? Am J Surg. 2018 Dec 1;216(6):1052–5. https://doi.org/10.1016/j.amjsurg.2018.04.005.

25. Hiller KM, Waterbrook A, Waters K. Timing of emergency medicine student evaluation does not affect scoring. J Emerg Med. 2016 Feb 1;50(2):302–7. https://doi.org/10.1016/j.jemermed.2015.09.010.

26. Traser CJ. “Do I really have to complete another evaluation?” exploring relationships among physicians’ evaluative load, evaluative strain, and the quality of clinical clerkship evaluations. Indiana University–Purdue University Indianapolis; 2017. http://dx.doi.org/10.7912/C2/2110

27. Abdou H, Kidd-Romero S, Brown RF, Kavic SM, Kubicki NS. Keep it SIMPL: improved feedback after implementation of an app-based feedback tool. Am Surg. 2022 Jul 1;88(7):1475–8. https://doi.org/10.1177/00031348221082279.

28. Faiella W, Navjot S, Ramer S. Competency-based cardiology training: a simple approach to improve supervisor completion of entrustable professional activities. CJC Open. 2024 Oct 1;6(10):1248–53. https://doi.org/10.1016/j.cjco.2024.07.007.

29. Elentra https://elentra.com

30. Chen HC, van den Broek WES, ten Cate O. The case for use of entrustable professional activities in undergraduate medical education. Acad Med. 2015 Apr;90(4):431–6. https://doi.org/10.1097/ACM.0000000000000586

31. Ryan MS, Khan AR, Park YS, et al. Workplace-based entrustment scales for the core EPAs: a multisite comparison of validity evidence for two proposed instruments using structured vignettes and trained raters. Acad Med. 2022 Apr 1;97(4):544–51. https://doi.org/10.1097/ACM.0000000000004222

32. ten Cate O. A primer on entrustable professional activities. Korean J Med Educ. 2018 Mar;30(1):1–10. https://doi.org/10.3946/kjme.2018.76.

33. Sanh V, Debut L, Chaumond J, Wolf T. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv; 2020. https://doi.org/10.48550/arXiv.1910.01108.

34. Pang B, Lee L. Seeing stars: exploiting class relationships for sentiment categorization with respect to rating scales. ACL. 2005:115–24. https://doi.org/10.3115/1219840.1219855.

35. DistilBERT fine-tuned SST2. Hugging Face; 2024.

36. Chan TM, Sebok-Syer SS, Sampson C, Monteiro S. The quality of assessment of learning (Qual) score: validity evidence for a scoring system aimed at rating short, workplace-based comments on trainee performance. Teach Learn Med. 2020;32(3). https://doi.org/10.1080/10401334.2019.1708365.

37. Choo EK, Woods R, Walker ME, O’Brien JM, Chan TM. The quality of assessment for learning score for evaluating written feedback in anesthesiology postgraduate medical education: a generalizability and decision study. Can Med Educ J. 2023;14(6):78–85. https://doi.org/10.36834/cmej.75876.

38. Johnson AEW, Pollard T, Shen L et al. MIMIC-III, a freely accessible critical care database. Sci Data. 2016;3:160035. https://doi.org/10.1038/sdata.2016.35.

39. Alsentzer E, Murphy JR, Boag W, et al. Publicly available clinical BERT embeddings. arXiv; 2019. https://doi.org/10.48550/arXiv.1904.03323.

40. Spadafore M, Yilmaz Y, Rally V, et al. Using natural language processing to evaluate the quality of supervisor narrative comments in competency-based medical education. Acad Med. 2024;99(5):534. https://doi.org/10.1097/ACM.0000000000005634

41. Harris CR, Millman KJ, van der Walt SJ, et al. Array programming with NumPy. Nature. 2020;585:357–62. https://doi.org/10.1038/s41586-020-2649-2.

42. Pandas Development Team. Pandas. Zenodo; 2025. https://doi.org/10.5281/zenodo.17229934.

43. Hunter JD. Matplotlib: a 2D graphics environment. Comput Sci Eng. 2007;9(3):90–5.

44. Waskom ML. Seaborn: statistical data visualization. J Open Source Softw. 2021;6(60):3021. https://doi.org/10.21105/joss.03021.

45. Virtanen P, Gommers R, Oliphant TE et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17:261–72. https://doi.org/10.1038/s41592-019-0686-2.

46. Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825–30.

47. Seabold S, Perktold J. Statsmodels: econometric and statistical modeling with Python. 2010.

48. Massie J, Ali JM. Workplace-based assessment: a review of user perceptions and strategies to address the identified shortcomings. Adv Health Sci Educ Theory Pract. 2016 May;21(2). https://doi.org/10.1007/s10459-015-9614-0.

49. Poeppelman RS, Cho J, Nachbor K, et al. “But why?”: explanatory feedback is a reliable marker of high-quality narrative assessment of surgical performance. Acad Med. 2025 May;100(5):614–20. https://doi.org/10.1097/ACM.0000000000005985

50. Dudek NL, Marks MB, Regehr G. Failure to fail: the perspectives of clinical supervisors. Acad Med. 2005 Oct;80(10 Suppl):S84–7. https://doi.org/10.1097/00001888-200510001-00023.

51. Dexter F, Vasilopoulos T, Fahy BG. Adjusting for resident rater leniency or severity improves the reliability of routine resident evaluations of faculty anesthesiologists. Cureus. 2025 Jun 19;17(6). https://doi.org/10.7759/cureus.86366.

52. Curtis R, Moon CC, Hanmore T, Hopman WM, Baxter S. Use the right words: evaluating the effect of word choice and word count on quality of narrative feedback in ophthalmology competency-based medical education assessments. Can Med Educ J. 2024 Apr 29. https://doi.org/10.36834/cmej.76671.

53. Page M, Gardner J, Booth J. Validating written feedback in clinical formative assessment. Assess Eval High Educ. 2020 Jul 3;45(5):697–713. https://doi.org/10.1080/02602938.2019.1691974.

54. Marty AP, Linsenmeyer M, George B, et al. Mobile technologies to support workplace-based assessment for entrustment decisions: guidelines for programs and educators: AMEE guide no. 154. Med Teach. 2023 Nov 2;45(11):1203–13. https://doi.org/10.1080/0142159X.2023.2168527.

55. Kit Y, Mokji MM. Sentiment analysis using pre-trained language model with no fine-tuning and less resource. IEEE Access. 2022;10:107056–65. https://doi.org/10.1109/ACCESS.2022.3212367.

56. Gin BC, ten Cate O, O’Sullivan PS, Boscardin C. Assessing supervisor versus trainee viewpoints of entrustment through cognitive and affective lenses: an artificial intelligence investigation of bias in feedback. Adv Health Sci Educ. 2024;29(5):1571–92. https://doi.org/10.1007/s10459-024-10311-9

Téléchargements

Publié

2026-06-30

Numéro

Rubrique

Recherche originale

Comment citer

1.
Quantification du délai : modélisation de l’impact de la rapidité de la rétroaction narrative dans les évaluations des activités professionnelles confiables (EPA) au cours de la formation en médecine interne à l’aide de l’intelligence artificielle. Can. Med. Ed. J [Internet]. 30 juin 2026 [cité 24 août 2026];. Disponible sur: https://cmej.ca/index.php/cmej/article/view/82276

Articles les plus lus du,de la,des même-s auteur-e-s

1 2 3 > >>