Structured Nursing Handover Report Generation from Clinical Speech using Fine-Tuned XLSR-53 and T5: A Benchmarking Study
Abstract
Accurate nursing handovers are critical for patient safety, as miscommunication during shift transitions leads to irreversible clinical errors. This work proposed an end-to-end pipeline that converts unstructured clinical nursing speech into standardized handover reports using a fine-tuned XLSR-53 acoustic model and T5-base text-to-text transformer. An Australian English clinical corpus of 200 synthetic nursing handover recordings from the CSIRO data access portal was utilised in this work. This benchmarking study was conducted within the CSIRO synthetic Australian English nursing handover corpus and does not represent a broad cross-domain clinical ASR benchmark. A domain-specific benchmarking study across seven state-of-the-art ASR architectures (Whisper Tiny/Base/Small, Wav2Vec2 Base/Large, HuBERT Large, XLSR-53) was conducted using this corpus. The experimental results further revealed XLSR-53 as the optimal architecture for clinical nursing speech recognition. A partial layer-freeze strategy was adopted in XLSR-53 by freezing the first 12 of 24 encoder layers, empirically validated through an ablation study with five freeze configurations (L=0, 6, 12, 18, 24). XLSR-53 preserves cross-lingual phonetic representations while enabling clinical vocabulary adaptation. A clinically motivated evaluation framework using curated medical vocabulary terms computes Medical Precision, Recall, and F1-Score along with standard WER, CER, and PER to assess reliability in clinical term recognition. Benchmarking against Google Health AI's MedASR zero-shot revealed that the proposed system XLSR-53 (L=12) achieved 17.15% WER against MedASR's 28.87% (p<0.001) with a Medical F1-Score of 0.98 and ROUGE-L of 0.92. Although results were obtained on synthetic Australian English speech, performance under real clinical conditions with background noise, overlapping speakers, and spontaneous interruptions requires further validation.
Downloads
References
M. I. Alhosani, F. R. Ahmed, N. Al-Yateem, H. S. Mobarak, and M. E. AbuRuz, "Assessment of nursing workload and adverse events reporting among critical care nurses in the United Arab Emirates," The Open Nursing Journal, vol. 17, e18744346281511. ISSN: 1874-4346, 2023. DOI: 10.2174/0118744346281511231120054125.
Z. Banda, M. Simbota, and C. Mula, "Nurses' perceptions on the effects of high nursing workload on patient care in an intensive care unit of a referral hospital in Malawi: A qualitative study," BMC Nursing, vol. 21(1), 136, 2022. DOI: 10.1186/s12912-022-00918-x.
M. Shohani and H. Tavan, "Factors affecting medication errors from the perspective of nursing staff," Journal of Clinical and Diagnostic Research, vol.12(3), pp. IC01–IC04, 2018. DOI: 10.7860/JCDR/2018/28447.11336.
E. Gesner, P. C. Dykes, L. Zhang, and P. Gazarian, "Documentation burden in nursing and its role in clinician burnout syndrome," Applied Clinical Informatics, vol. 13(5), pp. 983–990, 2022. DOI: 10.1055/s-0042-1757157.
R. A. Atinga, M. N. Gmaligan, A. Ayawine, and J. K. Yambah, "It's the patient that suffers from poor communication: Analyzing communication gaps and associated consequences in handover events from nurses' experiences," SSM - Qualitative Research in Health, vol. 6, Art. no. 100482, 2024. DOI: 10.1016/j.ssmqr.2024.100482.
J. Y. Uhm, E. Y. Lim, and J. Hyeong, "The impact of a standardized inter-department handover on nurses' perceptions and performance in Republic of Korea," Journal of Nursing Management, vol. 26, pp. 933–944, 2018. DOI: 10.1111/jonm.12608.
S. Ghosh, L. Ramamoorthy, and B. Pottakat, "Impact of structured clinical handover protocol on communication and patient satisfaction," Journal of Patient Experience, vol. 8, Art. no. 2374373521997733, 2021. DOI: 10.1177/2374373521997733.
A. L. Cooper, J. A. Brown, S. P. Eccles, N. Cooper, and M. A. Albrecht, "Is nursing and midwifery clinical documentation a burden? An empirical study of perception versus reality," Journal of Clinical Nursing, vol. 30, pp. 1645–1652, 2021. DOI: 10.1111/jocn.15718.
I. Köse Tosunöz and A. Aydınlı, "Nurses' perceptions of the effectiveness of handover practices, influencing factors and perceived barriers: A descriptive cross-sectional study of medical and surgical nurses," Collegian, vol. 32, pp. 242–249, 2025. DOI: 10.1016/j.colegn.2025.06.002.
J. A. Brown, A. L. Cooper, and M. A. Albrecht, "Development and content validation of the Burden of Documentation for Nurses and Midwives (BurDoNsaM) survey," Journal of Advanced Nursing, vol. 76, pp. 1273–1281, 2020. DOI: 10.1111/jan.14320.
X. Luo, L. Zhou, K. Adelgais, et al., "Assessing the effectiveness of automatic speech recognition technology in emergency medicine settings: A comparative study of four AI-powered engines," Journal of Healthcare Informatics Research, vol. 9, pp. 494–512, 2025. DOI: 10.1007/s41666-025-00193-w.
K. Denecke, L. Meier, J. G. Bauer, M. Bender, and C. Lueg, "Information capturing in pre-hospital emergency medical settings (EMS)," Studies in Health Technology and Informatics, vol. 270, pp. 613–617, 2020. DOI: 10.3233/shti200233.
C. C. Chiu, A. Tripathi, K. Chou, C. Co, N. Jaitly, D. Jaunzeikare, A. Kannan, P. Nguyen, H. Sak, A. Sankar, et al., "Speech recognition for medical conversations," in Proc. Interspeech 2018, pp.2972–2976, 2018. DOI: 10.21437/Interspeech.2018-40.
E. Edwards, W. Salloum, G. P. Finley, J. Fone, G. Cardiff, M. Miller, and D. Suendermann-Oeft, "Medical speech recognition: Reaching parity with humans," in Speech and Computer, A. Karpov, R. Potapova, and I. Mporas, Eds. Cham, Switzerland, pp.512–524, 2017. DOI: 10.1007/978-3-319-66429-3_51.
T. Hodgson, F. Magrabi, and E. Coiera, "Efficiency and safety of speech recognition for documentation in the electronic health record," Journal of the American Medical Informatics Association, vol. 24, pp. 1127–1133, 2017. DOI: 10.1093/jamia/ocx073.
Angel, Maricel; Suominen, Hanna; Zhou, Liyuan; & Hanlen, Leif (2014): Synthetic nursing handover training and development data set - audio files. v1. CSIRO. Data Collection. DOI: 10.4225/08/58d0977ab4888.
Sasikala D, S. Siva Sathya, S. Baghavathi Priya, D. Niranjan Kumar, and S. Vignesh, "A review of automatic speech recognition approaches for bridging communication gaps in clinical handover," in Proc. 2nd IEEE International Conference for Women in Computing (InCoWoCo), pp.1–7, 2025. DOI: 10.1109/InCoWoCo68239.2025.11407097.
A. J. Nashwan, A. Abujaber, and S. K. Ahmed, "Charting the future: The role of AI in transforming nursing documentation," Cureus, vol. 16, no. 3, 2024. DOI: 10.7759/cureus.57304.
S. Yadav, "Embracing artificial intelligence: Revolutionizing nursing documentation for a better future," Cureus, vol. 16, no. 4, 2024. DOI: 10.7759/cureus.57725.
J. Kodish-Wachs, E. Agassi, P. I. Kenny, and J. M. Overhage, "A systematic comparison of contemporary automatic speech recognition engines for conversational clinical speech," AMIA Annual Symposium Proceedings, vol. 2018, pp. 683–689, 2018. PMCID: PMC6371385.
A. Femi-Abodunde, K. Olinger, L. M. B. Burke, T. Benefield, E. R. Lee, K. McGinty, and B. M. Mervak, "Radiology dictation errors with COVID-19 protective equipment: Does wearing a surgical mask increase the dictation error rate?", Journal of Digital Imaging, vol. 34, pp. 1294–1301, 2021. DOI: 10.1007/s10278-021-00502-w.
K. Le-Duc, "VietMed: A dataset and benchmark for automatic speech recognition of Vietnamese in the medical domain",arXiv, arXiv:2404.05659, 2024. DOI: 10.48550/arXiv.2404.05659.
S. Banerjee, A. Agarwal, and P. Ghosh, "High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR," arXiv preprint, arXiv:2412.00055, 2024. DOI: 10.48550/arXiv.2412.00055.
T. Afonja, T. Olatunji, S. Ogun, N. A. Etori, A. Owodunni, and M. Yekini, "Performant ASR models for medical entities in accented speech," in Proc. Interspeech 2024, pp. 2315–2319, 2024. DOI: 10.21437/Interspeech.2024-2261.
A. Prasad and P. Jyothi, "How accents confound: Probing for accent information in end-to-end speech recognition systems," in Proc. 58th Annual Meeting of the Association for Computational Linguistics (ACL), 2020, pp. 3739–3753. DOI: 10.18653/v1/2020.acl-main.345.
M. Najafian, A. DeMarco, S. Cox, and M. Russell, "Unsupervised model selection for recognition of regional accented speech," in Proc. Interspeech 2014, pp.2967-2971, 2014. DOI: 10.21437/Interspeech.2014-495.
C. T. Do, S. Imai, R. S. Doddipatla, and T. Hain, "Improving accented speech recognition using data augmentation based on unsupervised text-to-speech synthesis," in Proceedings of the 2024 32nd European Signal Processing Conference (EUSIPCO), Lyon, France, 2024, pp. 136–140, 2024. DOI: 10.23919/EUSIPCO63174.2024.10715166.
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, "Unsupervised cross-lingual representation learning for speech recognition," in Proc. Interspeech 2021, Brno, Czech Republic, pp. 2426–2430, 2021. DOI: 10.21437/Interspeech.2021-329.
X. Zhang and L. He, "End-to-end cross-lingual spoken language understanding model with multilingual pretraining," in Proc. Interspeech 2021, pp.4728–4732, 2021. DOI: 10.21437/Interspeech.2021-818.
T. Pekarek Rosin and S. Wermter, "Replay to remember: Continual layer-specific fine-tuning for German speech recognition," in Proc. 32nd International Conference on Artificial Neural Networks (ICANN 2023), pp.489–500, 2023. DOI: 10.1007/978-3-031-44195-0_40.
Y. Li, K. Harrigian, A. Zirikly, and M. Dredze, "Are clinical T5 models better for clinical text?" ,in Proc. 4th Machine Learning for Health Symposium, Proceedings of Machine Learning Research, vol. 259, pp.636-667, 2025. DOI: 10.48550/arXiv.2412.05845.
P. Manakul, Y. Fathullah, A. Liusie, V. Raina, and M. Gales, "CUED at ProbSum 2023: Hierarchical ensemble of summarization models," in Proc. 22nd Workshop on Biomedical Natural Language Processing and BioNLP Shared Tasks, Toronto, Canada, pp. 516–523, 2023. DOI: 10.18653/v1/2023.bionlp-1.51.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, "Exploring the limits of transfer learning with a unified text-to-text transformer," arXiv preprint, arXiv:1910.10683, 2020. DOI: 10.48550/arXiv.1910.10683.
J. Giorgi et al., "Clinical note generation from doctor-patient conversations using large language models," in Proc. ClinicalNLP Workshop, pp.323-334, 2023. DOI: 10.18653/v1/2023.clinicalnlp-1.36.
A. Toma, P. R. Lawler, J. Ba, R. G. Krishnan, B. B. Rubin, and B. Wang, "Clinical Camel: An open expert-level medical language model with dialogue-based knowledge encoding," arXiv preprint, arXiv:2305.12031, 2023. DOI: 10.48550/arXiv.2305.12031.
Y. Labrak, A. Bazoge, E. Morin, P.-A. Gourraud, M. Rouvier, and R. Dufour, "BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains," in Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, pp. 5848–5864, 2024. DOI: 10.18653/v1/2024.findings-acl.348.
P. C. Nair, D. Gupta, and B. I. Devi, "Extracting clinical relationships from discharge summaries of suprasellar lesion patients using Gemini LLM", Procedia Computer Science, vol. 258, pp.2391-2404, 2025. DOI: 10.1016/j.procs.2025.04.502.
D. Sasikala, R. Sudarshan, and S. Sivasathya, "Harnessing LLMs for medical insights: NER extraction from summarized medical text," in Proc. 15th International Conference on Computing Communication and Networking Technologies(ICCCNT), pp.1–6, 2024. DOI: 10.1109/ICCCNT61001.2024.10724860.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, "Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks," in Proc. 23rd International Conference on Machine Learning (ICML), pp. 369–376, 2006. DOI: 10.1145/1143844.1143891.
D. Klakow and J. Peters, "Testing the correlation of word error rate and perplexity," Speech Communication, vol. 38, no. 1–2, pp. 19–28, 2002. DOI: 10.1016/S0167-6393(01)00041-3.
C. J. van Rijsbergen, Information Retrieval, 2nd ed. London, UK: Butterworths, 1979. Available: https://dl.acm.org/doi/book/10.5555/539927.
C.-Y. Lin, "ROUGE: A package for automatic evaluation of summaries," in Proc. ACL Workshop on Text Summarization Branches Out, Barcelona, Spain, pp. 74–81, 2004. ACL Anthology ID: W04-1013 . Available: https://aclanthology.org/W04-1013/.
A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing, 3rd ed. Upper Saddle River, NJ: Prentice Hall, 2009. ISBN: 978-0-13-198842-2.
A. Radford et al., "Robust speech recognition via large-scale weak supervision," in Proc. International Conference on Machine Learning (ICML), vol. 202, pp. 28492–28518, 2023. DOI: 10.48550/arXiv.2212.04356.
W. N. Hsu, B. Bolte, Y. H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, "HuBERT: Self-supervised speech representation learning by masked prediction of hidden units," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021. DOI: 10.1109/TASLP.2021.3122291.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, "wav2vec 2.0: A framework for self-supervised learning of speech representations," in Proc. Advances in Neural Information Processing Systems (NeurIPS), 2020, pp. 12449–12460. DOI: 10.48550/arXiv.2006.11477.
K. Wu, E. Variani, T. Bagby, S. Reddy, and R. Pilgrim, "MedASR: An Open-Source Model for High-Accuracy Medical Dictation,"
arXiv, arXiv:2605.16555, 2026. DOI: 10.48550/arXiv.2605.16555.
S. S. Shapiro and M. B. Wilk, "An analysis of variance test for normality (complete samples)," Biometrika, vol. 52, no. 3/4, pp. 591–611, 1965. DOI: 10.2307/2333709.
F. Wilcoxon, "Individual comparisons by ranking methods," Biometrics Bulletin, vol. 1, no. 6, pp. 80–83, 1945. DOI: 10.2307/3001968.
W. H. Kruskal and W. A. Wallis, "Use of ranks in one-criterion variance analysis," Journal of the American Statistical Association, vol. 47, no. 260, pp. 583–621, 1952. DOI: 10.1080/01621459.1952.10483441.
J. Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates, 1988. DOI: 10.4324/9780203771587.
B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York, NY: Chapman and all/CRC, 1994. DOI: 10.1201/9780429246593.
Suominen H, Zhou L, Hanlen L, Ferraro G, "Benchmarking clinical speech recognition and information extraction: New data, methods, and evaluations," JMIR Medical Informatics, vol. 3, no. 2, Art. no. e19, 2015.
DOI: 10.2196/medinform.4321.
Copyright (c) 2026 Sasikala D, Siva Sathya S, Niranjan Kumar D, Vignesh S

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-ShareAlikel 4.0 International (CC BY-SA 4.0) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).


.png)
.png)
.png)
.png)
.png)