MULTI-TASK DEEP LEARNING SEQUENCE PREDICTION FOR BILINGUAL HANDWRITTEN CLINICAL PRESCRIPTIONS

Authors

  • Ayesha Ahmad Author
  • Hafiz Haris Ali Author
  • Hafiz Muhammad Hussnain Author
  • Muhammad Hussnain Butt Author
  • Nadeem Iqbal Author
  • Abdul Razzaq Author
  • Fatima Bukhari Author

Keywords:

CNN, TrOCR, HTR, VGG Image Annotator, CER, WER

Abstract

Handwritten Text Recognition (HTR) in medical domain documents, particularly doctors’ prescriptions, remain a major challenge due to unconstrained handwriting, domain-specific terminology, and multi-script variations. To address these complexities, this study presents a specialized recognition workflow evaluated on a custom doctors’ prescription dataset categorized into seven distinct structural classes (Personal Information, Medicine Name, Dosage, Diagnostic, Symptoms, Numeric Data, and Text) across English, Urdu, and Numeric scripts. Ground-truth annotation was performed using the VGG Image Annotator (VIA) and exported to JSON format for pipeline integration. The proposed methodology employs a hybrid pipeline CNN - TrOCR comprising and EfficientNet-B0 backbone couples with a Transformer encoder for structural classification, followed by Transformer based OCR (TrOCR) for end-to-end word level text transcription. Experimental evaluation demonstrates that the classification network achieves an accuracy of 87.77% for language identification and 69.69% for structural category classification. For sequence recognition using TrOCR, the system achieves a Character Error Rate (CER) of 23.51% and a Word Error Rate (WER) of 37.91%. These findings highlights the potential  of combining convolution-transformer feature extractors with dedicated sequence decoders to transcribe complex, multi-lingual medical handwriting, offering a foundation for automated clinical documentation workflows.

Downloads

Published

2026-09-22