EVALUATING LARGE LANGUAGE MODELS FOR CLINICAL TERMINOLOGY TRANSLATION IN LOW-RESOURCE LANGUAGES: EVIDENCE FROM ENGLISH-TO-SINDHI NEURODEVELOPMENTAL DISORDER TERMS
Keywords:
large language models; technology assessment; Sindhi language; neurodevelopmental disorders; clinical terminology; low-resource languages; machine translation; healthcare technology managementAbstract
Diagnostic labels for neurodevelopmental disorders exist in English but have no agreed Sindhi equivalents, yet large language models are increasingly used to fill that gap without systematic evaluation of their handling of clinical vocabulary. This study evaluated four widely available systems (ChatGPT, Gemini, Microsoft Copilot and DeepSeek) on eighteen DSM-5-TR neurodevelopmental disorder terms. Model output was compared with a benchmark constructed by three independent Sindhi-language experts, who translated each term separately before reconciling discrepancies through consensus. Eight metrics were computed: Exact Match, Normalized Levenshtein Similarity, chrF2, BLEU, Translation Edit Rate, Jaccard Similarity, Cosine Similarity and ROUGE-L. Differences across systems were assessed using Friedman's test with Holm-Bonferroni correction. Leadership was distributed rather than concentrated: Copilot led on chrF2, Jaccard Similarity and Cosine Similarity; ChatGPT on BLEU, Translation Edit Rate and ROUGE-L; Gemini on Normalized Levenshtein Similarity; DeepSeek on none. No system reproduced an expert translation exactly for any term. Of the seven testable metrics, only chrF2 differed significantly after correction (χ² = 14.33, df = 3, p = .002, adjusted p = .017, Kendall's W = 0.265); post-hoc comparisons resolved no specific pair, and the composite ranking spanned only 0.35 of a rank, from ChatGPT (2.34) to DeepSeek (2.69). For healthcare organizations considering LLM-based translation, these systems can generate a usable first draft of specialized terminology in an under-resourced language, but none currently produces output that a clinical service could deploy without bilingual expert review.


