MINDMIRROR: DIAGNOSING AND REPAIRING FAILURE MODES IN A HYBRID DENSE-EMBEDDING AND LLM PIPELINE FOR DOCUMENT COMPARISON AND ORIGINALITY ASSESSMENT

Authors

  • Sohail Khan Author
  • Muhammad Javaid Iqbal Author
  • Arfan Jaffar Author
  • Nouman Aziz Author
  • Saba Ramzan Author
  • Inam Ul Hassan Author

Keywords:

Cosine similarity, document comparison, information retrieval, large language models, paraphrase detection, plagiarism detection, semantic textual similarity.

Abstract

Combining dense retrieval with a large language model (LLM) makes the creation of such systems straightforward, but it remains difficult to believe that they are accurate. Claims of success on carefully curated benchmarks are reported in the literature, but problems that make-or-break performance in production are not measured. The latter is not the case in this paper, but this research study develops a deployed document-comparison system, MindMirror, consisting of a dense-embedding stage and an LLM stage for qualitative analysis of the comparison results, and systematically identifies and addresses its failure cases in a hand-annotated set of 20 documents over 5 topics, 30 pairs, and a 5-level ordinal similarity scale for diagnosis. Four findings are reported. First, the lexical fallback that is used when the embedding service fails is not only less accurate, but ordinally invalid, showing a minimum adjacent-tier gap of -0.24 points (Spearman 0.842 to 0.955) with the dense score, an issue we attribute to unweighted term frequency, and demonstrate that by suppressing the stop words a sublinear term frequency can recover monotonicity (6.63 points, Spearman 0.842 to 0.955), and by repairing the lexical score using isotonic calibration brings it within 9.44 points ME from the dense score. Interestingly, comparing the two documents directly instead of the two documents under evaluation results in a worse score (-1.63 points), which has direct consequences for small collection retrieval. Second, embedding dimensionality does not correlate to accuracy: in five models on the provider's free tier, the 384-dimension multilingual-light-v3.0 does the best separation of tiers (5.92 points) whereas the 1024-dimension model does the worst (1.40). Third, chunk aggregation is not relevant to single-chunk documents but is crucial for multi-chunk documents, where the widely used max-pooling aggregation collapses (Spearman 0.548) when the documents share boilerplate. Fourthly, an ensemble of eight LLMs achieves the best single model performance (Spearman 0.9938) with no prior knowledge of the best-performing model, and using the five independent prompts simultaneously reduces latency by 1.37 times (still short of the five-fold optimum provider does not parallelise across a single account). This research study also reports a provider-entitlement mismatch, where only 8 of 18 models advertised were invocable. We believe that reporting a score that resulted from a degraded path, without revealing the degradation, is not a quality trade-off, but a correctness error, and the methodology for creating the diagnostic system used here, the ordinal corpora with adjacent-gap analysis, can be extended to any similarity system that is formed from third-party inference services.

Downloads

Published

2026-09-08