Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint reports a systematic engineering study of how to combine synthetic and real historical word images to train optical character recognition models for digitizing Manchu documents from the Qing empire. Adding real training data raised accuracy from 87.4–87.92% (synthetic-only) to 95.09–96.28%, with ensemble voting reaching 98.27%; this represents a practical advance for a critically endangered language's archival digitization, but is not peer-reviewed and carries no clinical or biomedical implications.
Methodological comparison of four training regimes (synthetic-only, real-only, joint, sequential) across three vision-language models and one CRNN. Historical Manchu word images from digitized Qing empire archival records (1636–1912), including both manuscripts and prints.. Intervention: Joint synthetic-real training and sequential synthetic-to-real training regimes for optical character recognition models.. Compared with: Synthetic-only, real-only, and ensemble voting approaches..
Synthetic-only configurations reached 87.4% to 87.92% word accuracy on real Manchu manuscripts Joint synthetic-real training raised leading configurations to 95.09% to 96.28% word accuracy Synthetic supplementation improved all three vision-language models; marginal effect for CRNN was training-objective-sensitive
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a methodological machine-learning study on a specialized historical document recognition task with no clinical or direct human health relevance; it reports technical performance metrics on a low-resource language digitization problem without peer review.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Manchu, now critically endangered, was one of the principal languages of the Qing empire (1636-1912), and its extensive archival record is increasingly digitized but remains difficult to search and analyze at scale. Previous work showed that vision-language models (VLMs) trained only on synthetic Manchu word images can reach 87.4% word accuracy on real Qing manuscripts and prints, leaving a substantial synthetic-to-real gap. This study examines how synthetic and real historical training data should be combined for low-resource OCR. Using 60,000 synthetic and 20,306 real historical word images, we evaluate three pretrained VLMs and a compact convolutional recurrent neural network (CRNN) under four regimes: synthetic-only, real-only, joint synthetic-real, and sequential synthetic-to-real training, following a common checkpoint-selection and archival evaluation protocol. Introducing real training images raises the leading configurations to between 95.09% and 96.28% word accuracy, while no synthetic-only configuration exceeds 87.92%. Synthetic supplementation substantially improves all three VLMs, whereas its marginal effect for the CRNN is sensitive to the training objective. Joint and sequential training yield broadly similar archival accuracy under the tested practical pipelines. A compact CRNN also reaches the leading performance range once real images are available, showing that model scale alone does not determine recognition accuracy. Finally, complementary errors among strong recognizers allow voting to raise accuracy to 98.27% without additional training, while an eighteenth-century Manchu dictionary provides a principled rule for adjudicating disagreements.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.