SEP 10, 2026 · PREPRINT
Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study
arXiv
This is a methodological machine-learning study on a specialized historical document recognition task with no clinical or direct human health relevance; it reports technical performance metrics on a low-resource language digitization problem without peer review.
Reported
Synthetic-only baseline word accu…87.4%
Synthetic-only configurations upp…87.92%
Joint/sequential training word ac…95.09% to 96.28%