[Submitted on 30 Jun 2026 (v1), last revised 23 Jul 2026 (this version, v4)]
Abstract:Maltese OCR is constrained by the absence of a public, reusable paragraph-scale training corpus. We address this by generating synthetic Maltese line images, fine-tuning the Tesseract 5 LSTM, and combining five deterministic Tesseract configurations through anchor-preserving, lexicon-gated word-level arbitration. The method uses a fixed anchor stream, a longest-stream fallback, a confusion-based anchor corrector, and a Maltese-specific diacritic-restoration gate. Unlike canonical ROVER, candidate streams cannot restructure the anchor through insertions or deletions; they propose only eligible substitutions at aligned anchor positions.
On the 422-paragraph development set of the DocEng 2026 Maltese OCR competition, the organizers' fine-tuned Tesseract baseline obtains CER 0.0234. Our pre-convention pipeline reaches CER 0.01317, a 44% reduction. Synthetic fine-tuning provides the largest single gain, while multi-stream arbitration contributes a further material reduction beyond the selected anchor, reaching CER 0.01220 in the current replay with paired-resampling support. A development-tuned label-convention normalization chain further reduces CER to 0.00700. We report recognition gains separately from benchmark-specific quote and dash normalization.
We also evaluate portability on Hungarian and Luxembourgish. Luxembourgish improves significantly over our stock baseline, while the Hungarian result is inconclusive. Finally, we release a 36,803-pair Maltese OCR corpus derived from EUR-Lex and Wikipedia. The held-out competition result remains under organizer embargo and is not reported
Submission history
From: Adam Darmanin [view email]
[v1]
Tue, 30 Jun 2026 22:58:41 UTC (87 KB)
[v2]
Thu, 2 Jul 2026 11:05:21 UTC (141 KB)
[v3]
Sun, 19 Jul 2026 22:38:38 UTC (1,263 KB)
[v4]
Thu, 23 Jul 2026 08:28:18 UTC (1,263 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.