Qinwu Xu, Yifan Jiang, Haoyu Ren
A multimodal LLM training framework combining synthetic data, OCR-aware fine-tuning, and visual chain-of-thought prompting to improve OCR and multilingual text understanding.
Multimodal LLMs fail on OCR and multilingual text understanding in real-world images with small fonts, blur, occlusion, and complex typography.
1) Large-scale synthetic OCR-to-translation data generation, 2) OCR-aware supervised fine-tuning with LoRA, 3) Structured visual chain-of-thought prompting for reasoning under uncertain visual conditions.
Significantly improved visual-text grounding on multilingual receipts, menus, posters, signs, handwritten text, and document images. Qualitative comparisons with GPT-5-class and Gemini-family models show improved OCR grounding and reduced hallucination.