dataqbs

Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation

· Source: arXiv cs.AI

Vision‑text models can automatically “correct” anomalous words that appear in images, producing transcriptions that sound natural but do not preserve the original content, thereby compromising OCR fidelity. The authors note that as a student progresses through training, the guidance from a static teacher becomes less effective, both across different checkpoints and among answer groups with varying task rewards. Based on this observation, they introduce GAD‑RL, an online distillation method that dynamically adjusts the teacher’s supervision according to the student’s current performance and the local distribution of its predictions. The frozen teacher receives input from both the reference transcription and the prefixes generated by the student. GAD‑RL disables distillation for groups whose output achieves a task reward of at least 0.95 and gradually reduces distillation intensity as the group’s average reward rises. It also weights forward KL divergence by the probability the student assigns to the teacher’s top‑1 token, limiting auxiliary updates when the student shows low confidence in that token. In experiments with the Qwen3.5‑2B model, GAD‑RL attains a 59.92 % Micro Recall on CHAOS‑Bench, outperforming GRPO and GRPO+OPD by 8.45 and 4.43 percentage points respectively, and achieves an overall score of 91.18 on OmniDocBench v1.6. This improvement is significant because it boosts text‑extraction accuracy in visual documents, benefiting applications that rely on reliable OCR, such as file digitization and information accessibility.

Read the original article on arXiv cs.AI

This summary is an informational synthesis produced by dataqbs.com. All rights to the original content belong to its author and the cited media outlet. We act solely as curators of technology news and claim no authorship.

Read this in Español · Deutsch