Visual document retrieval (VDR) has been dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Existing compression methods either train a smaller multi-vector encoder from scratch or distill only the query side, failing to produce a compact end-to-end single-vector retriever.
In a new paper, Hugging Face researchers present DistilVDR, a 524M-parameter end-to-end VDR system distilled bilaterally from a single 8B vision-language teacher using a pointwise cosine alignment loss. All supervision comes from the frozen teacher's embedding space, which was itself trained with relevance supervision, so the student objective requires no relevance labels, negative sampling, or contrastive terms.
The model mirrors VDR's text-query and image-document input asymmetry with an asymmetric encoder-only student that concentrates visual capacity on the document side (457M parameters) while keeping the query side at 70M parameters. Two variants share the same encoders and training but differ in the document encoder's visual-tile budget: DistilVDR-HiRes attains 61.74 average NDCG@5 on ViDoRe v1+v2+v3 (86.9% of the 8B teacher) and leads every reproduced sub-1B baseline on the high-resolution-sensitive v3 benchmark, while DistilVDR-Fast attains 59.98 at a 3x smaller visual-token budget.
Both variants store one million documents in a 15.6x smaller index than the strongest sub-1B multi-vector baseline and index the corpus an order of magnitude faster. The code is available on GitHub, and training details and models are released on the NanoVDR Hugging Face space.