ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
Embedding models designed for visual document retrieval. Trained on a large synthetic dataset using a DSE approach, improving retrieval quality, in cross-lingual scenarios and for visual-heavy documents, and support Matryoshka Representation Learning for reduced vector size with minimal performance impact.
Visual document retrieval—finding documents based on their visual content (e.g., scanned forms, infographics, slides)—remains a challenging problem. Traditional text-based embeddings often fail to capture layout, fonts, and graphical elements. vdr Embeddings directly addresses this gap by training a dedicated embedding model on synthetic data, making it a practical tool for real-world document retrieval systems.
The paper's focus on cross-lingual scenarios is particularly significant. Many documents contain mixed languages or non-English text, and existing models often degrade in such settings. By training on synthetic data that likely includes diverse languages, vdr Embeddings promises more robust performance globally.
The abstract reports improved retrieval quality in cross-lingual scenarios and for visual-heavy documents, but no concrete metrics (e.g., recall@k, mAP) or comparisons to baselines (e.g., CLIP, Dense Passage Retrieval) are provided. The Matryoshka Representation Learning claim of minimal performance impact with reduced vector size is also stated without quantitative evidence. This lack of empirical detail limits the ability to assess the model's practical advantages.
vdr Embeddings could lower the barrier for building visual document retrieval systems, especially in multilingual contexts. The Matryoshka feature is valuable for deployment on resource-constrained devices or large-scale databases. However, without rigorous benchmarking, the true impact remains uncertain. Future work should include comparisons on standard datasets like DocVQA or Visual Document Retrieval benchmarks.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba