LiLT
Unknown
Introduces Bi-directional attention complementation mechanism (BiACM) for cross-modal interaction of text and layout in document AI.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Introduces Bi-directional attention complementation mechanism (BiACM) for cross-modal interaction of text and layout in document AI.
Unknown
LayoutLMv2 integrates text, layout, and image in a single multi-modal Transformer pre-training framework with new cross-modal tasks.
Unknown
mmE5 is a multimodal multilingual embedding model trained on synthetic data via a framework that ensures broad task/language coverage, robust cross-modal alignment, and high fidelity through self-evaluation and refinement.
Saurabh Dash, Yiyang Nan, John Dang, et al.
Aya Vision introduces open-weight multilingual VLMs for 23 languages using synthetic annotation and cross-modal model merging to retain text skills.