ERNIE Layout
Unknown
ERNIE-Layout enhances document understanding by reorganizing tokens using layout knowledge and applying spatial-aware disentangled attention in a multi-modal transformer.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
ERNIE-Layout enhances document understanding by reorganizing tokens using layout knowledge and applying spatial-aware disentangled attention in a multi-modal transformer.
Yupan Huang, Tengchao Lv, Lei Cui, et al.
LayoutLMv3 introduces a unified text-image multimodal Transformer with masked language and image modeling for document AI.
Unknown
Introduces Bi-directional attention complementation mechanism (BiACM) for cross-modal interaction of text and layout in document AI.
Unknown
LayoutLMv2 integrates text, layout, and image in a single multi-modal Transformer pre-training framework with new cross-modal tasks.
Unknown
LAMBERT enhances RoBERTa with layout embeddings and relative bias for layout-aware document understanding without raw images.
Xinkun Huang, Jinyan Sang, Changrong Jin, et al.
Extends BERT with 2-D position and image embeddings for layout-aware document understanding.
Unknown
DocLLM extends LLMs to reason over visual documents using textual semantics and spatial layout, avoiding expensive image encoders.
Unknown
LMDX adapts arbitrary LLMs for document information extraction with layout encoding and grounding to prevent hallucinations.
Chuwei Luo, Changxu Cheng, Qi Zheng, et al.
GeoLayoutLM explicitly models geometric relations in pre-training to enhance feature representation for document layout analysis.
Erik Blasch
UDoP integrates text, image, and layout information through a Vision-Text-Layout Transformer for unified multimodal representation.