olmOCR
Unknown
olmOCR is an open-source Python toolkit that converts PDFs into linearized plain text while preserving structured content using a fine-tuned 7B VLM.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
olmOCR is an open-source Python toolkit that converts PDFs into linearized plain text while preserving structured content using a fine-tuned 7B VLM.
Unknown
Nougat uses a Visual Transformer to convert scientific document images into markup language via OCR.
Unknown
Donut is an OCR-free encoder-decoder Transformer that directly processes images and prompts to generate text, eliminating the need for optical character recognition.
Unknown
H2OVL-Mississippi introduces small, efficient vision-language models optimized for OCR and document analysis while maintaining strong general VLM performance.
Unknown
Idefics2 improves upon Idefics1 with enhanced OCR, simplified architecture, and better pre-trained backbones using open datasets and task-oriented fine-tuning.
Unknown
LLaVA 1.6 enhances multimodal AI with dynamic high-resolution processing, improved OCR, and scalable LLM integration.
Unknown
BLOOM is a 176B-parameter open-access decoder-only transformer collaboratively developed to democratize large language model technology.
Unknown
Neural operators use deep learning to accelerate scientific simulations and democratize access to computational science.
Pouria Mahdi, Haq Nawaz Malik
Persian Pixel is a large-scale synthetic OCR dataset with over 343,000 image-text pairs for Persian, generated via a rendering pipeline that models script complexities and applies stochastic degradations to bridge the synthetic-to-real gap.