Native and compact structured latents for 3d generation
Unknown
This paper introduces native and compact structured latents for efficient high-resolution 3D generation, including an image-to-3D model with about 4 billion parameters.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper introduces native and compact structured latents for efficient high-resolution 3D generation, including an image-to-3D model with about 4 billion parameters.
Unknown
Latent Consistency Models enable few-step inference for high-resolution image synthesis by distilling consistency into pre-trained latent diffusion models.
Paul Bergmann, Kilian Batzner, Michael Fauser, et al.
Introduces MVTec AD, a comprehensive dataset with 5354 high-resolution images across 15 categories for unsupervised anomaly detection, and benchmarks state-of-the-art methods.
Unknown
Eagle 2.5 introduces a generalist vision-language model family with Automatic Degrade Sampling and Image Area Preservation for long-context video and high-resolution image understanding.
Unknown
NVLM introduces a family of multimodal LLMs with a hybrid architecture and 1-D tile-tagging for dynamic high-resolution images, improving text-only and multimodal performance.
Unknown
LLaVA 1.6 enhances multimodal AI with dynamic high-resolution processing, improved OCR, and scalable LLM integration.