SigLIP 2
Unknown
SigLIP 2 improves multilingual vision-language encoders with captioning pretraining, self-supervised losses, and online data curation, offering native aspect ratio preservation.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
SigLIP 2 improves multilingual vision-language encoders with captioning pretraining, self-supervised losses, and online data curation, offering native aspect ratio preservation.
Unknown
SimCLRv2 presents a semi-supervised learning framework combining unsupervised pretraining, supervised fine-tuning, and distillation with unlabeled data.
Unknown
Introduces a training method for developing ultra-long context LLMs with context windows extending up to 4 million tokens via efficient continued pretraining with YaRN-based scaling and instruction tuning.
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, et al.
Llemma is an open-source LLM for mathematics, built by continued pretraining of Code Llama on a curated dataset of scientific papers, web math, and code, enabling tool use and formal theorem proving.
Unknown
LLaMA 2 Long extends LLaMA 2 to handle up to 32,768 tokens via continual pretraining with modified RoPE and synthetic instruction data.
Zhilin Yang, Zihang Dai, Yiming Yang, et al.
XLNet combines autoregressive and autoencoding pretraining via permutation language modeling, outperforming BERT on NLP tasks.
Unknown
Proves that in-context learning is guaranteed to emerge from multi-task pretraining under specific theoretical conditions.
Laila Rasmy, Yang Xiang, Ziqian Xie, et al.
Med-BERT adapts BERT to structured EHR data via pretraining on 28.5M patients, boosting disease prediction AUC by 1.21-6.14% and enabling small-dataset models to match those trained on tenfold larger data.