Stronger Normalization-Free Transformers
Mingzhi Chen, Taiming Lu, Jiacheng Zhu, et al.
Introduces Derf, a normalization-free pointwise function that outperforms LayerNorm, RMSNorm, and DyT across vision, speech, and DNA tasks.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Mingzhi Chen, Taiming Lu, Jiacheng Zhu, et al.
Introduces Derf, a normalization-free pointwise function that outperforms LayerNorm, RMSNorm, and DyT across vision, speech, and DNA tasks.
Unknown
Dynamic Tanh (DyT) replaces normalization layers in Transformers with a simple element-wise operation, achieving comparable or better performance across vision, language, and speech tasks.
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, et al.
ConvNeXt V2 enhances pure convolutional networks with a fully convolutional masked autoencoder and Global Response Normalization for superior self-supervised learning.
Unknown
SigLip introduces a pairwise sigmoid loss for language-image pre-training, enabling larger batch sizes and better performance at smaller batch sizes without global normalization.