Beyond IID: How General Are Tabular Foundation Models, Really?
Unknown
This paper critically evaluates tabular foundation models, showing they excel on tiny datasets but fail to generalize beyond IID assumptions.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper critically evaluates tabular foundation models, showing they excel on tiny datasets but fail to generalize beyond IID assumptions.
Catalin Ionescu, Dragos Papava, Vlad Olaru, et al.
Introduces Human3.6M, a large-scale dataset of 3.6 million accurate 3D human poses with synchronized images, motion capture, and depth data, along with statistical models and evaluation baselines for 3D human sensing.
Todor Ivanov, V. Penchev
This paper surveys AI benchmarks and datasets for evaluating large language models, providing a structured overview of existing resources and their applications.
Yin Dai, Yifan Gao, Fayu Liu
TransMed combines CNN and transformer architectures for multi-modal medical image classification, achieving significant accuracy improvements on parotid gland and knee injury datasets.
Ali Borji, Ming-Ming Cheng, Qibin Hou, et al.
A comprehensive survey of salient object detection covering 228 publications, including roots, core techniques, datasets, evaluation metrics, and future directions.
Oliver Wieder, Stefan M. Kohlbacher, Mélaine A. Kuenemann, et al.
This review structures the dynamic field of GNNs for molecular property prediction by classifying 80 GNNs used across 20+ properties and 48 datasets.
Yijia Xiao, Wanjia Zhao, Junkai Zhang, et al.
This paper provides the first comprehensive overview of Protein LLMs, covering architectures, training datasets, evaluation metrics, and applications, with a structured taxonomy from over 100 articles.
Unknown
Mitra introduces mixed synthetic priors to enhance tabular foundation models, improving in-context learning performance across diverse tabular datasets.
Unknown
This paper introduces in-context fine-tuning for time-series foundation models, enabling adaptation to new datasets without updating model weights.
Unknown
A comprehensive review of protein language models, datasets, and tools with a curated GitHub repository of resources.
Unknown
ProGen2 introduces a suite of protein language models scaled to 6.4B parameters, trained on diverse protein sequence datasets to explore scaling boundaries.
Unknown
The paper introduces a Common Task Framework (CTF) for scientific machine learning, featuring curated datasets and task-specific metrics to enable critical evaluation of algorithms.