Beyond IID: How General Are Tabular Foundation Models, Really?
Unknown
This paper critically evaluates tabular foundation models, showing they excel on tiny datasets but fail to generalize beyond IID assumptions.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper critically evaluates tabular foundation models, showing they excel on tiny datasets but fail to generalize beyond IID assumptions.
Mike A. Merrill, Alexander G Shaw, Nicholas Carlini, et al.
Terminal-Bench 2.0 is a hard benchmark of 89 real-world-inspired CLI tasks where frontier agents score below 65%, with error analysis and public dataset/harness.
Catalin Ionescu, Dragos Papava, Vlad Olaru, et al.
Introduces Human3.6M, a large-scale dataset of 3.6 million accurate 3D human poses with synchronized images, motion capture, and depth data, along with statistical models and evaluation baselines for 3D human sensing.
Zhouyuan Ma, Yutao Wu, Hanxun Huang, et al.
HarmProfile is a content-centric benchmark dataset of over 80,000 validated harmful artifacts from 23 frontier LLMs, defining model-level risk profiles and revealing that harmfulness and diversity grow with model capability.
Xu Cao, Houze Yang, Vipin Gunda, et al.
Introduces Promptable Gaze Target Estimation (PGE), a concept-driven paradigm using text or visual prompts for end-to-end gaze analysis, with a new dataset Gaze-Co and model GazeAnywhere.
Todor Ivanov, V. Penchev
This paper surveys AI benchmarks and datasets for evaluating large language models, providing a structured overview of existing resources and their applications.
Yin Dai, Yifan Gao, Fayu Liu
TransMed combines CNN and transformer architectures for multi-modal medical image classification, achieving significant accuracy improvements on parotid gland and knee injury datasets.
Ali Borji, Ming-Ming Cheng, Qibin Hou, et al.
A comprehensive survey of salient object detection covering 228 publications, including roots, core techniques, datasets, evaluation metrics, and future directions.
Oliver Wieder, Stefan M. Kohlbacher, Mélaine A. Kuenemann, et al.
This review structures the dynamic field of GNNs for molecular property prediction by classifying 80 GNNs used across 20+ properties and 48 datasets.
Bajaj, Payal, Campos, Daniel, Craswell, Nick, et al.
This paper presents a dataset for the Advertisement in Retrieval-Augmented Generation task at Touché 2025, using segments from MS MARCO V2.1 and queries from Webis Generated Native Ads 2024.
Yijia Xiao, Wanjia Zhao, Junkai Zhang, et al.
This paper provides the first comprehensive overview of Protein LLMs, covering architectures, training datasets, evaluation metrics, and applications, with a structured taxonomy from over 100 articles.
Unknown
Mitra introduces mixed synthetic priors to enhance tabular foundation models, improving in-context learning performance across diverse tabular datasets.