Can LLMs Reason and Plan?
Subbarao Kambhampati
Argues that LLMs lack genuine reasoning and planning capabilities, despite their impressive language generation, and that their apparent success is due to memorization and pattern matching.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Subbarao Kambhampati
Argues that LLMs lack genuine reasoning and planning capabilities, despite their impressive language generation, and that their apparent success is due to memorization and pattern matching.
Yijia Shao, Humishka Zope, Yucheng Jiang, et al.
Introduces an auditing framework with the Human Agency Scale to assess worker desires vs. AI capabilities for automating or augmenting occupational tasks.
Florian Wiesner, Matthias Wessling, Stephen Baek
A single transformer model trained on diverse simulation data demonstrates foundation model capabilities for physics, generalizing across domains without retraining.
Jacob T. Shreve, Sadia A. Khanani, Tufia C. Haddad
Reviews AI applications in oncology, including FDA-approved computer vision, multi-cancer detection, and ethical challenges like bias and interpretability.
Unknown
GPT-4V is a multimodal model that integrates text and vision capabilities for analyzing image inputs.
Unknown
Llama 3.1 adds multimodal capabilities via compositional cross-attention adapters for vision and speech, achieving competitive performance with GPT-4V and strong video reasoning.
Unknown
BloombergGPT is a 50B parameter language model trained on a mix of general and financial data, achieving top performance on financial NLP tasks without sacrificing general capabilities.
Xiaohan Xu, Ming Li, Chongyang Tao, et al.
This survey systematically reviews knowledge distillation techniques for transferring capabilities from large proprietary LLMs to smaller models.
Unknown
A comprehensive survey of in-context learning in large language models, covering its mechanisms, capabilities, and open challenges.
Unknown
This paper unlocks referring and grounding capabilities in multimodal large language models to create a more flexible human-computer interface.
Unknown
SEED-Bench is a large-scale benchmark for evaluating multimodal large language models across hierarchical capabilities.
Unknown
This survey reviews reasoning capabilities in foundation models, covering concepts, methodologies, and future outlook.