ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.
OmniScientist addresses a critical gap in AI-driven scientific discovery: existing systems typically reason over text, code, labels, or precomputed summaries, missing the rich spatial, temporal, cross-channel, and procedural information inherent in raw scientific data. By enabling an AI scientist to perceive and reason over heterogeneous raw evidence—such as images, signals, audio, video, 3D structures, trajectories, tables, formulae, and graphs—OmniScientist moves beyond workflow automation to true evidence-grounded discovery. This is a significant step toward AI systems that can autonomously conduct research across disciplines, not just within narrow domains.
The paper's emphasis on lifecycle-wide perception—where observations shape research questions, experimental decisions, and final claims—is a paradigm shift. It contrasts with prior work that treats perception as a preprocessing step, instead integrating it throughout the research pipeline. This holistic approach is likely to inspire future AI scientist frameworks to prioritize raw data access and multimodal reasoning.
OmniScientist achieved a mean overall paper score of 6.3 with the reference reasoning backbone, and completed all 36 cases from raw data to compiled manuscript. In paired comparisons against a blind variant that only receives precomputed scalar features, direct perception improved all 7 evaluation dimensions and won 85% of head-to-head judgments. These results strongly support the hypothesis that perception is essential for evidence-grounded scientific discovery.
The 85% win rate is particularly compelling, as it quantifies the benefit of raw data access over simplified feature representations. The improvement across all evaluation dimensions—likely including novelty, rigour, and clarity—indicates that perception enhances not just data interpretation but the entire research process.
OmniScientist sets a new benchmark for AI scientists by demonstrating that omni-modal perception is not just a nice-to-have but a necessity for evidence-grounded discovery. This work could accelerate research in fields where data is inherently multimodal, such as biomedical imaging, climate science, and materials science. By automating the full research pipeline with rigorous checks, it also has the potential to improve reproducibility and reduce human bias in scientific workflows.
The broader impact extends to the AI community: it challenges the assumption that precomputed features are sufficient for complex reasoning tasks and encourages the development of models that can directly consume raw data. As foundation models continue to advance, OmniScientist provides a blueprint for building AI scientists that can autonomously explore new frontiers, potentially leading to discoveries that would be difficult for humans alone. However, the reliance on a reference reasoning backbone and the limited number of cases suggest that further scaling and validation are needed before widespread adoption.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba