LangChain and NVIDIA today announced the NemoClaw for LangChain Deep Agents blueprint, an open reference architecture that combines LangChain’s agent harness, NVIDIA’s Nemotron 3 Ultra model, and NVIDIA’s OpenShell runtime. The goal is to help enterprises build, tune, and govern agent systems without locking themselves into a closed ecosystem.
Announced on July 8, 2026, the blueprint gives teams control over the full agent stack: an open model layer, a tuned agent harness, and a governed runtime. It is available immediately.
The Blueprint: Three Layers for Open Agent Systems
The NemoClaw for LangChain Deep Agents blueprint brings together three components. LangChain Deep Agents Code (dcode) provides the harness layer for long-running agents, handling planning, tool use, memory, and task execution. NVIDIA Nemotron 3 Ultra serves as the open model layer, offering advanced agent performance at lower inference cost. NVIDIA OpenShell provides a secure runtime with policies for tool, system, and data interaction.
The blueprint includes a Deep Agents harness profile that is specifically tuned for Nemotron 3 Ultra. This matters because agent performance improves when the model, harness, evals, and runtime are tuned together, rather than treated as separate pieces.
LangChain argues that closed ecosystems limit enterprise control over valuable intellectual property. Agent memory, workflows, traces, eval datasets, harness configuration, and tuning data are all proprietary IP that enterprises generate as they move agents to production. The open blueprint is designed to give enterprises ownership over that IP.
Cost Advantage: A 10x Reduction in Inference Cost
One of the most striking figures in the announcement is the inference cost comparison. In LangChain’s agent eval suite, Nemotron 3 Ultra with LangChain Deep Agents achieved an aggregate eval score of 0.86 at an inference cost of $4.48. The next closest performing model cost $43.48 for the same benchmark — roughly 10x higher.
The article states that lower inference cost is one of the biggest benefits of using open models. It also argues that lower inference cost changes how teams build and improve agents. When each iteration is expensive, teams run fewer evals, compare fewer variants, and avoid specialized agents. A more cost-efficient open stack makes it practical to run larger eval suites, compare more variants, and evaluate specialized agents.
Governance and Security for Regulated Industries
EY, the professional services firm, is building an implementation practice around the software stack. Geoff Vickrey, Global Chief Commercial Officer, NVIDIA, EY, said: " – Geoff Vickrey, Global Chief Commercial Officer, NVIDIA, EY"
Vickrey noted that EY clients in regulated industries are ready to move agentic AI into production but are constrained by governance, security, and control. Open agent architectures give enterprises transparency, control, and the freedom to deploy without committing to a closed stack. EY teams help give clients a secure, sandboxed foundation for always-on agents that meet enterprise standards.
Ecosystem Partners Supporting Production Deployment
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
The blueprint is supported by a range of infrastructure partners that help serve Nemotron models in production. Baseten, Fireworks, Nebius, Crusoe, DeepInfra, and Together AI are all ecosystem partners.
Philip Kiely, Head of Developer Relations at Baseten, said production agents need inference that is fast, reliable, and cost-efficient at scale. Baseten optimized Nemotron models for high throughput and low latency on NVIDIA hardware.
Lin Qiao, CEO and Cofounder of Fireworks AI, noted that agentic workloads make many model calls per task, so inference speed and cost determine viability. Fireworks serves Nemotron models with throughput and price-performance for high-volume agent systems.
Roman Chernin, Chief Business Officer at Nebius, said the next challenge for enterprise AI is running complex agentic workloads economically at production scale. Nebius gives dedicated infrastructure optimized for high-performance inference and cost-efficient scaling.
Tuning the Harness, Not Just the Model
The announcement includes two additional technical articles: "Deep Agents Code on NemoClaw: a governed blueprint for your most sensitive code" and "Tuning the harness, not the model: a Nemotron 3 Ultra playbook." These articles go deeper into how teams can tune the agent harness for their own workloads.
LangChain’s agent engineering platform, LangSmith, provides debugging, evaluation, and deployment capabilities. The company argues that evals work best when run throughout the agent development lifecycle — both before deployment and after deployment for monitoring.
For production agents, model choice is only one part of improving performance. Teams also need to control tools, context, evaluation, runtime, and policies. The blueprint is designed to give enterprises that control, allowing them to decide when, how, and why the system changes.
The Bottom Line for Enterprise Agents
The NemoClaw for LangChain Deep Agents blueprint represents a bet that open, governed agent systems will win over closed ecosystems in the enterprise. By combining an open model, a tuned harness, and a secure runtime, the partners aim to give enterprises the transparency and control they need to put agentic AI into production at scale.
With a 10x cost advantage on a key benchmark and a growing ecosystem of infrastructure partners, the blueprint offers a concrete path for enterprises that want to own their agent intelligence rather than rent it from a closed platform.

