Neura News

AI News

News reporting focused on AI and machine learning, covering the companies behind these technologies, their real-world applications, and the ethical concerns they raise. This includes areas like generative AI (large language models, text-to-image and video), speech tech, and predictive analytics.

Latest News

16 articles
Funding

Microsoft and Mistral in Multi-Billion-Dollar AI Infrastructure Deal

Microsoft and Mistral have expanded their strategic partnership with a multi-billion-dollar agreement to build AI infrastructure across Europe. Mistral will deploy thousands of Nvidia Vera Rubin GPUs, and Microsoft will use that capacity for its cloud and AI services. Mistral's Medium 3.5 and OCR 4 models are now available in Microsoft Foundry, with Medium 3.5 also in Copilot Studio. The deal targets regulated industries like finance, healthcare, and manufacturing, allowing companies to run Mistral models through Azure Local in the cloud, on-premises, or offline. Microsoft President Brad Smith emphasized Europe's need for powerful AI without sacrificing data control.

Jul 213 minNeura News
AI Tools

NVIDIA NeMo Automodel and Hugging Face Unite for Diffusion Model Fine-Tuning

NVIDIA and Hugging Face have collaborated to bring production-grade distributed diffusion training to any Diffusers-format model on the Hugging Face Hub through the NVIDIA NeMo Automodel open-source library. The integration supports models like FLUX.1-dev, Wan 2.1, and HunyuanVideo, enabling full fine-tuning and LoRA adaptation without checkpoint conversion or model rewrites. The workflow uses YAML-driven configuration, pre-encoded datasets, and scalable parallelism from one GPU to hundreds.

Jul 1710 minNeura News
Industry

Enterprises buy AI compute faster than they can measure costs

A new VentureBeat Pulse Research survey of 107 enterprises reveals a significant gap between AI infrastructure spending and the ability to track its economics. Most organizations run AI on hyperscalers and model APIs, but the next dollar is aimed at specialized compute they rarely use today. GPU utilization is at 50% or less for 83% of respondents, and fewer than half rigorously track compute costs. A majority plan to switch or add providers within the year.

Jul 166 minNeura News
AI Tools

Disaggregated Prefill and Decode for LLM Inference on SageMaker HyperPod

AWS has introduced a method to separate the prefill and decode phases of large language model inference on SageMaker HyperPod. By running each phase on dedicated GPU pools connected via Elastic Fabric Adapter with RDMA, the approach eliminates interference that occurs when the phases share a single GPU. This results in more consistent per-token latency and independent scaling for long-context, high-concurrency workloads.

Jul 106 minNeura News
AI Models

NVIDIA Confidential Computing Powers Apple Private Cloud Compute Expansion

NVIDIA announced that its GPUs with Confidential Computing are now used for confidential inference in Apple's Private Cloud Compute (PCC) as Apple expands PCC beyond its own data centers to Google Cloud. The integration, unveiled at Apple's WWDC 2026, uses NVIDIA Blackwell GPUs to support server-side inference for Apple Foundation Models, built by Apple and Google using technologies behind the Gemini family of models. The hardware-based security layer ensures that no one, including system builders, can access user data during processing.

Jun 93 minNeura News
Industry

Google and Nvidia eye Intel as backup chipmaker amid TSMC capacity crunch

Google has placed an order with Intel to manufacture over three million AI chips for 2028, while Nvidia is testing Intel's manufacturing process for its next-generation Feynman GPU architecture. The moves come as TSMC struggles to keep pace with surging demand for AI chips, giving Intel's foundry business a rare second chance after years of losses. Memory maker SK Hynix is also evaluating Intel's packaging technology.

Jun 84 minNeura News
AI Models

NVIDIA and LG Group Build AI Factory for Physical AI, Mobility

NVIDIA and LG Group are jointly building an AI factory to accelerate AI-driven businesses across robotics, autonomous driving, data center technologies, and GPU cloud services. The facility will provide LG with accelerated computing infrastructure for training, simulating, and deploying AI applications. The collaboration integrates NVIDIA's full-stack AI factory platform with LG's global leadership in consumer electronics, robotics, mobility components, smart spaces, and data center technologies.

Jun 85 minNeura News
AI Tools

AWS details observability strategy for SageMaker AI LLM inference

AWS published a technical guide on comprehensive observability for large language models deployed on Amazon SageMaker AI. The approach separates monitoring into infrastructure quantity and LLM quality dimensions, using CloudWatch and Amazon Managed Grafana to visualize GPU utilization, cost, and response quality signals. The solution includes alert thresholds for metrics such as safety scores and composite quality scores, evaluated by an LLM-as-judge setup.

May 304 minNeura News