Neura News

AI News

News reporting focused on AI and machine learning, covering the companies behind these technologies, their real-world applications, and the ethical concerns they raise. This includes areas like generative AI (large language models, text-to-image and video), speech tech, and predictive analytics.

Latest News

28 articles
AI Models

Anthropic's Fable 5 Hits a Corporate Spending Ceiling, Ramp Data Shows

New spending data from Ramp shows Anthropic's Fable 5, the most capable AI model on the market, captured only 6% of tokens and 11.4% of spending in its first month. The high price tag of $10 per million input tokens and $50 per million output tokens has deterred corporate buyers, signaling a ceiling on willingness to pay for frontier AI. Meanwhile, Anthropic leads OpenAI in U.S. adoption, but growth is slowing as open-source models close the gap.

Aug 135 minNeura News
AI Models

SpaceXAI launches Grok 4.6 with claims of top-tier reasoning

SpaceXAI has released Grok 4.6, a large language model that the company claims can outperform Anthropic's Claude Fable 5 in some areas. The model scored 61 on the Artificial Analysis Intelligence Index, placing it on par with OpenAI's GPT-5.6 Sol and one point behind the leader. Priced at $2 per million input tokens and $6 per million output tokens, Grok 4.6 is available through Cursor and Grok Build, with a faster edition costing double.

Aug 135 minNeura News
Developer

LangChain's Deep Agents v0.7 Cuts Token Use by 65% With Leaner Prompts

LangChain released Deep Agents v0.7 on July 29, 2026, an open-source agent harness update that cuts base input tokens by 65% on a default-agent turn, from roughly 6,000 to about 2,000 tokens. The release trims the built-in prompt, shortens tool descriptions, and makes the todo list middleware optional. It also adds middleware configurability and filesystem improvements that users have requested for months.

Aug 36 minNeura News
Industry

OpenAI Slashes GPT-5.6 Luna Price by 80% as AI Market Shifts to Infrastructure Economics

OpenAI has slashed the price of its GPT-5.6 Luna API model by 80%, reducing input costs from $1 to $0.20 per million tokens and output costs from $6 to $1.20 per million. The move responds to competition from open-weight models and signals a shift toward infrastructure-like economics in the AI market. The company also cut GPT-5.6 Terra by 20%, while the flagship Sol model remains unchanged.

Aug 111 minNeura News
AI Models

Petals Lets Users Run Large AI Models at Home Like BitTorrent

Petals is a decentralized platform that allows users to run large language models such as Llama 3.1, Mixtral, Falcon, and BLOOM on consumer-grade hardware by sharing computational resources in a peer-to-peer network. Users load only a portion of a model and join a network of others serving the remaining parts, enabling inference speeds of up to 6 tokens per second for Llama 2 70B and 4 tokens per second for Falcon 180B. The platform supports fine-tuning, custom sampling methods, and access to hidden states, combining the convenience of an API with the flexibility of PyTorch and Hugging Face Transformers.

Jul 232 minNeura News
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 215 minNeura News
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 214 minNeura News
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 216 minNeura News