AI Tools

Osaurus Runs Local and Cloud AI Models on Macs

Osaurus, an open-source Apple-only LLM server, lets Mac users switch between local and cloud AI models while keeping files and tools on their hardware. It evolved from a desktop AI companion called Dinoki and supports models like Llama, DeepSeek V4, and cloud providers such as OpenAI and Anthropic. The tool offers a secure sandbox, over 20 plugins, and has seen more than 112,000 downloads since launch.

Neura News

Neura News

Neura Market Editorial

May 15, 20264 min read

Originally reported by techcrunch.com

Osaurus Runs Local and Cloud AI Models on Macs

Osaurus Runs Local and Cloud AI Models on Macs

Startups now focus on software that works above basic AI models as those models turn common. Osaurus enters this area as an open-source server for Apple devices only. It helps users shift among various local AI models, running them either on their machine or in the cloud. All the while, files and tools stay on the user's own hardware.

Roots in a Desktop AI Idea

Osaurus grew from a concept for Dinoki, a desktop AI helper. Co-founder Terence Pae called it an "AI-powered Clippy." Dinoki users wondered why buy the app when they still paid for tokens. Those are the units AI firms charge to handle prompts and create replies.

This pushed Pae to consider local AI runs. Pae, a former software engineer at Tesla and Netflix, shared this in a TechCrunch call. He aimed to build an AI assistant that operates on a Mac. "You can do pretty much everything on your Mac locally, like browsing your files, accessing your browser, accessing your system configurations. I figured this would be a great way to position Osaurus as a personal AI for individuals."

Pae developed the tool openly as an open-source project. He added features and fixed issues step by step.

Flexible Connections and User Benefits

Osaurus now links to AI models hosted locally or cloud services like OpenAI and Anthropic. Users pick any model they want. They keep memory, files, and tools on their hardware.

AI models vary in strengths. Users gain by picking the best model for each task.

Osaurus acts as a "harness." That means a control layer tying models, tools, and workflows via one interface. Tools like OpenClaw or Hermes do similar work. Yet those target developers familiar with terminals. Some, like OpenClaw, raise security risks.

Osaurus offers a simple interface for everyday users. It runs in a hardware-isolated virtual sandbox for safety. This bounds the AI's reach and protects the computer and data.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Hardware Demands and Progress

Local AI runs demand heavy resources and depend on hardware. Systems need at least 64 GB of RAM for local models. For bigger ones like DeepSeek v4, Pae suggests 128 GB of RAM.

Pae expects these needs to drop over time. "I can see the potential of it, because the intelligence per wattage , which is like the metric for local AI, has been going up significantly. It's on its own curve of innovation. Last year, local AI could barely finish sentences, but today it can actually run tools, write code, access your browser, and order stuff from Amazon. It's just getting better and better," he said.

Osaurus handles MiniMax M2.5, Gemma 4, Qwen3.6, GPT-OSS, Llama, DeepSeek V4, and more. It works with Apple's on-device foundation models and Liquid AI's LFM family for devices. In the cloud, it ties to OpenAI, Anthropic, Gemini, xAI/Grok, Venice AI, OpenRouter, Ollama, and LM Studio.

As a full MCP server, it lets MCP-compatible clients use tools. It comes with over 20 native plugins for Mail, Calendar, Vision, macOS Use, XLSX, PPTX, Browser, Music, Git, Filesystem, Search, Fetch, and others. A recent update added voice features.

Downloads, Accelerator, and Next Steps

The project launched almost a year ago. Its site reports over 112,000 downloads.

Founders Terence Pae and Sam Yoo join the New York accelerator Alliance. They plan business uses, such as in legal or healthcare fields. Local LLMs there solve privacy issues.

Local AI power rises, which may cut needs for AI data centers. "We're seeing this explosive growth in the AI space where cloud AI providers have to scale up using data centers and infrastructure, but we feel like people haven't really seen the value of the local AI yet," Pae said. "Instead of relying on the cloud, they can actually deploy a Mac Studio on-prem, and it should use substantially less power. You still have the capabilities of the cloud, but you will not be dependent on a data center to be able to run that AI."

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read