
Ollama can now run super fast on Apple devices using the MLX framework. It improves time to first token and tokens per second to make local AI run faster. Apple Silicon uses a shared memory pool for CPU and GPU. With MLX, Ollama uses that memory without copying data back and forth, which means you get faster inference.
Ollama is now updated to run the fastest on Apple silicon, powered by MLX, Apple’s machine learning framework.
This change unlocks much faster performance to accelerate demanding work on macOS:
– Personal assistants like OpenClaw
– Coding agents like Claude Code, OpenCode,… pic.twitter.com/WImO0lyYnp— ollama (@ollama) March 31, 2026
This means you can now power OpenClaw with local models and get faster performance. It is also great for Claude Code, OpenCode, and Codex. With NVFP4 support, you get higher quality responses. There is also improved caching for better responsiveness.
[HT]

