Running AI models locally on a Mac has always sounded appealing because it gives users more control, better privacy, and freedom from subscription limits, but it has also come with a clear hardware challenge since even smaller large language models can consume a lot of memory and push Mac hardware hard during everyday use.
Ollama is trying to fix that problem with its new preview release, Ollama 0.19, which now uses Apple’s open source MLX framework to improve local AI performance on Apple silicon Macs, and that change matters because MLX is built to take advantage of Apple’s unified memory design, where the CPU and GPU share the same memory pool more efficiently.
Don’t miss the best of The Mac Observer
Set us as a preferred source and our Apple reporting ranks higher in your Google Search results and Discover feed — one tap, no account changes.
In its announcement, Ollama said:
“This results in a large speedup of Ollama on all Apple Silicon devices. On Apple’s M5, M5 Pro and M5 Max chips, Ollama leverages the new GPU Neural Accelerators to accelerate both time to first token (TTFT) and generation speed (tokens per second).”
That “large speedup” is the key takeaway here, especially for users who want faster response times from local assistants and coding tools, because lower time to first token and stronger token generation speed make local models feel much more practical during real work.
Ollama also says the update helps with personal assistants such as OpenClaw and coding agents like Claude Code, OpenCode, and Codex, while adding better caching performance and support for Nvidia’s NVFP4 compression format, which improves memory efficiency in supported setups.
There is still a catch, and it is a big one for many Mac users, because Ollama recommends a Mac with more than 32GB of unified memory, and the current MLX preview only supports one model, the 35 billion parameter version of Alibaba’s Qwen3.5.
Even so, this update shows why local AI on Macs is getting more attention, since users want more private workflows, fewer cloud limits, and better value from the hardware they already own, and Ollama’s MLX move pushes that shift forward in a meaningful way.
Discussion