Developer Tools

Ollama v0.32.3 fixes downloads, adds GPU support, improves integrations

GPU support on Windows ARM64 and lower memory usage for integrated GPUs are now live.

Deep Dive

Ollama's latest release, v0.32.3, addresses a critical bug where model downloads would stall before sending data, ensuring smoother setup for users. Integration improvements include restoring Claude Code channels, fixing Anthropic thinking streams, and making the Hermes Desktop client respect the --force-build flag. These changes streamline workflows for developers using Ollama with popular AI tools.

GPU support has been expanded significantly: CUDA now works on Windows ARM64 devices, B200 support is available via CUDA 12, and memory usage is reduced on Linux iGPUs using both CUDA and ROCm. The update also adds chat, thinking, and tool calling capabilities for Laguna 2.1 models, along with a Metal inference fix for better performance on Apple Silicon. Additionally, GLM tool calls that were silently dropped at the end of generation are now fixed. The underlying MLX and llama.cpp engines have been updated to their latest versions.

Key Points
  • Fixed model downloads stalling before data transfer begins.
  • Expanded GPU support: CUDA on Windows ARM64, B200 via CUDA 12, lower memory on Linux iGPUs.
  • Added chat, thinking, and tool calling for Laguna 2.1 models with a Metal inference fix.

Why It Matters

Ollama's update improves reliability and hardware compatibility, making local LLM deployment smoother for more users.

📬 Get the top 10 AI stories daily