Open Source

Unsloth launches AMD support for local LLM training with 70% less VRAM

Unsloth cuts VRAM usage by 70% for training and 80% for reinforcement learning on AMD GPUs

Deep Dive

Unsloth, the popular open-source framework for efficient LLM fine-tuning and inference, has officially added support for AMD GPUs. The update enables local inference, fine-tuning, reinforcement learning, and deployment on AMD hardware including Radeon RX 9000 and 7000 series, Instinct MI350 and MI300 accelerators, and Strix Halo / Ryzen AI Max systems. Users can expect up to 70% less VRAM usage during training and up to 80% less VRAM for reinforcement learning, thanks to optimized ROCm, Triton, bitsandbytes, PyTorch, and llama.cpp builds that install automatically. The tool also supports CPU-only inference on AMD processors.

The release includes a one-liner installer for Linux, WSL, and macOS, and a PowerShell script for Windows. Unsloth supports nearly all major model families, including Qwen, Gemma, DeepSeek, GLM, Kimi, MiniMax, and DiffusionGemma. Users can export trained models as GGUF, safetensors, or LoRA adapters, connect local models to coding agents like Claude Code and Codex, and track RAM/VRAM usage remotely. The framework also provides daily updated ROCm prebuilts for llama.cpp and a secure Cloudflare HTTPS tunneling feature called "LM Link" for remote access. This expansion significantly broadens the hardware options for local LLM development, previously dominated by NVIDIA GPUs.

Key Points
  • Supports Radeon RX 9000/7000, Instinct MI350/MI300, Strix Halo, and Ryzen AI Max GPUs
  • Up to 70% less VRAM for fine-tuning and 80% less for reinforcement learning
  • Automated installation with optimized ROCm, Triton, bitsandbytes, and llama.cpp builds

Why It Matters

AMD users now get NVIDIA-level efficiency for local LLM training, democratizing access to AI development hardware.

📬 Get the top 10 AI stories daily