
Description
Running an open model on your own machine means picking between GGUF and MLX, choosing a quant and setting up llama.cpp. Fine-tuning one means a pile of Python and watching VRAM run out. Usually those are two completely separate toolchains.
Unsloth puts both in one desktop app. Find Qwen, DeepSeek, Gemma or Kimi in the built-in model hub and start chatting with one click; then, in the same window, fine-tune on your own data 2× faster with 70% less VRAM and no accuracy loss. There are native installers for Windows, macOS and Linux, and it runs on NVIDIA, AMD and Intel GPUs or plain CPU.
It started as one of the most popular fine-tuning libraries on GitHub (over 76k stars) and now has a GUI, so you don't need to write code. Command-line fans can still run the Unsloth Studio web UI or call it from Python. It's Apache-2.0 licensed.
Model hub: browse Discover and On Device tabs showing each model's size, quant and download count, and run GGUF, MLX, diffusion, embedding and audio models.
Chat with search and deep research: local models can search the web, run deep research, read files and execute code, and long chats auto-compact so the context window doesn't overflow.
Local models for Claude Code:
No-code fine-tuning: LoRA, QLoRA, full fine-tuning, pretraining and RL methods like GRPO and DPO, for LLMs, diffusion, TTS and embedding models.
Datasets from your documents: drop in PDFs, CSVs or Word files and Data Recipes turns them into training data, skipping the manual cleanup.
Export when done: save to GGUF, FP8, NVFP4 and more, ready for Ollama or llama.cpp.
Images and video locally: run and train image and video diffusion and multimodal models on your own machine.
Serve it as an API: an OpenAI-compatible endpoint that other devices on your LAN can reach, plus secure remote access over Cloudflare HTTPS.
Unsloth puts both in one desktop app. Find Qwen, DeepSeek, Gemma or Kimi in the built-in model hub and start chatting with one click; then, in the same window, fine-tune on your own data 2× faster with 70% less VRAM and no accuracy loss. There are native installers for Windows, macOS and Linux, and it runs on NVIDIA, AMD and Intel GPUs or plain CPU.
It started as one of the most popular fine-tuning libraries on GitHub (over 76k stars) and now has a GUI, so you don't need to write code. Command-line fans can still run the Unsloth Studio web UI or call it from Python. It's Apache-2.0 licensed.
Features
Model hub: browse Discover and On Device tabs showing each model's size, quant and download count, and run GGUF, MLX, diffusion, embedding and audio models.
Chat with search and deep research: local models can search the web, run deep research, read files and execute code, and long chats auto-compact so the context window doesn't overflow.
Local models for Claude Code:
unsloth start claude points Claude Code at a local model in one command, and Codex, OpenCode and other agents work the same way, with tool calling and MCP.No-code fine-tuning: LoRA, QLoRA, full fine-tuning, pretraining and RL methods like GRPO and DPO, for LLMs, diffusion, TTS and embedding models.
Datasets from your documents: drop in PDFs, CSVs or Word files and Data Recipes turns them into training data, skipping the manual cleanup.
Export when done: save to GGUF, FP8, NVFP4 and more, ready for Ollama or llama.cpp.
Images and video locally: run and train image and video diffusion and multimodal models on your own machine.
Serve it as an API: an OpenAI-compatible endpoint that other devices on your LAN can reach, plus secure remote access over Cloudflare HTTPS.

