Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
-
Updated
Aug 15, 2026 - Python
Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
Zero-friction LLM fine-tuning skill for Claude Code, Gemini CLI & any ACP agent. Unsloth on NVIDIA · TRL+MPS/MLX on Apple Silicon. Automates env setup, LoRA training (SFT, DPO, GRPO, vision), post-hoc GRPO log diagnostics, evaluation, and export end-to-end. Part of the Gaslamp AI platform.
A framework for agentic tool use training with reinforcement learning
Upload your data → Get a fine-tuned SLM. Free.
LoRA fine-tune Gemma 4 31B to speak caveman-mode natively. Style: github.com/JuliusBrussee/caveman
Build your own offline AI from any documents. Free. No coding. LoRA fine-tuning + RAG + GGUF export.
An implementation of GRPO for Unsloth's VLMs training
从对话数据到训练:数字分身 + 模型蒸馏 From Dialogue Data to Training Closed-Loop: Digital Twin + Model Distillation
本项目利用医学领域的 CoT 数据对 Deepseek-R1-Distill-Qwen-7B 进行微调,通过 QLoRA 量化和 Unsloth 加速训练,显著提升模型在复杂医学推理任务中的慢思考能力。知识蒸馏技术使轻量级模型获得大模型的推理优势,实现高效、准确且具有解释性的医学问答系统。
Natural language control for Python CLI tools using locally-trained SLMs (CPU inference)
Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF export · vLLM inference · BLEU/ROUGE/BERTScore · full CLI · Heretic Mode to unlock full model potential
Local-first desktop assistant for the whole job hunt — scrape listings, match with a local LLM, generate applications in your own voice, and track everything on a Kanban board.
LLM finetuning for Sudoku solving
Finetune Web UI is a user-interface for training and deploying pre-trained models.
Adversarial LLM arena using GRPO + SFT to train a model that generates questions hard LLMs can't solve
7.67× LoRA / 8.35× Full FT speedup for Qwen3.5 (0.8B–27B) on NVIDIA DGX Spark — wall-clock parity with rented H100. Lossless within BF16. Three-command interactive wizard handles model picker, data validator, training, and merge.
Add a description, image, and links to the unsloth topic page so that developers can more easily learn about it.
To associate your repository with the unsloth topic, visit your repo's landing page and select "manage topics."