A native C++ harness for local AI.

LlamaBoss runs GGUF models on your own GPU through llama.cpp, in a Windows app written in C++. No Electron, no Docker. Its agent shows its work in the chat, checks native file changes against shared folders, and asks you to review risky actions.

Beta 0.1.19 for Windows 10 and 11. About 97,000 lines of C++ on wxWidgets and llama.cpp, MIT licensed.

Try a theme

A desktop app, not a black box.

Everything the model does shows up in the chat as it happens: what it searched, what it read, what it wants to run. You stay in charge of anything that changes your machine.

Review risky actions

Each chat gets its own workspace folder. Inside it, and inside any Project you attach, the agent can read, search and write freely, and every step shows up in the chat.

With normal approval settings, deleting files, PowerShell beyond simple read-only commands, arbitrary Python code and package installs ask first. Native writes outside shared folders need a folder grant. Allow once, trust the chat for this app session, or deny. Existing scripts in supported folders can run without a separate card; review imported Skills before using them.

read, list, grepruns, shown in the chat
write, editruns in shared folders, asks anywhere else
delete, PowerShell, Python code, installsasks by default, shows the exact command
⚠ Approval Required · pending · PowerShell
> Remove-Item .\build\obj -Recurse
[ Allow Once ] [ Allow Always ] [ Deny ]

Big files don't flood the conversation

When a file or tool result is too large to paste into the prompt, LlamaBoss saves it to disk and hands the model a short card instead: size, line count, structure, and a preview.

The model then searches inside it and reads only the lines it needs. Small local models can work through files many times larger than their context window.

📄 Read · stored
[LARGE OUTPUT -> stored as variable]
Vars\read_manifest.txt
286 KB · 5,412 lines · JSON, 14 top-level keys
head: {"name": "llamaboss", "files": [ …
🔍 Grep · 1 match
> "timeout" Vars\read_manifest.txt
📄 Read Range · lines 4,990–5,010
The first-byte timeout is 900 seconds, set on line 5,001.

Many windows, one model in memory

Open as many chat windows as you like. They share a single local model service, so your GPU loads the model once instead of once per window.

Each window can also talk to a different remote model at the same time without getting in each other's way.

Update checker review
You: read the diff
Permit report macro
Qwen: done, 3 sheets
Release notes
You: shorter please
llama-server · Qwen3.8-27Bloaded once

Your GPU first, the cloud when you want it

LlamaBoss ships with llama.cpp for CPU and NVIDIA CUDA, so local chat needs no Ollama, Docker or Python. Download a curated model or point it at your own GGUF folder.

When you want a bigger model, add OpenRouter, OpenAI or any OpenAI-compatible server. Local and remote models sit in the same picker.

Local
● Qwen3.8-27B-Q4_K_M16.4 GB
  gemma-4-12B-it7.3 GB
OpenRouter
  glm-5.3remote
OpenAI
  gpt-5.6-lunaremote
  + Add model…

Projects, Workflows and Skills

A Project keeps sources and notes for a line of work and can be attached to any chat. Skills are reusable instructions you can import, export and share as a zip. Project Workflows keep repeatable plans close to the files they use.

Projects and Skills live in ordinary Windows folders you can open, edit and back up.

%USERPROFILE%\LlamaBoss\
├─ Projects\
│ └─ llamaboss\ notes, sources, workflows
├─ Skills\
│ ├─ webSearch\SKILL.md
│ └─ runPod\SKILL.md + scripts\
└─ Chats\ one workspace per chat

What touches the network, and when.

Local chat never leaves your PC. Downloads, providers and tools can use the network. Approved code and existing Skill scripts can also connect to services.

  • Chatting with a local modelnever
  • Remote modelswhen you pick one
  • Model downloadswhen you start one
  • Updateswhen you check or install an update
  • Fetching a web pagewhen the agent asks for one
  • Installing Python packagesafter you approve

Download LlamaBoss.

Run the installer, pick a model, and start chatting. No account or sign-up needed.

Prefer to build it yourself? The full C++ source is on GitHub.

Download the installer
  • VersionBeta 0.1.19
  • SystemWindows 10 or 11, x64
  • GPUNVIDIA CUDA or CPU only
  • LicenseMIT
Seeing a SmartScreen warning? The installer isn't code-signed yet. If you downloaded it from llamaboss.com, choose More info, then Run anyway.