projects
Local LLM Experiments

Local LLM Experiments

completed· 2026

How far agentic coding goes with local models on a 32 GB MacBook Pro M5, measured against cloud models.

Why

The scope of validity and what I did not measure are written before the results.

I wanted to know whether a local model can really act as a coding agent, not in a demo but in a harness that writes a whole app. The task is always the same: a CRUD app with a FastAPI and SQLite backend and a React frontend, generated from scratch and scored against a fixed rubric.

  • On a small scaffold, Qwen3.6-35B-A3B quantized to 4 bits with MLX scored 18 out of 18 in 3.5 to 7 minutes, with no human fixes. Replicated twice.
  • On the same task Opus 4.7 scored 17 out of 18 at 0.76 dollars per run, DeepSeek V4 Flash 17 out of 18 at 0.002: cost moves by four orders of magnitude, the result stays within the rubric’s noise.
  • Local tool calling needs three things aligned: the model’s training, the chat template and the server’s parser. If one link is broken, the model fails out of the box.
  • Splitting the work into a roadmap rescues fragile models but breaks capable ones, because resetting context between tasks loses coherence.

The scope is narrow and the repo says so: one task, most cells with N=1, no data on refactors or long contexts. It is a qualitative starting point, not a benchmark. Logs, generated code, scores and costs for every run are public.

LLM Dash

For daily use there is LLM Dash, the control panel for local models on my Mac: four ready profiles (build, coder, plan and a fast one for when memory is tight), served with mlx_lm.server, a dashboard to switch models and watch memory and speed, and opencode integration. Only models whose tool calling was actually verified, with the same harness as the experiments, make it into a profile.