The best LLM to run locally is MiMo: 3 practical picks

MiMo, gpt-oss and real hardware tests explain model fit, memory bandwidth, tool calls, privacy and GPU…

The best LLM to run locally is MiMo: 3 practical picks

The short version

  • MiMo-V2.6-Distill-Qwen-9B in Q4_K_M is the default local coding pick for ordinary modern computers.
  • On an M3 Max with 128 GB, gpt-oss:20b generated 74 tokens per second and gpt-oss:120b reached 51.
  • Readers should validate tool calls, memory behavior, privacy controls and workloads before trusting any local model deployment.

I watched a smart local model turn a valid tool call into XML-flavored soup because its bundled chat template chose the wrong parser. The weights worked. My coding agent didn’t.

For most people, the best LLM to run locally is MiMo-V2.6-Distill-Qwen-9B in Q4_K_M. Its file is about 5.8 GB, reported coding results are strong for its size, and it leaves memory for the surrounding application. On bigger Apple Silicon machines, I would also consider gpt-oss:20b and gpt-oss:120b. I have run both, and their speed surprised me.

There is no universal best local LLM. Nobody has published a controlled comparison using identical prompts, tools, hardware, runtime and privacy setup. A definitive leaderboard grades the football team, stadium and referee under one name.

MiMo is my default local coding model

MiMo-V2.6-Distill-Qwen-9B Q4_K_M is my first recommendation for a normal modern computer. People searching for the best open source coding LLM usually mean downloadable open weights, and MiMo offers the best balance I have seen between file size and reported agent performance.

Bar chart comparing current figures against their baselines: Annualized revenue run rate 65 $ versus 9 $, Forecast annualized revenue 120 $ versus 65 $, remaining performance obligations, or… 664 $ versus 209 $.

Xiaomi reports a SWE Pro avg@3 score of about 45% for MiMo versus 32% for Qwen3.5-9B. That gap matters under the vendor’s test conditions, though vendor charts belong in the “interesting, please verify” drawer, not carved into marble outside Rome.

On AutomationBench, Xiaomi reports 30% for MiMo against 5% for the same Qwen baseline. We lack controlled evidence that either advantage survives laptop-friendly quantization, so I test MiMo on my repositories before handing over the keys.

The Q4_K_M GGUF is about 5.8 GB, versus roughly 18 GB for the listed BF16 weights. That puts MiMo on ordinary hardware, but the runtime still needs memory for its KV cache and application overhead.

Here’s the launch chain. A GGUF file bundles model weights with tokenizer metadata and may include a chat template. Quantization stores many weights at lower precision, reducing memory use and data movement during generation. A compatible runtime executes those packed weights through low-level kernels. Hugging Face warns that an incompatible path can dequantize the model, consuming the memory I thought I had saved. The chat template converts my messages and tool definitions into MiMo’s expected token pattern; a parser converts its response back into a structured tool call. Break either translation and my expensive artificial intelligence mumbles brackets at VS Code.

MiMo’s model discussion includes user reports of malformed tool calls and missing reasoning with the bundled llama.cpp template. Those users supplied a replacement and reported cleaner tool round trips, but that is anecdotal community testing. I verify each deployment with a read, a write and a follow-up containing the tool result. If one fails, the model stays away from my repository.

That boring test has saved me more time than any leaderboard.

Bigger local LLMs need memory bandwidth, not optimism

My second pick is gpt-oss:20b for a Mac with plenty of unified memory. On my M3 Max with 128 GB, it generated about 74 tokens per second, processed prompts at roughly 756 tokens per second and produced its first token in around four seconds, with the MXFP4 model fully resident in memory.

My third is gpt-oss:120b on the same class of machine. On that M3 Max with 128 GB, it generated about 51 tokens per second, processed prompts at around 215 tokens per second and produced its first token in roughly six seconds while staying fully resident. That matters more than screenshot bragging rights.

I measured both on August 25, 2026.

Those results establish speed on one machine, not whether either gpt-oss model writes better code than MiMo in your repository. Claiming otherwise is benchmark cosplay.

A model can fit in memory yet become miserable during long sessions. Each processed token adds key and value state to the KV cache, letting the model reuse earlier attention work. As the conversation grows, that cache’s memory traffic can bottleneck performance even when the weights fit comfortably in RAM. Cache compression adds room but has poorly measured, task-specific quality costs. Partial offload adds another tax as layers and state move among CPU, GPU and system memory. Thus two people can say “it runs” while one gets a fluid assistant and the other grows a beard between tokens.

Nvidia’s MLPerf Edge Agentic submission shows the impact of cache behavior. Its Qwen3.6-27B setup completed the workload 6.4 times faster than the llama.cpp reference, which took two hours and 37 minutes. Nvidia reported that hot cache supplied 96% of prompt tokens, avoiding repeated processing of shared conversation history. This was one Jetson AGX Thor configuration using Nvidia’s optimization stack, so I would not paste that multiplier onto a Mac buying guide. But the mechanism applies everywhere: agent loops repeatedly feed old context into the model, and cached history eliminates plenty of duplicate work.

When readers ask whether they need a NAS or server for home AI, I separate storage from inference. A NAS keeps GGUF files available across machines. Fast inference needs high-bandwidth memory and a runtime that properly supports the accelerator. I would buy a dedicated server for shared access, constant availability or hardware isolation. For one person, a desktop or high-memory laptop usually brings fewer headaches.

The Raspberry Pi is a great low-power homelab controller. In separate test series cited by a Pi-versus-mini-PC comparison, an 8 GB Pi 5 idled at 2.8 W while an ASUS NUC averaged 5.4 W. The Pi loses flexibility once I need Proxmox, virtual machines, transcoding or faster networking. Samulczyk eventually moved his setup to an x86 small-form-factor PC for that reason. I love the Pi, but affection is a poor hypervisor.

Orange workstation radiating heat beside a blue NAS and a small controller board on a mint platform.

Local coding agents earn trust through checks

I let a local model propose code. The compiler gets veto power.

The useful pattern pairs a candidate answer with a cheap deterministic check. A coding agent can format, compile or run a focused test. Network automation can reject configurations violating known constraints. Failed inputs, and tasks without reliable checks, move to a stronger remote model with explicit permission. Routine work stays local, with an escape hatch for ambiguity. Remote escalation belongs in the design; no shame required. After two decades shipping software, I trust rejection paths more than confident prose.

This matters because local-model benchmarks often measure the serving stack too. Research comparing Ollama, llama.cpp, vLLM and SGLang found that stacks could reject requests before inference or handle tool calls differently. Changing the aggregation method moved reported tool-use fidelity by as much as 55 points on the same evaluation. A model leaderboard can secretly become a parser leaderboard, returning us to the opening XML soup.

Production tests are just as ruthless. On SWE-Serve tasks with end-to-end coverage, agents passed 69% when those tests were excluded from scoring and 46% when the checks counted. Hidden serving tests exposed patches that looked correct locally but failed after deployment. Every founder learns this, usually late on Friday while takeaway cools beside the keyboard.

Tool authorization also requires enforcement at execution. In one MCP evaluation, a server checking permission only inside the tool body exposed forbidden tools in 21% of attempts. Permission-aware visibility combined with per-tool invocation enforcement produced zero exposures in the reported trials. Hiding a tool from the model interface cannot stop a scripted client from naming it directly, and an agent guessing a hidden name is hardly science fiction.

Local deployment also offers weaker privacy guarantees than many assume. A security preprint found plaintext prompt remnants in allocator-managed memory and persistence inside wrapper software. Its researchers exploited a saved-conversation-state authorization flaw in every controlled attempt on one consumer serving interface. That result covers the tested implementation, not every local server, but destroys the comforting idea that staying inside the house automatically makes traffic secure.

EDPB Deputy Chair Jelena Virant Burnik said:

The new EDPB guidelines are a major step in further aligning how Data Protection Authorities decide whether an administrative fine should be imposed, either on its own or alongside other corrective measures. The GDPR significantly increased the corrective powers of DPAs, with fines serving as an important instrument for effective enforcement. The guidelines reaffirm our commitment to providing greater clarity and ensuring the consistent application of the GDPR across Europe.

I want tenant isolation, memory cleanup and encrypted persistence wherever prompts are saved. The green “localhost” badge can keep its emotional-support duties.

Self-hosted image generation changes the hardware plan

My RTX 5060 Ti has 16 GB of VRAM, which sounds generous until ComfyUI occupies it for self hosted image generation AI. With that process loaded, Ollama saw only 150 MB of available VRAM in my setup. A 20B language model then ran entirely on the CPU.

The mechanism is simple. Image generation loads its model and working state onto the GPU. Ollama places language-model layers in whatever capacity remains. With almost no free VRAM, those layers stay in system memory and the CPU does the work. The model still launches, making a quick status check look successful; generation speed reveals the truth. Shared hardware therefore needs scheduling or separate accelerators if both workloads must remain responsive. My GPU cannot run a diffusion kitchen and LLM dining room during the same service.

The supplied evidence does not identify the best self-hosted image-generation model. Visual-language models discussed around local AI generally analyze images; creation uses a separate generation stack. I would rather leave the recommendation open than crown something using unrelated vision benchmarks.

The same honesty applies to local LLM rankings. No controlled cross-hardware test names one winner across chat, coding, tool use and long-context work. MiMo’s advantage after common laptop quantization remains unknown. Aggressive quantization and KV-cache compression also cause workload-specific losses that generic hardware specifications cannot predict.

Clausebench observed:

A later date is more time to do the same amount of work, not less work.

My specific bet: by the end of 2028, the winning local coding product will ship a tested model file, runtime and chat template as one signed bundle, with deterministic checks enabled by default. Until then, I am keeping one hand on the test suite and the other away from `sudo`.

Frequently asked questions

What is the best LLM to run locally for coding?

The best LLM to run locally for most people is MiMo-V2.6-Distill-Qwen-9B in Q4_K_M. Its GGUF is about 5.8 GB, reported coding and automation results are strong for its size, and it preserves memory for the application, KV cache and runtime overhead.

Do you need a NAS or server for a local LLM?

Local LLM inference needs high-bandwidth memory and accelerator support, while a NAS mainly stores and shares model files. A dedicated server makes sense for shared access, constant availability or isolation. For one person, a desktop or high-memory laptop usually provides a simpler setup with fewer operational headaches.

Can self-hosted image generation and a local LLM share one GPU?

Self-hosted image generation can consume nearly all available GPU memory, forcing a local language model onto the CPU. In the reported setup, ComfyUI left only 150 MB free on a 16 GB RTX 5060 Ti. Responsive simultaneous workloads therefore require scheduling or separate accelerators.

Sources

Related reading

Luca

Luca

Luca by the way is the personal blog of Los Angeles based entrepreneur Luca Capula. A true Italian who lives between Torino and LA.

More posts →