Kimi K3 Local Setup: Run the 2M-Context Model On Your PC
A 2M-token context window sounds like a spec-sheet flex. It isn't — it's the difference between "paste the relevant file" and "paste the whole repository." Kimi K3 brought that to the table, and running it locally means the long-context experiments don't cost you a cent per token.
Companion guide to the video — every command written out, copy along.
What 2M context actually buys you
Rough scale, so the number means something:
- ~8K tokens: one long file — the old normal
- ~200K tokens: a medium project, or a thick book
- 2M tokens: a large monorepo, your entire document archive, or months of chat logs
The workflow change: instead of curating what the model sees, you point it at everything and ask the question. "Where is auth handled in this codebase?" stops being a search problem.
Reality check: full-context prompts are slower and the middle of very long contexts is weaker than the edges (true for every model today). Use the big window when you need it, not as a habit.
Hardware: be honest here first
Long context = long KV cache. The model weights are one cost; feeding it 2M tokens is another.
- Comfortable: 24 GB+ VRAM (or Mac unified memory) for meaningful context lengths
- Workable: 12–16 GB with smaller quants and moderate contexts (think 100–200K, still huge)
- Full 2M locally: workstation territory, or quantized hard enough that you trade quality for it
If your machine is in the middle band: you still get a locally-run model with a massively larger usable context than last year's locals. That's the win.
1. Install Ollama
Same as any local setup:
# Windows/macOS: installer from ollama.com
# Linux:
curl -fsSL https://ollama.com/install.sh | sh
2. Pull Kimi K3
ollama pull kimi-k3
If the tag has moved, ollama search kimi shows the current listing — pull that. Size on disk depends on the quantization variant, so check your free space for a multi-GB download.
3. Load a whole project and ask about it
ollama run kimi-k3
Now the fun part. In a new terminal, feed it your codebase:
cat src/**/*.ts | ollama run kimi-k3 "Summarize the architecture of this codebase and list every external service it talks to."
Watch the context fill. This is the moment the 2M window stops being a number.
4. Wire it into your workflow
http://localhost:11434 is your OpenAI-compatible endpoint, same as any Ollama model:
- Point Codex CLI (or any compatible client) at it — the provider config is in Use Codex CLI with Any AI Model.
- Repo-scale questions become one prompt instead of a RAG pipeline you have to maintain.
Pro tip: for code questions, include your directory tree (tree src/) with the files. The model navigates a map better than a pile.
Recap + exercise
You checked whether your hardware honestly fits the context length you need, installed Ollama, pulled Kimi K3, and fed it a whole project in one prompt.
Exercise (15 min): take a repo you know well, ask Kimi K3 to find a subtle bug or explain the trickiest module, and grade the answer against what you know to be true. That's the real benchmark — not leaderboards.
The full video walkthrough is on the DevKingOv channel, and the Local AI Setup course is free on the portal — track your progress lesson by lesson.
FAQ
Can I really use the full 2M context locally?
Only with serious VRAM/ unified memory, and quality at extreme depths varies. Most practical local use sits in the 100K–500K range — which is still enormous.
Kimi K3 vs GLM-5.3 — which should I run locally?
Keep both if disk allows. Kimi K3 when the task is "understand all of this at once"; GLM-5.3 for snappier general coding chat. The Codex CLI guide shows switching providers per task.
Does long context replace RAG?
For personal-scale archives and single repos, increasingly yes. For shared, constantly-updating knowledge bases at work, RAG still earns its keep.
Prefer watching?
Every post here is a lesson in a free video course — follow along on YouTube and track your progress on the portal.
Keep reading
GLM-5.3 Local Setup (2026): Install & Run It On Your PC
The step-by-step local setup for GLM-5.3 — hardware requirements, Ollama install, the exact pull commands, and how to wire it into your dev tools. Copy along in 15 minutes.
What AI Coding Actually Costs in 2026
Cursor at $0.07/task vs Claude Code at ~$4 — the real numbers, the token-efficiency axis, the 65% output cut, and the free local floor. A cost playbook for AI-assisted development.
Use Codex CLI with Any AI Model (Not Just OpenAI)
Codex CLI is an OpenAI-compatible client, not an OpenAI-only one. One config file points it at GLM, Claude via OpenRouter, or a local model.
Want help applying this? Book a 1-on-1 with a consultant.
Find a Consultant