Meet Capuchin. Try it free.
Capuchin is the AI that powers MonkeysCode — a coding model built for one thing: shipping software. It plans, writes, and verifies across your whole codebase, thinks hard when the problem is hard, and stays instant when it isn't. Try it free with the promo — no card — and get it unlimited on every paid plan.
Free to try, unlimited on every paid plan. A managed service, so there is nothing to provision and nothing to maintain.
A model that does one job, exceptionally.
General chat models spread their capacity across everything. Capuchin spends it all on software engineering — reading large codebases, calling tools, editing across files, and checking its own work against your tests. That focus is why it's both sharper on real coding and lean enough to run fast and cheap, wherever you put it.
Capuchin Flash and Capuchin Reason.
One model family, two profiles — so you never trade speed for depth or depth for speed.
Capuchin Flash
instantBest for: autocomplete, apply, quick edits
Feel: sub-second
Where it shines: the 80% of everyday work
Capuchin Reason
deliberateBest for: planning, multi-file agents, debugging
Feel: step-by-step reasoning
Where it shines: the 20% that's actually hard
MonkeysCode routes between them automatically — or you can pin one.
What Capuchin does.
Instant when it's easy. Deliberate when it's hard.
Flash mode answers and edits land immediately for the everyday work — completions, small changes, quick questions.
Reason mode for the hard stuff, Capuchin thinks step by step before it acts — planning multi-file changes, untangling bugs.
Automatic routing MonkeysCode picks the mode by the task, so you get speed by default and depth on demand.
Budgeted thinking reasoning is capped per task, so 'thinking' never means 'spinning.'
Two minds in one model
Instant when it's easy. Deliberate when it's hard.
Built to use tools, not just talk about them.
Reliable tool use Capuchin calls your terminal, git, test runner, browser, and MCP servers — accurately, call after call.
Long-horizon stability it holds the thread across dozens of tool calls and many turns.
Plans then executes capuchin produces a plan you can approve, then carries it out across files in parallel.
Self-checking it runs your tests and fixes what it broke before handing back.
Agent-native
Built to use tools, not just talk about them.
Reads your codebase, not just the open file.
Massive context window hundreds of thousands of tokens, so whole services and long histories fit at once.
Repo-aware paired with MonkeysCode's on-device index and dependency graph, it reasons across files, services, and repos.
Keeps the thread long agent runs are summarized and carried forward without losing important context.
Sees the whole picture
Reads your codebase, not just the open file.
Fluent across your stack — not just one language.
Many languages, first-class strong across PHP, Python, Go, Rust, JavaScript/TypeScript, Java, and more.
Built for MonkeysLegion extra-sharp on PHP 8.4 and the MonkeysLegion framework, where most tools are weakest.
Real-world tasks tuned on the kind of multi-file, multi-tool work engineers actually do.
Polyglot
Fluent across your stack — not just one language.
Cheap to run is why it is unlimited.
Unlimited because it is efficient Capuchin is specialised for software engineering rather than general conversation, which is why serving it costs a fraction of a frontier model and why we include it without metering.
90% context cache hit rate long agent runs reuse cached context instead of resending it, cutting input cost dramatically on exactly the workloads that would otherwise be expensive.
Your frontier allowance lasts longer too the same context engine that makes Capuchin cheap makes Claude, Gemini and ChatGPT cheap. A $20 allowance stretches to 300+ frontier agent runs here.
Efficient by design
Cheap to run is why it is unlimited.
We benchmark Capuchin on the endpoint you actually use.
Most model benchmarks are run on a research configuration that nobody serves in production. Ours are run against the same endpoint your editor calls, at the same quantization and context settings, so the number reflects what you get.
| Benchmark | Capuchin Reason | What it measures |
|---|---|---|
| SWE-bench Pro | 62.1% | Resolving real GitHub issues end to end |
| Terminal-Bench 2.1 | 81.0 | Agentic command-line execution |
| FrontierSWE | 74.4% | Long-horizon engineering tasks |
| GPQA Diamond | 91.2% | Advanced scientific reasoning |
| Cost per completed task | ~$0.06 | What a finished unit of work actually costs |
Measured on our own production serving endpoint, at the quantization and context configuration we actually serve. Methodology published alongside every figure. Re-measured on each release.
What we can prove today.
run on Capuchin — not a projection, that is the measured share across real user sessions.
after the context engine rewrite, sessions that previously stalled at eight exchanges now run past a hundred.
the on-device context engine sends relevant code slices rather than whole files, keeping even the priciest frontier model cheap.
not planned — implemented, permission-gated and callable today.
A model is only as useful as what it can reach.
Reasoning quality matters less than people think if the model cannot act. Capuchin has 50+ implemented tools covering the filesystem, version control, the terminal, testing, the browser, search and Model Context Protocol servers — each behind an explicit permission.
That is why it completes work rather than describing it. See the full tool list →
A managed service, the same way you already use AI.
Capuchin runs on GPU infrastructure we own and operate, delivered the same way Claude and ChatGPT are delivered. There is nothing to install, nothing to provision, and no weights to manage. You get the current version automatically.
Always current
New checkpoints roll out continuously
Built for latency
Speculative decoding + prefix caching
Included, not metered
Unlimited on every paid plan
Need fully offline?
Use Ollama, llama.cpp or vLLM
Capuchin gets better every release.
Specialized by us
Tuned relentlessly for real software engineering inside MonkeysCode.
Learns from real work
With your permission, Capuchin improves from opt-in signals — never from private code.
Ships often
New checkpoints roll out on a steady cadence, gated against our benchmarks.
Open where it counts
Your data stays yours; improvement is consent-first by design.
Unlimited Capuchin + the frontier in one plan.
Upgrade to Pro and get unlimited Capuchin plus a monthly allowance of Claude, Gemini, and ChatGPT — one subscription instead of three. And your allowance goes further here, by design.
The on-device context engine sends only the relevant slices of your codebase — not whole files — so even the priciest frontier model stays cheap per request.
Repeated agent steps reuse cached context instead of re-sending it, cutting input token cost dramatically on long agentic runs.
Everyday work runs on unlimited Capuchin; frontier tokens are reserved for the hard problems. A $20 allowance stretches to hundreds of agent runs.
Measured on real agentic sessions with the highest-cost frontier model. Live consumption dashboard, per-task budgets, and a hard spend cap included — you see and control every token.
Free
The full IDE with 100% local models or your own key (BYOK). Try Capuchin free with the promo.
Try Capuchin freeThe default that respects your choice.
Capuchin is what powers MonkeysCode out of the box. But the moment you want a different model — Claude, Gemini, ChatGPT, or your own, local or remote — it's one switch away. Capuchin is the AI we stand behind; model freedom is the promise underneath it.
Start free — no cardCode with Capuchin. Free to start.
MonkeysCode's own AI — fast, reasoning, agent-native. Try it free, and it's one plan away from the whole frontier.