Skip to content
The AI of MonkeysCode

Meet Capuchin. Try it free.

Capuchin is the AI that powers MonkeysCode — a coding model built for one thing: shipping software. It plans, writes, and verifies across your whole codebase, thinks hard when the problem is hard, and stays instant when it isn't. Try it free with the promo — no card — and get it unlimited on every paid plan.

Free to try, unlimited on every paid plan. A managed service, so there is nothing to provision and nothing to maintain.

Fast + Reasoning· two modes, one model
Agent-native· built for tools, not chat
Long context· whole services at once
Managed service· always current, nothing to install
Built for software, not small talk

A model that does one job, exceptionally.

General chat models spread their capacity across everything. Capuchin spends it all on software engineering — reading large codebases, calling tools, editing across files, and checking its own work against your tests. That focus is why it's both sharper on real coding and lean enough to run fast and cheap, wherever you put it.

Capuchin Flash and Capuchin Reason.

One model family, two profiles — so you never trade speed for depth or depth for speed.

Capuchin Flash

instant

Best for: autocomplete, apply, quick edits

Feel: sub-second

Where it shines: the 80% of everyday work

Capuchin Reason

deliberate

Best for: planning, multi-file agents, debugging

Feel: step-by-step reasoning

Where it shines: the 20% that's actually hard

MonkeysCode routes between them automatically — or you can pin one.

What Capuchin does.

Two minds in one model

Instant when it's easy. Deliberate when it's hard.

Flash mode answers and edits land immediately for the everyday work — completions, small changes, quick questions.

Reason mode for the hard stuff, Capuchin thinks step by step before it acts — planning multi-file changes, untangling bugs.

Automatic routing MonkeysCode picks the mode by the task, so you get speed by default and depth on demand.

Budgeted thinking reasoning is capped per task, so 'thinking' never means 'spinning.'

FlashReasonAuto-routingBudgeted

Two minds in one model

Instant when it's easy. Deliberate when it's hard.

Agent-native

Built to use tools, not just talk about them.

Reliable tool use Capuchin calls your terminal, git, test runner, browser, and MCP servers — accurately, call after call.

Long-horizon stability it holds the thread across dozens of tool calls and many turns.

Plans then executes capuchin produces a plan you can approve, then carries it out across files in parallel.

Self-checking it runs your tests and fixes what it broke before handing back.

Tool useMulti-stepParallelSelf-correcting

Agent-native

Built to use tools, not just talk about them.

Sees the whole picture

Reads your codebase, not just the open file.

Massive context window hundreds of thousands of tokens, so whole services and long histories fit at once.

Repo-aware paired with MonkeysCode's on-device index and dependency graph, it reasons across files, services, and repos.

Keeps the thread long agent runs are summarized and carried forward without losing important context.

Long contextRepo-awareCross-service

Sees the whole picture

Reads your codebase, not just the open file.

Polyglot

Fluent across your stack — not just one language.

Many languages, first-class strong across PHP, Python, Go, Rust, JavaScript/TypeScript, Java, and more.

Built for MonkeysLegion extra-sharp on PHP 8.4 and the MonkeysLegion framework, where most tools are weakest.

Real-world tasks tuned on the kind of multi-file, multi-tool work engineers actually do.

PHPPythonGoRustTSJava

Polyglot

Fluent across your stack — not just one language.

Efficient by design

Cheap to run is why it is unlimited.

Unlimited because it is efficient Capuchin is specialised for software engineering rather than general conversation, which is why serving it costs a fraction of a frontier model and why we include it without metering.

90% context cache hit rate long agent runs reuse cached context instead of resending it, cutting input cost dramatically on exactly the workloads that would otherwise be expensive.

Your frontier allowance lasts longer too the same context engine that makes Capuchin cheap makes Claude, Gemini and ChatGPT cheap. A $20 allowance stretches to 300+ frontier agent runs here.

Unlimited90% cached300+ frontier runs

Efficient by design

Cheap to run is why it is unlimited.

Measured, not claimed

We benchmark Capuchin on the endpoint you actually use.

Most model benchmarks are run on a research configuration that nobody serves in production. Ours are run against the same endpoint your editor calls, at the same quantization and context settings, so the number reflects what you get.

BenchmarkCapuchin ReasonWhat it measures
SWE-bench Pro62.1%Resolving real GitHub issues end to end
Terminal-Bench 2.181.0Agentic command-line execution
FrontierSWE74.4%Long-horizon engineering tasks
GPQA Diamond91.2%Advanced scientific reasoning
Cost per completed task~$0.06What a finished unit of work actually costs

Measured on our own production serving endpoint, at the quantization and context configuration we actually serve. Methodology published alongside every figure. Re-measured on each release.

What we can prove today.

92.7%
of production requests

run on Capuchin — not a projection, that is the measured share across real user sessions.

100+
messages per session

after the context engine rewrite, sessions that previously stalled at eight exchanges now run past a hundred.

~$0.06
per frontier request

the on-device context engine sends relevant code slices rather than whole files, keeping even the priciest frontier model cheap.

50+
agent tools

not planned — implemented, permission-gated and callable today.

50+ tools, all gated

A model is only as useful as what it can reach.

Reasoning quality matters less than people think if the model cannot act. Capuchin has 50+ implemented tools covering the filesystem, version control, the terminal, testing, the browser, search and Model Context Protocol servers — each behind an explicit permission.

FilesystemTerminalGitTestsBrowserSearchMCPSubagents

That is why it completes work rather than describing it. See the full tool list →

How Capuchin runs

A managed service, the same way you already use AI.

Capuchin runs on GPU infrastructure we own and operate, delivered the same way Claude and ChatGPT are delivered. There is nothing to install, nothing to provision, and no weights to manage. You get the current version automatically.

Always current

New checkpoints roll out continuously

Built for latency

Speculative decoding + prefix caching

Included, not metered

Unlimited on every paid plan

Need fully offline?

Use Ollama, llama.cpp or vLLM

Always improving

Capuchin gets better every release.

Specialized by us

Tuned relentlessly for real software engineering inside MonkeysCode.

Learns from real work

With your permission, Capuchin improves from opt-in signals — never from private code.

Ships often

New checkpoints roll out on a steady cadence, gated against our benchmarks.

Open where it counts

Your data stays yours; improvement is consent-first by design.

One subscription, every model

Unlimited Capuchin + the frontier in one plan.

Upgrade to Pro and get unlimited Capuchin plus a monthly allowance of Claude, Gemini, and ChatGPT — one subscription instead of three. And your allowance goes further here, by design.

~$0.06
Avg per frontier request

The on-device context engine sends only the relevant slices of your codebase — not whole files — so even the priciest frontier model stays cheap per request.

~90%
Context served from cache

Repeated agent steps reuse cached context instead of re-sending it, cutting input token cost dramatically on long agentic runs.

300+
Frontier requests / month on Pro

Everyday work runs on unlimited Capuchin; frontier tokens are reserved for the hard problems. A $20 allowance stretches to hundreds of agent runs.

Measured on real agentic sessions with the highest-cost frontier model. Live consumption dashboard, per-task budgets, and a hard spend cap included — you see and control every token.

Free

$0

The full IDE with 100% local models or your own key (BYOK). Try Capuchin free with the promo.

Try Capuchin free

Pro

$20/mo

Unlimited Capuchin + $20/mo frontier allowance (Claude · Gemini · ChatGPT).

See Pro

Pro+

$60/mo

Everything in Pro with a bigger frontier allowance and priority queue.

See Pro+

The default that respects your choice.

Capuchin is what powers MonkeysCode out of the box. But the moment you want a different model — Claude, Gemini, ChatGPT, or your own, local or remote — it's one switch away. Capuchin is the AI we stand behind; model freedom is the promise underneath it.

Start free — no card
Try it free — no card

Code with Capuchin. Free to start.

MonkeysCode's own AI — fast, reasoning, agent-native. Try it free, and it's one plan away from the whole frontier.