Everything an agentic IDE should do — with Capuchin built in.
Agents that plan, edit, and verify — powered by Capuchin, free to try. Claude, Gemini, and ChatGPT included in your plan and used token-efficiently. Fully local or BYOK when you want it.
Agents that plan, edit, run, and check.
50+ real agent tools file edits, terminal, git, test runner, browser, search, subagents — all implemented, all gated by permissions.
Plan vs Fast approve a plan for complex work, or fire instant edits.
Multi-file edits coherent changes across many files, reviewed as one diff set.
Parallel sub-agents split by file cluster or service, then merged.
Agentic editing
Agents that plan, edit, run, and check.
Capuchin built in — the big three in your plan.
Capuchin — free to try fast-and-reasoning, two tiers (Flash + Reason), tuned for coding. Try it free with the promo; unlimited on paid plans.
Claude · Gemini · ChatGPT included a monthly frontier allowance on paid plans — one subscription instead of three. BYOK also supported.
Automatic routing everyday work runs on Capuchin; frontier tokens are reserved for the problems that deserve them.
Your own, local or remote Ollama/llama.cpp/vLLM on your machine, or any OpenAI-compatible endpoint. Same agent features either way.
Capuchin built in
Capuchin built in — the big three in your plan.
Your frontier allowance goes further here.
~$0.06 avg per frontier request the on-device context engine sends only relevant code slices — not whole files — even on the priciest frontier model.
~90% cached context long agentic runs reuse cached context instead of re-sending it, slashing input token cost.
Per-task budgets + hard cap agents stop instead of surprising you. Live consumption dashboard included.
Capuchin absorbs the routine unlimited Capuchin on Pro handles the everyday 80% — a $20 allowance stretches to 300+ frontier agent runs.
Efficiency is a quality argument a model that wastes context on irrelevant files has less room for the code that matters. Sending precise slices is why sessions here run past a hundred messages instead of degrading after eight.
Every token pulls its weight
Your frontier allowance goes further here.
Your codebase never has to leave your machine.
On-device index Merkle-tracked, tree-sitter-chunked, embedded and searched locally.
Air-gapped mode local index + local models + internal registry. Zero outbound calls.
Optional team sync share an encrypted index when you want; turn it off and you're fully local.
Monorepo-aware a symbol/dependency graph so agents reason across services, not just nearby files.
Local-first
Your codebase never has to leave your machine.
Every agent run, signed and replayable.
Signed event logs a tamper-evident, replayable trail of plan, tools, and diffs.
Test-gated apply agents can't commit code that fails your tests.
Content-signed diffs verify exactly what changed and reproduce it byte for byte.
Audit-ready hand a run to a reviewer or compliance officer to check themselves.
Verifiable by design
Every agent run, signed and replayable.
Run the agent where the code actually is.
Local full capability on your machine — semantic search, interactive shell sessions, background processes, checkpoints, OS-level sandboxing.
SSH point the agent at a remote dev box. It translates every command over SSH and runs there, not here.
Cloud execute inside a managed container workspace, sandboxed by default.
Remote Control connect over WebSocket to a peer machine running the remote host agent — useful for a build server, a test rig, or a colleague's environment.
Four backends
Run the agent where the code actually is.
Undo anything the agent did.
Pre-edit snapshots captured automatically before the agent's first mutation of any file.
Selective revert roll back one file or an entire run.
Revert planning see exactly what will be restored or deleted before it happens.
Honest about limits unrevertable operations flagged rather than silently skipped.
Reversible by default
Undo anything the agent did.
Baseline-aware type checking.
Baseline capture the agent captures diagnostics at run start and reports only new ones.
Per-language detection detects the right checker per language automatically.
Monorepo scoping scopes to affected packages instead of re-checking the entire tree.
Transient retry a single retry on transient environment failures rather than reporting a false problem.
Only what it broke
Baseline-aware type checking.
Research threads with real files.
Long-lived threads scoped to a workspace, linking many runs. Findings accumulate across days.
Real markdown output stored under .monkeyscode/docs/, git-tracked and editable in any editor.
Lifecycle management open → active → resolved → archived.
Longer than a session
Research threads with real files.
It runs inside a box, on every platform.
Platform sandboxing Seatbelt on macOS, bubblewrap on Linux, AppContainer on Windows. Workspace-write mode constrains writes and denies network unless approved.
Secret scanning pre-commit detection with blocking findings and redaction in tool output.
Prompt injection defence untrusted content is fenced and classified, with injection detection before it reaches the model.
Policy engine per-tool permission tiers with allow, deny and ask decisions, configurable at workspace and organisation level.
Sandboxed by default
It runs inside a box, on every platform.
The extensions you want, from a registry you can trust.
Open VSX by default vendor-neutral, legally clean for an independent IDE.
Sideload anything install any .vsix directly.
Self-hostable registry mirror and curate your own catalog.
Org governance admins set allowlists and pin versions.
Extensions, your way
The extensions you want, from a registry you can trust.
Native on all three. Same engine everywhere.
Windows · macOS · Linux signed, notarized installers; AppImage/deb/rpm/Flatpak on Linux.
One engine editor, CLI, and CI share the same core.
Background & scheduled agents long-running and cron-style, headless.
Cross-platform
Native on all three. Same engine everywhere.
Everything the agent can actually do.
Not a roadmap. Every tool below is implemented, permission-gated, and callable today. Each one is subject to the permission mode you set, from read-only planning through fully autonomous execution inside a sandbox.
Filesystem
Read file · Write file · Structured edit · Multi-file edit · Create · Delete · Move · Directory listing
Search & navigation
Ripgrep search · Glob matching · Semantic code search · Symbol lookup · Find references · Dependency graph traversal
Terminal & execution
Run command · Streaming PTY session · Background process · Environment inspection
Version control
Git status · Diff · Stage · Commit · Branch · Log · Blame · Worktree management
Testing & verification
Detect test command · Run suite · Run targeted tests · Parse failures · Test-gated apply
Browser
Navigate · Interact · Screenshot · Console capture
Extensibility
MCP server tools · Subagent spawn · Custom command execution
Every tool is permission-gated. Nothing with a side effect runs without either your explicit approval or a mode you deliberately enabled. Nothing runs outside the sandbox.
Why we publish a number. Tool coverage is the difference between an agent that suggests and an agent that finishes. Most tools in this category describe their capabilities in prose. We count them, list them, and gate every one, because that is what makes the list verifiable rather than marketing.
One engine, every surface — included in your plan.
Your subscription grows with the platform at no extra cost.
Code Agent
A desktop app to launch, watch, and orchestrate parallel agents across your projects — without the full editor. Same engine, same 50+ tools.
Learn moreMonkeysCode CLI
Agentic coding from the command line and CI — headless, same engine. Runs Capuchin, Claude, Gemini, ChatGPT, or your own model.
Start free with Capuchin.
Try Capuchin free, no card. One plan adds unlimited Capuchin plus Claude, Gemini, and ChatGPT — with every token counted.