fold is one gateway with one job — govern the path between MCP clients and MCP
servers — and that single position covers a family of problems. Most
deployments start with one of these and grow into the others, because they all
ride the same config.
Unify a federation
Acquisitions, child orgs, and regional teams each build their own MCP
servers — Python from the ML org, TypeScript from platform, Go from last
year’s acquisition. fold presents them as one virtual server with
namespaced tools (gh__create_pr, search__query), per-upstream
owner: { org, team } metadata, and graceful degradation when one team’s
server is down. No team rewrites anything.
Draw the security boundary
One choke point instead of N: OAuth 2.0 resource server with enterprise
IdP integration (EMA/ID-JAG), deny-by-default tool allowlists, list
filtering so callers never see tools they can’t call, and an audit event
for every request — including the 401s, denials, and 429s your SOC team
actually hunts for.
Broker credentials
Agents hold exactly one token, audience-bound to fold. Per upstream, fold
exchanges it (RFC 8693) for an upstream-audience token preserving user
identity, or injects service credentials. Upstream API keys live in the
gateway’s environment and never reach a model’s context window.
Protect fragile services
Agent traffic is bursty and retry-happy. List caching, global and
per-upstream rate limits with Retry-After, and per-upstream circuit
breakers stand between an agent storm and the twenty-year-old ERP that
must not fall over. Where a rate limit smooths a burst and forgets it,
a budget accumulates across a calendar period, so a
busy month meets a ceiling.
Govern each customer separately
A tenant is a named set of principals carrying its own allowance, its own
rate-limit bucket, and its own view of the federation — resolved from
claims the IdP already asserts. Ten agents on one team share one bucket
rather than holding ten between them, and every audit event they produce
carries the tenant’s name. Tenancy →
Federate local servers
Most MCP servers are local processes speaking stdio, not HTTP endpoints.
The fold-stdio shim runs one and serves it over streamable HTTP, so it
joins the federation as an ordinary url upstream with every credential
strategy, guard, and policy rule applied unchanged. Stdio servers →
Govern vendor MCP servers
Third-party and SaaS MCP endpoints go behind your auth, your allowlists,
and your audit trail — instead of every employee pasting a personal API
key into every client. Swap a vendor without touching a single client
configuration.
Discover upstreams automatically
A team ships an MCP server, the registry lists it, and it appears behind
the gateway — no fold config change. On Kubernetes, label a Service
fold.run/upstream: "true" and fold-discovery does the rest, with
gateway-side bounds on what a discovered upstream may claim.
Discovery →
Expose tools outward, carefully
Give partners or customers a curated, policy-scoped subset of internal
tools on one hardened endpoint — DNS-rebinding protection, host
allowlists, and per-principal rate limits included. A single static
binary deploys anywhere in your VPC.
Conformant, provably
The official MCP conformance suite runs through fold on every merge —
40/40 checks, including sampling, elicitation, and subscriptions bridged
through the gateway — alongside a latency gate on the proxy path.
Conformance →
Every use case above is the same binary and the same config file — an
upstreams list plus optional auth, policy, audit, and tenants
blocks. Begin with a single passthrough upstream in front of your most-used
server, then add governance as the surface grows. Nothing
here is a separate product or a separate deployment: declare no tenants and
no budgets, and the gateway behaves exactly as it did before either existed.