Skip to content

Security model

What fold trusts, what it enforces, and how the pieces compose.

Four configuration values are trust anchors — compromising any of them compromises the gateway. Validation forces each onto https (loopback exempt for development):

AnchorWhy it is one
auth.issuers[].issuer / jwksUriThe inbound identity root: forging a principal only requires substituting the key set. Issuers are allowlisted and checked before any network I/O; JWKS fetches are single-flighted, size-bounded, and timeout-bounded so unknown-kid floods can’t be amplified against the IdP.
auth.ema.idpIssuer / idpJwksUriThe same, for the ID-JAG exchange path.
Upstream tokenEndpointsCarry client secrets (client-credentials, token-exchange).
discovery.urlDecides where traffic routes and where credentials attach: whoever controls the document can add upstreams. Documents are strictly parsed, size-capped (4 MiB), and validated whole against the running config — a collision with a static upstream rejects the entire document.

Every /mcp request passes, in order: host/origin allowlist (DNS-rebinding protection) → body-size cap → Bearer verification (trusted issuer, JWKS signature, exact audience per RFC 8707, non-empty sub, asymmetric algorithms only — RS/ES/EdDSA) → global and per-principal rate limits → routing → policy → per-upstream guards → the upstream. Every terminal response — including the refusals — produces exactly one audit event; audit is the single exit door. See Architecture for the full pipeline diagram.

With auth.mode: "required", failed verification answers 401 with a WWW-Authenticate challenge pointing at /.well-known/oauth-protected-resource (RFC 9728), which the gateway publishes.

With EMA configured, fold additionally acts as a deliberately one-grant-wide authorization server:

{
"ema": {
"idpIssuer": "https://acme.okta.com",
"idpJwksUri": "https://acme.okta.com/oauth2/v1/keys",
"signingKeyRef": "FOLD_EMA_KEY",
"tokenTtlSec": 600,
"tokenRateLimitPerMinute": 600
}
}

POST /oauth/token exchanges an enterprise-IdP ID-JAG (Identity Assertion JWT Authorization Grant, RFC 7523 jwt-bearer) for a short-lived fold-signed access token. Each assertion’s jti is single-use — recorded fleet-wide via Redis when configured, so a captured ID-JAG can’t be redeemed twice — the token endpoint is rate-limited against amplification, and issuers marked mode: "exchange" are never accepted as direct bearer issuers. Everything fold accepts afterward has aud = fold, which keeps upstream token exchange coherent.

Policy is deny-by-default and enforced twice: named invocations (tools/call, prompts/get, resources/read, and the completions and subscriptions derived from them) are denied outright, and list results are filtered per principal — a caller never sees a tool it can’t call. Protocol plumbing (ping, the lists themselves) is not policy-gated; invisibility plus call-denial is the enforcement pair.

Rules match subjects, groups, issuers, and verified token claims (ABAC). Subjects, groups, and claim names are only meaningful within an issuer, so rules should pin issuers whenever more than one IdP is trusted — otherwise a lower-assurance IdP could mint a principal that satisfies a rule written for another.

Task ownership follows the same principle: a task minted through the gateway is bound to the minting principal, and another caller’s requests for it answer exactly like an unknown id — no existence leak.

That binding is an authorization record rather than a routing hint, so it lives in shared state: with Redis configured every instance of a fleet reads the same ownership, and it survives a rolling restart — a caller can’t reach another principal’s task by landing on an instance that didn’t serve the mint. Records key on a digest of the task id and hold a digest of the owning principal, so neither the id a caller names nor its subject claims enter shared state verbatim, and a Redis outage falls back to the records that instance mirrored locally rather than to none. Two limits worth knowing: records expire after 24 hours, and a gateway without Redis is still per-instance. In both cases the task falls through to the locate-by-probe path and becomes reachable by any caller, exactly as one fold never saw minted always was.

Credentials never travel further than configured

Section titled “Credentials never travel further than configured”

Upstream credentials (API keys, exchanged tokens, passthrough bearers) attach per outgoing request and only to requests bound for a configured endpoint host of that upstream. Two layers enforce it: the HTTP client refuses cross-host redirects outright, and the transport re-checks the destination host before attaching anything — a hostile upstream answering 3xx can’t capture a credential. The token-endpoint client refuses redirects entirely, not just cross-host ones: Go replays POST bodies on 307/308, and those requests carry the client secret and, under token-exchange, the caller’s own bearer token as subject_token — so a redirecting token endpoint would otherwise hand both to the host it names. Its response body is size-bounded, and concurrent first-time callers for one identity share a single grant request instead of becoming a burst of them against the IdP.

The same redirect rule covers fold’s other two credentialed outbound clients. The discovery poller refuses every redirect — the document decides where traffic routes, and Go only strips a bearer credential when a redirect leaves the domain, so a sibling host or a plain-http same-host target would still have received it. The audit webhook refuses them too: its POST carries the sink’s configured headers and a batch of records naming principals and tools.

Each upstream picks one credential strategy:

StrategyIdentity at the upstreamWhen
none—Trusted network, no upstream auth.
staticA shared API key.API-key upstreams.
passthroughThe caller’s raw Bearer token, forwarded as-is.Upstreams doing strict RFC 8707 audience checks will reject it — prefer token-exchange.
client-credentialsfold’s own service identity for that upstream.Tokens cached until 60s before expiry.
token-exchange (RFC 8693)The end user, in a token minted for that upstream’s audience.Recommended enterprise default — preserves user identity end-to-end. Cached per (upstream, subject).

passthrough and token-exchange derive per-principal credentials, so both require auth.mode: "required" — without a verified caller identity there’s no subject to exchange for, and passthrough would forward whatever header an anonymous caller supplied. Exchanged tokens cache per (upstream, issuer, subject); per-caller strategies disable list caching, so one caller’s per-user list can never serve another.

Secrets never appear in the config document — secretRef fields name environment variables — and /health withholds URLs, owners, and error text (which can name env vars or internal hosts) unless auth is disabled, i.e. on deployments already private by posture.

Audit: every terminal response, one exit door

Section titled “Audit: every terminal response, one exit door”

One JSON event is emitted per terminal response — including 401s, 403-equivalents, and 429s, the events a SOC team actually hunts for — carrying the principal, the upstream, the authorization decision plus the matching rule id, the outcome, and latency. Sinks are configurable (stdout, webhook); webhook delivery is asynchronous and batched, so shipping audit events never adds request latency. Because audit sits at the single exit door in the pipeline (see Architecture), nothing that reaches a terminal response can bypass it — a denial is exactly as auditable as a success.

Discovery moves an authorization boundary — treat it that way

Section titled “Discovery moves an authorization boundary — treat it that way”

With dynamic discovery, who can register an upstream becomes a security decision made outside fold — on Kubernetes with fold-discovery, it’s whoever can create or label a Service in a watched namespace (see Discovery for the mechanics). Three consequences, each with a control:

  • Credential references are the sharp edge. A registered upstream chooses both its secretRef names and its destination URL — ungated, that’s an exfiltration path for any gateway-held secret, and passthrough would forward caller tokens to a URL of the registrant’s choosing. Two independent gates close it: the producer refuses credentialed strategies and secret references by default, and the gateway enforces its own discovery.allowedAuthStrategies / allowedSecretRefs / allowedCredentialHosts allowlists as a backstop, rejecting a violating document whole. Set the gateway-side allowlists whenever the discovery source isn’t operated by the gateway’s own operators.
  • Identity claims need bounds. A registration colliding with a static upstream id makes the gateway reject every future document — fail-safe, but a freeze an attacker can cause, so alert on fold_discovery_syncs_total{outcome="rejected"}. Among discovered entries, namespace prefixing requires both the upstream id and the MCP namespace to carry the registering Kubernetes namespace’s prefix; contested claims drop every claimant, so list order can’t hand an identity to whoever sorts earlier.
  • Policy is the exposure gate, and wildcards defeat it. A discovered upstream is inert until a policy rule grants its tools — unless rules use "server": "*", which makes every future registration instantly callable. With discovery enabled, name servers explicitly in allow rules.

The global rate limit protects the gateway; perPrincipalPerMinute gives each authenticated principal its own bucket so one tenant’s flood can’t 429 the rest. Per-upstream limits and circuit breakers protect fragile backends; with Redis, all of this state — plus EMA replay protection — is fleet-wide. Redis outages fail open, bounded at 500 ms per operation: the gateway degrades to per-instance enforcement rather than going down.

Content inspection — DLP, PII filtering, prompt-injection detection — is out of scope by design. Inspecting request and response bodies means buffering and rewriting traffic, which conflicts with fold’s invisibility rule (behavior through the gateway matches hitting the upstream directly) and its latency gate. fold’s security model is structural instead: deny-by-default allowlists, per-principal invisibility, claim-gated (ABAC) rules, credential brokering so agents never hold upstream keys, and a complete audit trail feeding the SIEM that does the detecting.

For vulnerability reporting and the supported-versions policy, see SECURITY.md at the repo root — this page is the architecture, that one is the process.

  • Architecture — the full request pipeline these controls sit inside.
  • Configuration — auth, policy, audit, and discovery field references.
  • Conformance — how the official MCP conformance suite verifies the gateway stays invisible.
  • Defaults — every default reviewed as a deliberate security decision.