security / threat model

Assume the model is compromised. Prove it doesn't matter.

The LLM, its context window, and everything it reads are classified untrusted, and the adversary is assumed to steer them at will. MCPIP's guarantee is architectural: compromised reasoning is not enough to produce a protected state change. This is the full adversary model — public here, not a gated PDF.

prompt injectionpoisoned documenttool-output trapUNTRUSTEDthe modelits context windowevery tool + doc it readsthe only pathMCPIPdeterministicfail-closedTRUSTEDpayments · treasuryproduction infrastructuresystems of recordtool callauthorized onlyno route around the gate — bypass is not a policy, it is a path that does not exist

01 / the adversary

Grant the attacker the whole model. It still cannot cross.

what the adversary controls

  • every prompt, and the full context window
  • every document, page, and tool output the agent reads
  • the exact wording, structure, and timing of each request
  • full knowledge of the protocol — it has read the source

what it never obtains

  • a verified identity it did not earn
  • the real name or address of any protected system
  • an approval that outlives its exact payload
  • an execution path that skips the gate
  • an explanation for any denial it can mine

02 / the guarantees

Seven guarantees. Non-negotiable.

Each holds on every request — or the request does not run. They are enforced by the architecture, not asked of the operator.

1

Timing-safe by construction

Every comparison of a token, PIN, signature, or secret is timing-safe. There is no fast path for an attacker to measure, so a valid secret can never be teased out one probe at a time.

timing-safe

2

Bound to the exact request

High-risk actions are cryptographically bound to the exact parameters they were approved for. Change a single value between approval and execution — a recipient, an amount, a path — and the action is denied. Approval for one request can never be replayed against another.

bound to the payload

3

Strict by default

Every input is strictly validated before it can reach your systems. Malformed, oversized, or smuggled payloads — unexpected fields, hidden control characters, structures built to exhaust a parser — are rejected at the door rather than interpreted.

strict by default

4

Verified identity only

Who an agent is comes only from a cryptographically verified token — never from anything the agent can set for itself. Identity smuggled into a request payload is a hard denial, not a value that is quietly trusted or silently stripped.

verified identity only

5

Capabilities, never roles

Privileged actions gate on explicit, granted capabilities and hard isolation boundaries — not on a role label. The word “admin” in a token authorizes nothing on its own, and an agent cannot even see the actions belonging to a compartment it was never granted.

capabilities, not roles

6

Fail closed, stay opaque

Any failure — a bad parse, a failed check, a missing record, a lost race — denies immediately. The caller sees only a generic denial and a correlation_id, while the real reason is sealed privately into a tamper-evident record before the action would ever run. The gate never leaks why, so it can never be turned into an oracle.

MCPIPDenied + correlation_id

7

Scales without shared trust

Authorization keeps no shared state inside any single node. Gates can be added, removed, or restarted independently, so the system scales horizontally under load without weakening a single guarantee or trusting one node’s memory.

no shared auth state

✓ all passing

Verify it yourself

974 tests, a self-verifying adversarial demo, and an independently re-checkable audit export. Every guarantee here is falsifiable on your own machine in minutes — not asserted in a datasheet.

./scripts/quickstart_demo.sh

03 / the attack catalog

Every known move against an agent — and where it dies.

The outcome column is the shape of the denial recorded internally; the caller only ever sees a generic MCPIPDenied and a correlation id.

AttackWhat the gate doesOutcome
Prompt injectionThe ask changes; the authorization does not.DENY · sealed
Poisoned document / tool descriptionSteers what the agent requests, never what the gate allows.DENY · sealed
Target enumerationReal systems are never named; unknown aliases deny opaquely.unknown alias
Approval replayStep-up approvals are consume-once; a second use finds nothing.not found
Payload tampering after approvalApproval binds to the exact payload; one byte of drift kills it.payload mismatch
Forged or expired tokenIdentity comes only from cryptographic verification.invalid token
Identity smuggled in argumentsIdentity-shaped input is a hard deny, never a value to trust.hard deny
Role-claim escalationRole labels authorize nothing; capabilities are granted, not claimed.no capability
Cross-tenant / cross-compartment reachHard isolation — the agent cannot even see across the boundary.hard deny
Error-oracle probingDenials are uniform and opaque; there is no signal to mine.opaque deny
Audit tamperingRecords are signed and hash-linked; any edit breaks the chain.chain break

04 / the two hard guarantees, drawn

Payload-bound approval. Evidence before execution.

AGENTMCPIPSYSTEMinvoke skill_payroll_run · full payloadstage approvalbound to the exact payload · consume-oncestep-up challenge · one-timeconfirm + one-time PINverify · consume · sealverdict sealed before executionexecute — only nowΔ 1 byte of payload driftpayload_mismatch → DENYthe approval never transfers to a different action

the step-up ceremony — approval binds to the exact payload and is consumed once; drift of a single byte denies.

THE RECORD IS WRITTEN BEFORE THE ACTION RUNSany edit, drop, or reorder breaks the chainrecord 12verdict · sealedsigned ✓record 13verdict · sealedsigned ✓record 14verdict · sealedsigned ✓record 15verdict · sealedsigned ✓h(prev)h(prev)h(prev)t₀ · verdict sealedt₁ · side effect runsguaranteed orderingexport any epoch · verify the chain independently — no trust in us requireda system that logs after execution can lose the one record that mattered · this ordering cannot

write-before-execute — the sealed, hash-linked record exists before the side effect runs, in guaranteed order.

05 / non-goals

The guarantee is strong because the scope is honest.

MCPIP limits the blast radius of a compromised agent to exactly its granted envelope, and records everything inside it. What it deliberately does not attempt:

Making the model harder to fool

Prompt hardening is worthwhile, but the guarantee here does not depend on it. MCPIP assumes the model is already fooled.

Judging content

What the model says is out of scope. What it does is the whole scope — governed action, not filtered speech.

Anomaly detection

The gate is deterministic, not probabilistic. Pair it with your monitoring and SIEM; it does not replace them.

Replacing IAM or your vault

MCPIP consumes verified identities and composes in front of your secrets manager. It holds neither role and neither set of keys.