MCP Runtime Decision Evaluator

Why this page exists. A policy pack is valuable when reviewers can review it. It becomes enterprise infrastructure when an agent host, MCP gateway, or CI admission check can execute the same policy before a tool call happens.

Rechecked September 29, 2026: MCP 2026-07-28 is still current and stateless. There is no negotiation handshake. Each request carries protocol version and capabilities. Servers MUST implement server/discover. kill_session here is a host-session kill switch, not Mcp-Session-Id. Streamable HTTP revisions through 2025-11-25 could assign that header; 2026-07-28 ignores it and does not mint session IDs. Streamable HTTP Security & Endpoint requires servers to validate the Origin header on every incoming connection. If Origin is present and invalid, the server MUST respond with HTTP 403 Forbidden. Without that check, a remote page can use DNS rebinding to call a local MCP server. Streamable HTTP request metadata mirrors MCP-Protocol-Version, Mcp-Method, and Mcp-Name so a gateway can route without parsing the body. Servers that process the body MUST reject disagreement with HTTP 400 and JSON-RPC HeaderMismatch (-32020). Rechecked September 12, 2026 against the OWASP Agent Control Standard (ACS) v0.1.0 §6.4: the Observed Agent MUST wait for and apply a Guardian decision; on timeout, transport failure, or an error without a decision, ACS defaults to on_decision_failure: proceed (fail-open). This evaluator kills that fail-open path instead of allowing the tool call.

The product bet

Agentic security programs will not scale because a prompt says “stay in scope.” They scale when the platform can make boring, repeatable decisions about what an agent may do right now.

The runtime decision evaluator is the enforcement bridge. It consumes the generated MCP Gateway Policy Pack and one runtime request, then returns a structured decision:

  • allow for declared read-only context.
  • allow_scoped_branch for branch writes that satisfy branch, path, file count, and diff limits.
  • allow_scoped_ticket for declared ticket or incident writes.
  • hold_for_approval when a tool scope or change class needs a typed human approval record.
  • deny when the workflow, agent class, namespace, gate phase, branch, path, or diff size drifts from policy.
  • kill_session when a runtime kill signal fires.

This makes AI easier for the enterprise operator: agents do not need to interpret policy text, and reviewers do not need to reconstruct why a tool call was allowed.

Workflow at a glance

MCP Runtime Decision Evaluator workflow

Normalize an MCP runtime event and return one deterministic allow, hold, deny, or kill decision with reasons and required evidence.

mcp-governance
  1. Signal

    Load the runtime event

    Capture identity, tenant, workflow, run, server, capability, arguments, data, destination, approvals, and session history.

  2. Scope

    Resolve decision inputs

    Validate schemas and load authorization, connector trust, tool risk, policy, entitlement, context, egress, and telemetry state.

  3. Decision

    Apply ordered rules

    Evaluate hard stops, ACS Guardian fail-open decision failures, Streamable HTTP Origin DNS-rebinding (HTTP 403), HeaderMismatch (-32020) header/body disagreement, explicit denies, missing evidence, approval holds, argument constraints, rate/sequence, and allow conditions.

  4. Action

    Return one decision

    Emit allow, hold, deny, or kill with stable reason codes, transformations, obligations, and next action.

  5. Proof

    Persist the decision

    Record input/evidence hashes, policy version, latency, response status, correlation, and linked incident or approval.

Decision gate

Do validated inputs satisfy a single allow path with no hard-stop, deny, or hold condition?

Proceed

Return allow with explicit obligations and expiry.

Hold or stop

Return hold, deny, or kill with missing evidence and containment guidance.

Evidence to retain

  • normalized runtime event
  • ordered rule trace
  • decision/reason receipt

Expected outputs

  • runtime decision JSON
  • approval/evidence request
  • containment signal

What was added

The evaluator lives in the runtime surface, not just the docs:

  • CI checks that exercise allow, deny, hold, and kill decisions against the checked-in gateway policy.

Evaluate a declared read-only tool call when ACS Guardian evidence is unspecified:

python3 scripts/evaluate_mcp_gateway_decision.py \
  --workflow-id vulnerable-dependency-remediation \
  --agent-id sr-agent::vulnerable-dependency-remediation::codex \
  --agent-class codex \
  --run-id run-ci \
  --tool-namespace advisories.vulnerability \
  --tool-access-mode read \
  --gate-phase tool_call \
  --expect-decision allow

Evaluate an ACS Guardian timeout that proceeded fail-open:

python3 scripts/evaluate_mcp_gateway_decision.py \
  --workflow-id vulnerable-dependency-remediation \
  --agent-id sr-agent::vulnerable-dependency-remediation::codex \
  --agent-class codex \
  --run-id run-acs-fail-open \
  --tool-namespace advisories.vulnerability \
  --tool-access-mode read \
  --gate-phase tool_call \
  --guardian-decision-status timeout \
  --on-decision-failure proceed \
  --expect-decision kill_session

Evaluate the same timeout with ACS fail-closed posture:

python3 scripts/evaluate_mcp_gateway_decision.py \
  --workflow-id vulnerable-dependency-remediation \
  --agent-id sr-agent::vulnerable-dependency-remediation::codex \
  --agent-class codex \
  --run-id run-acs-fail-closed \
  --tool-namespace advisories.vulnerability \
  --tool-access-mode read \
  --gate-phase tool_call \
  --guardian-decision-status timeout \
  --on-decision-failure deny \
  --expect-decision deny

Evaluate a Streamable HTTP tools/call whose headers match the body:

python3 scripts/evaluate_mcp_gateway_decision.py \
  --workflow-id vulnerable-dependency-remediation \
  --agent-id sr-agent::vulnerable-dependency-remediation::codex \
  --agent-class codex \
  --run-id run-header-match \
  --tool-namespace advisories.vulnerability \
  --tool-access-mode read \
  --gate-phase tool_call \
  --mcp-method tools/call \
  --mcp-name advisories.vulnerability.read \
  --mcp-protocol-version 2026-07-28 \
  --jsonrpc-method tools/call \
  --jsonrpc-name advisories.vulnerability.read \
  --jsonrpc-protocol-version 2026-07-28 \
  --expect-decision allow

Evaluate a gateway-routable tools/list header whose body is tools/call:

python3 scripts/evaluate_mcp_gateway_decision.py \
  --workflow-id vulnerable-dependency-remediation \
  --agent-id sr-agent::vulnerable-dependency-remediation::codex \
  --agent-class codex \
  --run-id run-header-mismatch \
  --tool-namespace advisories.vulnerability \
  --tool-access-mode read \
  --gate-phase tool_call \
  --mcp-method tools/list \
  --jsonrpc-method tools/call \
  --expect-decision deny

Evaluate a host-app Origin that matches the Streamable HTTP allowlist:

python3 scripts/evaluate_mcp_gateway_decision.py \
  --workflow-id vulnerable-dependency-remediation \
  --agent-id sr-agent::vulnerable-dependency-remediation::codex \
  --agent-class codex \
  --run-id run-origin-allow \
  --tool-namespace advisories.vulnerability \
  --tool-access-mode read \
  --gate-phase tool_call \
  --origin https://mcp.security-recipes.ai \
  --allowed-origin https://mcp.security-recipes.ai \
  --expect-decision allow

Evaluate a DNS-rebinding Origin that does not match the allowlist:

python3 scripts/evaluate_mcp_gateway_decision.py \
  --workflow-id vulnerable-dependency-remediation \
  --agent-id sr-agent::vulnerable-dependency-remediation::codex \
  --agent-class codex \
  --run-id run-origin-deny \
  --tool-namespace advisories.vulnerability \
  --tool-access-mode read \
  --gate-phase tool_call \
  --origin https://attacker.example \
  --allowed-origin https://mcp.security-recipes.ai \
  --expect-decision deny

The response includes the decision, matched workflow, matched scope, violations, approval state, source manifest hash, and observed runtime attributes. That output can be attached to PR evidence, MCP gateway logs, or red-team transcripts.

Decision model

The evaluator fails closed:

  1. Unknown workflow IDs return deny.
  2. Inactive workflows return deny.
  3. Missing runtime attributes return deny.
  4. Undeclared agent classes return deny.
  5. Undeclared namespace and access-mode pairs return deny.
  6. Unknown gate phases return deny.
  7. Branch writes must use the workflow branch prefix and declared file scope.
  8. Forbidden paths beat allowed paths.
  9. Approval-required scopes return hold_for_approval until a typed approval record is present.
  10. Runtime kill signals return kill_session before ordinary allow or deny checks.
  11. An ACS Guardian decision failure (timeout, transport_failure, or error_without_decision) with fail-open on_decision_failure (proceed, or the ACS default when the posture is omitted) returns kill_session. The same failure with on_decision_failure=deny returns deny so the tool is not executed. Unspecified Guardian evidence stays on the prior allow path.
  12. Observed Streamable HTTP MCP-Protocol-Version, Mcp-Method, or Mcp-Name values that disagree with the JSON-RPC body return deny (HeaderMismatch, -32020). Unspecified headers stay on the prior allow path so stdio and CI admission checks stay valid.
  13. Observed Streamable HTTP Origin that is null, a non-http(s) scheme, a URL with a path, missing from the allowlist, or present without an allowlist returns deny (HTTP 403). Unspecified Origin stays on the prior allow path so stdio and CI admission checks stay valid.

The important design choice is that the evaluator does not ask the model to decide whether a call is safe. The model requests a tool call; the policy layer decides.

Industry alignment

This is the practical enforcement layer implied by current AI security guidance:

  • MCP Streamable HTTP requires servers to validate Origin on every incoming connection and to reject an invalid header with HTTP 403. A browser page whose DNS rebinds onto a local MCP server otherwise inherits the server’s network position. The same page requires MCP-Protocol-Version, Mcp-Method, and Mcp-Name on every POST so intermediaries can route without parsing the body. Servers that process the body MUST reject header/body disagreement with HTTP 400 and JSON-RPC HeaderMismatch (-32020). Otherwise a gateway can authorize tools/list from the header while the server executes tools/call from the body. This evaluator denies both the invalid Origin and the header/body split.
  • OWASP Agent Control Standard (ACS-Core §6.4, donated 2026-09-01) requires the Observed Agent to wait for and apply a Guardian decision. The ACS default on_decision_failure posture is proceed. An adversary who can disrupt that channel converts control into audit unless the runtime fails closed. This evaluator kills fail-open proceeds and denies fail-closed failures instead of executing the tool.
  • OWASP MCP Top 10 calls out MCP authentication, authorization, audit telemetry, command execution, shadow servers, and over-sharing risks. A deterministic evaluator gives gateways a repeatable authorization and audit decision for each call.
  • OWASP Top 10 for Agentic Applications centers tool misuse, goal hijacking, identity abuse, and rogue agent behavior. The evaluator restricts each action to workflow, identity, namespace, gate phase, and scope.
  • NIST AI RMF frames trustworthy AI around govern, map, measure, and manage. Runtime decisions turn mapped policy into measurable control evidence.
  • NIST AI RMF Generative AI Profile emphasizes lifecycle governance and measurement for generative AI systems. The evaluator creates a repeatable measurement point for agentic tool use.
  • CISA Secure by Design pushes secure defaults, transparency, and customer security outcomes. Default deny plus auditable reasons is the secure default for agent actions.

Enterprise checklist

Before using the evaluator in production:

  • Pin the gateway policy artifact by source_manifest.sha256.
  • Require every tool call to include workflow_id, agent_id, run_id, namespace, access mode, gate phase, and any write scope.
  • Store every non-allow decision in the MCP gateway audit log.
  • Treat hold_for_approval records as typed approvals, not chat messages.
  • Attach the evaluator response to PR evidence for branch writes.
  • Alert on repeated denials, forbidden path attempts, and kill-session events.

See also