Context Poisoning Guard
Why this page exists. A secure context layer cannot only hash context. It has to inspect context for instruction-like payloads before that context is returned to an agent through MCP.
Rechecked October 4, 2026 against OWASP
LLM01:2026 Prompt Injection
in the
GenAI LLM Top 10 2026
edition (published 2026-08-04). Common Example #5 is invisible-character
injection and exfiltration. Mitigation #5 says to strip tag-block
(U+E0000 to U+E007F), variation-selector (U+FE00 to U+FE0F),
and zero-width (U+200B, U+200C, U+200D, U+2060) characters at
every ingest and render boundary. This pack already scanned zero-width
and bidi controls; it now also flags tag-block characters and runs of
three or more variation selectors. A single variation selector used as
an emoji modifier stays on the prior allow path. This pass does not
claim human review of the pack.
The product bet
SecurityRecipes is positioned as the secure context layer for agentic AI. The strongest enterprise version of that idea is not a recipe catalog. It is a controlled context supply chain:
- registered source roots,
- owners and trust tiers,
- retrieval decisions,
- source hashes,
- poisoning controls,
- and deterministic inspection before context reaches an agent.
The Context Poisoning Guard adds that inspection layer. It scans every registered context root from the Secure Context Registry and produces a generated evidence pack that says whether a source passes, contains only documented adversarial examples, should hold for review, or should be blocked until fixed. Leftover CVE drafts that invent a next version or copy unrelated floors are untrusted context: cite NVD or GHSA, and do not treat that text as a retrieval-approved floor.
Workflow at a glance
Context Poisoning Guard workflow
Detect hidden instructions, invisible Unicode smuggling, exfiltration cues, authority spoofing, malicious markup, and trust-boundary violations before context reaches an agent.
Signal
Receive candidate context
Capture source, owner, retrieval path, hash, content type, workflow, trust tier, transformations, and destination.
Scope
Normalize safely
Decode supported containers and markup, expose hidden text/links/metadata, bound size and recursion, and preserve original hashes.
Decision
Scan poisoning signals
Detect instruction override, secret requests, tool abuse, authority claims, data exfiltration, encoded payloads, invisible Unicode tag-block or variation-selector smuggling, and cross-boundary redirects.
Action
Apply trust policy
Allow clean context, strip untrusted instructions, restrict to quoted evidence, quarantine, hold, deny, or kill active attacks.
Proof
Record guard evidence
Write source/normalized hashes, matched signals, transformations, trust decision, affected workflows, and owner action.
Decision gate
Is context provenance intact and free of instructions or payloads that violate its declared evidence role?
Return only the approved normalized context with citations.
Quarantine, deny, or kill malicious, unprovenanced, secret-seeking, or boundary-breaking context.
Evidence to retain
- source provenance and hashes
- normalized scan findings
- guard decision and transformations
Expected outputs
- approved context object
- quarantine report
- runtime kill signal
What was added
- Source profile:
data/assurance/context-poisoning-guard-profile.json - Generator:
scripts/generate_context_poisoning_guard_pack.py - Evidence pack:
data/evidence/context-poisoning-guard-pack.json - MCP tool:
recipes_context_poisoning_guard_pack
Regenerate and validate the pack:
python3 scripts/generate_context_poisoning_guard_pack.py
python3 scripts/generate_context_poisoning_guard_pack.py --check
What it scans
| Rule | Severity | Why it matters |
|---|---|---|
| Direct instruction override | Critical | Detects text that asks an agent to ignore or override higher-priority instructions. |
| Secret exfiltration request | Critical | Detects transfer language near secrets, tokens, credentials, private keys, or environment dumps. |
| Approval bypass request | High | Detects requests to skip, bypass, remove, or disable review, approval, policy, CI, or guardrails. |
| Hidden HTML instruction | High | Detects hidden HTML/comment patterns that may evade human review but remain visible to models. |
| External callback instruction | High | Detects send/post/upload/callback language near external URLs. |
| Encoded payload | Medium | Detects long base64-like strings that may hide instructions or data. |
| Zero-width control | Medium | Detects zero-width and bidirectional controls that can hide or reorder text. |
| Invisible Unicode smuggling | High | Detects tag-block characters (U+E0000 to U+E007F) and runs of three or more variation selectors (U+FE00 to U+FE0F) that reviewers never see. |
The guard is intentionally conservative. It does not pretend regexes can solve prompt injection. It creates evidence and routing:
passwhen no markers are detected.allow_with_adversarial_exampleswhen markers appear only in documented red-team, threat-model, or defensive examples.hold_for_context_reviewwhen normal guidance contains high-risk markers.block_until_removedwhen critical actionable findings appear outside approved examples.
Why this is enterprise-grade
This feature makes AI easier for reviewers because it turns a hard question into a simple artifact:
Can this context be returned to an agent?
An MCP server, AI platform intake workflow, or procurement reviewer can ask the guard pack for source-level decisions and findings instead of reading every page manually. The answer carries source ID, path, line, rule ID, severity, disposition, and source hash.
The generated pack supports:
- recipe publication review,
- MCP server intake,
- quarterly secure-context recertification,
- red-team replay planning,
- trust review diligence,
- and future hosted context monitoring.
MCP examples
Get the portfolio-level summary:
{}
Get all sources held for context review:
{
"decision": "hold_for_context_review"
}
Get actionable critical findings for one source:
{
"source_id": "recipes",
"severity": "critical",
"actionable_only": true
}
Get all direct instruction override matches:
{
"rule_id": "direct-instruction-override"
}
Get invisible Unicode smuggling matches:
{
"rule_id": "invisible-unicode-smuggling"
}
Industry alignment
The guard follows current agentic AI and MCP security guidance:
- OWASP LLM01:2026 Prompt Injection for invisible-character injection, including tag-block ASCII smuggling and variation-selector runs that stay hidden in review UIs.
- OpenAI guidance on prompt injection resistance for treating prompt injection as an impact-limiting problem, not only a string-filtering problem.
- OWASP MCP Tool Poisoning for the risk of hidden or malicious instructions in MCP tool metadata and runtime context.
- OWASP Agentic AI Threats and Mitigations for agent threat models around autonomy, tools, delegation, and retrieved context.
- MCP Security Best Practices for scoped access, token-safety, confused-deputy prevention, and auditability.
- NIST AI RMF Generative AI Profile and CISA AI Data Security guidance for AI data provenance, integrity, monitoring, and lifecycle controls.
See also
- Secure Context Trust Pack for registered context roots and hashes.
- Secure Context Firewall for runtime retrieval decisions.
- Context Egress Boundary for outbound data-boundary decisions after retrieval.
- Agentic Red-Team Drill Pack for adversarial examples that should stay labeled as test payloads.
- Agentic Threat Radar for source-backed prioritization.