Guardrails outside the model
An AI agent that can read files, write files and send messages has capabilities. A prompt injection is text that tries to talk the agent into using those capabilities for someone else. If the only thing between the injected text and the action is the model's own judgement, the attacker is arguing with the guard in the guard's own language.
The problem with guardrails inside the model
A rule such as "never read the secrets folder" written into a system prompt lives in the same channel as the attacker's text. The model weighs both. Sometimes it holds, sometimes it does not, and you cannot prove which. A defence that is probabilistic by construction is a poor fit for a question that has a yes-or-no answer: is this actor allowed to do this action on this resource, right now?
Move the decision out
The idea is small. The agent may request anything. A separate function, which never sees the agent's reasoning, decides whether the request runs. That function reads only trusted inputs: the actor, the action, the target, and a store of capabilities that a human granted. It returns one of three answers:
- ALLOW: the request matches a grant.
- DENY: it does not.
- REQUIRE_APPROVAL: it needs a human in the loop first.
Prompt injection changes the request. It does not change the grant, so it does not change the decision.
What that looks like
renker-core-authz is a standard-library-only Python implementation of this idea, with a hash-chained audit log. A least-privilege grant is one action verb plus one path scope, and it expires:
from datetime import datetime, timedelta, timezone
from renker_core_authz import (
Actor, Capability, CapabilityStore, PathScope, GuardedFilesystem,
)
from renker_core_authz.audit import AuditLog
now = datetime.now(timezone.utc)
store = CapabilityStore()
store.grant(
Capability(
capability="filesystem.write",
scope=PathScope(base="~/Documents/drafts"),
granted_to="agent:session-1",
granted_by="human:owner",
issued_at=now,
expires_at=now + timedelta(hours=1),
)
)
guard = GuardedFilesystem(store, AuditLog("audit.log"))
result = guard.write(Actor("agent", "session-1"), "~/Documents/drafts/note.txt", "hi")
print(result.decision.value, result.reason) # ALLOW / DENY with an explainable reason
Watching an injection fail
The companion repository renker-agent-demo runs a scripted scenario with no LLM, no network and no API key. A file the agent reads contains hidden instructions to exfiltrate a secret. The gateway still decides from the grants:
| Step | Request | Decision |
|---|---|---|
| legit read task | workspace/task.txt | ALLOW |
| legit write draft | workspace/drafts/summary.txt | ALLOW |
| injected read secret | secrets/api_key.txt (out of scope) | DENY |
| injected traversal write | drafts/../../secrets/stolen.txt | DENY |
| injected send exfil | outbox/exfil.eml | REQUIRE_APPROVAL, human declines |
| legit send reply | outbox/reply.eml | REQUIRE_APPROVAL, human approves |
None of those outcomes depended on the agent behaving well. The demo also edits one audit entry on disk and shows that verify() catches it.
What this does not do
This is a prototype-grade library, not an externally audited product. It does not provide cryptographic authentication of actors (identity is validated, not authenticated), protection once the host account is fully compromised, enforcement for actions that are not routed through it, an externally anchored audit, or multi-process audit coordination. The audit log is tamper-evident, not immutable: an attacker who can rewrite both the log and its head anchor can still forge history.
The boundary only helps if every side effect goes through it. A tool that touches the filesystem directly, bypassing the gateway, is not covered. That is a property of the design, and it is stated up front rather than discovered later.
Try it
pip install -e .
python scripts/run_demo.py
Source and issues: renker-core-authz (Apache-2.0) and renker-agent-demo. Feedback on the threat model is the most useful thing you can send.