RENKER

Guardrails outside the model

Sebastian Renker · 11 October 2026 · 5 min read

An AI agent that can read files, write files and send messages has capabilities. A prompt injection is text that tries to talk the agent into using those capabilities for someone else. If the only thing between the injected text and the action is the model's own judgement, the attacker is arguing with the guard in the guard's own language.

The problem with guardrails inside the model

A rule such as "never read the secrets folder" written into a system prompt lives in the same channel as the attacker's text. The model weighs both. Sometimes it holds, sometimes it does not, and you cannot prove which. A defence that is probabilistic by construction is a poor fit for a question that has a yes-or-no answer: is this actor allowed to do this action on this resource, right now?

Move the decision out

The idea is small. The agent may request anything. A separate function, which never sees the agent's reasoning, decides whether the request runs. That function reads only trusted inputs: the actor, the action, the target, and a store of capabilities that a human granted. It returns one of three answers:

Prompt injection changes the request. It does not change the grant, so it does not change the decision.

What that looks like

renker-core-authz is a standard-library-only Python implementation of this idea, with a hash-chained audit log. A least-privilege grant is one action verb plus one path scope, and it expires:

from datetime import datetime, timedelta, timezone
from renker_core_authz import (
    Actor, Capability, CapabilityStore, PathScope, GuardedFilesystem,
)
from renker_core_authz.audit import AuditLog

now = datetime.now(timezone.utc)
store = CapabilityStore()
store.grant(
    Capability(
        capability="filesystem.write",
        scope=PathScope(base="~/Documents/drafts"),
        granted_to="agent:session-1",
        granted_by="human:owner",
        issued_at=now,
        expires_at=now + timedelta(hours=1),
    )
)

guard = GuardedFilesystem(store, AuditLog("audit.log"))
result = guard.write(Actor("agent", "session-1"), "~/Documents/drafts/note.txt", "hi")
print(result.decision.value, result.reason)  # ALLOW / DENY with an explainable reason

Watching an injection fail

The companion repository renker-agent-demo runs a scripted scenario with no LLM, no network and no API key. A file the agent reads contains hidden instructions to exfiltrate a secret. The gateway still decides from the grants:

StepRequestDecision
legit read taskworkspace/task.txtALLOW
legit write draftworkspace/drafts/summary.txtALLOW
injected read secretsecrets/api_key.txt (out of scope)DENY
injected traversal writedrafts/../../secrets/stolen.txtDENY
injected send exfiloutbox/exfil.emlREQUIRE_APPROVAL, human declines
legit send replyoutbox/reply.emlREQUIRE_APPROVAL, human approves

None of those outcomes depended on the agent behaving well. The demo also edits one audit entry on disk and shows that verify() catches it.

What this does not do

This is a prototype-grade library, not an externally audited product. It does not provide cryptographic authentication of actors (identity is validated, not authenticated), protection once the host account is fully compromised, enforcement for actions that are not routed through it, an externally anchored audit, or multi-process audit coordination. The audit log is tamper-evident, not immutable: an attacker who can rewrite both the log and its head anchor can still forge history.

The boundary only helps if every side effect goes through it. A tool that touches the filesystem directly, bypassing the gateway, is not covered. That is a property of the design, and it is stated up front rather than discovered later.

Try it

pip install -e .
python scripts/run_demo.py

Source and issues: renker-core-authz (Apache-2.0) and renker-agent-demo. Feedback on the threat model is the most useful thing you can send.