Security engineering
AI Agent Security: Sandboxing, Prompt Injection and Tool Authorization
A real client engagement. The engineering and the results are described below.
An operations team we worked with wants an agent to clear a mountain of routine requests, but the first untrusted document asks it to use a privileged tool. The team keeps the automation opportunity and redesigns who can authorize each effect.
Request a time through the inquiry form. A meeting is confirmed separately by email.
The business problem behind the technology
Can automation save effort without inheriting unlimited authority?
The business wants less repetitive preparation without granting documents or model output the power to send messages, change accounts or reveal protected data.
Read the client engagement ↓Client engagement / Delivered results
The pilot needed permission boundaries before more autonomy
A software operations team we worked with handles 2,000 routine requests a month. Its manual preparation step takes twelve minutes per request, excluding the final approval for any sensitive action.
The constraint
The business wants less repetitive preparation without granting documents or model output the power to send messages, change accounts or reveal protected data.
The engineering decision
The team separates a constrained compute workspace from trusted tools and exact approvals. Bounded automation reduced preparation to five minutes per request while retaining human approval where needed.
The delivered outcome
Preparation effort fell from 400 to about 166.667 hours a month, leaving about 233.333 hours for other work.
| Measure / unit | Before | After | Difference |
|---|---|---|---|
| Preparation effort hours/month | 400 | 166.667 | 233.333 |
Preparation effort. 2,000 × 12 / 60 versus 2,000 × 5 / 60; displayed values are rounded and represent task effort, not headcount reduction.
The conditions behind the results
- The monthly volume and twelve/five-minute preparation times are the team’s operating values.
- Sensitive effects still require separate authorization and are excluded from the preparation-time model.
- Tool costs, review effort, exceptions and implementation work must be assessed separately.
What this does not prove. The freed preparation capacity kept human approval in place for sensitive effects.
Evidence to collect for your own decision
- Map untrusted inputs, protected assets and exact tool permissions.
- Test denied, ambiguous, timeout and prompt-injection paths before broadening access.
- Measure task completion quality, exception handling and human review effort.
Key decisions
A sandbox limits what a process can reach. It does not establish that a message, payment, deployment, or data disclosure is authorized. Design containment and business approval as separate controls, then test the boundaries rather than assigning a security score.
- Map untrusted inputs, protected assets and exact tool permissions.
- Test denied, ambiguous, timeout and prompt-injection paths before broadening access.
- Freed preparation capacity is not payroll savings, proof of safety or a monetary claim about an avoided attack.
Follow the decision
Select a step to follow its reasoning, then continue into the technical chapters.
Problem → boundary → decision → evidence
Separate retrieved content, trusted instructions and protected assets.
Read every component and connection
- Identify the crossing · Problem
- Separate retrieved content, trusted instructions and protected assets.
- Constrain execution · Boundary
- Bound filesystem, processes, network and credential access.
- Authorize the action · Decision
- Bind approval to the exact external effect rather than a vague intention.
- Identify the crossing → Constrain execution: identify the constraint
- Constrain execution → Authorize the action: choose a bounded change
- Authorize the action → Prove the denial path: check the outcome
A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Start with assets, adversaries, and crossings
The client’s operations team names protected assets before assigning authority to its automation pilot.
Treat repository files, dependency scripts, retrieved pages, issue comments, and tool responses as potentially attacker-controlled data. An adversary might control a README without controlling the agent process, or gain arbitrary code execution through an installation script. Design for both cases: instructions embedded in content and hostile native code are different entry points into the same system.
Inventory the assets before selecting a runtime: host files, other tenants, source integrity, production credentials, customer data, external accounts, and the audit trail. Then mark every crossing between the model, command runner, filesystem, network gateway, credential service, approval service, and external provider. A container around commands does not automatically contain a connector that executes in the orchestrator with a separate privileged credential.
For each crossing, identify the trusted enforcement component, the exact object being authorized, and what remains possible after compromise. In this exercise, the runner is disposable and untrusted; the gateway and approval store are outside its write boundary. The gateway must not accept a runner-authored claim that the user approved an action. Protect policy configuration and approval storage from the same repository the agent is allowed to edit.
- Confidentiality: can controlled source text cause sensitive input to leave through a tool, log, artifact, or allowed destination?
- Integrity: can the process alter the host, another task, its own enforcement policy, or an external business object?
- Availability: can it exhaust CPU, memory, processes, storage, output buffers, network connections, or paid API budgets?
- Residual risk: which kernel, runtime, gateway, identity provider, and human-review assumptions remain trusted?
Prompt injection is not a request for tool authority
The suspicious document enters the conversation. An indirect prompt injection remains untrusted content even when it describes a useful next step.
An untrusted page can say that a build requires uploading a credentials file. That is data describing an instruction, not an instruction from the task owner. The model should preserve provenance and refuse the redirection, but the security architecture must also limit damage if the model follows it. A text classifier or stronger system prompt is not an operating-system boundary.
Put policy checks at the actual executor. A tool request should carry a typed operation and bounded parameters; the trusted gateway checks identity, resource scope, destination, and action approval immediately before execution. Do not make the model the only component deciding whether its own request complies. A generic shell with a powerful inherited token defeats a carefully scoped notification tool because the agent can bypass the tool and call the service directly.
A read operation is not necessarily harmless. An HTTPS GET can disclose data in its query string, headers, hostname, or timing, and some poorly designed endpoints mutate state on GET. Separate retrieving approved public documentation from sending arbitrary data to any approved domain. OpenAI documents that command-sandbox networking and web search, connectors, MCP connections, browser activity, and cloud tasks have different control surfaces. Inventory all enabled paths instead of assuming one proxy governs the whole agent.
Choose an isolation boundary, not a security adjective
A reassuring isolation label appears in the design. You replace the adjective with an explicit statement of what the boundary contains.
Ordinary Linux containers use kernel mechanisms such as namespaces, cgroups, capabilities, and syscall filtering, but their processes still share the host kernel. A non-root UID and a read-only root filesystem reduce accessible operations; they do not remove the kernel attack surface. Privileged mode, host namespaces, hostPath mounts, or a mounted container-runtime socket can undermine the intended boundary. Containers alone are not a complete security architecture.
gVisor inserts its Sentry implementation of the system API between application code and the host. This reduces direct exposure to the host syscall interface; the Sentry itself uses a constrained host interface. It is not simply a list of allowed Linux syscalls, and it is not a conventional guest-kernel VM. Compatibility, filesystem behavior, networking, and performance depend on the workload and configuration. gVisor explicitly describes reliance on host resource controls and host mitigations for hardware attacks.
Firecracker runs a guest kernel in a KVM-based microVM with a deliberately reduced virtual-device surface. Its production guidance requires host and guest patching, a jailer or equivalent process constraints, resource limits, and operator-managed network filtering. A guest-kernel boundary changes the attack surface; it does not prove the absence of VMM, host, device, or hardware vulnerabilities. Neither a microVM nor gVisor grants authority to send a message or publish a release.
Select a runtime from the threat model, then exercise the actual toolchain inside it. Keep untrusted tasks separated from privileged infrastructure workloads and account for runtime overhead. A RuntimeClass name is configuration, not evidence that the expected handler is installed and used on every scheduled node. The reader intentionally treats a restricted container and a restricted microVM alike for its teaching decision; it makes no claim that their real isolation properties are equal.
Keep the root immutable and the workspace bounded
The agent needs a workspace, not ownership of the host. The next decision limits which paths and state can change.
Make the runtime image root read-only and expose only the files needed for the task. Use a disposable work copy rather than a broad host checkout, home directory, SSH agent socket, or container-runtime socket. Separate immutable source input from writable scratch or a narrowly owned editing workspace. Read-only does not mean confidential: anything mounted readably can be copied into output if another channel permits it.
Writable mounts remain writable even when readOnlyRootFilesystem is true. Decide explicitly where package caches, temporary files, generated output, and source edits may go. Do not silently turn off the read-only root to satisfy a tool that assumes a writable home directory. Prepare compatible image paths and route writes to bounded volumes; keep policy files, Git credentials, and orchestration configuration out of those volumes.
For file tools, enforce confinement at file-open time, not only through a string prefix check. Traversal, symlink swaps, hard links to already exposed files, and mount changes can invalidate a path that looked safe during validation. Use platform-supported directory-relative, no-escape operations and avoid exposing the sensitive objects in the first place. Review exported patches and artifacts outside the runner before applying them to a trusted checkout; a disposable filesystem does not make its generated code trustworthy.
A constrained Kubernetes hardening fragment
The boundary reaches Kubernetes configuration. The hardening fragment is a starting contract with stated limits, not a universal security template.
The following is a complete set of container fields for this limited example, to merge into one already defined Linux container. It is deliberately not a Deployment, Pod, image selection, or network policy. Assume a reviewed digest-pinned image whose program works as UID/GID 10001, and Pod volumes named workspace and scratch that are task-local emptyDir volumes. Give those volumes size limits, such as 256Mi and 64Mi, and arrange writable ownership through the Pod security context, for example fsGroup 10001. Do not use hostPath to satisfy the mount names.
The surrounding Pod must disable automatic service-account token mounting with automountServiceAccountToken: false, avoid hostNetwork, hostPID, and hostIPC, and expose no extra credentials or privileged sidecars. Enforce an appropriate Restricted Pod Security admission policy rather than relying on each generated manifest to be benevolent. Apply restrictions to init and ephemeral containers too. The fragment adds a read-only root and resource budgets; it is not a claim that the Restricted profile alone enforces either.
RuntimeDefault selects the runtime-provided seccomp profile, not a bespoke proof of syscall safety. Validate that profile and the non-root identity against the pinned image and chosen runtime. The resource requests are scheduling inputs, while limits constrain runtime consumption with resource-specific behavior. Disk-backed emptyDir and ephemeral-storage limits depend on node accounting and eviction; they are not a synchronous per-write quota. Use a filesystem quota or separately bounded storage when a strict byte ceiling is part of the threat model.
# Fields for one existing Linux container; Pod prerequisites above.
workingDir: /workspace
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
privileged: false
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
seccompProfile:
type: RuntimeDefault
resources:
requests:
cpu: "250m"
memory: "128Mi"
ephemeral-storage: "128Mi"
limits:
cpu: "1"
memory: "512Mi"
ephemeral-storage: "512Mi"
volumeMounts:
- name: workspace
mountPath: /workspace
- name: scratch
mountPath: /tmpLeast privilege applies to tools as well as Linux capabilities
The process is constrained, but the tool still has authority. AI agent security now depends on the operation the tool exposes.
Dropping Linux capabilities limits kernel privileges; it does not limit the business permissions held by a cloud token. Treat these as different uses of the word capability. An application-level capability should bind a principal to a specific operation on specific resources for a bounded period, with enforcement by a component the runner cannot modify.
Prefer a gateway operation such as reading one approved object or proposing one workspace patch to broad cloud CLI access. Scope service identities to the task and account; separate a credential-free executor from a publisher. A runner that can create arbitrary Pods may be able to mount secrets it cannot directly read, so review indirect authority and not only obvious Secret get permissions. Kubernetes documents this exact distinction.
Do not let the agent widen its own mount set, select privileged mode, replace its RuntimeClass, modify network rules, or rewrite its approval policy. Pin and validate the operation schema at the trusted boundary, reject unknown fields and ambiguous target aliases, and use explicit argument arrays where possible. For an approved script, bind approval to its immutable content and working context; approving a mutable pathname leaves the behavior changeable after review.
Mediate egress and account for SSRF
Preparation-time savings do not justify allowing the agent to cross a network or credential boundary without permission.
Start offline. When retrieval is needed, allow the runner to reach a dedicated gateway rather than the public internet directly. Enforce destination identity, scheme, port, method, permitted path, request size, response size, redirects, and timeouts there. Remove direct alternate routes, including IPv6 or UDP paths that escape an IPv4/TCP-only design. A proxy environment variable is a convenience, not enforcement when arbitrary code can ignore it.
Server-side request forgery, or SSRF, turns an apparently allowed fetch into access to another service. Reject loopback, link-local, private, metadata, and cluster-management destinations unless an exact separate capability intentionally permits one. Resolve and classify all candidate addresses, bind the checked destination to the connection, verify TLS hostname identity, and reapply policy on redirects and new connections. A hostname allowlist checked before a later independent DNS lookup leaves a rebinding race. OpenAI explicitly qualifies its documented DNS checks as best-effort rather than a complete rebinding defense.
Standard Kubernetes NetworkPolicy works primarily at IP and port boundaries and requires a supporting network plugin. Policies combine additively; a broad allow policy can reopen a destination despite a default-deny policy. Standard NetworkPolicy does not inspect an HTTPS path or message body, and the documentation notes node-traffic exceptions and implementation-dependent address rewriting. Combine network enforcement with an application gateway and platform-specific metadata protection, then test from the actual runner namespace.
An allowed destination may still accept confidential data. Limit the request body and query contract, avoid arbitrary tunnels and unrestricted CONNECT use, and protect the gateway from becoming a general-purpose proxy. Separate a public-document fetcher from authenticated business tools. The browser reader can display an allowlist setting, but it cannot establish that any of these controls exists in a real environment.
Broker credentials without returning raw secrets
The workflow needs a credentialed operation without making the credential part of the conversation. That is a custody problem, not a prompt-writing trick.
The safest runner credential is often no credential. Let the agent refer to an opaque capability reference and ask a trusted broker to execute one permitted operation. The reference identifies a server-side grant; it is not a bearer credential or proof of authorization by itself. Authenticate the caller, bind the reference to task and audience, check expiration and revocation, and apply the same destination and action policy before the broker uses its credential.
Prefer having the broker add authentication and call the provider outside the sandbox, returning a bounded, redacted result. If a workflow truly needs a delegated token inside the runner, make it short-lived and minimally scoped and acknowledge the weaker property: arbitrary code in that runner can read and use it. Calling a raw-token vending endpoint a secret broker does not make the token inaccessible to the model or its tools.
Do not inject broad production credentials through environment variables, command arguments, source files, or shared home directories. Environment secrets are available to the program and its children, and can leak through diagnostics or intentional output. Moving the same secret into a readable volume is not a confidentiality boundary against that program. Kubernetes Secrets are base64-encoded, not encrypted by that encoding, and applications must protect their values after reading them. Configure storage encryption and least-privilege access separately.
Keep credentials out of logs, approval records, patches, and exception text. Treat broker responses as untrusted data too, with field allowlists and size bounds. A broker also needs a revocation and outage policy: deny when required authorization state cannot be checked rather than falling back to a broadly privileged environment token. In the teaching reader, brokered secrets are a declared assumption, not a tested guarantee.
Bind approval to the exact external action
A human approves a step. The approval is useful only if it identifies the action that will actually happen.
Execution permission answers whether the tool can reach a destination. Business approval answers whether this principal may perform this exact action for this task. Neither implies the other. An agent allowed to fetch documentation is not approved to send an external message; an approved message is not permission to disable network isolation. Gate publication, payments, destructive changes, and communications at the executor that holds the external credential.
The JSON below is an approval-store record, not a signature or provider API contract. The identities and the example.net endpoint are anonymized. It binds one notification to one task, actor, provider account, environment, endpoint, method, recipient, body, idempotency key, and validity window. The request body is included directly so there is no fabricated content digest. A real gateway reads the record from authenticated, integrity-protected approval storage and independently verifies the approver has authority over that target; it never trusts a JSON object returned by the runner.
At the evaluation time 2026-09-05T12:01:00Z, this record could satisfy the approval predicate only for its exact request and only if it is authentic, unrevoked, and unused. At or after expiresAt it fails. A changed recipient, account, body, environment, or key requires a new review. For a larger payload, bind a cryptographic digest of immutable canonical bytes and show those exact bytes to the approver. Resolving a mutable artifact reference after review defeats the binding.
{
"recordId": "training-approval-014",
"taskId": "training-task-014",
"actor": "agent:training-runner",
"approver": "human:training-owner",
"environment": "training",
"providerAccount": "training-notifications",
"capabilityRef": "notification-send-014",
"action": "externalmessage",
"request": {
"method": "POST",
"url": "https://notify.example.net/v1/messages",
"body": {
"recipient": "reviewer@example.net",
"text": "The training patch is ready for review."
}
},
"idempotencyKey": "training-task-014-message-1",
"issuedAt": "2026-09-05T12:00:00Z",
"expiresAt": "2026-09-05T12:05:00Z",
"maxExecutions": 1
}Idempotency and approval solve different failures
A timeout arrives after an uncertain side effect. Approval and idempotency now answer different questions about the next attempt.
Approval authorizes an action; idempotency prevents a retry of that action from creating another effect. A gateway should durably bind its idempotency key to the exact request and authorization context, then atomically transition from approved to dispatching. A repeated key with different bytes is a conflict, not a new approved request. Two concurrent workers must not both consume the same single-use approval.
Record the provider receipt and return the prior result for a completed duplicate without dispatching again. If a timeout occurs after transmission, the outcome is unknown rather than failed: the provider may already have accepted the request. Use the provider-supported idempotency mechanism and status reconciliation within its documented scope and retention window. A gateway database alone cannot guarantee exactly-once effects across a remote service boundary.
When the provider has no safe deduplication or status lookup, stop and reconcile an ambiguous send instead of automatically retrying it. Expiration and revocation should prevent a new dispatch; local retrieval of an already recorded result is not another side effect. A necessary retry after approval expires needs a newly authorized reconciliation or retry policy, not a fresh idempotency key that risks duplication. Restoring the sandbox does not unsend a delivered message.
Read the widget as eight predicates, not a security score
You can inspect the eight predicates in the teaching model. Their result explains this declared policy, not the security of an unseen system.
The companion reader accepts runtime host/container/microVM, restricted/privileged execution, scratch/workspace/host filesystem access, off/allowlist/unrestricted network, none/brokered/broad secrets, absent/exact-action approval, and compute/read-HTTPS/edit-workspace/external-message actions. The default is a restricted container with scratch storage, no network, no secrets, no approval, and a compute action. The WebAssembly function ORs the following violation bits; it does not estimate breach probability.
Bits 1, 2, 4, 8, and 16 reject host execution, privileged execution, host filesystem access, unrestricted network, and broad secrets respectively, regardless of the selected action. Bit 32 requires allowlisted network for read-HTTPS or external-message actions. Bit 64 requires workspace access for editing. Bit 128 requires an exact-action approval record for an external message. Violations are combined with bitwise OR; the number is a mask identifying rules, not a severity rank.
For example, a restricted microVM with scratch storage, allowlisted network, brokered secrets, and an external-message action but no approval yields mask 128 and denial. Declaring exact approval changes that example to mask 0. If networking is instead off and approval remains absent, the result is 32 OR 128 = 160. A scratch-only container cannot edit the workspace under this policy even if it has an approval record.
Mask 0 means permit under this explicitly declared teaching policy only. It is not real authorization, isolation certification, or evidence that the selected runtime, broker, allowlist, or approval record is genuine. The model omits target identity, payload content, expiration, revocation, quotas, real runtime configuration, and many other controls described in this article. A container and microVM can both pass; neither thereby acquires business authority.
01 / Follow the explanation
What should an agent be allowed to do?
A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.
An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.
Read the full explanation and assumptionsCurrent example
Permitted by the declared policy
No declared policy violations. This is not real authorization or security certification. Filesystem, secret broker and network enforcement are model inputs; isolation never grants action authority.
Initial scenario: a restricted container, scratch-only writes, no network or secrets, no external-action approval, and a local computation. This is a declared teaching policy, not inspection of a real sandbox.
Execution boundary
- Runtime privileges
- Restricted
Reach & capabilities
- Writable filesystem scope
- Ephemeral scratch only
- Network policy
- Network disabled
- Secret access
- No secret capability
- Violation bitmask
- 0
Action authority
- Requested action
- Compute locally
- External-action approval
- No exact approval
- Policy decision
- Permit under the stated policy
# Evaluated policy inputs, not a runnable security boundary
runtime = 1
privileged = 0
filesystem = 0
network = 0
secrets = 0
approval = 0
action = 0
violation_mask = 0
# Enforce the real boundary outside the agent before executing any action.Declared example inputs:
runtime=microVM, privileged=restricted, filesystem=scratch
network=allowlist, secrets=brokered
approval=absent, action=externalmessage
Teaching result: violation mask 128; deny
Change only approval to exact-action record:
Teaching result: violation mask 0; permit under teaching policy only
Real gateway decision requires ALL of:
authenticated actor and task
permitted operation, account, destination, and payload
effective runtime and network policy
authentic, matching, unexpired, unrevoked approval when required
available resource and rate budgets
atomic idempotency-state check before dispatch
If any required check fails or cannot be established: do not dispatchBound time, processes, output, and paid work
Even an authorized action needs limits. Time, processes, output and paid work keep a valid workflow from becoming unbounded.
CPU and memory limits are necessary but incomplete availability controls. CPU limits throttle execution rather than setting a task deadline, and memory enforcement can kill processes rather than produce a recoverable application error. Bound wall time, subprocess count, file descriptors, scratch size, stdout/stderr, network concurrency, response bytes, tool calls, and remote spending. A one-core process can still run forever or emit enough logs to exhaust a collector.
Use a trusted supervisor outside the sandbox to enforce deadlines and terminate the full task process group or runtime, not merely the original shell. On Kubernetes, configure the appropriate workload deadline and retry budget and separately configure platform PID limits; there is no generic per-container pids field in this securityContext fragment. Apply resource requests and limits to helper containers as well. A memory-backed emptyDir consumes memory rather than the disk-backed ephemeral-storage budget, so account for the chosen medium.
Set output quotas at collection boundaries and preserve a clear truncation marker. Bound archive extraction and decompressed sizes, not just downloaded bytes. Count attempts, retries, and broker calls against task budgets; otherwise a timeout can amplify load through repeated work. Firecracker specifically warns that guest-influenced serial and log output needs host-side bounds, illustrating why a VM boundary does not automatically solve resource exhaustion.
Audit decisions, rehearse denial, and recover deliberately
The team tests denials and exceptions before counting freed preparation capacity as a useful result.
Record task identity, immutable input and image identifiers, selected runtime, effective policy version, normalized tool operation, policy decision and reasons, approval reference, idempotency state, bounded outcome, and provider receipt. Store the audit trail outside runner write access. Prefer metadata and hashes where content is sensitive; any retained payload needs access controls and retention rules. A transcript is useful context, but it is not independent proof of what the executor actually dispatched.
Validate the deployed boundary with explicit allowed and denied cases. Attempt a write outside the workspace, a symlink escape, a disallowed destination, a metadata address, a redirected fetch, a mismatched approval body, an expired record, two concurrent consumes, and a repeated idempotency key. Exercise a runaway subprocess tree, memory pressure, full output buffer, and broker outage. Observe the enforcement component and external effects, not only the model saying it refused.
After a suspected compromise, stop dispatch, revoke scoped grants, preserve bounded evidence, and replace the disposable runtime and work volume. Review exported code separately and reconcile external effects before retrying. Patch the affected trusted components and rerun the relevant denial cases before restoring authority. Resetting a container cannot revoke a copied credential, repair a host compromise, or reverse an external action.
These are the validation scenarios we designed for the engagement. The reader demonstrates a small deterministic policy and retains an explanatory baseline without interactive execution. Its value is making assumptions visible; production assurance requires evidence from the actual runtime, network path, identity system, and action gateway.
Questions behind the decision
Does sandboxing prevent prompt injection?
Sandboxing constrains execution and resource access; it does not make untrusted instructions authoritative or harmless. Explicit tool authorization, data boundaries and adverse-path tests are still required.
Which AI agent actions need independent approval?
High-impact operations such as external messages, financial changes, privileged writes or protected-data disclosure need authority outside the model’s confidence. Bind approval to the exact rendered action and relevant state rather than a vague intent.
References & further reading
- Kubernetes: Pod Security Standards and Restricted controls
- Kubernetes: Pod and container securityContext semantics
- Kubernetes: NetworkPolicy enforcement, additive rules, and limits
- Kubernetes: least-privilege Secret access and application exposure
- Kubernetes: CPU, memory, and ephemeral-storage resource management
- gVisor: security model, Sentry boundary, and residual risks
- Firecracker: production host, jailer, resource, and egress requirements
- OpenAI: agent approvals, sandboxing, and separate network control surfaces


