CLI reference
Exit codes, store health, the Claude Code hook, argv parsing, compound commands, dry runs, and
library usage for the aegis CLI.
← back to the README
Argv forms
Global options in front of the verb (kubectl -n prod delete …, git -C /repo push …,
helm --kube-context prod uninstall …), glued short flags (-nprod), label selectors
(-l role=worker) and comma-separated kinds (nodes,pods) all parse to the same intents as
their canonical forms; a leading option Aegis does not recognise is an error (exit 65), never a
guess at where the verb starts. kubectl delete namespace prod additionally emits a synthetic
*/* delete intent scoped to that namespace, so namespace-scoped deletion rules fire on it.
Exit codes
The exit code is the worst verdict across all evaluated intents, or a tool error. Verdict codes
depend on --exit-style; tool errors are identical in every style and never print a verdict.
| exit code | --exit-style aegis (default) |
--exit-style claude-hook |
--exit-style ci |
|---|---|---|---|
0 |
ALLOW | ALLOW (prints nothing) | ALLOW |
1 |
— | — | ESCALATE or BLOCK |
2 |
ESCALATE | ESCALATE or BLOCK, printing {"decision": "block", "reason": "<verdict>: <citations>"} |
— |
3 |
BLOCK | — | — |
64 |
usage error (unknown flag, no argv after --, compound argv without --split-compound, a command string Aegis refuses to evaluate statically) |
same | same |
65 |
bad data: unparseable argv, --now, YAML or JSON; degraded store (see below); no signing key, or a missing/bad signature |
same | same |
66 |
a constraints/authority/environments/plan/key file does not exist or is unreadable | same | same |
70 |
internal error; the exception class name is on stderr | same | same |
Every error is one aegis: error: … line on stderr, never a traceback. claude-hook exists
because Claude Code PreToolUse hooks treat exit 2 as block and any other non-zero code as a
non-blocking error, which would invert the default 3 = BLOCK contract.
Store health
Every output carries the state of the constraint store, so a degraded store can never be
mistaken for a clean allow. Each per-intent JSON line and the plan summary include a
store_health object — loaded, quarantined: [{id, reason}], principals,
constraints_sha256 (of the raw file bytes), warnings — and --pretty ends with a
STORE: loaded=N quarantined=M principals=P line plus one line per quarantined rule, preceded
by one WARNING: … line per store warning (using example signing key, insecure: signatures
not verified, ledger: chain-broken, rate-limit key typos). Quarantines also go to stderr as
aegis: WARNING … lines.
An untrustworthy constraint gets no vote. A constraint that was quarantined at load
(tampered, forged, principal-mismatch, invalid) or discarded at decision time
(tampered, unauthorized) contributes nothing to the verdict — but it is never silently
dropped. It is named in discarded, counted in STORE: … quarantined=N, and logged to stderr:
aegis check kubectl --constraints /tmp/oneflip.yaml --key file:data/example-signing.key \
--plan-constraints '' --pretty -- kubectl delete node/x --context kind-local
aegis: WARNING Quarantined constraint no-delete-nodes: provenance hash mismatch
ALLOW: kubernetes delete node/x
discarded: [{'id': 'no-delete-nodes', 'reason': 'tampered'}]
covered: True latency_ms: 0.25
STORE: loaded=20 quarantined=1 principals=3
quarantined: no-delete-nodes (tampered)
warning: Quarantined constraint no-delete-nodes: provenance hash mismatch
Why not stop the line instead? Because if an untrustworthy rule could force ESCALATE, anyone able to write a rule could stall every matching action — a denial of service on the guardrail, and the usual reason guardrails get switched off. Measured on the benchmark, that policy scores the same poison-susceptibility and over-block rate as a verifier with no trust model at all. Visibility, not obedience, is what makes a poisoned rule safe.
If you would rather an edited rule stopped the line, and accept that anyone who can write a
rule can then stop it, opt in with --on-untrusted-match escalate. The same command then
exits 2:
ESCALATE: kubernetes delete node/x
discarded: [{'id': 'no-delete-nodes', 'reason': 'tampered'}]
note: fail-closed: no-delete-nodes (tampered)
Even then, an untrustworthy rule can never produce BLOCK.
(/tmp/oneflip.yaml here is data/constraints.example.yaml with the last hex digit of
no-delete-nodes’s provenance_hash flipped and re-signed with the example key — a valid
signature, invalid content, exactly what a hand-edited rule or a one-bit storage error looks
like. --key is explicit because the example key normally auto-resolves next to
--constraints, and /tmp has no example-signing.key of its own. --context kind-local
resolves the environment and --plan-constraints '' switches off plan rules, so the tampered
rule is the only thing deciding this intent.)
Aegis refuses to decide at all (exit 65, one-line message, no verdict) when the store loaded
zero constraints, the authority map grants nothing to anyone, or more than
--max-quarantine-ratio (default 0.10) of the constraints were quarantined. --fail-closed
additionally turns an uncovered intent (no rule matched) into ESCALATE with the note
fail-closed: uncovered.
Some clauses do fail closed, because they cover missing information rather than an
attacker’s choice: a rule that scopes on env when the
intent’s environment could not be resolved contributes ESCALATE with the note
env-unresolved: <id> (see Environment mapping), and an
intent whose target the parser could not pin down (git push -f with no refspec) contributes
ESCALATE with the note unknown-target, so it cannot slip past a ref/main rule as ref/*.
Claude Code hook
For Claude Code, Codex, Copilot, VS Code and Cursor, use aegis hook <agent> and
aegis install <agent> instead; see Coding agents. The script below predates
them and is kept for existing setups: unlike aegis hook, it gates every project it is
registered in, whether or not the project has a policy.
examples/claude-code-hook.sh is a PreToolUse hook: it reads the hook JSON from stdin and
hands tool_input.command — the raw string — to aegis check command --exit-style claude-hook.
ALLOW lets the tool call proceed; ESCALATE and BLOCK exit 2 with the reason on stderr (which
Claude Code shows to the model). Compound commands are split and launchers unwrapped (see
“Compound commands” below); binaries Aegis has no parser for are not gated unless you add
--fail-closed via AEGIS_ARGS; any tool error — including a command string Aegis refuses to
evaluate statically, such as kubectl delete $(cat x) — is converted into a block, so the hook
never fails open. It runs the aegis in the venv next to it (AEGIS_BIN overrides).
{"hooks": {"PreToolUse": [{"matcher": "Bash", "hooks": [
{"type": "command", "command": "/path/to/aegis-devops/examples/claude-code-hook.sh"}]}]}}
Verified snapshot
aegis snapshot loads and verifies the policy as aegis check does and prints what may
actually vote:
aegis snapshot --pretty
SNAPSHOT eac55f5b2b2163b6…
aegis 0.2.1; verified=30 excluded=0
input: environments 9d970db5…
input: plan_constraints c2f595cd…
input: sources_manifest 2fbfa3b2…
verified counts constraints that are loaded and, at this moment, intact, authorized
(authority is otherwise checked per decision) and source-verified at load — a signed
constraint with a self-consistent hash is not proof that its cited source backs it. excluded
lists everything else with its reason (tampered, forged, unauthorized, source-unverified,
invalid, …). The digest covers every input — the verified constraints, the authority map, the
environment and plan-constraint maps, the sources manifest, repos.yaml/signers.yaml/
agents.yaml when present, the commit each Git source ref points at, store settings that change
evaluation (the default time zone, tzdata availability) and the Aegis version — so anything
built from the policy can be traced to exactly this state.
It refuses --insecure, --sources '', and any decision-shaping file that does not verify: the
environment and plan-constraint maps are loaded through their loaders (signature and shape),
agents.yaml through the identity-model loader (below). The snapshot is detached from the loaded store: it keeps canonical
records and hands out fresh copies, so nothing can change what it describes after the fact.
Without --pretty it prints one JSON object. aegis sources is the per-constraint view of the
same load.
Identity model
aegis agents verifies agents.yaml (signature, shape, and that its principal holds the
identity class) and prints the identity model server-side enforcement is compiled for:
aegis agents --pretty
AGENTS /home/me/.config/aegis/agents.yaml
mode=deny-by-default enforcement=report-only principal=admin
kubernetes:
break-glass group aegis:break-glass
trusted group platform-admins
trusted serviceaccount argocd:argocd-application-controller
(every other identity is treated as an agent)
aws:
break-glass role arn:aws:iam::111122223333:role/BreakGlass
trusted role arn:aws:iam::111122223333:role/Deploy
(every other identity is treated as an agent)
Exit 0, 1 when the model has a warning (workflow-unbound, no-break-glass), 65 when it does
not verify, 66 when there is no agents.yaml (--agents PATH points elsewhere). Without
--pretty it prints one JSON object. See Configuration.
Server-side compile (preview)
aegis compile aws --account <12-digit id> --out DIR compiles the verified snapshot to AWS
Service Control Policies scoped to the agent identities in agents.yaml; --check DIR exits 1
on drift. It refuses --insecure and unsigned policy like aegis snapshot, and needs
agents.yaml. Options: --partition, --escalate deny|omit, --max-policies. See
Server-side enforcement.
aegis compile kubernetes --cluster <name> --out DIR compiles it to Kubernetes
ValidatingAdmissionPolicies (policies.yaml, for kubectl apply -f) scoped to the same
agents, with the same coverage report, manifest and --check DIR; --escalate deny|omit.
Identity audit
aegis audit-identity aws checks agents.yaml against the IAM roles and users that actually
exist in an AWS account. It is read-only: it runs aws iam get-account-authorization-details
with your credentials (--profile), or reads a saved copy of that output (--inventory FILE).
aegis audit-identity aws --profile sandbox-admin --pretty
AUDIT aws account 701331084529 (deny-by-default, report-only)
would restrict (2):
role arn:aws:iam::701331084529:role/coding-agent last used 2026-09-29T03:12:50+00:00
role arn:aws:iam::701331084529:role/nightly-backup last used 2026-09-28T02:00:00+00:00
exempt (3):
role arn:aws:iam::701331084529:role/BreakGlass
...
service-linked, never restricted by SCPs: 5
- would restrict: every role and user a compiled policy would restrict — in
deny-by-default, everything not trusted or break-glass. This is the report-only review: add the legitimate automation here (backup jobs, cleanup functions, deploy roles) totrustedbefore settingenforcement: enforce.--would-restrictprints only this list. The last-used date helps tell live automation from leftovers; Identity Center andOrganizationAccountAccessRoleroles carry a hint. - missing (a problem): an identity
agents.yamllists for this account that does not exist. A mistyped break-glass role means there is no break-glass. - findings, on trust policies:
workflow-unbound,oidc-wildcard-subject,oidc-no-subject— GitHub OIDC trust that any job in a repository (or any repository) can use. On an exempt role this is a problem: an agent job can assume it withAssumeRoleWithWebIdentity, which no SCP condition on the caller can see.exempt-trusts-restricted,exempt-trusts-account— an exempt role that a restricted identity (or the whole account) may assume. The compiled self-protection block closes this once enforced, so it is a problem only whilereport-only.
Exit 0, 1 on a problem, 65 if agents.yaml does not verify, 66 if it is missing.
aegis audit-identity kubernetes [--context CTX] does the same for a cluster, read-only (get
and SubjectAccessReviews through kubectl with your credentials; --inventory FILE reads a
copy saved with --save-inventory):
- would restrict: every ServiceAccount, and every user and group named in an RBAC binding
(the API has no user objects), that the compiled policies would restrict, with its bindings
and how many pods run as it. The control plane is always exempt, as in the compiler, and groups
every identity carries (
system:authenticated,system:serviceaccounts) are not listed.system:masters, the API server’s kubelet client and node bootstrap tokens carry a hint. - missing (a problem): a listed ServiceAccount that does not exist, or a break-glass user or group no binding grants anything.
- findings (problems): an identity the policies restrict that RBAC still lets impersonate
users, groups, ServiceAccounts or UIDs; update, patch or delete ValidatingAdmissionPolicies,
their bindings or webhook configurations; create mutating webhooks or policies; or escalate,
bind or write cluster role bindings. Admission cannot see impersonation or protect admission
objects, so RBAC is the only control there. Each is asked as the identity with the groups it
really carries, by name for every trusted and break-glass user, group and ServiceAccount
(RBAC can grant impersonation by
resourceNames), and per namespace for ServiceAccounts (a RoleBinding grants it in one namespace). Note that the built-ineditandadminroles include impersonating ServiceAccounts in their namespace. A review that fails (no right to run it) is an error, never a “no”.
Compound commands
Agent frameworks hand over a shell string, not an argv. aegis check command -- "<string>"
(and aegis check argv --split-compound) turns that string into the simple commands it would
run — splitting on ;, &&, ||, |, & and newlines, unwrapping sudo, env VAR=…,
timeout, nice, nohup, command, time, sh -c "…" and the k/tf/g aliases — and
checks every one whose binary Aegis knows. The exit code is the worst verdict across them:
aegis check command -- "kubectl get pods; kubectl delete node/w1" # BLOCK, exit 3
aegis check command -- "sudo kubectl delete node/w1" # BLOCK, exit 3
aegis check command -- "kubectl get pods | grep x" # ALLOW, exit 0
aegis check command -- 'kubectl delete $(cat x)' # exit 64: command rejected
Inside a pipeline an unknown binary (| grep x) produces nothing; on its own it produces a
synthetic shell / binary/<name> / exec intent that --fail-closed escalates. KUBECONFIG=…
and env AWS_PROFILE=… prefixes land in the intent’s metadata (kubeconfig, profile, …) so
the environment map sees them. Fail closed: anything whose argv cannot be known without
running it is refused as a usage error (exit 64, command rejected: <reason>) — command
substitution ($(…), backticks), $VAR outside single quotes, process substitution,
subshells, here-docs, eval/exec/source/./xargs, and unbalanced quotes. Without the
flag, aegis check <target> still treats any shell metacharacter in an argv as a usage error.
Dry runs
A rehearsal (kubectl --dry-run=client|server, aws --dry-run, az --what-if|--dry-run,
gcloud --dry-run) can’t change infrastructure, so Aegis never blocks or escalates it — the
verdict is always ALLOW, with dry_run: true and would_be reporting what a real run would
have gotten:
ALLOW: kubernetes scale deployment/api-server (dry-run; would be BLOCK)
citations: no-scale-prod-peak
covered: True latency_ms: 0.12
PLAN ALLOW: 1 intent(s)
note: would_be: BLOCK
terraform/tofu plans are evaluated normally — the plan JSON is the proposed change, not a
rehearsal of one.
Library usage
from aegis_core.authority import load_authority_map
from aegis_core.store import ConstraintStore
from aegis_core.interceptor import AegisInterceptor
from aegis_core.parser import from_kubectl
authority_map = load_authority_map("data/authority.example.yaml")
store = ConstraintStore.load("data/constraints.example.yaml", authority_map=authority_map)
interceptor = AegisInterceptor(store)
intent = from_kubectl(["kubectl", "delete", "node/worker-1"])
decision = interceptor.intercept(intent)
print(decision.verdict, decision.citations, decision.covered)
# BLOCK ['no-delete-nodes'] True
Decision also carries discarded (constraints that matched but were thrown out, with why),
notes (fail-closed: …, rate-limit: …), latency_ms, dry_run, and would_be.
store.health is the StoreHealth the CLI prints; AegisInterceptor(store, fail_closed=True)
is the library form of --fail-closed.