← All categories

AI Agent Safety & Governance

AI agent safety and governance turns operational boundaries into written, reviewable controls: permissions, defensive handling, audit evidence, safety cases, and readiness gates.

Who it is for: Agent operators, safety leads, engineering managers, and reviewers who need explicit controls and evidence before or after agent actions.

Problems this area addresses

  • Tool access without a documented risk and approval model
  • Untrusted content reaching an agent without defensive handling rules
  • Agent actions that cannot be reconstructed or reviewed
  • Deployment decisions made without evidence-backed safety and readiness gates

Recommended evaluation path

Start with the permission and safety-gate products to define authority. Add prompt-injection defense for agents that read untrusted content, then use audit trails, safety cases, and deployment-readiness records to review evidence.

Start here

Start with the permission and safety-gate products to define authority. Add prompt-injection defense for agents that read untrusted content, then use audit trails, safety cases, and deployment-readiness records to review evidence.

  • Start with: Agent Tool-Use Permission Matrix A written permission system for agent tool use: risk tiers, deny-by-default matrices, allowlist/blocklist policies, and an audit-log schema.
  • Add: AI Agent Permission and Safety Gate Kit A written safety system for agents acting on your behalf: permission tiers, human-approval checkpoint patterns, risk classification, do-not-touch boundaries, and rollback playbooks.
  • Choose this specialist option: Agent Prompt Injection Defense Kit Defensive patterns against prompt injection: a content trust ladder, instruction hierarchy model, untrusted-content policies, detection heuristics, and regression test suites.

Choose a bundle when: several exact members independently fit. Start with Wave 2 Agent Trust & Safety Stack, or use Agent Tool Permissions vs Safety Gates for a guided comparison.

Products in this category

The product type and summary below show how each option differs. Open a product page for exact scope, tiers, buyer fit, and bundle membership.

  1. 01 · policy kit (risk ladder, matrices, machine-readable contracts)

    Agent Tool-Use Permission Matrix

    A written permission system for agent tool use: risk tiers, deny-by-default matrices, allowlist/blocklist policies, and an audit-log schema.

  2. 02 · safety kit (permission tiers, approval checkpoints, rollback playbooks)

    AI Agent Permission and Safety Gate Kit

    A written safety system for agents acting on your behalf: permission tiers, human-approval checkpoint patterns, risk classification, do-not-touch boundaries, and rollback playbooks.

  3. 03 · defensive security kit (trust taxonomy, quarantine policies, detection heuristics, regression suites)

    Agent Prompt Injection Defense Kit

    Defensive patterns against prompt injection: a content trust ladder, instruction hierarchy model, untrusted-content policies, detection heuristics, and regression test suites.

  4. 04 · schema + method kit (action taxonomy, entry schemas, integrity and redaction rules)

    Agent Audit Trail Generator

    A written audit-trail standard for agent actions: an 8-class action taxonomy, per-class JSON entry schemas, append-only integrity rules (tamper-evident, honestly not tamper-proof), redaction rules, and sampling plans.

  5. 05 · governance kit (claim-argument-evidence safety cases, evidence registry, review records)

    Agent Safety Case Builder

    A written safety-case system for agent deployments: claim-argument-evidence framing, safety-case and evidence-registry JSON schemas, hazard analysis, and review records.

  6. 06 · checklist system (evidence-gated readiness gates, sign-off records)

    Agent Deployment Readiness Checklist System

    An evidence-gated go/no-go system for agent deployments: capability/safety/rollback/monitoring/ownership gate families, evidence-required checklists, and sign-off records.

  7. 07 · governance kit (classification model, retention schedules, scorecard)

    Agent Memory Governance Kit

    A written governance system for what agents may remember: memory classification, a binding never-store list, retention schedules, and audit scorecards.

Relevant bundles

  • Wave 2 Agent Trust & Safety StackThe five-product trust stack for proving and protecting agent behavior: evaluation harnesses, prompt-injection defense, audit trails, safety cases, and incident response - designed to work as one system.
  • Agent Operator Core BundleThe reliability stack for anyone running AI agents: harnesses to structure work, scorecards to judge output, and safety gates to prevent damage.
  • Wave 1 Agent Infrastructure StackThe five-product governance stack for running agents like infrastructure: memory governance, tool permissions, context engineering, error recovery, and workflow compilation - designed to work as one system.

Decision guides for this category

Compatibility considerations

These products are document, schema, policy, and checklist systems. Your own agent runtime or workflow platform must apply the controls and produce the evidence.

Limitations

They do not provide runtime enforcement, monitoring, certification, legal advice, or a guarantee that every failure or attack will be prevented.