Topic route / 07

AI and agent security

Production AI systems assessed through their data, deployment identities, tools, and real-world authority—not prompt lists alone.

Recommended order

Start here, then go deeper.

Assess the production pipeline, keep pentest authority outside the model, implement a deterministic execution broker, then measure AI researchers by verified coverage.

  1. 01
    Assessment methodThe Model Is Not the Target. The Pipeline Is.

    A field methodology for using MITRE ATLAS without turning an AI assessment into matrix theatre: map the production system, follow authority into tools and data, test reachable attack paths, and label the evidence only after impact is proven.

    18 min ↗
  2. 02
    Pentest operationsThe Model Found the Vulnerability. The Tool Call Became the Incident.

    A balanced operating model for AI-assisted pentesting: where models improve coverage and evidence work, where excessive agency turns a valid test into a destructive action, and how to keep cloud, shell, and Domain Admin authority outside the model.

    18 min ↗
  3. 03
    Controlled labThe Model Proposed the Action. The Broker Decided Whether It Could Exist.

    A practical architecture for AI-assisted pentest execution: resolve scope outside the model, classify side effects, issue short-lived capabilities, deny high-impact authority, and preserve a decision record that can be independently verified.

    16 min ↗
  4. 04
    Research benchmarkAI Vulnerability Discovery: One Frontier Model or Three Specialists?

    A reproducible benchmark design for the decision security teams actually face: spend the same research budget on repeated runs of one strong model, or on a diverse model team—and count only vulnerabilities that survive root-cause review, reproduction, and a fixed-version negative control.

    22 min ↗
Full topic archive

Every matching record.

Methods and named-vulnerability research remain visually and editorially separate.