Three composable scanners that audit the entire AI-agent supply chain — the parts, the assembly, and the governable whole.
Bulwark is a defensive security suite for agentic AI systems. A modern AI agent is assembled from third-party parts nobody wrote and can't see inside — a model off the Hub, an MCP server from a gist, a pile of tools wired into an autonomous loop, an unaudited requirements.txt. Each is a trust boundary, and a system built entirely from individually benign parts can still be dangerous because of how they're wired together. Bulwark answers the three questions that actually matter, with three tools that share one engine, one finding taxonomy, and one report format. Packaged for PyPI under the bulwark-* namespace (bulwark-airlock / bulwark-warden / bulwark-manifest / bulwark-suite), with a MkDocs docs site, CodeQL CI, and a CITATION.cff for citable releases.
The suite is three composable scanners. Airlock scans the parts — is this model / MCP server / tool-spec itself malicious or unsafe? Warden scans the assembly — given how the agent is wired, does it have more power than its job needs? Manifest inventories the whole system — what is my AI made of, and is it governable? They compose literally: `bulwark scan ./project --scan-risk --govern` builds a CycloneDX AI-BOM of a project, calls Airlock on each model/MCP component and Warden on each agent assembly, and folds their findings inline into one governance report mapped to NIST AI RMF and the EU AI Act. Every tool is deterministic-first (fully useful with zero AI), CI-friendly (--fail-on exit codes + SARIF for GitHub code scanning), and defensive-only — it detects and reports, and never executes or imports the artifacts it scans.
Python 3.11+, uv workspace monorepo, YAML rule engine, CycloneDX ML-BOM, SPDX, SARIF 2.1.0, OSV, Ollama (optional AI), ruff, mypy, pytest, nox, GitHub Actions CI