xorcise-ai
xorcise
Pythonโœจ New

Run cyber-AI agents against real missions in isolated environments. Record everything as OpenTelemetry evidence. Grade the evidence, not the claim.

Last updated Aug 9, 2026
10
Stars
0
Forks
11
Issues
0
Stars/day
Attention Score
32
Language breakdown
Python 61.5%
TypeScript 37.7%
CSS 0.6%
JavaScript 0.1%
Shell 0.1%
Dockerfile 0.0%
โ–ธ Files click to expand
README

XORCISE.AI โ€” Trust Evidence, not Claims.

PyPI Python 3.12+ Tested on Ubuntu CI License Apache-2.0

Quickstart ยท Output ยท Agents ยท Missions ยท Open source ยท Website ยท Docs ยท Contributing ยท Security


Run your cyber-AI agent against a real mission. Watch everything it does. Grade the evidence.

AI can take action. It cannot bear consequences.

A benchmark score tells you an agent finished. It says nothing about the destructive commands it tried on the way there. XORCISE runs the agent against a live target inside a contained environment, records every command, tool call and dead end as OpenTelemetry evidence, and grades that evidence against the mission's own criteria.

Trust is not declared. It is demonstrated.

Quickstart

Tested on Ubuntu. Needs Python 3.12+ and Docker Engine.

pip install xorcise
xorcise doctor                            # checks the host first
xorcise up                                # boots the stack, prints the console URL
xorcise config set-model --name <model> --key <key>          # the judge โ€” half the score
xorcise agent register --name my-agent --kind claude-code
xorcise mission list
xorcise mission pull aviary-access
xorcise run create --agent my-agent --mission aviary-access
xorcise run launch-cmd <run_id>           # paste into your agent's terminal, then run it
xorcise run status <run_id>               # score, breakdown, evidence

xorcise down stops it all. No Docker on the box? xorcise up --stub is the self-contained demo. xorcise --help has the rest, and docs.xorcise.ai walks through a first run end to end.

Prefer to work from source? See Contributing โ†’ Setup.

What comes out

| | | |---|---| | Live trace | every command, tool call and message, streaming into the console as it happens | | Score | deterministic checks plus a bring-your-own-model judge | | Report | the full run record, exportable โ€” Markdown, HTML, JSONL | | Leaderboard | agents ranked across recorded results |

Every run gets its own private network and a fresh environment, created for the run and destroyed after it. An agent under evaluation cannot reach the host, or another run.

Bring your agent

XORCISE evaluates the agent you already use.

| | | |---|---| | OpenHands | full trace + tool-call capture | | Claude Code | via OTLP telemetry | | Codex CLI | via OTLP telemetry | | Anything custom | register it, drive it with the connect prompt, submit over REST |

Activity is normalised into one event model, so the trace, the grading and the report read the same whichever harness produced the run.

Missions

A mission is a self-contained target: services, a network, and the criteria an agent is graded against. Packaged as bundles, pulled on demand.

Missions are deliberately vulnerable โ€” SQL injection, IDOR, network pivots. That is the point: they exist so an agent has something real to find.

Run XORCISE on infrastructure you are willing to lose โ€” a dedicated VM or an isolated cloud
environment, never a workstation holding credentials you care about. It executes untrusted
agent code against vulnerable targets by design.
>
**Only point XORCISE at systems you own, or that you have specific written authorisation to
test.** See Acceptable use.

Open source

XORCISE goes public in parts, not whole. This repository is the engine โ€” the CLI, harness adapters, isolation, grading and console โ€” under Apache-2.0, with issues and pull requests open.

The evaluation technology is open source. The commercial layer โ€” managed deployment, runtime, command and sovereign hosting โ€” is not. The agent skills and the documentation source are published separately as they are readied.

Documentation & help

| | | |---|---| | xorcise.ai | the project website โ€” what XORCISE is and who it is for | | Documentation | first run, missions, grading, traces, the full CLI and API reference | | Contributing | dev setup, the test lanes, the PR process, versioning | | Security | what's in scope, and how to report privately | | Acceptable use | what to point XORCISE at, export control, sanctions | | Maintainers ยท Code of Conduct | who to ask, and how we work |

Found a vulnerability? Do not open a public issue โ€” report it privately. Flaws in the harness, the isolation boundary or the supply chain are in scope; flaws inside a mission are the content.

License

Apache-2.0 ยฉ 2026 Fifth Domain Pty Ltd (ACN 606 251 585)

XORCISE.AI is a business name of Fifth Domain Pty Ltd. Contributions are accepted under the Contributor License Agreement. "XORCISE" and "XORCISE.AI", the associated logos and wordmarks, and the xorcise.ai domain are trademarks โ€” see the trademark policy and NOTICE.


XORCISE.AI โ€” Trust Evidence, not Claims.

๐Ÿ”— More in this category

ยฉ 2026 GitRepoTrend ยท xorcise-ai/xorcise ยท Updated daily from GitHub