elengtis
A configurable MCP prompt-injection benchmark. Declarative YAML scenarios, complete agent trajectories, and effects verified independently of the model that was driving.
The pitch
An MCP server hands a model text it did not write. Sometimes that text contains instructions, and sometimes the model follows them. Testing that by hand means reading transcripts and deciding, by eye, whether anything happened. elengtis makes the scenario a YAML file, runs it a fixed number of times against a real server, and then checks the world instead of the transcript.
It is a scenario-based research tool, and it finds the behavior its rules describe. That is not the same as telling you a server is safe.
Quick start
$ pip install elengtis $ elengtis example --out results/demo $ cat results/demo/summary.txt
The packaged example needs no API key, no network service and no Docker. It drives a synthetic stdio fixture with a fresh fictional canary per trial, so the first run tells you the harness works before you point it at anything real.
A scenario is five phases
- setup Trusted MCP or HTTP actions plant the controlled state, including the injected text.
- exercise The model gets a prompt and an explicit tool allowlist. Setup and verification tools stay hidden unless you list them.
- proposal_rules Structured tool calls and arguments are matched. Every rule carries positive and negative examples, and validation runs them like unit tests for the YAML.
- verify Trusted checks inspect the resulting state. A verifier that calls back into the audited server is not an independent boundary, so container scenarios have to use HTTP or a SQLite snapshot.
- cleanup Best effort, and it runs even after a failure.
References inside a scenario are data, never executable templates:
tool: {binding: submit_credential}
arguments:
credential: {runner: canary}
A campaign binds those logical names to concrete tool names on each
target, and expands into a stable target × scenario × trial
matrix. Missing bindings are rejected before a result directory exists.
The model never grades itself
Two independent bits per attempt. proposed comes from
deterministic matching against the recorded tool calls;
completed comes from the trusted runner checking state
afterward. An attempt that was made and failed is a different result
from one that was never made, and both differ from a verifier that could
not produce a verdict at all, which records null rather than
a pass.
Each output directory keeps manifest.json,
runs.jsonl, one evidence document per attempt, and a summary.
Requests, messages, tool calls, matches and lifecycle records are all
retained, so a surprising row can be read back rather than re-run.
--resume picks up between trials, and a retry gets a new
attempt ID instead of inflating the denominator.
Targets
stdio targets get a fresh child process per trial. Streamable HTTP
targets get a fresh client session and nothing more, which is why those
runs are labelled externally_managed in the evidence and warn
during preflight. elengtis does not claim that an externally managed
server was reset between trials.
For a local image there is isolated_container: a fresh
hardened container per trial on an internal no-egress network, read-only
root, dropped capabilities, no image pull. Preflight starts it, checks the
allowlists and tears it down without spending a model call.
$ elengtis validate campaign.yaml $ elengtis preflight campaign.yaml $ elengtis run --config campaign.yaml --out results/research
Tracked bundles for the MCP Everything server and DBHub 1.2.3 live under
experiments/real-targets/. Only audit systems you are
authorized to test.
Four engines, one evaluator
create_agent, graph, langchain and
reference all receive identical prompts and identical allowed
tools, and all return orchestration observations only. One scenario
evaluator assigns proposal and completion meaning afterward, which
prevents four separate implementations of the experiment's semantics.
create_agent is the default. graph is for
workflows where sequential dispatch, turn budgets, branching or recovery
have to be visible and programmable. The framework capability experiment
found no universal winner between them, so the distinction stayed a role
distinction rather than becoming a recommendation.
Running it without a model
Set model_free: true with the reference engine
and a campaign runs only its declarative setup, verification and cleanup.
No model is constructed, zero model turns and tool calls are recorded, and
model or routing settings are rejected outright. It is how you debug a
scenario's plumbing, and how the CI suite stays free.