elengtis

A configurable MCP prompt-injection benchmark. Declarative YAML scenarios, complete agent trajectories, and effects verified independently of the model that was driving.

  • Python 3.12+
  • LangChain + LangGraph
  • MCP
  • Apache-2.0
  • CLI + workbench

The pitch

An MCP server hands a model text it did not write. Sometimes that text contains instructions, and sometimes the model follows them. Testing that by hand means reading transcripts and deciding, by eye, whether anything happened. elengtis makes the scenario a YAML file, runs it a fixed number of times against a real server, and then checks the world instead of the transcript.

It is a scenario-based research tool, and it finds the behavior its rules describe. That is not the same as telling you a server is safe.

Quick start

$ pip install elengtis
$ elengtis example --out results/demo
$ cat results/demo/summary.txt

The packaged example needs no API key, no network service and no Docker. It drives a synthetic stdio fixture with a fresh fictional canary per trial, so the first run tells you the harness works before you point it at anything real.

A scenario is five phases

  • setup Trusted MCP or HTTP actions plant the controlled state, including the injected text.
  • exercise The model gets a prompt and an explicit tool allowlist. Setup and verification tools stay hidden unless you list them.
  • proposal_rules Structured tool calls and arguments are matched. Every rule carries positive and negative examples, and validation runs them like unit tests for the YAML.
  • verify Trusted checks inspect the resulting state. A verifier that calls back into the audited server is not an independent boundary, so container scenarios have to use HTTP or a SQLite snapshot.
  • cleanup Best effort, and it runs even after a failure.

References inside a scenario are data, never executable templates:

tool: {binding: submit_credential}
arguments:
  credential: {runner: canary}

A campaign binds those logical names to concrete tool names on each target, and expands into a stable target × scenario × trial matrix. Missing bindings are rejected before a result directory exists.

The model never grades itself

Two independent bits per attempt. proposed comes from deterministic matching against the recorded tool calls; completed comes from the trusted runner checking state afterward. An attempt that was made and failed is a different result from one that was never made, and both differ from a verifier that could not produce a verdict at all, which records null rather than a pass.

Each output directory keeps manifest.json, runs.jsonl, one evidence document per attempt, and a summary. Requests, messages, tool calls, matches and lifecycle records are all retained, so a surprising row can be read back rather than re-run. --resume picks up between trials, and a retry gets a new attempt ID instead of inflating the denominator.

Targets

stdio targets get a fresh child process per trial. Streamable HTTP targets get a fresh client session and nothing more, which is why those runs are labelled externally_managed in the evidence and warn during preflight. elengtis does not claim that an externally managed server was reset between trials.

For a local image there is isolated_container: a fresh hardened container per trial on an internal no-egress network, read-only root, dropped capabilities, no image pull. Preflight starts it, checks the allowlists and tears it down without spending a model call.

$ elengtis validate campaign.yaml
$ elengtis preflight campaign.yaml
$ elengtis run --config campaign.yaml --out results/research

Tracked bundles for the MCP Everything server and DBHub 1.2.3 live under experiments/real-targets/. Only audit systems you are authorized to test.

Four engines, one evaluator

create_agent, graph, langchain and reference all receive identical prompts and identical allowed tools, and all return orchestration observations only. One scenario evaluator assigns proposal and completion meaning afterward, which prevents four separate implementations of the experiment's semantics.

create_agent is the default. graph is for workflows where sequential dispatch, turn budgets, branching or recovery have to be visible and programmable. The framework capability experiment found no universal winner between them, so the distinction stayed a role distinction rather than becoming a recommendation.

Running it without a model

Set model_free: true with the reference engine and a campaign runs only its declarative setup, verification and cleanup. No model is constructed, zero model turns and tool calls are recorded, and model or routing settings are rejected outright. It is how you debug a scenario's plumbing, and how the CI suite stays free.