CreditDesk

A provider-neutral evaluation environment. Run one credit-underwriting episode and watch the tool-call trajectory and rubric score render.

Proof of architecture — two tasks, two seeded worlds. Replay is the public default (offline, deterministic). Live models are gated off unless a server operator enables them. Not a benchmark; not a task library.

Draft status. Both tasks and their rubrics are AI-authored drafts pending expert sign-off. They are not difficulty-calibrated, carry no inter-rater agreement figure, and no expert has signed them. No live frontier model has been run on the current reward.

What you are looking at. A correct analyst (the gold trajectory) scores 1.0 on both worlds; a plausible-but-wrong one (the weak trajectory) is capped by the critical gate however fluent its memo reads. The guilty world is the one where the trap is present; in the innocent world the trap is absent and the permissive decision is the correct one — so an agent that reflexively cries fraud fails it too. Both trajectories are fixed recordings, not live models.

Drive it over MCP

These environments run as Model Context Protocol servers. One server process binds to one environment and exposes that task's tools to any MCP client, so a lab can point its own agent at the world without importing our code. The reward stays behind an operator gate — the agent never sees its own score, and the score is frozen when the episode closes.

Covers all seven tasks across the three engines, one process per environment. Install steps, a client config snippet, the wire contracts, an example session and the known limitations are in creditdesk-mcp/README.md.

Scoreboard

The full cross-engine scoreboard — every task on both worlds, gold against weak — is generated by eval/run_eval.py and read from the repository at request time, so what you see here is what the file says. The live-model columns are empty by design.