Batch-test prompts against CSV test cases, score responses automatically with configurable judge rubrics, and put a mandatory human in the loop before anything counts as "final." Juri replaces ad-hoc console testing with a repeatable, auditable evaluation workflow — now extended with agent tool-gateway auditing, reusable Skills, and platform-enforced safety hooks.
Manually eyeballing whether an LLM's response is "good enough" doesn't scale past a handful of examples. Juri turns that judgment call into a structured pipeline: define a prompt once as a versioned Script, attach one or more Judge Configurations (rubric + scoring model), run it against a CSV of test cases, and get consistent, comparable, human-reviewed results every time.
Every prompt, every generation parameter, and every judge rubric is versioned. A run always points back to the exact script version that produced it — no more "which prompt was this?"
Uses the Claude Agent SDK (or a direct Messages API path) to both generate candidate responses and score them against configurable rubrics — accuracy, relevance, completeness, clarity, or fully custom dimensions.
No run is ever "final" purely from automated judging. Every required row is reviewed, overridable, and requires a reviewer comment — a durable, resumable, auditable trail.
From prompt to finalized, exportable report.
Create a Script — system prompt, generation model, temperature/max tokens — and attach one or more Judge Configurations with rubric + pass threshold, or start from the org-wide Default Judge.
Upload test cases: input prompt, optional expected answer, test case ID, category/tag.
Juri generates a candidate response per row and scores it against every attached judge config. Partial row failures don't abort the batch.
Rows required by the run's snapshotted policy enter mandatory pending review. Reviewers claim, accept, or override scores/verdicts, and must comment.
Finalize the run, view the multi-judge comparison report, and export as CSV, JSON, or HTML.
Real screens from the running Juri application — the Atlas theme, live.
Workload snapshot and trends: pending review, in-progress, finalized runs, latest pass rate, judge disagreement, review turnaround, judge drift, and recent activity/comparisons at a glance.
An agent's "Generate test cases" tab — drafts a run-ready CSV from the agent's design doc, with configurable case types (happy path, edge, ambiguity, hallucination, safety, RAG, tool use, format) and should_pass / should_fail tagging.
Side-by-side input, candidate response, and every attached judge's verdict with dimension scores and justification — keyboard shortcuts for accept, override, and navigation keep review moving fast.
Baseline vs. candidate diffed by match key — LLM usage deltas, agent trace deltas, and a full performance review with resolved/new findings, exportable as CSV, JSON, or HTML.
What's shipped and usable today, grouped by area. Newer capabilities are flagged.
Named, versioned prompt bundles with generation params (model, system prompt, temperature, max tokens) and full version history per run.
1..N judge configs per script, each with its own judge model, rubric text, scoring dimensions, scoring scale (1–5, 1–10, pass/fail), and optional custom output schema. Inherits from an org-level Default Judge unless overridden.
Row-by-row generate → judge loop with live progress ("87/200 processed") and per-row failure isolation. Includes a dedicated flow for grounded/RAG test sets.
Side-by-side prompt, candidate response, and all judge scores/justifications. Accept, override any dimension or the overall verdict, comment required.
Per-row side-by-side scores across judge configs, disagreement rows surfaced prominently, and aggregate averages to spot systematically harsh or lenient configs.
CSV and JSON for raw data, HTML for shareable comparison/full reports. Print-to-PDF from the browser covers PDF needs — no dedicated PDF pipeline.
Run-to-run comparison and judge drift tracking to see how scoring shifts as prompts or rubrics evolve over time.
Pending runs start in a shared team queue. Reviewers can claim; admins can assign or unassign. Assigned runs are visible only to the assignee and admins.
Per-script review policy: full, random sample, borderline-margin, or failures-only — snapshotted per run so mid-run edits don't change that batch.
One deployment, multiple teams within a single organization. Tool admins manage the org-level default judge; teams manage their own scripts and runs.
A workspace for iterating on prompts and judge rubrics before committing them to a versioned script.
Version diffs across a judge configuration's history, with pass-rate and override-rate tracking per version — see exactly how a rubric change shifted outcomes.
A narrow, scoped, evaluate-only HTTP API (/api/v1/integrations/{app_id}/…,
max 100 rows/batch) for callers who supply their own candidate_response and
just want consistent judging — with revealable bearer tokens and a signed-callback delivery
log in Admin. The legacy free-form CI API and general org webhooks have been retired in
favor of this narrower surface.
Admin-managed registry of outbound MCP servers (slug, transport, config JSON, secrets via
env:VAR_NAME) plus built-in no-network verification tools — compare_to_expected,
validate_json_schema, check_required_fields, regex_match,
contains_all — usable directly inside judge rubrics.
A dedicated per-call view of token usage and timing, separate from run-level reports — useful for cost and latency investigation across generation and judging calls.
A built-in /guide help page inside the app itself, so reviewers and admins
don't need to leave the tool to understand a workflow.
MCP Servers point Juri outward at tools it can call. Tool Gateways flip that: Juri
itself becomes the MCP server (POST /api/v1/gateways/{id}/mcp) that an external
agent under test — Cursor, Claude Desktop, or a custom runner — calls into. Every tool call
the agent makes is audited, and gateways support record/replay and fault injection against
that traffic, with full trajectory spans for debugging multi-step agent behavior.
Every tool call an external agent makes through the gateway is logged with arguments, results, and timing — a complete trajectory span per run.
Capture a real agent session once, then replay recorded tool responses deterministically for regression testing of agent behavior.
Inject errors, latency, or malformed responses into specific tool calls to see how an agent under test recovers.
*delete*,
*drop*, *purge*, *truncate*, *rm*),
blocking of destructive arguments (SQL DROP/DELETE/TRUNCATE,
HTTP DELETE), run/team binding checks, denylist-aware tools/list
filtering, and secret redaction in audit spans. Teams can layer on additional denylist globs
(Admin → Users & Teams) but cannot remove or disable the platform defaults.
Configured under Admin → Skills. Skills improve prompt cost and quality and let teams attach advisory playbooks — they're versioned and attach directly to agent generation or judge capabilities, up to five per attachment.
prompt_cache — an ephemeral cache for long system prompts (roughly 3,300+
characters) to cut repeated-generation cost. token_budget — truncates
oversized prompts, candidates, or tool results before they blow a run's budget.
context_memory — compacts multi-turn Agent SDK tool input/output into a
rolling summary so long agent traces stay within context.
Upload or paste a SKILL.md playbook at Admin → Skills → New
skill; each edit creates a new version, and attaching a skill pins that version
id. Skills are advisory only — they cannot grant tools, widen allowlists, or disable
product safety hooks.
There's no single "Logs" screen — instead, every layer of the system gets its own purpose-built log, each with its own retention and access controls.
Admin → LLM Calls. Every AI call leaves a row: kind (judge, generation, Prompt Lab, cache prewarm, run analysis), transport (Agent SDK, Messages API, mock, message batch), status, latency, and token counts including cache tokens. Full request/response text is viewable per call, and admins can purge rows older than a configurable window.
Admin → Tool Gateways → Audit. A timeline view groups the last 200 calls by run, then run-only, then trace, then unlinked — with a table view, a "linked to runs only" filter, and a cassette-coverage report for record/replay hit and miss rates. Denied calls are shown as such, and payloads are redacted for secrets automatically.
Every Integration App with a signed callback configured gets a delivery log tracking success and error status for each callback fired on a run's status changes — pending review, finalized, failed, or cancelled.
A tab on any existing agent that drafts a run-compatible CSV directly from that agent's design/solution document, name, description, and system prompt — instead of hand-writing every row.
Choose how many rows to generate and which natures to mix in — happy path, edge cases,
RAG-grounded, tool use, and more — and optionally tag rows should_pass or
should_fail so judge calibration has known-answer rows to check against.
Output is a downloadable CSV, not an auto-started run — nothing is saved to the agent and no run begins until a reviewer or admin uploads it via New Run, keeping the mandatory human-in-the-loop guarantee intact even for generated test sets.
Concrete scenarios Juri is built to support for QA and Dev teams working with LLM-backed features and agents.
A team updates a customer-support system prompt. Before merging, they re-run the existing CSV of 200 support questions against both the old and new script version, and use the run-to-run comparison to confirm the new prompt doesn't quietly regress accuracy or tone.
Two judge configurations — one scoring strictly on factual accuracy, one on a broader helpfulness rubric — are attached to the same script. The comparison report highlights rows where they disagree, so the team can tune thresholds or retire the rubric that's miscalibrated. Judge config History shows whether the last rubric edit actually moved the pass rate.
Before moving a production prompt from one model to another, a script is duplicated with the new generation model and run against the same test CSV, letting the team compare pass rates, judge scores, and per-call token/latency telemetry side by side prior to cutover.
Every reviewed row carries a mandatory comment and a persisted human decision. For teams in regulated environments, the finalized run plus HTML export becomes the evidence that a human confirmed AI-generated output before it was trusted.
A team already generates candidate responses in their own CI job and just needs consistent
scoring. They register an Integration App, post candidate_response values (up to
100 rows per batch) to Juri's evaluate-only API, and get judged results back with an optional
signed callback tracked in a delivery log.
A team building a Cursor-style coding agent points it at a Juri Tool Gateway instead of its real tool backend. Every tool call is audited and platform safety hooks block destructive calls outright, while fault injection confirms the agent degrades gracefully when a tool returns an error mid-task.
A 500-row run lands in the shared pending-review queue. Reviewers claim batches of rows, an admin reassigns a stuck batch to a teammate going on leave, and the run only finalizes once every row required by its snapshotted review policy has a human decision.
Not real customer data — shown to illustrate the shape of a run.
support_qa_testset.csv| Test ID | Category | Judge: Accuracy | Judge: Tone | Verdict | Review status |
|---|---|---|---|---|---|
| TC-001 | billing | 9 / 10 | 9 / 10 | Pass | Reviewed — accepted |
| TC-002 | billing | 6 / 10 | 8 / 10 | Borderline | Reviewed — overridden to Fail |
| TC-003 | technical | 8 / 10 | 7 / 10 | Pass | Reviewed — accepted |
| TC-004 | account | 4 / 10 | 6 / 10 | Fail | Reviewed — accepted |
| TC-005 | technical | 9 / 10 | 9 / 10 | Pass | Pending |
Rows with a copper left border are disagreement rows — flagged automatically wherever attached judge configs diverge, matching the Atlas app's report view.
Fields shown as they appear in the app's agent editor.
| Field | Example value |
|---|---|
| generation_mode | llm_prompt |
| generation_model | claude-sonnet |
| temperature | 0.20 |
| system_prompt | "You are a billing support assistant. Be concise, cite the relevant account setting, and never promise a refund." |
| skills attached | prompt_cache, token_budget |
| primary judge | strict-accuracy-v2 (pass_threshold 7, type: rubric) |
Aggregate view for one finished run — report.html's per-judge summary and flaky assertions.
finalized
2 / 200
91.5%
| Judge | avg_score | pass_rate | pass_threshold |
|---|---|---|---|
| strict-accuracy-v2 | 8.1 / 10 | 89% | 7 |
| tone-check | 8.6 / 10 | 94% | 6 |
| Flaky assertion | fail_count | total | fail_rate |
|---|---|---|---|
| contains_all(required_terms) | 6 | 200 | 3.0% |
| regex_match(order_id_format) | 3 | 200 | 1.5% |
Baseline vs. candidate run, matched by match_key — from compare.html.
24 rows
6 rows
3 pass→fail · 9 fail→pass
| Key | Baseline score | Candidate score | Delta | Verdict flip |
|---|---|---|---|---|
| TC-014 | 6 / 10 | 9 / 10 | +3 | fail → pass |
| TC-057 | 8 / 10 | 5 / 10 | −3 | pass → fail |
| TC-091 | 7 / 10 | 7 / 10 | 0 | — |
Judge calibration over time and by category — from drift.html.
7.2%
4.8%
tone-check · 11.3%
| Judge | N | Override % | Pass flip % | Harsh % | Lenient % |
|---|---|---|---|---|---|
| strict-accuracy-v2 (#14) | 640 | 5.9% | 3.1% | 2.2% | 3.7% |
| tone-check (#9) | 640 | 9.4% | 11.3% | 8.0% | 3.3% |
All values on this page are illustrative examples matching the app's real field and column names — not figures from a live deployment.
Set per script, snapshotted per run, so a policy change mid-run doesn't affect a batch already in flight.
| Mode | Behavior |
|---|---|
| full | Every row requires human review before the run can finalize. |
| sample | A random percentage of rows require review; failed and rate-limited rows are always included. |
| borderline | Only rows scoring within a configurable margin of the judge's pass threshold require review. |
| failures_only | Only rows a judge marked as fail (plus failed/rate-limited rows) require review. |
Runs start unassigned in a shared team queue. Reviewers can claim rows for themselves; admins can assign or unassign a run to a specific reviewer. Once assigned, a run is visible only to that reviewer and to admins — keeping review work distributed without duplicate effort.
A pragmatic, single-organization internal tool — not built for multi-tenant SaaS scale, and deliberately so.
ANTHROPIC_API_KEY) for generation + judging
First boot seeds MySQL from scripts/schema_mysql.sql; run
alembic stamp head as a known first-boot fixup before further migrations.
For a full blank-machine runbook (Docker install, git clone, MySQL, host Claude OAuth
mounts), see docs/INSTALL.md in the repository.
Smoke test path once running: Admin → Default Judge → create a Script and
attach a judge → Runs → upload a CSV with a prompt column →
Review → view the Report and export.