Juri Internal Product · LLM-as-Judge

The LLM prompt evaluation platform for QA & Dev teams

Batch-test prompts against CSV test cases, score responses automatically with configurable judge rubrics, and put a mandatory human in the loop before anything counts as "final." Juri replaces ad-hoc console testing with a repeatable, auditable evaluation workflow — now extended with agent tool-gateway auditing, reusable Skills, and platform-enforced safety hooks.

Overview

What Juri does

Manually eyeballing whether an LLM's response is "good enough" doesn't scale past a handful of examples. Juri turns that judgment call into a structured pipeline: define a prompt once as a versioned Script, attach one or more Judge Configurations (rubric + scoring model), run it against a CSV of test cases, and get consistent, comparable, human-reviewed results every time.

🧪

Repeatable testing

Every prompt, every generation parameter, and every judge rubric is versioned. A run always points back to the exact script version that produced it — no more "which prompt was this?"

⚖️

LLM-as-judge scoring

Uses the Claude Agent SDK (or a direct Messages API path) to both generate candidate responses and score them against configurable rubrics — accuracy, relevance, completeness, clarity, or fully custom dimensions.

👤

Human-in-the-loop, always

No run is ever "final" purely from automated judging. Every required row is reviewed, overridable, and requires a reviewer comment — a durable, resumable, auditable trail.

How it works

The core loop

From prompt to finalized, exportable report.

1. Configure

Create a Script — system prompt, generation model, temperature/max tokens — and attach one or more Judge Configurations with rubric + pass threshold, or start from the org-wide Default Judge.

2. Upload CSV

Upload test cases: input prompt, optional expected answer, test case ID, category/tag.

3. Run

Juri generates a candidate response per row and scores it against every attached judge config. Partial row failures don't abort the batch.

4. Review

Rows required by the run's snapshotted policy enter mandatory pending review. Reviewers claim, accept, or override scores/verdicts, and must comment.

5. Compare & export

Finalize the run, view the multi-judge comparison report, and export as CSV, JSON, or HTML.

Product tour

Straight from the app

Real screens from the running Juri application — the Atlas theme, live.

Dashboard

Workload snapshot and trends: pending review, in-progress, finalized runs, latest pass rate, judge disagreement, review turnaround, judge drift, and recent activity/comparisons at a glance.

Juri app dashboard screenshot

Generate test cases

An agent's "Generate test cases" tab — drafts a run-ready CSV from the agent's design doc, with configurable case types (happy path, edge, ambiguity, hallucination, safety, RAG, tool use, format) and should_pass / should_fail tagging.

Juri app generate test cases screenshot

Review workspace

Side-by-side input, candidate response, and every attached judge's verdict with dimension scores and justification — keyboard shortcuts for accept, override, and navigation keep review moving fast.

Juri app review workspace screenshot

Run-to-run compare

Baseline vs. candidate diffed by match key — LLM usage deltas, agent trace deltas, and a full performance review with resolved/new findings, exportable as CSV, JSON, or HTML.

Juri app run-to-run compare screenshot
Features

Functionality supported

What's shipped and usable today, grouped by area. Newer capabilities are flagged.

Scripts & versioning

Named, versioned prompt bundles with generation params (model, system prompt, temperature, max tokens) and full version history per run.

Judge configurations

1..N judge configs per script, each with its own judge model, rubric text, scoring dimensions, scoring scale (1–5, 1–10, pass/fail), and optional custom output schema. Inherits from an org-level Default Judge unless overridden.

CSV batch runs

Row-by-row generate → judge loop with live progress ("87/200 processed") and per-row failure isolation. Includes a dedicated flow for grounded/RAG test sets.

Mandatory review workspace

Side-by-side prompt, candidate response, and all judge scores/justifications. Accept, override any dimension or the overall verdict, comment required.

Comparison reports

Per-row side-by-side scores across judge configs, disagreement rows surfaced prominently, and aggregate averages to spot systematically harsh or lenient configs.

Export

CSV and JSON for raw data, HTML for shareable comparison/full reports. Print-to-PDF from the browser covers PDF needs — no dedicated PDF pipeline.

Dashboards & drift

Run-to-run comparison and judge drift tracking to see how scoring shifts as prompts or rubrics evolve over time.

Review assignment

Pending runs start in a shared team queue. Reviewers can claim; admins can assign or unassign. Assigned runs are visible only to the assignee and admins.

Sampling review modes

Per-script review policy: full, random sample, borderline-margin, or failures-only — snapshotted per run so mid-run edits don't change that batch.

Teams & roles

One deployment, multiple teams within a single organization. Tool admins manage the org-level default judge; teams manage their own scripts and runs.

Prompt Lab

A workspace for iterating on prompts and judge rubrics before committing them to a versioned script.

Judge config History New

Version diffs across a judge configuration's history, with pass-rate and override-rate tracking per version — see exactly how a rubric change shifted outcomes.

Integration Apps API

A narrow, scoped, evaluate-only HTTP API (/api/v1/integrations/{app_id}/…, max 100 rows/batch) for callers who supply their own candidate_response and just want consistent judging — with revealable bearer tokens and a signed-callback delivery log in Admin. The legacy free-form CI API and general org webhooks have been retired in favor of this narrower surface.

MCP Servers registry New

Admin-managed registry of outbound MCP servers (slug, transport, config JSON, secrets via env:VAR_NAME) plus built-in no-network verification tools — compare_to_expected, validate_json_schema, check_required_fields, regex_match, contains_all — usable directly inside judge rubrics.

LLM Calls telemetry New

A dedicated per-call view of token usage and timing, separate from run-level reports — useful for cost and latency investigation across generation and judging calls.

In-app field guide

A built-in /guide help page inside the app itself, so reviewers and admins don't need to leave the tool to understand a workflow.

New · Agent evaluation

Tool Gateways New

MCP Servers point Juri outward at tools it can call. Tool Gateways flip that: Juri itself becomes the MCP server (POST /api/v1/gateways/{id}/mcp) that an external agent under test — Cursor, Claude Desktop, or a custom runner — calls into. Every tool call the agent makes is audited, and gateways support record/replay and fault injection against that traffic, with full trajectory spans for debugging multi-step agent behavior.

Audit

Every tool call an external agent makes through the gateway is logged with arguments, results, and timing — a complete trajectory span per run.

Record & replay

Capture a real agent session once, then replay recorded tool responses deterministically for regression testing of agent behavior.

Fault injection

Inject errors, latency, or malformed responses into specific tool calls to see how an agent under test recovers.

🛡️
Platform-enforced safety hooks apply to every Tool Gateway call, with no opt-out: denylist globs on destructive-looking tool names (*delete*, *drop*, *purge*, *truncate*, *rm*), blocking of destructive arguments (SQL DROP/DELETE/TRUNCATE, HTTP DELETE), run/team binding checks, denylist-aware tools/list filtering, and secret redaction in audit spans. Teams can layer on additional denylist globs (Admin → Users & Teams) but cannot remove or disable the platform defaults.
New · Reusable behavior

Product Skills New

Configured under Admin → Skills. Skills improve prompt cost and quality and let teams attach advisory playbooks — they're versioned and attach directly to agent generation or judge capabilities, up to five per attachment.

Platform skills

prompt_cache — an ephemeral cache for long system prompts (roughly 3,300+ characters) to cut repeated-generation cost. token_budget — truncates oversized prompts, candidates, or tool results before they blow a run's budget. context_memory — compacts multi-turn Agent SDK tool input/output into a rolling summary so long agent traces stay within context.

Custom skills

Upload or paste a SKILL.md playbook at Admin → Skills → New skill; each edit creates a new version, and attaching a skill pins that version id. Skills are advisory only — they cannot grant tools, widen allowlists, or disable product safety hooks.

New · Observability

Logs & audit trails New

There's no single "Logs" screen — instead, every layer of the system gets its own purpose-built log, each with its own retention and access controls.

LLM Calls

Admin → LLM Calls. Every AI call leaves a row: kind (judge, generation, Prompt Lab, cache prewarm, run analysis), transport (Agent SDK, Messages API, mock, message batch), status, latency, and token counts including cache tokens. Full request/response text is viewable per call, and admins can purge rows older than a configurable window.

Tool Gateway audit

Admin → Tool Gateways → Audit. A timeline view groups the last 200 calls by run, then run-only, then trace, then unlinked — with a table view, a "linked to runs only" filter, and a cassette-coverage report for record/replay hit and miss rates. Denied calls are shown as such, and payloads are redacted for secrets automatically.

Callback deliveries

Every Integration App with a signed callback configured gets a delivery log tracking success and error status for each callback fired on a run's status changes — pending review, finalized, failed, or cancelled.

New · Faster test authoring

Generate test cases New

A tab on any existing agent that drafts a run-compatible CSV directly from that agent's design/solution document, name, description, and system prompt — instead of hand-writing every row.

Configurable generation

Choose how many rows to generate and which natures to mix in — happy path, edge cases, RAG-grounded, tool use, and more — and optionally tag rows should_pass or should_fail so judge calibration has known-answer rows to check against.

Review before it counts

Output is a downloadable CSV, not an auto-started run — nothing is saved to the agent and no run begins until a reviewer or admin uploads it via New Run, keeping the mandatory human-in-the-loop guarantee intact even for generated test sets.

Use cases

What teams use Juri for

Concrete scenarios Juri is built to support for QA and Dev teams working with LLM-backed features and agents.

Regression testing for prompts

Catch quality regressions before a prompt change ships

A team updates a customer-support system prompt. Before merging, they re-run the existing CSV of 200 support questions against both the old and new script version, and use the run-to-run comparison to confirm the new prompt doesn't quietly regress accuracy or tone.

Judge calibration

Decide which judge rubric to trust

Two judge configurations — one scoring strictly on factual accuracy, one on a broader helpfulness rubric — are attached to the same script. The comparison report highlights rows where they disagree, so the team can tune thresholds or retire the rubric that's miscalibrated. Judge config History shows whether the last rubric edit actually moved the pass rate.

Model migration validation

Validate a switch to a new or cheaper model

Before moving a production prompt from one model to another, a script is duplicated with the new generation model and run against the same test CSV, letting the team compare pass rates, judge scores, and per-call token/latency telemetry side by side prior to cutover.

Compliance & audit trail

Keep an auditable record of human-approved outputs

Every reviewed row carries a mandatory comment and a persisted human decision. For teams in regulated environments, the finalized run plus HTML export becomes the evidence that a human confirmed AI-generated output before it was trusted.

Third-party integration scoring

Score responses generated by an external pipeline

A team already generates candidate responses in their own CI job and just needs consistent scoring. They register an Integration App, post candidate_response values (up to 100 rows per batch) to Juri's evaluate-only API, and get judged results back with an optional signed callback tracked in a delivery log.

Agent tool-call auditing

Verify an autonomous agent doesn't misuse its tools

A team building a Cursor-style coding agent points it at a Juri Tool Gateway instead of its real tool backend. Every tool call is audited and platform safety hooks block destructive calls outright, while fault injection confirms the agent degrades gracefully when a tool returns an error mid-task.

Review workload management

Distribute review work across a QA team

A 500-row run lands in the shared pending-review queue. Reviewers claim batches of rows, an admin reassigns a stuck batch to a teammate going on leave, and the run only finalizes once every row required by its snapshotted review policy has a human decision.

Sample data

Illustrative input and output

Not real customer data — shown to illustrate the shape of a run.

Sample input CSV — support_qa_testset.csv

test_id,category,input_prompt,expected_answer
TC-001,billing,"How do I update my payment method?","Guide the user to Account > Billing > Payment Methods"
TC-002,billing,"Why was I charged twice this month?","Explain proration and offer to check the invoice"
TC-003,technical,"The app crashes when I export a report","Ask for steps to reproduce, offer workaround, escalate if unresolved"
TC-004,account,"How do I delete my account permanently?","Explain the deletion process and 30-day recovery window"
TC-005,technical,"API returns a 429 error intermittently","Explain rate limits and suggest exponential backoff"

Sample run output

Test IDCategoryJudge: AccuracyJudge: ToneVerdictReview status
TC-001billing9 / 109 / 10PassReviewed — accepted
TC-002billing6 / 108 / 10BorderlineReviewed — overridden to Fail
TC-003technical8 / 107 / 10PassReviewed — accepted
TC-004account4 / 106 / 10FailReviewed — accepted
TC-005technical9 / 109 / 10PassPending

Rows with a copper left border are disagreement rows — flagged automatically wherever attached judge configs diverge, matching the Atlas app's report view.

Sample judge configuration (JSON)

{
  "name": "strict-accuracy-v2",
  "judge_model": "claude-sonnet",
  "scoring_scale": "1-10",
  "dimensions": ["accuracy", "completeness", "tone"],
  "pass_threshold": 7,
  "rubric_prompt": "Score the candidate response strictly on factual correctness against the expected answer. Penalize hedging."
}

Sample agent configuration

Fields shown as they appear in the app's agent editor.

FieldExample value
generation_modellm_prompt
generation_modelclaude-sonnet
temperature0.20
system_prompt"You are a billing support assistant. Be concise, cite the relevant account setting, and never promise a refund."
skills attachedprompt_cache, token_budget
primary judgestrict-accuracy-v2 (pass_threshold 7, type: rubric)

Sample run report

Aggregate view for one finished run — report.html's per-judge summary and flaky assertions.

Run status

finalized

Disagreement count

2 / 200

Avg pass rate

91.5%

Judgeavg_scorepass_ratepass_threshold
strict-accuracy-v28.1 / 1089%7
tone-check8.6 / 1094%6
Flaky assertionfail_counttotalfail_rate
contains_all(required_terms)62003.0%
regex_match(order_id_format)32001.5%

Sample comparison report

Baseline vs. candidate run, matched by match_key — from compare.html.

Improved

24 rows

Regressed

6 rows

Verdict flips

3 pass→fail · 9 fail→pass

KeyBaseline scoreCandidate scoreDeltaVerdict flip
TC-0146 / 109 / 10+3fail → pass
TC-0578 / 105 / 10−3pass → fail
TC-0917 / 107 / 100—

Sample drift dashboard

Judge calibration over time and by category — from drift.html.

Overall override %

7.2%

Overall pass-flip %

4.8%

Highest flip rate

tone-check · 11.3%

JudgeNOverride %Pass flip %Harsh %Lenient %
strict-accuracy-v2 (#14)6405.9%3.1%2.2%3.7%
tone-check (#9)6409.4%11.3%8.0%3.3%

All values on this page are illustrative examples matching the app's real field and column names — not figures from a live deployment.

Review policy

Sampling review modes

Set per script, snapshotted per run, so a policy change mid-run doesn't affect a batch already in flight.

ModeBehavior
fullEvery row requires human review before the run can finalize.
sampleA random percentage of rows require review; failed and rate-limited rows are always included.
borderlineOnly rows scoring within a configurable margin of the judge's pass threshold require review.
failures_onlyOnly rows a judge marked as fail (plus failed/rate-limited rows) require review.

Runs start unassigned in a shared team queue. Reviewers can claim rows for themselves; admins can assign or unassign a run to a specific reviewer. Once assigned, a run is visible only to that reviewer and to admins — keeping review work distributed without duplicate effort.

Architecture

At a glance

A pragmatic, single-organization internal tool — not built for multi-tenant SaaS scale, and deliberately so.

Getting started

Local smoke test via Docker Compose

cp .env.example .env
# set SECRET_KEY (openssl rand -hex 32) and MySQL passwords
# with host Claude OAuth (claude auth login on the host):
docker compose -f docker-compose.yml -f docker-compose.host-oauth.yml up -d --build
# — or, with a console API key instead, drop the second -f —
docker compose exec app python -m app.cli create-admin
# open http://localhost:8000

First boot seeds MySQL from scripts/schema_mysql.sql; run alembic stamp head as a known first-boot fixup before further migrations. For a full blank-machine runbook (Docker install, git clone, MySQL, host Claude OAuth mounts), see docs/INSTALL.md in the repository.

Smoke test path once running: Admin → Default Judge → create a Script and attach a judge → Runs → upload a CSV with a prompt column → Review → view the Report and export.