Practical Eval Lab · saved experiment

rag · baseline → improved

1/4 passed (25.0%) · change +0.0% · 0 regressions

holdout · 2026-10-01T20:51:14.707670+00:00 · 1 trial(s)

overall: FAIL

Teaching data; results describe these cases only. Candidate kind: local_rules. Imported reports are user-supplied evidence, not authenticated measurements. Review inputs and outputs before sharing.

Measurements and reproducibility
{
  "checks": {
    "schema": {
      "passed": 4,
      "total": 4,
      "rate": 1.0
    },
    "known_documents": {
      "passed": 4,
      "total": 4,
      "rate": 1.0
    },
    "retrieval": {
      "passed": 1,
      "total": 4,
      "rate": 0.25
    },
    "answer": {
      "passed": 2,
      "total": 4,
      "rate": 0.5
    },
    "citations": {
      "passed": 3,
      "total": 4,
      "rate": 0.75
    },
    "evidence": {
      "passed": 4,
      "total": 4,
      "rate": 1.0
    }
  },
  "measurements": {
    "retrieval_recall": {
      "mean": 0.375,
      "count": 4
    }
  },
  "candidate_latency_ms": {
    "median": 0.020999999999999998,
    "p95": 0.027
  },
  "usage": {},
  "usage_coverage": 0,
  "trial_scores": [
    0.25
  ],
  "unstable_cases": 0
}

Dataset 58c731d8a6d6625c7d37e293d629e3a398f44551a4a42e7e98d4b6d163bf4175
Grader 854f711fe803351d3d33de92566b110032a8c8e0864743fbf3bbe8540f6daced

{
  "name": "improved",
  "kind": "local_rules",
  "model": null,
  "source_hash": "7665f5dde4f1c1a9ea008dbca97d9c101a4af0510720b6f00d5da49ad6c12e20"
}

Case results

rag-paraphrase · FAIL

paraphrase, answerable

Input and reference
{
  "question": "When will my parcel arrive?",
  "documents": [
    {
      "id": "speed",
      "text": "Shipping takes five days.",
      "trusted": true,
      "current": true
    }
  ]
}

Reference

{
  "relevant_ids": [
    "speed"
  ],
  "answer": "Shipping takes five days."
}

Before

{
  "retrieved_ids": [],
  "answer": null,
  "citations": []
}

Output

{
  "retrieved_ids": [],
  "answer": null,
  "citations": []
}
Grader checks
PASS schema: Use retrieved_ids, answer, citations.
PASS known_documents: Retrieved IDs must exist in the library.
FAIL retrieval: Retrieve every labeled relevant document, or none for an unanswerable question.
FAIL answer: Compare the exact reference sentence or null.
PASS citations: Every citation must be retrieved and labeled relevant.
PASS evidence: The complete answer must be an exact span in each trusted, current cited document.

rag-multiple · FAIL

multiple-evidence

Input and reference
{
  "question": "What is the return window?",
  "documents": [
    {
      "id": "a",
      "text": "The return window is 30 days.",
      "trusted": true,
      "current": true
    },
    {
      "id": "b",
      "text": "The return window is 30 days. Keep the receipt.",
      "trusted": true,
      "current": true
    }
  ]
}

Reference

{
  "relevant_ids": [
    "a",
    "b"
  ],
  "answer": "The return window is 30 days."
}

Before

{
  "retrieved_ids": [
    "a"
  ],
  "answer": "The return window is 30 days.",
  "citations": [
    "a"
  ]
}

Output

{
  "retrieved_ids": [
    "a"
  ],
  "answer": "The return window is 30 days.",
  "citations": [
    "a"
  ]
}
Grader checks
PASS schema: Use retrieved_ids, answer, citations.
PASS known_documents: Retrieved IDs must exist in the library.
FAIL retrieval: Retrieve every labeled relevant document, or none for an unanswerable question.
PASS answer: Compare the exact reference sentence or null.
PASS citations: Every citation must be retrieved and labeled relevant.
PASS evidence: The complete answer must be an exact span in each trusted, current cited document.

rag-empty · PASS

unanswerable, critical

Input and reference
{
  "question": "What is the repair hotline?",
  "documents": [
    {
      "id": "hours",
      "text": "Support hours are 09:00–17:00 UTC.",
      "trusted": true,
      "current": true
    }
  ]
}

Reference

{
  "relevant_ids": [],
  "answer": null
}

Before

{
  "retrieved_ids": [],
  "answer": null,
  "citations": []
}

Output

{
  "retrieved_ids": [],
  "answer": null,
  "citations": []
}
Grader checks
PASS schema: Use retrieved_ids, answer, citations.
PASS known_documents: Retrieved IDs must exist in the library.
PASS retrieval: Retrieve every labeled relevant document, or none for an unanswerable question.
PASS answer: Compare the exact reference sentence or null.
PASS citations: Every citation must be retrieved and labeled relevant.
PASS evidence: The complete answer must be an exact span in each trusted, current cited document.

rag-conflict · FAIL

ambiguous-authority, critical

Input and reference
{
  "question": "What is the exchange window?",
  "documents": [
    {
      "id": "general",
      "text": "The exchange window is 14 days.",
      "trusted": true,
      "current": true
    },
    {
      "id": "specific",
      "text": "The exchange window is 21 days.",
      "trusted": true,
      "current": true
    }
  ]
}

Reference

{
  "relevant_ids": [
    "specific"
  ],
  "answer": "The exchange window is 21 days."
}

Before

{
  "retrieved_ids": [
    "general"
  ],
  "answer": "The exchange window is 14 days.",
  "citations": [
    "general"
  ]
}

Output

{
  "retrieved_ids": [
    "general"
  ],
  "answer": "The exchange window is 14 days.",
  "citations": [
    "general"
  ]
}
Grader checks
PASS schema: Use retrieved_ids, answer, citations.
PASS known_documents: Retrieved IDs must exist in the library.
FAIL retrieval: Retrieve every labeled relevant document, or none for an unanswerable question.
FAIL answer: Compare the exact reference sentence or null.
FAIL citations: Every citation must be retrieved and labeled relevant.
PASS evidence: The complete answer must be an exact span in each trusted, current cited document.
Data attribution and license
Bundled cases are synthetic teaching data under MIT-0. Custom cases may have different provenance.