Teaching data; results describe these cases only. Candidate kind: local_rules.
Imported reports are user-supplied evidence, not authenticated measurements. Review inputs and outputs before sharing.
PASS schema: Use retrieved_ids, answer, citations.
PASS known_documents: Retrieved IDs must exist in the library.
FAIL retrieval: Retrieve every labeled relevant document, or none for an unanswerable question.
FAIL answer: Compare the exact reference sentence or null.
PASS citations: Every citation must be retrieved and labeled relevant.
PASS evidence: The complete answer must be an exact span in each trusted, current cited document.
rag-multiple · FAIL
multiple-evidence
Input and reference
{
"question": "What is the return window?",
"documents": [
{
"id": "a",
"text": "The return window is 30 days.",
"trusted": true,
"current": true
},
{
"id": "b",
"text": "The return window is 30 days. Keep the receipt.",
"trusted": true,
"current": true
}
]
}
Reference
{
"relevant_ids": [
"a",
"b"
],
"answer": "The return window is 30 days."
}
Before
{
"retrieved_ids": [
"a"
],
"answer": "The return window is 30 days.",
"citations": [
"a"
]
}
Output
{
"retrieved_ids": [
"a"
],
"answer": "The return window is 30 days.",
"citations": [
"a"
]
}
Grader checks
PASS schema: Use retrieved_ids, answer, citations.
PASS known_documents: Retrieved IDs must exist in the library.
FAIL retrieval: Retrieve every labeled relevant document, or none for an unanswerable question.
PASS answer: Compare the exact reference sentence or null.
PASS citations: Every citation must be retrieved and labeled relevant.
PASS evidence: The complete answer must be an exact span in each trusted, current cited document.
rag-empty · PASS
unanswerable, critical
Input and reference
{
"question": "What is the repair hotline?",
"documents": [
{
"id": "hours",
"text": "Support hours are 09:00–17:00 UTC.",
"trusted": true,
"current": true
}
]
}
PASS schema: Use retrieved_ids, answer, citations.
PASS known_documents: Retrieved IDs must exist in the library.
PASS retrieval: Retrieve every labeled relevant document, or none for an unanswerable question.
PASS answer: Compare the exact reference sentence or null.
PASS citations: Every citation must be retrieved and labeled relevant.
PASS evidence: The complete answer must be an exact span in each trusted, current cited document.
rag-conflict · FAIL
ambiguous-authority, critical
Input and reference
{
"question": "What is the exchange window?",
"documents": [
{
"id": "general",
"text": "The exchange window is 14 days.",
"trusted": true,
"current": true
},
{
"id": "specific",
"text": "The exchange window is 21 days.",
"trusted": true,
"current": true
}
]
}
Reference
{
"relevant_ids": [
"specific"
],
"answer": "The exchange window is 21 days."
}
Before
{
"retrieved_ids": [
"general"
],
"answer": "The exchange window is 14 days.",
"citations": [
"general"
]
}
Output
{
"retrieved_ids": [
"general"
],
"answer": "The exchange window is 14 days.",
"citations": [
"general"
]
}
Grader checks
PASS schema: Use retrieved_ids, answer, citations.
PASS known_documents: Retrieved IDs must exist in the library.
FAIL retrieval: Retrieve every labeled relevant document, or none for an unanswerable question.
FAIL answer: Compare the exact reference sentence or null.
FAIL citations: Every citation must be retrieved and labeled relevant.
PASS evidence: The complete answer must be an exact span in each trusted, current cited document.
Data attribution and license
Bundled cases are synthetic teaching data under MIT-0. Custom cases may have different provenance.