Recipe catalog / explanation-level
Grade shown reasoning
How much reasoning does answer show for its conclusion to question, on a five-level rubric from bare conclusion to rigorous chain?
You assess whether students or assistants showed their work, separately from whether the final answer is right.
Explore this recipe interactively ยท Source and implementation guide
Use explanation-level in TypeScript
Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.
import { explanationLevel } from 'jev-recipes/explanation-level';
const result = await explanationLevel({
"question": "A train leaves at 9:40 and travels 150 km at 60 km/h. When does it arrive?",
"answer": "Time is distance over speed: 150 / 60 = 2.5 hours, which is 2 hours 30 minutes. Adding that to 9:40 gives 12:10. Check: 60 km/h for 2 hours covers 120 km, and the remaining 30 km takes half an hour, so 2.5 hours is right. The train arrives at 12:10.",
"minConfidence": 0.8
});
console.log(result);
Input contract
| Field | Type | Needed |
|---|---|---|
| question | string | Required |
| answer | string | Required |
| minConfidence | number | Optional |
Full input and result schemas
{
"input": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"question": {
"type": "string"
},
"answer": {
"type": "string"
},
"minConfidence": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"required": [
"question",
"answer"
]
},
"result": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"model": {
"type": "string"
},
"usage": {
"type": "object",
"properties": {
"input_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"output_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"input_tokens",
"output_tokens"
],
"additionalProperties": false
},
"status": {
"type": "string",
"enum": [
"ready",
"review"
]
},
"score": {
"type": "number",
"minimum": 0
},
"level": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"probabilities": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"explanation": {
"type": "string",
"enum": [
"bare",
"asserted",
"partial",
"complete",
"rigorous"
]
}
},
"required": [
"model",
"usage",
"status",
"score",
"level",
"confidence",
"probabilities",
"explanation"
],
"additionalProperties": false
}
}Saved example result
This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo explanation-level.
{
"model": "demo-fixture",
"usage": {
"input_tokens": 0,
"output_tokens": 0
},
"status": "ready",
"score": 3.87,
"level": 4,
"confidence": 0.89,
"probabilities": {
"0": 0,
"1": 0,
"2": 0.02,
"3": 0.09,
"4": 0.89
},
"explanation": "rigorous"
}
Evaluation evidence
No verified live accuracy measurement is available. Evaluate representative cases before using this decision in your workflow.
Use the evaluation guide to measure this decision on your own labeled cases.
Limitations
- Grades the visible reasoning, not its correctness; a complete chain can rest on a false premise.
- Rewards shown steps, so a correct one-line answer to a trivial question grades bare.
- The score is an expected value over levels. Application code chooses cutoffs.
Related recipes
- answer-relevance: Use answer-relevance to check that the answer addresses the question at all.
- certainty-match: Use certainty-match to check whether the confidence expressed fits the reasoning given.