Recipe catalog / evidence-strength
Grade evidence strength
How strongly does evidence support the entire claim, on a five-level rubric?
You need a graded strength for weighting or ranking evidence, not just a supported or unsupported label.
Explore this recipe interactively ยท Source and implementation guide
Use evidence-strength in TypeScript
Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.
import { evidenceStrength } from 'jev-recipes/evidence-strength';
const result = await evidenceStrength({
"claim": "Guests can export reports as CSV.",
"evidence": "Guest accounts may export any report they can view. Export formats include CSV and PDF.",
"minConfidence": 0.8
});
console.log(result);
Input contract
| Field | Type | Needed |
|---|---|---|
| claim | string | Required |
| evidence | string | Required |
| minConfidence | number | Optional |
Full input and result schemas
{
"input": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"claim": {
"type": "string"
},
"evidence": {
"type": "string"
},
"minConfidence": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"required": [
"claim",
"evidence"
]
},
"result": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"model": {
"type": "string"
},
"usage": {
"type": "object",
"properties": {
"input_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"output_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"input_tokens",
"output_tokens"
],
"additionalProperties": false
},
"status": {
"type": "string",
"enum": [
"ready",
"review"
]
},
"score": {
"type": "number",
"minimum": 0
},
"level": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"probabilities": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"strength": {
"type": "string",
"enum": [
"none",
"weak",
"moderate",
"strong",
"conclusive"
]
}
},
"required": [
"model",
"usage",
"status",
"score",
"level",
"confidence",
"probabilities",
"strength"
],
"additionalProperties": false
}
}Saved example result
This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo evidence-strength.
{
"model": "demo-fixture",
"usage": {
"input_tokens": 0,
"output_tokens": 0
},
"status": "ready",
"score": 3.83,
"level": 4,
"confidence": 0.86,
"probabilities": {
"0": 0.01,
"1": 0.01,
"2": 0.03,
"3": 0.09,
"4": 0.86
},
"strength": "conclusive"
}
Evaluation evidence
No verified live accuracy measurement is available. Evaluate representative cases before using this decision in your workflow.
Use the evaluation guide to measure this decision on your own labeled cases.
Limitations
- Grades support only. Contradiction lands at the lowest level; use verify to distinguish it.
- The score is an expected value over rubric levels. Application code chooses cutoffs.
Related recipes
- verify: Use verify for a categorical supported, contradicted, or unsupported verdict per claim.
- answerability: Use answerability to decide whether evidence can answer a whole question.