Recipe catalog / completion-gate
Gate an agent claiming it is done
Did an agent finish task, judging report against evidence, with unproven claims, quietly narrowed scope, open questions, and unresolved errors flagged in the same call?
A coding agent says it is finished and you must decide, before accepting or before letting it stop, whether the work is actually complete.
Explore this recipe interactively ยท Source and implementation guide
Use completion-gate in TypeScript
Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.
import { completionGate } from 'jev-recipes/completion-gate';
const result = await completionGate({
"task": "Add input validation to the signup endpoint, write unit tests for the new validation, and make sure the full test suite passes.",
"report": "I added validation for email and password fields in signup.ts and wrote three unit tests covering the new checks. All tests pass. Let me know if you want me to also validate the username field.",
"evidence": "$ npm test\n\nTest Files 12 passed (12)\n Tests 87 passed (87)\n\n$ git diff --stat\n src/signup.ts | 18 ++++++++++++\n tests/signup.test.ts | 41 +++++++++++++++++++++++++",
"minConfidence": 0.8
});
console.log(result);
Input contract
| Field | Type | Needed |
|---|---|---|
| task | string | Required |
| report | string | Required |
| evidence | string | Optional |
| minConfidence | number | Optional |
Full input and result schemas
{
"input": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"task": {
"type": "string"
},
"report": {
"type": "string"
},
"evidence": {
"type": "string"
},
"minConfidence": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"required": [
"task",
"report"
]
},
"result": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"model": {
"type": "string"
},
"usage": {
"type": "object",
"properties": {
"input_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"output_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"input_tokens",
"output_tokens"
],
"additionalProperties": false
},
"status": {
"type": "string",
"enum": [
"ready",
"review"
]
},
"verdict": {
"type": "string",
"enum": [
"complete",
"incomplete",
"unverified",
"unclear"
]
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"probabilities": {
"type": "object",
"propertyNames": {
"type": "string",
"enum": [
"complete",
"incomplete",
"unverified",
"unclear"
]
},
"additionalProperties": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"required": [
"complete",
"incomplete",
"unverified",
"unclear"
]
},
"signals": {
"type": "object",
"properties": {
"claimsWithoutEvidence": {
"type": "object",
"properties": {
"status": {
"type": "string",
"enum": [
"ready",
"review"
]
},
"verdict": {
"type": "string",
"enum": [
"present",
"absent"
]
},
"probability": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"required": [
"status",
"verdict",
"probability",
"confidence"
],
"additionalProperties": false
},
"scopeNarrowed": {
"type": "object",
"properties": {
"status": {
"type": "string",
"enum": [
"ready",
"review"
]
},
"verdict": {
"type": "string",
"enum": [
"present",
"absent"
]
},
"probability": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"required": [
"status",
"verdict",
"probability",
"confidence"
],
"additionalProperties": false
},
"openQuestions": {
"type": "object",
"properties": {
"status": {
"type": "string",
"enum": [
"ready",
"review"
]
},
"verdict": {
"type": "string",
"enum": [
"present",
"absent"
]
},
"probability": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"required": [
"status",
"verdict",
"probability",
"confidence"
],
"additionalProperties": false
},
"unresolvedErrors": {
"type": "object",
"properties": {
"status": {
"type": "string",
"enum": [
"ready",
"review"
]
},
"verdict": {
"type": "string",
"enum": [
"present",
"absent"
]
},
"probability": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"required": [
"status",
"verdict",
"probability",
"confidence"
],
"additionalProperties": false
}
},
"required": [
"claimsWithoutEvidence",
"scopeNarrowed",
"openQuestions",
"unresolvedErrors"
],
"additionalProperties": false
},
"detected": {
"type": "array",
"items": {
"type": "string",
"enum": [
"claimsWithoutEvidence",
"scopeNarrowed",
"openQuestions",
"unresolvedErrors"
]
}
}
},
"required": [
"model",
"usage",
"status",
"verdict",
"confidence",
"probabilities",
"signals",
"detected"
],
"additionalProperties": false
}
}Saved example result
This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo completion-gate.
{
"model": "demo-fixture",
"usage": {
"input_tokens": 0,
"output_tokens": 0
},
"status": "ready",
"verdict": "complete",
"confidence": 0.9,
"probabilities": {
"complete": 0.9,
"incomplete": 0.05,
"unverified": 0.03,
"unclear": 0.02
},
"signals": {
"claimsWithoutEvidence": {
"status": "ready",
"verdict": "absent",
"probability": 0.06,
"confidence": 0.94
},
"scopeNarrowed": {
"status": "ready",
"verdict": "absent",
"probability": 0.08,
"confidence": 0.92
},
"openQuestions": {
"status": "ready",
"verdict": "absent",
"probability": 0.12,
"confidence": 0.88
},
"unresolvedErrors": {
"status": "ready",
"verdict": "absent",
"probability": 0.03,
"confidence": 0.97
}
},
"detected": []
}
Evaluation evidence
typesafe-ai/jev / 2026-09-27 / 40 held-out cases
Scoring revision 1.
30 ready decisions, with 100% accuracy among those decisions.
95% case-level interval: 91% to 100%. Related synthetic cases are correlated.
Measured on these synthetic cases
This measurement uses an earlier or unverified recipe or evaluator version. Rerun with the current recipe and evaluator before treating these numbers as current.
Use the evaluation guide to measure this decision on your own labeled cases.
Limitations
- Judges only the supplied report and evidence. It does not run tests or inspect the repository; supply that output in evidence.
- A complete verdict means the supplied material shows the task done, not that the work is correct or well made.
Related recipes
- step-complete: Use step-complete to check one explicit completion condition against evidence.
- goal-drift: Use goal-drift while the agent is still working to catch a step that wanders from the goal.
- result-plausibility: Use result-plausibility to check whether a single tool result is a real answer.