Recipe catalog / evaluation-mention
Label explicit mentions of model evaluation
Distinguish a response referring to its own evaluation from general evaluation discussion or no such mention.
You need to find explicit mentions of being tested, graded, or evaluated in saved model responses.
Explore this recipe interactively ยท Source and implementation guide
Use evaluation-mention in TypeScript
Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.
import { evaluationMention } from 'jev-recipes/evaluation-mention';
const result = await evaluationMention({
"response": "This looks like a benchmark that will grade my answer, though I cannot know for sure."
});
console.log(result);
Input contract
| Field | Type | Needed |
|---|---|---|
| response | string | Required |
| context | string | Optional |
| minConfidence | number | Optional |
Full input and result schemas
{
"input": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"response": {
"type": "string",
"description": "The response to inspect for explicit references to evaluation of an AI response or model."
},
"context": {
"type": "string",
"description": "Surrounding text used only to resolve who or what the response refers to."
},
"minConfidence": {
"type": "number",
"minimum": 0,
"maximum": 1
}
},
"required": [
"response"
]
},
"result": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"model": {
"type": "string"
},
"usage": {
"type": "object",
"properties": {
"input_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"output_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"input_tokens",
"output_tokens"
],
"additionalProperties": false
},
"status": {
"type": "string",
"enum": [
"ready",
"review"
]
},
"verdict": {
"type": "string",
"enum": [
"self_reference",
"discussion",
"none",
"unclear"
]
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"probabilities": {
"type": "object",
"propertyNames": {
"type": "string",
"enum": [
"self_reference",
"discussion",
"none",
"unclear"
]
},
"additionalProperties": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"required": [
"self_reference",
"discussion",
"none",
"unclear"
]
}
},
"required": [
"model",
"usage",
"status",
"verdict",
"confidence",
"probabilities"
],
"additionalProperties": false
}
}Saved example result
This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo evaluation-mention.
{
"model": "demo-fixture",
"usage": {
"input_tokens": 0,
"output_tokens": 0
},
"status": "ready",
"verdict": "self_reference",
"confidence": 0.97,
"probabilities": {
"self_reference": 0.97,
"discussion": 0.01,
"none": 0.01,
"unclear": 0.01
}
}
Evaluation evidence
No verified live accuracy measurement is available. Evaluate representative cases before using this decision in your workflow.
Use the evaluation guide to measure this decision on your own labeled cases.
Limitations
- Detects explicit wording only. It does not infer hidden evaluation awareness, strategic behavior, or internal goals.
- Self-reference includes uncertainty and denial. It does not establish that evaluation is occurring or that the model believes it is.
- A missing mention does not show absence of awareness. Quoted or hypothetical first-person text needs careful attribution.
Related recipes
- claim-stance: Use claim-stance to distinguish affirming from denying a specific evaluation claim.
- context-role: Use context-role to classify the role of supplied text rather than mentions inside a response.
- uncertainty-expression: Use uncertainty-expression to label how certain the response sounds about a specific evaluation claim.