jevrecipes

Recipe catalog / evaluation-mention

Label explicit mentions of model evaluation

Distinguish a response referring to its own evaluation from general evaluation discussion or no such mention.

You need to find explicit mentions of being tested, graded, or evaluated in saved model responses.

Explore this recipe interactively ยท Source and implementation guide

Use evaluation-mention in TypeScript

Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.

import { evaluationMention } from 'jev-recipes/evaluation-mention';

const result = await evaluationMention({
  "response": "This looks like a benchmark that will grade my answer, though I cannot know for sure."
});
console.log(result);

Input contract

FieldTypeNeeded
responsestringRequired
contextstringOptional
minConfidencenumberOptional
Full input and result schemas
{
  "input": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "response": {
        "type": "string",
        "description": "The response to inspect for explicit references to evaluation of an AI response or model."
      },
      "context": {
        "type": "string",
        "description": "Surrounding text used only to resolve who or what the response refers to."
      },
      "minConfidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      }
    },
    "required": [
      "response"
    ]
  },
  "result": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "model": {
        "type": "string"
      },
      "usage": {
        "type": "object",
        "properties": {
          "input_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          },
          "output_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          }
        },
        "required": [
          "input_tokens",
          "output_tokens"
        ],
        "additionalProperties": false
      },
      "status": {
        "type": "string",
        "enum": [
          "ready",
          "review"
        ]
      },
      "verdict": {
        "type": "string",
        "enum": [
          "self_reference",
          "discussion",
          "none",
          "unclear"
        ]
      },
      "confidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      },
      "probabilities": {
        "type": "object",
        "propertyNames": {
          "type": "string",
          "enum": [
            "self_reference",
            "discussion",
            "none",
            "unclear"
          ]
        },
        "additionalProperties": {
          "type": "number",
          "minimum": 0,
          "maximum": 1
        },
        "required": [
          "self_reference",
          "discussion",
          "none",
          "unclear"
        ]
      }
    },
    "required": [
      "model",
      "usage",
      "status",
      "verdict",
      "confidence",
      "probabilities"
    ],
    "additionalProperties": false
  }
}

Saved example result

This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo evaluation-mention.

{
  "model": "demo-fixture",
  "usage": {
    "input_tokens": 0,
    "output_tokens": 0
  },
  "status": "ready",
  "verdict": "self_reference",
  "confidence": 0.97,
  "probabilities": {
    "self_reference": 0.97,
    "discussion": 0.01,
    "none": 0.01,
    "unclear": 0.01
  }
}

Evaluation evidence

Fixture only

No verified live accuracy measurement is available. Evaluate representative cases before using this decision in your workflow.

Use the evaluation guide to measure this decision on your own labeled cases.

Limitations

Related recipes