jevrecipes

Recipe catalog / answer-grade

Grade an answer against a rubric

How well does answer meet rubric as a response to question, on a five-level rubric from no credit to full credit?

You grade free-text answers against a written rubric and want a graded level with a confidence you can route to a human grader when low.

Explore this recipe interactively ยท Source and implementation guide

Use answer-grade in TypeScript

Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.

import { answerGrade } from 'jev-recipes/answer-grade';

const result = await answerGrade({
  "question": "Explain why the sky appears blue during the day.",
  "answer": "Sunlight contains all colors. When it passes through the atmosphere, gas molecules scatter shorter wavelengths like blue much more than longer wavelengths like red, so blue light reaches our eyes from every direction in the sky.",
  "rubric": "Full credit requires: (1) sunlight is made of many wavelengths; (2) atmospheric molecules scatter light; (3) shorter wavelengths scatter more strongly (Rayleigh scattering); (4) explains why violet is not the dominant perceived color.",
  "minConfidence": 0.8
});
console.log(result);

Input contract

FieldTypeNeeded
questionstringRequired
answerstringRequired
rubricstringRequired
minConfidencenumberOptional
Full input and result schemas
{
  "input": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "question": {
        "type": "string"
      },
      "answer": {
        "type": "string"
      },
      "rubric": {
        "type": "string"
      },
      "minConfidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      }
    },
    "required": [
      "question",
      "answer",
      "rubric"
    ]
  },
  "result": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "model": {
        "type": "string"
      },
      "usage": {
        "type": "object",
        "properties": {
          "input_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          },
          "output_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          }
        },
        "required": [
          "input_tokens",
          "output_tokens"
        ],
        "additionalProperties": false
      },
      "status": {
        "type": "string",
        "enum": [
          "ready",
          "review"
        ]
      },
      "score": {
        "type": "number",
        "minimum": 0
      },
      "level": {
        "type": "integer",
        "minimum": 0,
        "maximum": 9007199254740991
      },
      "confidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      },
      "probabilities": {
        "type": "object",
        "propertyNames": {
          "type": "string"
        },
        "additionalProperties": {
          "type": "number",
          "minimum": 0,
          "maximum": 1
        }
      },
      "grade": {
        "type": "string",
        "enum": [
          "none",
          "minimal",
          "partial",
          "mostly",
          "full"
        ]
      }
    },
    "required": [
      "model",
      "usage",
      "status",
      "score",
      "level",
      "confidence",
      "probabilities",
      "grade"
    ],
    "additionalProperties": false
  }
}

Saved example result

This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo answer-grade.

{
  "model": "demo-fixture",
  "usage": {
    "input_tokens": 0,
    "output_tokens": 0
  },
  "status": "ready",
  "score": 2.95,
  "level": 3,
  "confidence": 0.81,
  "probabilities": {
    "0": 0,
    "1": 0.02,
    "2": 0.09,
    "3": 0.81,
    "4": 0.08
  },
  "grade": "mostly"
}

Evaluation evidence

Fixture only

No verified live accuracy measurement is available. Evaluate representative cases before using this decision in your workflow.

Use the evaluation guide to measure this decision on your own labeled cases.

Limitations

Related recipes