jevrecipes

Recipe catalog / extraction-fidelity

Grade extraction fidelity

How faithfully does extracted represent the facts in source, without invented, altered, or dropped values, on a five-level rubric?

You need to grade a structured extraction against its source document before trusting, storing, or acting on the values.

Explore this recipe interactively ยท Source and implementation guide

Use extraction-fidelity in TypeScript

Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.

import { extractionFidelity } from 'jev-recipes/extraction-fidelity';

const result = await extractionFidelity({
  "source": "Invoice INV-2041 issued 3 March 2026 to Harbor Lighting Ltd for 12 LED fixtures at $85.00 each. Subtotal $1,020.00, tax $81.60, total $1,101.60. Payment due 2 April 2026.",
  "extracted": "{\"invoiceNumber\":\"INV-2041\",\"issued\":\"2026-03-03\",\"customer\":\"Harbor Lighting Ltd\",\"total\":1101.6,\"due\":\"2026-04-02\"}",
  "minConfidence": 0.8
});
console.log(result);

Input contract

FieldTypeNeeded
sourcestringRequired
extractedstringRequired
minConfidencenumberOptional
Full input and result schemas
{
  "input": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "source": {
        "type": "string"
      },
      "extracted": {
        "type": "string"
      },
      "minConfidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      }
    },
    "required": [
      "source",
      "extracted"
    ]
  },
  "result": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "model": {
        "type": "string"
      },
      "usage": {
        "type": "object",
        "properties": {
          "input_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          },
          "output_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          }
        },
        "required": [
          "input_tokens",
          "output_tokens"
        ],
        "additionalProperties": false
      },
      "status": {
        "type": "string",
        "enum": [
          "ready",
          "review"
        ]
      },
      "score": {
        "type": "number",
        "minimum": 0
      },
      "level": {
        "type": "integer",
        "minimum": 0,
        "maximum": 9007199254740991
      },
      "confidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      },
      "probabilities": {
        "type": "object",
        "propertyNames": {
          "type": "string"
        },
        "additionalProperties": {
          "type": "number",
          "minimum": 0,
          "maximum": 1
        }
      },
      "fidelity": {
        "type": "string",
        "enum": [
          "poor",
          "low",
          "fair",
          "high",
          "exact"
        ]
      }
    },
    "required": [
      "model",
      "usage",
      "status",
      "score",
      "level",
      "confidence",
      "probabilities",
      "fidelity"
    ],
    "additionalProperties": false
  }
}

Saved example result

This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo extraction-fidelity.

{
  "model": "demo-fixture",
  "usage": {
    "input_tokens": 0,
    "output_tokens": 0
  },
  "status": "ready",
  "score": 3.04,
  "level": 3,
  "confidence": 0.81,
  "probabilities": {
    "0": 0,
    "1": 0.01,
    "2": 0.06,
    "3": 0.81,
    "4": 0.12
  },
  "fidelity": "high"
}

Evaluation evidence

Earlier-evaluator measurement

jev-1.13.0 / 2026-09-27 / 40 held-out cases

Scoring revision 1.

100%All-case accuracy
0Sent for review
0Failed calls / cases

40 ready decisions, with 100% accuracy among those decisions.

95% case-level interval: 91% to 100%. Related synthetic cases are correlated.

Measured on these synthetic cases

This measurement uses an earlier or unverified recipe or evaluator version. Rerun with the current recipe and evaluator before treating these numbers as current.

Use the evaluation guide to measure this decision on your own labeled cases.

Limitations

Related recipes