jevrecipes

Recipe catalog / context-prune

Prune agent context

Which of items, earlier tool results and messages in an agent session, are still needed to finish objective, so the rest can be dropped from context?

An agent session is growing long and you want to drop stale tool output and messages before the next model turn without summarizing what must stay verbatim.

Explore this recipe interactively ยท Source and implementation guide

Use context-prune in TypeScript

Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.

import { contextPrune } from 'jev-recipes/context-prune';

const result = await contextPrune({
  "objective": "Fix the failing test in tests/auth.test.ts, which expects a 401 for expired tokens but receives a 500, then run the suite and commit.",
  "items": [
    {
      "id": "ls-root",
      "text": "Tool result (ls): README.md package.json src tests docs node_modules"
    },
    {
      "id": "auth-source",
      "text": "Tool result (cat src/auth.ts): export function verify(token) { const payload = jwt.verify(token, SECRET); if (payload.exp < Date.now()) throw new Error('expired'); return payload; }"
    },
    {
      "id": "readme",
      "text": "Tool result (cat README.md): # Acme API. Install with npm ci. Run with npm start. See docs/ for endpoints."
    },
    {
      "id": "user-note",
      "text": "User: don't touch the token format, other services depend on it."
    },
    {
      "id": "test-output",
      "text": "Tool result (npm test): FAIL tests/auth.test.ts > returns 401 for expired token. Expected 401, received 500. TypeError: Cannot read properties of undefined (reading 'status') at src/auth.ts:9"
    }
  ],
  "recent": "The agent has read the failing test and the auth source and is about to edit src/auth.ts.",
  "minConfidence": 0.8
});
console.log(result);

Input contract

FieldTypeNeeded
objectivestringRequired
itemsarrayRequired
recentstringOptional
minConfidencenumberOptional
Full input and result schemas
{
  "input": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "objective": {
        "type": "string"
      },
      "items": {
        "minItems": 1,
        "maxItems": 50,
        "type": "array",
        "items": {
          "type": "object",
          "properties": {
            "id": {
              "type": "string"
            },
            "text": {
              "type": "string"
            }
          },
          "required": [
            "id",
            "text"
          ]
        }
      },
      "recent": {
        "type": "string"
      },
      "minConfidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      }
    },
    "required": [
      "objective",
      "items"
    ]
  },
  "result": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "model": {
        "type": "string"
      },
      "usage": {
        "type": "object",
        "properties": {
          "input_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          },
          "output_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          }
        },
        "required": [
          "input_tokens",
          "output_tokens"
        ],
        "additionalProperties": false
      },
      "status": {
        "type": "string",
        "enum": [
          "ready",
          "review"
        ]
      },
      "items": {
        "type": "array",
        "items": {
          "type": "object",
          "properties": {
            "id": {
              "type": "string"
            },
            "status": {
              "type": "string",
              "enum": [
                "ready",
                "review"
              ]
            },
            "verdict": {
              "type": "string",
              "enum": [
                "keep",
                "drop"
              ]
            },
            "probability": {
              "type": "number",
              "minimum": 0,
              "maximum": 1
            },
            "confidence": {
              "type": "number",
              "minimum": 0,
              "maximum": 1
            }
          },
          "required": [
            "id",
            "status",
            "verdict",
            "probability",
            "confidence"
          ],
          "additionalProperties": false
        }
      },
      "keep": {
        "type": "array",
        "items": {
          "type": "string"
        }
      },
      "drop": {
        "type": "array",
        "items": {
          "type": "string"
        }
      },
      "evaluated": {
        "type": "integer",
        "minimum": 0,
        "maximum": 9007199254740991
      }
    },
    "required": [
      "model",
      "usage",
      "status",
      "items",
      "keep",
      "drop",
      "evaluated"
    ],
    "additionalProperties": false
  }
}

Saved example result

This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo context-prune.

{
  "model": "demo-fixture",
  "usage": {
    "input_tokens": 0,
    "output_tokens": 0
  },
  "status": "ready",
  "items": [
    {
      "id": "ls-root",
      "status": "ready",
      "verdict": "drop",
      "probability": 0.06,
      "confidence": 0.94
    },
    {
      "id": "auth-source",
      "status": "ready",
      "verdict": "keep",
      "probability": 0.95,
      "confidence": 0.95
    },
    {
      "id": "readme",
      "status": "ready",
      "verdict": "drop",
      "probability": 0.08,
      "confidence": 0.92
    },
    {
      "id": "user-note",
      "status": "ready",
      "verdict": "keep",
      "probability": 0.93,
      "confidence": 0.93
    },
    {
      "id": "test-output",
      "status": "ready",
      "verdict": "keep",
      "probability": 0.96,
      "confidence": 0.96
    }
  ],
  "keep": [
    "auth-source",
    "user-note",
    "test-output"
  ],
  "drop": [
    "ls-root",
    "readme"
  ],
  "evaluated": 5
}

Evaluation evidence

Earlier-evaluator measurement

typesafe-ai/jev / 2026-09-27 / 20 held-out cases

Scoring revision 1.

100%All-case accuracy
9Sent for review
0Failed calls / cases

11 ready decisions, with 100% accuracy among those decisions.

95% case-level interval: 84% to 100%. Related synthetic cases are correlated.

Experimental: declared acceptance policy not met

This measurement uses an earlier or unverified recipe or evaluator version. Rerun with the current recipe and evaluator before treating these numbers as current.

Use the evaluation guide to measure this decision on your own labeled cases.

Limitations

Related recipes