jevrecipes

Recipe catalog / action-compare

Compare two candidate actions

Which of firstAction and secondAction better advances goal within constraints?

An agent has narrowed to two next steps and needs a head-to-head preference under the stated goal and constraints.

Explore this recipe interactively ยท Source and implementation guide

Use action-compare in TypeScript

Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.

import { actionCompare } from 'jev-recipes/action-compare';

const result = await actionCompare({
  "goal": "Find out why the nightly export job failed last night.",
  "firstAction": "Open the job's log for last night's run and read the last 200 lines.",
  "secondAction": "Re-run the export job now and see if it fails again.",
  "constraints": "Do not trigger production jobs during business hours.",
  "minConfidence": 0.8
});
console.log(result);

Input contract

FieldTypeNeeded
goalstringRequired
firstActionstringRequired
secondActionstringRequired
constraintsstringOptional
minConfidencenumberOptional
Full input and result schemas
{
  "input": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "goal": {
        "type": "string"
      },
      "firstAction": {
        "type": "string"
      },
      "secondAction": {
        "type": "string"
      },
      "constraints": {
        "type": "string"
      },
      "minConfidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      }
    },
    "required": [
      "goal",
      "firstAction",
      "secondAction"
    ]
  },
  "result": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "model": {
        "type": "string"
      },
      "usage": {
        "type": "object",
        "properties": {
          "input_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          },
          "output_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          }
        },
        "required": [
          "input_tokens",
          "output_tokens"
        ],
        "additionalProperties": false
      },
      "status": {
        "type": "string",
        "enum": [
          "ready",
          "review"
        ]
      },
      "verdict": {
        "type": "string",
        "enum": [
          "first",
          "second",
          "tie",
          "neither",
          "unclear"
        ]
      },
      "confidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      },
      "probabilities": {
        "type": "object",
        "propertyNames": {
          "type": "string",
          "enum": [
            "first",
            "second",
            "tie",
            "neither",
            "unclear"
          ]
        },
        "additionalProperties": {
          "type": "number",
          "minimum": 0,
          "maximum": 1
        },
        "required": [
          "first",
          "second",
          "tie",
          "neither",
          "unclear"
        ]
      }
    },
    "required": [
      "model",
      "usage",
      "status",
      "verdict",
      "confidence",
      "probabilities"
    ],
    "additionalProperties": false
  }
}

Saved example result

This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo action-compare.

{
  "model": "demo-fixture",
  "usage": {
    "input_tokens": 0,
    "output_tokens": 0
  },
  "status": "ready",
  "verdict": "first",
  "confidence": 0.88,
  "probabilities": {
    "first": 0.88,
    "second": 0.04,
    "tie": 0.04,
    "neither": 0.02,
    "unclear": 0.02
  }
}

Evaluation evidence

Fixture only

No verified live accuracy measurement is available. Evaluate representative cases before using this decision in your workflow.

Use the evaluation guide to measure this decision on your own labeled cases.

Limitations

Related recipes