jevrecipes

Recipe catalog / tool-compare

Compare two tools for a task

Which of firstTool and secondTool, as described by their stated capabilities, better fits task?

An agent has two candidate tools for one step and needs a head-to-head preference based on the capability descriptions it has been given.

Explore this recipe interactively ยท Source and implementation guide

Use tool-compare in TypeScript

Install with npm install jev-recipes. Requires Node.js 22.9 or newer and ES modules. Set TYPESAFE_API_KEY in your server environment for live calls, which send input to TypeSafe and use API quota. See the installation guide.

import { toolCompare } from 'jev-recipes/tool-compare';

const result = await toolCompare({
  "task": "Extract every line-item table from a batch of 40 scanned supplier invoices, delivered as image-only PDFs, into CSV rows with the page number each row came from.",
  "firstTool": "pdf_text_extract: returns the embedded text layer of a PDF as a single plain-text string per page. Does not perform OCR and returns empty output for scanned or image-only pages.",
  "secondTool": "document_ocr_tables: runs OCR on scanned or image-based PDFs and returns detected tables as structured rows, each with cell text, a table index, and the source page number.",
  "minConfidence": 0.8
});
console.log(result);

Input contract

FieldTypeNeeded
taskstringRequired
firstToolstringRequired
secondToolstringRequired
minConfidencenumberOptional
Full input and result schemas
{
  "input": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "task": {
        "type": "string"
      },
      "firstTool": {
        "type": "string"
      },
      "secondTool": {
        "type": "string"
      },
      "minConfidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      }
    },
    "required": [
      "task",
      "firstTool",
      "secondTool"
    ]
  },
  "result": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "model": {
        "type": "string"
      },
      "usage": {
        "type": "object",
        "properties": {
          "input_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          },
          "output_tokens": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          }
        },
        "required": [
          "input_tokens",
          "output_tokens"
        ],
        "additionalProperties": false
      },
      "status": {
        "type": "string",
        "enum": [
          "ready",
          "review"
        ]
      },
      "verdict": {
        "type": "string",
        "enum": [
          "first",
          "second",
          "tie",
          "neither",
          "unclear"
        ]
      },
      "confidence": {
        "type": "number",
        "minimum": 0,
        "maximum": 1
      },
      "probabilities": {
        "type": "object",
        "propertyNames": {
          "type": "string",
          "enum": [
            "first",
            "second",
            "tie",
            "neither",
            "unclear"
          ]
        },
        "additionalProperties": {
          "type": "number",
          "minimum": 0,
          "maximum": 1
        },
        "required": [
          "first",
          "second",
          "tie",
          "neither",
          "unclear"
        ]
      }
    },
    "required": [
      "model",
      "usage",
      "status",
      "verdict",
      "confidence",
      "probabilities"
    ],
    "additionalProperties": false
  }
}

Saved example result

This hand-authored response demonstrates the contract. It is not a model accuracy measurement. Run it without an API key: npx jev-recipes demo tool-compare.

{
  "model": "demo-fixture",
  "usage": {
    "input_tokens": 0,
    "output_tokens": 0
  },
  "status": "ready",
  "verdict": "second",
  "confidence": 0.93,
  "probabilities": {
    "first": 0.02,
    "second": 0.93,
    "tie": 0.02,
    "neither": 0.02,
    "unclear": 0.01
  }
}

Evaluation evidence

Fixture only

No verified live accuracy measurement is available. Evaluate representative cases before using this decision in your workflow.

Use the evaluation guide to measure this decision on your own labeled cases.

Limitations

Related recipes