Jev recipe catalog
Browse 248 TypeScript recipes for bounded decisions. Each guide includes an input contract, a saved example, evaluation evidence, and limitations.
Use the interactive explorer to search and compare recipes.
workflow
- Compare two candidate actions
Which of firstAction and secondAction better advances goal within constraints?
Fixture only - Label the side effects of an action
Which side effects does the described action involve: writing or modifying data, sending a message or notification, spending or moving money, deleting something, or calling an external service?
Earlier-evaluator measurement - Grade action reversibility
How reversible is action, given any context, from a trivial undo to an irreversible external effect?
Unknown response origin - Check proposed action scope
Is proposedAction within the work requested in request and constraints?
Fixture only - Grade how actionable an alert is
How actionable is alert for the on-call engineer who receives it, on a five-level rubric from pure noise to a guided first step?
Fixture only - Classify the ground an appeal asserts
What ground does appeal primarily assert: a factual error, a procedural error, new evidence, hardship, a misapplied rule, or something else?
Fixture only - Check argument meaning
Does proposedValue for argument express the intended value in request and context?
Fixture only - Detect breaking changes
Does change indicate a change that would break existing callers, integrations, stored data, or documented behavior?
Fixture only - Check plan against budget
Does plan, as described, plausibly fit within budget, the stated limits on steps, time, cost, or calls?
Fixture only - Grade bug report completeness
How complete is report for someone to reproduce and triage it, on a five-level rubric?
Fixture only - Grade change risk
How risky is change to ship, given context, on a five-level rubric?
Fixture only - Check a change against a change-window policy
Does the described change fall within the written change-window or freeze policy?
Fixture only - Choose a checkers move
Recommend a supplied legal checkers move from a structured board and player, with built-in American/English checkers instructions.
Unknown response origin - Choose a game action from supplied candidates
Recommend one eligible game action against a supplied goal, using rules, current state, and optional player history.
Unknown response origin - Check commit message accuracy
Does message accurately describe change, neither omitting a material part nor claiming work not present?
Fixture only - Gate an agent claiming it is done
Did an agent finish task, judging report against evidence, with unproven claims, quietly narrowed scope, open questions, and unresolved errors flagged in the same call?
Earlier-evaluator measurement - Compare two musical continuations
Which of firstContinuation and secondContinuation better follows the musical context in motif, contour, harmony, and any style supplied?
Fixture only - Check a corrective action against its root cause
Does action address the cause stated in rootCause rather than only the symptom or the affected items?
Fixture only - Classify a request to a music program
What does request ask a DAW or music program to do: record, edit, mix, apply an effect, arrange, export, or unclear?
Fixture only - Grade the risk of missing a deadline
How at risk is the work of missing deadline given progress, from on track to already missed or impossible?
Fixture only - Label the facets of a defect report
Which of these does report state: the part or lot identifier, a description of the defect, where it was detected, the quantity affected, a containment action?
Fixture only - Check delegation fit
Does subtask fall within the capabilities described for the delegate?
Fixture only - Flag hazards in a code diff
Which hazards does diff introduce: leaked secrets, destructive commands, debug leftovers, weakened tests, or dependency changes?
Unknown response origin - Steer the loudness
Should the upcoming phrase be softer, equally loud, or louder, judged from audience remarks about volume in context and the dynamics written in recentMaterial?
Fixture only - Label what an applicant statement addresses
Which of these does statement address: residency, income, household size, identity documents, prior benefits received?
Fixture only - Select the practice exercise that addresses feedback
Which candidate practice exercise in candidates best addresses the problems raised in feedback, given the player's goal?
Fixture only - Classify a described failure
Which supplied category best describes the observed failure?
Fixture only - Choose an action for every finger
On beat, which listed option should each finger in fingers take, given music and style, or rest?
Fixture only - Choose a game action from JSON
Send game state, player state, and available actions as JSON; receive the original selected action with Jev confidence and probabilities.
Fixture only - Classify game phase
Which phase of the game does state describe: opening, midgame, endgame, or game over?
Fixture only - Detect goal drift
Does step still serve goal, or has work drifted to something goal did not ask for?
Unknown response origin - Check handoff rules
Decide whether a request matches your human escalation rules.
Earlier-evaluator measurement - Grade handoff completeness
How ready is item, a work description being handed to another agent or person, to be picked up without asking questions, on a five-level rubric?
Fixture only - Grade the severity an incident report describes
What severity does the wording of report describe, on a five-level rubric from no user impact to total outage or data loss?
Fixture only - Detect agent-directed instructions
Does text contain instructions aimed at steering an AI system or agent?
Unknown response origin - Grade instruction clarity
How unambiguous is instruction for a delegate who has only context, on a five-level rubric?
Fixture only - Check instructions for conflict
Decide whether two instructions can both be followed under the supplied circumstances.
Fixture only - Check instruction applicability
Does the explicit scope of instruction cover task and context?
Fixture only - Decide which conflicting instruction wins
When firstInstruction and secondInstruction conflict, which should take precedence under the stated policy?
Fixture only - Check an intake question against its purpose
Does question ask only for what purpose needs, avoiding unrelated personal detail?
Fixture only - Check whether an itinerary is feasible in sequence
Can the consecutive items in itinerary be carried out in the order described: enough transfer time between them, each location reachable from the previous one, and no overlapping commitments?
Fixture only - Label what a job posting states
Which of these does posting state: salary or range, work location, remote policy, experience level, required qualifications?
Fixture only - Classify what a log line reports
What does line report: an error, a warning, a lifecycle event, a handled request, a metric, or debug output?
Fixture only - Route a request to a model tier
Which of the described models should serve request, and how much reasoning effort does it need, decided in one call so a cheap model handles easy turns and a capable one handles hard turns?
Earlier-evaluator measurement - Check whether now is a moment to change key
Is now a musically suitable moment to modulate away from key, judged on the phrase position, cadence, and stability described in recentMaterial?
Fixture only - Check a passage against a requested mood
Does the described passage deliver requestedMood?
Fixture only - Check move explanation fit
Does explanation give a reason for move that is consistent with the described game state?
Fixture only - Pick the next chord
Which candidate chord best continues progression, given any key and style supplied?
Fixture only - Pick the next rhythmic value
Which candidate rhythmic value best continues recentRhythm within meter, given any style supplied?
Fixture only - Pick the next melody note
Which candidate note best continues the melody in recentNotes, given any key and style supplied?
Fixture only - Check whether a phrase has closed
Does recentNotes form a complete musical phrase that is ready to cadence or rest, or is it still open, given any meter supplied?
Fixture only - Detect personal information
Does text contain information identifying a specific private individual?
Unknown response origin - Grade plan completeness
How completely does plan cover what task requires, on a five-level rubric?
Fixture only - Check an expense against a written policy
Does expense, as described, comply with the written policy?
Fixture only - Label the facets of a postmortem
Which of these does postmortem include: a timeline, a root cause, customer impact, contributing factors, action items?
Fixture only - Compare two tasks for priority
Which of firstTask and secondTask should be done first under criteria?
Fixture only - Detect a stalled agent
Does transcript, the recent agent steps, show the agent failing to make progress toward objective by repeating actions, circling, or reprocessing the same information?
Earlier-evaluator measurement - Check evidence for a qualification
Does profile contain concrete evidence that the candidate meets requirement, not just matching keywords?
Fixture only - Check interview question relevance
Does question ask only about matters relevant to the requirements of role, rather than personal circumstances unrelated to the work?
Fixture only - Compare attempted approaches
Does proposedAttempt use essentially the same approach as previousAttempt for objective?
Fixture only - Grade how repetitive recent material is
How repetitive is recentMaterial, from varied to stuck on one idea?
Fixture only - Label the facets of a progress report
Which of these does report include: an outcome statement, supporting evidence, blockers, a next step, open questions?
Fixture only - Interpret a reported tool outcome
What outcome does result report for task?
Fixture only - Check result plausibility
Is result a plausible, internally consistent answer to request rather than an error, placeholder, empty, or unrelated output dressed as data?
Fixture only - Check tool result usefulness
Does result provide information useful for task?
Fixture only - Decide whether to retry
Does failure describe a transient condition where an identical retry could succeed, given any attempt history?
Earlier-evaluator measurement - Classify review comment kind
What is the primary kind of comment?
Fixture only - Check whether symptoms implicate a recent change
Do symptoms plausibly point at the recently deployed change as the cause, judged on the timing and scope each describes?
Fixture only - Grade the depth of a root cause analysis
How deep does analysis go in explaining a failure, on a five-level rubric from restating the symptom to naming a verified systemic cause?
Fixture only - Route a request
Choose a named handler or return a review decision.
Earlier-evaluator measurement - Check rule compliance
Is the described action permitted by the written rules?
Fixture only - Check whether a runbook applies to an incident
Does runbook address the symptoms and component described in incident?
Fixture only - Classify a safety event
What kind of safety event does report describe: a near miss, a first-aid injury, a medical-treatment injury, property damage, an environmental release, or an unsafe condition?
Fixture only - Check a proposed slot against constraints
Does the time or slot proposed in proposal satisfy the availability constraints written in constraints?
Fixture only - Check deal stage against evidence
Does evidence support that the deal is at stage as described?
Fixture only - Check one completion condition
Does evidence establish that condition has been met?
Fixture only - Check step progress
How does observation change progress toward objective relative to previousState?
Fixture only - Assess whether a player can act now
Interpret game rules and current state to label a player's turn or reaction opportunity as act, wait, inactive, or unclear.
Fixture only - Grade task complexity
How complex is task, from a single lookup to open-ended work, on a five-level rubric?
Fixture only - Check the dependency between two tasks
Identify whether either of two tasks requires the other to finish before it can start.
Fixture only - Compare tasks for duplicate work
Decide whether two tasks request the same outcome, overlapping work, or distinct work.
Fixture only - Detect overlapping tasks
Do firstTask and secondTask cover overlapping work, such that two workers would duplicate or collide?
Fixture only - Steer the tempo
Should the beat slow down, hold, or speed up next, given what listeners said about pace in context and how fast recentMaterial moved?
Fixture only - Grade harmonic tension
How much harmonic tension does the current point in progression carry, from fully resolved to a peak demanding resolution, given any key supplied?
Fixture only - Gate an agent tool call
Should an agent run toolCall now, ask a person first, or refuse it, given request and policy, with irreversible, destructive, out-of-scope, exfiltration, and injection risks flagged in the same call?
Earlier-evaluator measurement - Compare two tools for a task
Which of firstTool and secondTool, as described by their stated capabilities, better fits task?
Fixture only - Check a tool fit
Can the capabilities explicitly described in tool perform task?
Fixture only - Decide whether an event should wake a waiting agent
Should event wake an agent that is paused until waitingFor happens: wake now, keep waiting, or ignore it as unrelated?
Unknown response origin
conversation
- Grade audience age suitability
What is the youngest general audience for which content is appropriate, on a five-level rubric?
Fixture only - Grade buying intent
How strong is the purchase intent expressed in message, on a five-level rubric?
Earlier-evaluator measurement - Identify who should initiate a callback
Which party is expected to initiate the next call?
Earlier-evaluator measurement - Check cancellation intent
Does message ask to cancel, pause, or continue task?
Fixture only - Identify personal and situational explanations
Label whether an explanation attributes a specified behavior to the person, circumstances, both, or gives no cause.
Fixture only - Grade the strength of causal language
How strong is the causal claim in the wording of statement, from no relationship claimed to causation stated as fact?
Fixture only - Check required information
Identify missing or ambiguous information before proceeding.
Unknown response origin - Grade clickbait level
How much does headline rely on curiosity gaps, exaggeration, or emotional bait instead of stating what the content is?
Fixture only - Grade commitment strength
How firmly does statement commit its speaker to an action or outcome, on a five-level rubric?
Fixture only - Interpret a confirmation
Does response clearly agree to or reject this exact proposal?
Fixture only - Detect an explicit consent request
Does text explicitly ask the reader for agreement or permission before something proceeds?
Fixture only - Distinguish requirements from preferences
Classify a stated constraint as required, preferred, optional, or unclear.
Fixture only - Recognize a contact opt-out
What scope of future contact does the sender ask to stop?
Earlier-evaluator measurement - Locate a correction target
Which supplied field or statement is message correcting?
Fixture only - Identify expressed emotion
What primary emotion does the wording of message express?
Fixture only - Check whether a message acknowledges a mistake
Does message explicitly acknowledge an earlier mistake and state a correction, rather than silently changing course or ignoring it?
Fixture only - Flag applicant-preference wording in a listing
Does listing express a preference for, or limitation on, applicants based on personal characteristics rather than describing the property?
Fixture only - Detect specific financial advice
Does text give a specific recommendation to buy, sell, hold, or allocate money, rather than general education?
Fixture only - Link a follow-up request
Which supplied earlier request does message follow up on?
Earlier-evaluator measurement - Recognize requested follow-up timing
When does the sender want another contact, if any?
Earlier-evaluator measurement - Grade the certainty of a financial projection
How certain is the wording of statement, a financial projection, from explicitly speculative to stated as fact?
Fixture only - Detect a change of intent
How does message change currentGoal?
Fixture only - Classify a listener's request
What does a listener's message ask of the performance?
Fixture only - Identify a requested mood
What mood does the listener's message ask to hear?
Fixture only - Identify a stated source of motivation
Label a stated reason for an activity as intrinsic enjoyment, a separate outcome, both, or not stated.
Fixture only - Grade how strongly a statement claims novelty
How strongly does the wording of statement claim novelty for a contribution, from no novelty claim to first-ever or unprecedented?
Fixture only - Classify sales objection kind
What primary sales objection does message raise?
Fixture only - Identify gain and loss framing
Label whether wording presents a specified outcome through gains, losses, both, or neither.
Fixture only - Identify a persuasion technique
Which persuasion technique, if any, does the wording of message primarily use?
Fixture only - Grade a policy violation
How severely does content violate the supplied policy, on a five-level rubric?
Fixture only - Grade expressed politeness
How polite is the wording of message toward its recipient, on a five-level rubric?
Fixture only - Check what a question takes for granted
Label whether a question takes a supplied claim for granted or leaves that claim open.
Fixture only - Check whether a question steers an answer
Label whether a question's wording favors, disfavors, or stays neutral toward a proposed answer.
Fixture only - Resolve a reference
Which supplied candidate does reference refer to in message and context?
Fixture only - Check reported resolution
Does message establish that the customer reports issue as resolved?
Fixture only - Check whether a reply is needed
Does message require a substantive reply in context?
Earlier-evaluator measurement - Detect spam content
Is message unsolicited promotional, scam, or bulk content rather than a genuine contribution?
Fixture only - Detect a topic shift
Does message stay with currentTopic, introduce a different topic, or contain both?
Fixture only - Identify the purpose of a trip
What purpose does the traveler's message express for the trip: business, leisure, a family visit, medical care, relocation, or attending an event?
Fixture only - Identify a conversation turn
What is the primary communicative purpose of message in context?
Fixture only
answer-quality
- Check answer consistency
Do firstStatement and secondStatement make compatible claims about the same subject and circumstances?
Fixture only - Check answer coverage
Check whether a draft answers each supplied question.
Fixture only - Label the disclosures in an answer
Which of these does draft include: stated uncertainty, stated limitations, cited sources, stated assumptions?
Fixture only - Grade an answer against a rubric
How well does answer meet rubric as a response to question, on a five-level rubric from no credit to full credit?
Fixture only - Check answer relevance
How directly does draft address request?
Fixture only - Check certainty wording
Does the certainty expressed in draft match assessment?
Fixture only - Match claims to citations
Find supplied passages that independently support an entire claim.
Fixture only - Check citation requirements
Do citationRules require evidence for statement?
Fixture only - Label a response's stance toward a claim
Label whether a response affirms, denies, mixes positions on, or does not address a supplied claim.
Unknown response origin - Compare two drafts
Which draft better satisfies request under rubric?
Fixture only - Label explicit mentions of model evaluation
Distinguish a response referring to its own evaluation from general evaluation discussion or no such mention.
Fixture only - Grade shown reasoning
How much reasoning does answer show for its conclusion to question, on a five-level rubric from bare conclusion to rigorous chain?
Fixture only - Grade feedback actionability
How actionable is feedback for its recipient, on a five-level rubric from no direction to a specific change with reason and example?
Fixture only - Check requested format compliance
Does response follow the structure or format request explicitly asks for, such as a list, table, JSON, item count, sections, or language?
Fixture only - Grade draft grounding
How much of the substantive content in draft is backed by evidence, on a five-level rubric?
Fixture only - Grade how easy patient instructions are to follow
How easy are instructions to follow for a general reader, from dense jargon to plain and stepwise?
Fixture only - Check response length fit
Is the length and detail of response proportionate to what request asks for?
Fixture only - Check question alignment to a learning objective
Does question assess the skill or knowledge stated in objective, rather than something adjacent?
Fixture only - Label the aspects performance feedback addresses
Which aspects of a musical performance does feedback address: rhythm, pitch or intonation, dynamics, technique, and expression or phrasing?
Fixture only - Check reply commitments
Does reply promise actions or outcomes beyond allowedCommitments?
Earlier-evaluator measurement - Label refusal behavior in a response
Distinguish an explicit refusal, a substantive attempt, mixed behavior, and a stated inability to fulfill a request.
Unknown response origin - Check summary coverage
Check whether a summary preserves each supplied point.
Fixture only - Check writing criteria
Check a draft against each supplied writing criterion.
Fixture only - Label expressed certainty about a claim
Label categorical, qualified, or unresolved wording about a supplied claim without inferring internal confidence.
Fixture only
knowledge
- Check an answer after a source change
Does updatedEvidence still support the entire claim that was based on previousEvidence?
Fixture only - Check a statement's attributed source
Check whether supplied source text attributes a statement to the claimed speaker or source.
Fixture only - Check audience fit
Does the level of explanation in document fit the knowledge and needs explicitly described in audience?
Fixture only - Check a budget narrative against its line items
Does narrative explain each line item in lineItems and nothing that lineItems does not list?
Fixture only - Check an item against its category
Does the product described by item belong under category as defined?
Fixture only - Assess a text revision
Does the revision from before to after change material meaning, conditions, or obligations?
Fixture only - Detect conflicting clauses
Do firstClause and secondClause impose requirements that cannot both be satisfied?
Fixture only - Classify contract clause kind
What does clause primarily do?
Fixture only - Check whether a comparable property fits the subject
Is comparable similar enough to subject in type, size, age, condition, and location wording to support a valuation comparison?
Fixture only - Check whether consent wording covers a use
Does the wording of consent cover the described use?
Fixture only - Label article facets
Which of these does article include: a clear thesis, supporting evidence, counterarguments, a call to action, and signals of author expertise?
Fixture only - Detect perishable content
Does content contain claims that are likely to go stale, such as prices, versions, dates, current events, or latest wording?
Fixture only - Label what a property disclosure addresses
Which of these does disclosure address: known defects, prior repairs, environmental hazards, boundary or easement issues, association rules?
Fixture only - Identify document purpose
What is the primary purpose of document?
Fixture only - Match records to one entity
Do firstRecord and secondRecord describe the same real-world entity despite formatting, abbreviation, or partial fields?
Earlier-evaluator measurement - Match an expense to a category list
Does expense clearly fall under one of the caller's categories, under several equally, under none, or is it unclear?
Fixture only - Grade extraction fidelity
How faithfully does extracted represent the facts in source, without invented, altered, or dropped values, on a five-level rubric?
Earlier-evaluator measurement - Select a field value
Which supplied candidate is the value of field in document?
Earlier-evaluator measurement - Compare two funding opportunities for a program
Which of firstOpportunity and secondOpportunity better fits program in purpose, eligibility wording, and scope?
Fixture only - Check headline fit
Does headline accurately represent what body says, without promising more than the body delivers?
Fixture only - Label what an invoice states
Which of these does invoice state: vendor identity, an invoice number, a due date, line items, a tax or total breakdown?
Fixture only - Compare two listings for a request
Which of firstListing and secondListing better satisfies request?
Fixture only - Check a listing description against its fact sheet
Does the free-text description contradict the structured facts about the property?
Fixture only - Identify the cause of loss in a claim narrative
What cause of loss does narrative describe: weather, fire, water, theft, collision, wear and tear, vandalism, or something else?
Fixture only - Check lyrics against the mood of the music
Do lyrics fit the mood and pacing of the music described in music?
Fixture only - Label what a methods section states
Which of these does methods state: sample size, data source, analysis method, limitations, and preregistration or a protocol?
Fixture only - Check a narrative for internal contradictions
Are the statements within narrative consistent with each other in timeline, cause, and extent, with no statement contradicting another in the same account?
Fixture only - Label what an official notice states
Which of these does notice state: an action the recipient must take, a deadline, the consequence of inaction, a contact for questions, a right to appeal?
Fixture only - Label what a property offer states
Which of these does offer state: a price, financing, contingencies, a closing or move-in date, an earnest money or security deposit?
Fixture only - Grade the difficulty of a musical passage
How hard is the passage described in passage for a player of instrument, from beginner to virtuoso?
Fixture only - Identify a passage's mood
What mood does the described musical passage express?
Fixture only - Label the statements in a privacy notice
Which of these does notice state: what data is collected, why it is used, how long it is kept, who it is shared with, how to contact the controller?
Fixture only - Check a listing against a request
Does listing describe a product that satisfies what request asks for, including stated must-have attributes?
Fixture only - Label the facets of a grant proposal
Which of these does proposal include: a statement of need, measurable objectives, planned activities, a budget, an evaluation plan?
Fixture only - Match a ledger record to a statement line
Do record and statementLine describe the same transaction, judging payee, purpose, and timing wording?
Fixture only - Check whether a requirement is testable
Decide whether a requirement defines an observable way to distinguish meeting it from failing it.
Fixture only - Label the aspects of a product review
Which of these does review comment on: quality, price, shipping, service, or a defect?
Fixture only - Classify a peer review's recommendation
What recommendation does the wording of review express: accept, minor revision, major revision, or reject?
Fixture only - Classify search intent
What is the searcher behind query trying to do: learn, reach a site, buy, compare before buying, or find something nearby?
Fixture only - Label the facets a settlement offer letter states
Which of these does offer state: the settlement amount, the basis for that amount, a deadline to respond, release terms, and how to dispute or appeal?
Fixture only - Check a patch against a requested sound
Does the patch or preset described in patch deliver the sound character asked for in request?
Fixture only - Identify a musical style
What style does the described passage or request evoke?
Fixture only
retrieval
- Check answerability
Decide whether supplied evidence can answer an entire question.
Unknown response origin - Check a cached answer
Does cachedAnswer address question with the same relevant meaning and conditions as originalQuestion?
Fixture only - Identify a passage role
What role does passage play in answering question?
Fixture only - Compare evidence for conflicts
Do firstPassage and secondPassage give incompatible evidence relevant to question under the same conditions?
Fixture only - Compare the origins of two pieces of evidence
Check whether supplied provenance shows shared or separate evidence origins for one claim, or leaves their relationship unresolved.
Fixture only - Check new evidence
Does passage add material information relevant to question beyond existingEvidence?
Fixture only - Grade evidence strength
How strongly does evidence support the entire claim, on a five-level rubric?
Fixture only - Check information freshness needs
Does question require a current or time-specific state that can change, or stable conceptual knowledge?
Fixture only - Recover a paragraph boundary
Do two adjacent extracted text fragments continue one paragraph or belong to separate blocks?
Earlier-evaluator measurement - Compare two passages for a question
Which of firstPassage and secondPassage better helps answer question?
Fixture only - Compare passages for duplication
How much material information do firstPassage and secondPassage share?
Fixture only - Check passage self-containment
Can passage be understood on its own, without unresolved references to surrounding text?
Fixture only - Compare question meaning
Do firstQuestion and secondQuestion request the same information under the same stated conditions?
Fixture only - Check question specificity
Does question, interpreted with context, identify a focused information need?
Fixture only - Rerank evidence
Select candidate passages by relevance to a query.
Unknown response origin - Check whether retrieval is needed
Does request require facts beyond context?
Fixture only - Check source applicability
Does the scope described in passage apply to scenario?
Fixture only - Identify a text block’s structure
Is extracted text a heading, body paragraph, list item, code, table, caption, formula, or other block?
Earlier-evaluator measurement - Verify claims
Check each supplied claim against its paired evidence.
Unknown response origin
support
- Classify a patient's appointment request
What does the patient primarily want from message: to schedule, reschedule, or cancel an appointment, get results, get a refill, or ask a question?
Fixture only - Check a previously attempted step
Does conversation establish whether the customer already performed step?
Fixture only - Grade the urgency a patient's wording asks for
How urgent is the care that message asks for, from routine to emergency, judged on what the patient's wording says rather than on medical assessment?
Fixture only - Label the facets an insurance claim narrative states
Which of these does claim state: when the incident happened, where it happened, what caused it, what was damaged or lost, and whether there are witnesses or evidence?
Fixture only - Classify a billing dispute
What kind of billing dispute does message raise?
Fixture only - Check whether a message asks for escalation
Does message explicitly ask for escalation to a higher tier, a manager, or on-call engineering?
Fixture only - Classify customer feedback
What is the primary kind of feedback in message?
Fixture only - Detect expressed frustration
Does message express frustration or dissatisfaction in its wording?
Fixture only - Label the booking facets a guest request states
Which of these does request state: travel dates, party size, accessibility needs, a budget, or a special occasion?
Fixture only - Match a known incident
Which supplied incident is supported as a match for ticket?
Fixture only - Classify an instrument problem report
What kind of problem does report describe with a musical instrument: tuning, buzz or rattle, no sound, intonation, mechanical, cosmetic, or unclear?
Fixture only - Label the facets of an instrument repair message
Which of these does message state about an instrument problem: the instrument and model, the symptom, when it started, recent changes, and the environment it is kept in?
Fixture only - Assess reported issue impact
What practical impact does message explicitly describe?
Fixture only - Check whether an issue returned
Label a reported issue as new, ongoing without resolution, or returned after reported recovery.
Fixture only - Flag a maintenance request that describes a safety hazard
Does request describe a safety hazard such as a gas smell, active water intrusion, exposed wiring, structural damage, no heat in cold weather, or a blocked exit?
Fixture only - Detect a medication or dosage mention
Does message mention a medication, supplement, or dosage?
Fixture only - Label the facets of a message
Which of these does message do: ask a question, report a problem, request an action, state a deadline, reference prior contact?
Fixture only - Select a reply template
Which supplied approved template applies to request and context?
Fixture only - Check whether a review response addresses the review
Does response engage with the specific complaints and praise raised in review, rather than offering a generic thank-you or apology?
Fixture only - Route many requests in batches
Which named route handles each of up to 500 requests, judged in batched Jev requests so a queue of tickets, emails, or events is routed in a handful of calls instead of one per item?
Earlier-version measurement - Grade expressed satisfaction
How much satisfaction with the outcome does message express at the close of an interaction, on a five-level rubric?
Fixture only - Compare sentiment across messages
How does the sentiment expressed in laterMessage compare with earlierMessage from the same person?
Fixture only - Classify a reported shipping problem
What shipping problem does message report?
Fixture only - Label the symptom facets a patient message states
Which of these does message state about a symptom: when it began, how bad it is, how long it has lasted, what makes it better or worse, what the patient has already tried?
Fixture only - Compare support tickets
Do firstTicket and secondTicket describe the same underlying reported issue?
Fixture only - Check troubleshooting applicability
Does procedure address symptoms under the described circumstances?
Fixture only - Detect an explicit urgency request
Does message explicitly request urgent attention?
Fixture only - Check workaround fit
Can workaround address issue without violating constraints?
Fixture only
memory
- Prune agent context
Which of items, earlier tool results and messages in an agent session, are still needed to finish objective, so the rest can be dropped from context?
Earlier-evaluator measurement - Assess fact stability
Is fact about an enduring or historical attribute, or a state that is expected to change?
Fixture only - Compare a new fact with memory
How does newFact relate to existingMemory?
Fixture only - Identify memory scope
What is the narrowest explicitly supported scope of fact in context?
Fixture only - Identify whom a memory describes
Label whether one candidate memory describes the user, someone else, or a group including the user.
Fixture only - Assess a candidate memory
How useful is fact for future work under purpose?
Fixture only - Identify a stated preference
Does statement express an ongoing preference, a factual assertion, or a temporary request?
Fixture only