LLM Answer Match

import langwatch

df = langwatch.dataset.get_dataset("dataset-id").to_pandas()

evaluation = langwatch.evaluation.init("my-incredible-evaluation")

for index, row in evaluation.loop(df.iterrows()):
    # your execution code here       
    evaluation.run(
        "langevals/llm_answer_match",
        index=index,
        data={
            "input": row["input"],
            "output": output,
            "expected_output": row["expected_output"],
        },
        settings={}
    )

[
  {
    "status": "processed",
    "score": 123,
    "passed": true,
    "label": "<string>",
    "details": "<string>",
    "cost": {
      "currency": "<string>",
      "amount": 123
    },
    "raw_response": {},
    "error_type": "<string>",
    "traceback": [
      "<string>"
    ]
  }
]

POST

langevals

llm_answer_match

evaluate

import langwatch

df = langwatch.dataset.get_dataset("dataset-id").to_pandas()

evaluation = langwatch.evaluation.init("my-incredible-evaluation")

for index, row in evaluation.loop(df.iterrows()):
    # your execution code here       
    evaluation.run(
        "langevals/llm_answer_match",
        index=index,
        data={
            "input": row["input"],
            "output": output,
            "expected_output": row["expected_output"],
        },
        settings={}
    )

[
  {
    "status": "processed",
    "score": 123,
    "passed": true,
    "label": "<string>",
    "details": "<string>",
    "cost": {
      "currency": "<string>",
      "amount": 123
    },
    "raw_response": {},
    "error_type": "<string>",
    "traceback": [
      "<string>"
    ]
  }
]

Authorizations

X-Auth-Token

string

header

required

Body

application/json

data

object

required

Show child attributes

data.input

string

The input text to evaluate

data.output

string

The output text to evaluate

data.expected_output

string

Expected output for comparison

settings

object

Evaluator settings

Show child attributes

settings.model

string

default:openai/gpt-5

The model to use for evaluation

settings.max_tokens

number

default:8192

Max tokens allowed for evaluation

settings.prompt

string

default:"Verify that the predicted answer matches the gold answer for the question. Style does not matter

Prompt for the comparison

Response

Successful evaluation

status

enum<string>

required

Available options:

processed,

skipped,

error

score

number

Evaluation score

passed

boolean

Whether the evaluation passed

label

string

Evaluation label

details

string

Additional details about the evaluation

cost

object

Show child attributes

cost.currency

string

required

cost.amount

number

required

raw_response

object

Raw response from the evaluator

error_type

string

Type of error if status is 'error'

traceback

string[]

Error traceback if status is 'error'

Exact Match Evaluator LLM-as-a-Judge Boolean Evaluator

⌘I

Get Started

Observability

Agent Simulations

Evaluation

Prompt Management

Platform

Examples & Cookbooks

Authorizations

Body

Response