LLM Factual Match

import langwatch

df = langwatch.dataset.get_dataset("dataset-id").to_pandas()

evaluation = langwatch.evaluation.init("my-incredible-evaluation")

for index, row in evaluation.loop(df.iterrows()):
    # your execution code here       
    evaluation.run(
        "ragas/factual_correctness",
        index=index,
        data={
            "output": output,
            "expected_output": row["expected_output"],
        },
        settings={}
    )

[
  {
    "status": "processed",
    "score": 123,
    "passed": true,
    "label": "<string>",
    "details": "<string>",
    "cost": {
      "currency": "<string>",
      "amount": 123
    },
    "raw_response": {},
    "error_type": "<string>",
    "traceback": [
      "<string>"
    ]
  }
]

POST

ragas

factual_correctness

evaluate

import langwatch

df = langwatch.dataset.get_dataset("dataset-id").to_pandas()

evaluation = langwatch.evaluation.init("my-incredible-evaluation")

for index, row in evaluation.loop(df.iterrows()):
    # your execution code here       
    evaluation.run(
        "ragas/factual_correctness",
        index=index,
        data={
            "output": output,
            "expected_output": row["expected_output"],
        },
        settings={}
    )

[
  {
    "status": "processed",
    "score": 123,
    "passed": true,
    "label": "<string>",
    "details": "<string>",
    "cost": {
      "currency": "<string>",
      "amount": 123
    },
    "raw_response": {},
    "error_type": "<string>",
    "traceback": [
      "<string>"
    ]
  }
]

Authorizations

X-Auth-Token

string

header

required

Body

application/json

data

object

required

Show child attributes

data.output

string

required

The output text to evaluate

data.expected_output

string

required

Expected output for comparison

settings

object

Evaluator settings

Show child attributes

settings.model

string

default:openai/gpt-5

The model to use for evaluation.

settings.max_tokens

number

default:2048

The maximum number of tokens allowed for evaluation, a too high number can be costly. Entries above this amount will be skipped.

settings.mode

string

default:f1

The mode to use for the factual correctness metric.

settings.atomicity

string

default:low

The level of atomicity for claim decomposition.

settings.coverage

string

default:low

The level of coverage for claim decomposition.

Response

Successful evaluation

status

enum<string>

required

Available options:

processed,

skipped,

error

score

number

Evaluation score

passed

boolean

Whether the evaluation passed

label

string

Evaluation label

details

string

Additional details about the evaluation

cost

object

Show child attributes

cost.currency

string

required

cost.amount

number

required

raw_response

object

Raw response from the evaluator

error_type

string

Type of error if status is 'error'

traceback

string[]

Error traceback if status is 'error'

Context Recall Ragas Faithfulness

⌘I

Get Started

Observability

Agent Simulations

Evaluation

Prompt Management

Platform

Examples & Cookbooks

Authorizations

Body

Response