LLM evaluation

Use semantic similarity to a reference

Hard120 pts~45 min
  • Semantic similarity
  • Thresholds
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-use-semantic-similarity-to-a-reference

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Score sampled answers against a reference answer with a similarity metric and a threshold.

Your task

  1. 1Reference: "You can request a refund within 30 days of delivery."
  2. 2Ask "How long do I have to request a refund?" at temperature 1 with seeds 1–3.
  3. 3Score each output by reference-word recall: the share of the reference's words that appear in the output.
  4. 4Assert every score ≥ 0.4 while at least one output differs from the reference text.

Acceptance criteria

  • POST /ai/chat returns 200
  • At least 3 assertions pass

LLM evaluation · AI Testing · Output validation