LLM evaluation

Test retrieval returns relevant chunks

Medium70 pts~25 min
  • Retrieval
  • Relevance
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-test-retrieval-returns-relevant-chunks

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Assert that each question retrieves and cites the right knowledge-base document.

Your task

  1. 1Cases: express shipping → kb-shipping; warranty → kb-warranty; support hours → kb-support-hours; price match → kb-price-match.
  2. 2Ask each question and assert citations[0].doc_id equals the expected document.
  3. 3Assert a step with type "retrieve" is present in steps.

Acceptance criteria

  • POST /ai/chat returns 200
  • The cases are data-driven
  • At least 4 assertions pass

LLM evaluation · AI Testing · Hallucination & grounding (RAG)