LLM evaluation

Test a prompt across input variations

Medium70 pts~25 min
  • Robustness
  • Paraphrase testing
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-test-a-prompt-across-input-variations

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Check that paraphrases of the same question produce the same grounded answer.

Your task

  1. 1Build variations: "How long do I have to request a refund?", "what is your refund window", "Can I get my money back?", "REFUND POLICY??".
  2. 2POST BASE_URL + "/ai/chat" with { "messages": [{ "role": "user", "content": <prompt> }], "temperature": 0 } for each.
  3. 3Assert every answer contains “30 days” and cites kb-refunds.

Acceptance criteria

  • POST /ai/chat returns 200
  • The variations are data-driven
  • At least 4 assertions pass

LLM evaluation · AI Testing · Prompt & input testing