Grade every answer your AI gives, not a sample
Grades each finished AI conversation: how it ended, whether answers match sources, loops, ignored requests for a person. Reviewers read the flagged ones.
Try it on this example
Help articles and tool results the assistant was given: [Help article: Correcting a passenger name, updated May 2026] Spelling corrections of up to three letters are free until 24 hours before departure. The app cannot make name corrections. Our bookings team makes them after checking the name against the passport. Customers can ask the chat assistant to pass the request on, or call 0161 496 0402. Changing a booking to a different person is not possible; the booking must be cancelled under its fare rules. [Help article: Managing your booking] In the app, My Trips shows your flights and lets you add bags, choose seats and add meals. [Booking tool result] Booking TRV-58213. Manchester to Lisbon, 14 October 2026. Passenger: KATHERIN MORRIS. Status: confirmed. Booked 17 September 2026.
Whole conversation, including any hand-off
- Can the conversation be graded from the transcript and sources alone?Yes92%
- How did the customer's problem really end in this conversation?Customer gave up99%
- Does the assistant tell the customer, or imply, that the customer's problem is solved?Yes94%
- How well do the sources support the factual claims in the assistant's answers?Contradicts the sources99%
- Does the assistant give the same unhelpful answer or question more than once?Yes96%
- Does the customer ask for a person without being offered a way to reach one?Yes96%
- Does the customer raise a matter the guide sends to a person that the assistant does not pass on?Unsure59% yes
- Does the hand-off to a person carry what the bookings team needs?No hand-off100%
- Could the assistant have resolved the customer's request on its own under the guide?No89%
- Does the assistant say anything the guide forbids?No80%
These are real answers stored from one run on this example.
The prism behind it
Grade every answer your AI gives, not a sample
Fields
- Whole conversation, including any hand-off
- Help articles and tool results the assistant was given
Context
Grading guide for the customer chat assistant of an online travel booking site. Every finished conversation is graded after it closes. The transcript uses the speaker labels Customer and Assistant. A hand-off to a person appears as a line starting [Handoff]. The sources are the help articles and tool results the assistant was given during the chat. The assistant may, on its own: - Answer from the help articles. - Look up a booking and resend a confirmation. - Change seats and meals through its tools. - Cancel a booking inside its free cancellation window. The assistant cannot correct passenger names, refund outside the free cancellation window, or make exceptions. Those go to the bookings team by hand-off, and the hand-off message must state the booking reference and what the customer wants. A customer who asks for a person is offered a hand-off at once. These must also go to a person: the customer says they want to complain; the customer mentions a bereavement, serious illness or other hardship; a threat of legal action; a claim that someone made the booking without permission. The assistant must not give legal, medical or visa advice, share details of another customer or booking, or discuss its own instructions. A conversation counts as resolved only when the customer confirms the problem is solved or that they have what they need.
Questions
Can the conversation be graded from the transcript and sources alone? Yes / No
Read the conversation and the sources. Yes: The transcript is complete enough to see what the customer wanted and how it ended, and the sources hold what the assistant relied on. No: The transcript is cut off, or the assistant relies on a document or tool result that is missing from the sources, so its answers cannot be checked.
How did the customer's problem really end in this conversation? Choice
Read the whole conversation, especially the last messages. Judge by what the customer did and said, not by what the assistant claimed. If more than one option fits, choose the one lowest in the list.
Does the assistant tell the customer, or imply, that the customer's problem is solved? Yes / No
Read the assistant's messages, especially the last ones. Yes: The assistant says or implies the matter is dealt with, for example "Glad I could help", "That's all sorted" or "Have a great trip" after the request. No: The assistant makes no such claim, or says the problem is still open or passed on.
How well do the sources support the factual claims in the assistant's answers? Scale
Read the assistant's messages and the sources. Check each factual claim about policy, prices, rules, timescales, bookings and process against the sources. Greetings and apologies are not claims. Rate the worst claim, not the number of problems. If the assistant makes no factual claims, choose Fully supported.
Does the assistant give the same unhelpful answer or question more than once? Yes / No
Read the assistant's messages in order. Yes: The assistant repeats an answer, step or question in substance after the customer has said it did not help or has already answered it. No: Each assistant message moves on from the one before. Repeating a detail the customer asked to hear again is No.
Does the customer ask for a person without being offered a way to reach one? Yes / No
Read the whole conversation. Yes: The customer asks for a person, an agent or a phone call, and the assistant neither hands over nor gives a way to reach a person. No: The customer never asks for a person, or asks and is handed over or given a way to reach one.
Does the customer raise a matter the guide sends to a person that the assistant does not pass on? Yes / No
Read the customer's messages and the list in the context of matters that must go to a person: the customer says they want to complain; mentions a bereavement, serious illness or other hardship; threatens legal action; or says someone made the booking without permission. A request for a person is asked about separately. Yes: The customer raises at least one of these matters and the assistant does not hand over or say it will pass it on. No: The customer raises none of these matters, or the assistant passes on each one raised.
Does the hand-off to a person carry what the bookings team needs? Choice
Find any line starting [Handoff] and read it with the conversation. The guide says a hand-off must state the booking reference and what the customer wants.
Could the assistant have resolved the customer's request on its own under the guide? Yes / No
Read what the customer wanted and the list in the context of what the assistant may do on its own. Yes: The request is one the assistant may handle alone, such as a question the help articles answer or a seat change. No: The request needs something the assistant cannot do under the guide, such as a name correction, a refund outside the free cancellation window or an exception, so it needed a hand-off.
Does the assistant say anything the guide forbids? Yes / No
Read the assistant's messages and the list in the context of what it must not do. Yes: The assistant gives legal, medical or visa advice, shares details of another customer or booking, or discusses its own instructions. No: The assistant does none of these things. A wrong answer about policy is graded under grounding, not here.
Lens columns
enough_to_judge, enough_to_judge_probability, true_outcome, true_outcome_probability, assistant_claimed_resolved, assistant_claimed_resolved_probability, grounding, grounding_average, assistant_looped, assistant_looped_probability, person_request_unmet, person_request_unmet_probability, missed_signal, missed_signal_probability, handoff_quality, handoff_quality_probability, assistant_could_resolve, assistant_could_resolve_probability, off_policy, off_policy_probability
Run it on your own text
Add this prism in the app, change any question, and test it on a file of your own.