Article

A Good Demo Is Not an AI Product: How I Evaluate AI Features Before Shipping

A practical way to evaluate AI features before release: real test cases, grounded answers, useful graders, human review, regression checks, and the product signals that matter after launch.

A Good Demo Is Not an AI Product: How I Evaluate AI Features Before Shipping cover

AI demos are easy to fall in love with.

You prepare a clean input. You ask a clear question. The model returns a sharp answer in a few seconds. Someone on the team says, “This is actually really good,” and for a moment the feature feels almost finished.

Then a real user arrives.

They write half a sentence. They use a term that never appeared in the prompt. The source is missing the answer. Two documents disagree. The relevant passage is buried near the end of a long transcript. The model gives a confident answer anyway, and the interface makes it look more trustworthy than it is.

This is usually the point where an AI feature stops being a demo and starts becoming a product.

The difficult part is not making the model produce an impressive response. The difficult part is knowing whether the whole feature behaves well across the messy situations users will actually create.

That is what evals are for.

I do not think about evals as a research ceremony or a dashboard full of decimal scores. I think of them as a practical agreement with the product:

these are the jobs the feature should handle, these are the failures we care about, and this is how we will notice when a change makes things worse.

That agreement can start small. In many products, it should.

The demo is the friendliest possible user

A demo usually contains a hidden advantage: the person running it already understands the system.

They know which question will work. They know which document contains the answer. They avoid the awkward wording that confuses retrieval. If the first response is weak, they quietly try again with a better prompt.

None of that is dishonest. It is how prototypes are explored.

The problem begins when we treat that exploration as evidence that the product is reliable.

If I test an AI feature by chatting with it until I get a good answer, I am not really testing the feature. I am testing my ability to find inputs it handles well.

Real users do not know the hidden rules. They should not need to.

They will ask broad questions, underspecified questions, questions with typos, and questions whose answers are not in the available material. They will paste too much content. They will click twice. They will expect the second request to remember the first. They will read a polished answer as a promise, even when the model was guessing.

A useful evaluation starts by leaving the demo path on purpose.

An eval is a test with room for uncertainty

Traditional software gives us many clean assertions.

The API should return 401 when the session is missing. The calculation should produce the expected number. The button should stay disabled while the request is running. The database should contain one record, not two.

AI output is often less tidy.

Two answers can use different words and both be correct. A concise answer may be better than a longer one even if it contains fewer details. A useful summary does not have one perfect reference string. The same model can produce slightly different results from the same input.

That does not make testing impossible. It means the test has to match the kind of decision we are making.

Some things remain exact:

  • did the system use the correct source?
  • did it return a real timestamp?
  • did it call the right tool with valid arguments?
  • did it avoid an action that required confirmation?
  • did the structured output match the schema?
  • did it stay inside the latency and cost budget?

Other things need judgment:

  • is the answer supported by the source?
  • does it answer the user's actual question?
  • is an important qualification missing?
  • is the tone appropriate for the situation?
  • should the system have said it did not know?

An eval is simply a repeatable way to make those judgments.

The OpenAI evaluation guide describes evals as structured tests for measuring an AI system despite the variability of generative models. I like that definition because it keeps the idea grounded. We are not trying to remove uncertainty from the model. We are trying to make the uncertainty visible enough to manage.

I would not tie that process to one vendor dashboard. Platforms and APIs change. The valuable part is the set of cases, expectations, and failure definitions that belongs to the product team.

Test the product, not only the model

When an AI answer is wrong, the model is an easy suspect.

Sometimes it is the model.

But an AI feature is usually a pipeline, and the failure may have happened long before the final response appeared.

Imagine a feature that answers questions about a long video and returns a timestamp.

The product may need to:

  1. obtain the transcript;
  2. divide it into useful passages;
  3. find the passages relevant to the question;
  4. give those passages to the model with clear instructions;
  5. generate an answer grounded in that material;
  6. attach the correct timestamp;
  7. show uncertainty when the source does not support an answer.

If the final answer is wrong, several different things may have happened.

The transcript could be inaccurate. The right passage may never have reached the model. The context selection step may have selected a similar but irrelevant section. The model may have ignored the evidence. The answer may be correct but linked to the wrong timestamp. The backend may have done everything correctly while the interface hid the source.

One score called “answer quality” will not tell me which problem I have.

This is why I try to evaluate the layers separately.

For a source-grounded answer, I want to know at least:

  • whether the system found the right evidence;
  • whether the answer stayed inside that evidence;
  • whether the citation or timestamp points to the supporting part;
  • whether the system handled a missing answer honestly;
  • whether the result arrived quickly enough to be useful.

That is a product evaluation, not just a model evaluation.

It tests the thing the user experiences.

Start with the failure you cannot accept

It is tempting to begin with a metric.

Should groundedness be above 0.9? Should the judge score answers from one to five? How many examples do we need before the result is statistically useful?

Those questions matter later. First I want a more human conversation:

what kind of failure would make us regret shipping this feature?

For a video question-answering tool, my list might include:

  • inventing a claim that does not appear in the video;
  • attaching a timestamp that does not support the answer;
  • blending statements from different speakers as if they were one opinion;
  • answering confidently when the transcript contains no answer;
  • exposing text from content the user should not be able to access;
  • taking so long that searching manually would have been faster.

For an agent with tool access, the unacceptable failures are different:

  • choosing the wrong account or environment;
  • modifying data when the user asked only for an explanation;
  • sending a draft without approval;
  • calling the right tool with the wrong parameters;
  • reporting success when the external system did not change;
  • repeating an action after a retry and creating a duplicate.

This list gives the evaluation a reason to exist.

It also stops the team from optimizing a convenient score while missing the failure users will remember.

I would start with twenty useful cases

You do not need a giant evaluation platform to begin.

I would rather have twenty carefully chosen cases that reflect the product than two thousand synthetic prompts nobody has read.

The first set can come from work the team already does manually:

  • the obvious happy paths used in demos;
  • questions people naturally ask during user testing;
  • failures found while building the feature;
  • examples from support conversations or bug reports;
  • edge cases the team is worried about;
  • cases where the correct behavior is to refuse, ask, or say “I don't know.”

Anthropic's practical guide to agent evals similarly recommends starting early with a small set of real tasks rather than waiting for a perfect benchmark. The important part is not the exact number. It is that each case represents a behavior the product genuinely needs.

For the video example, a small starting set could contain:

  • a direct factual question with one clear supporting passage;
  • a question whose answer is spread across two nearby passages;
  • a question using different words from the transcript;
  • a question about something mentioned several times;
  • a question with a plausible but unsupported premise;
  • a question the video does not answer;
  • a transcript with a name or technical term transcribed incorrectly;
  • a long video where the relevant moment appears near the end;
  • a follow-up question that depends on the previous answer;
  • an instruction that tries to make the system ignore its source.

That set is already more informative than repeatedly asking the nicest question in the demo.

A useful case needs a clear reason to pass

An evaluation case should not be a mysterious prompt with one preferred paragraph attached to it.

It should explain what matters.

For example:

type EvalCase = {
  id: string;
  question: string;
  source: TranscriptChunk[];
  expected: {
    answerable: boolean;
    requiredFacts?: string[];
    supportingChunkIds?: string[];
    forbiddenClaims?: string[];
  };
};

This is deliberately simple.

The expected result does not prescribe the exact sentence the model must write. It captures the facts that have to be present, the evidence that should support them, and the claims the system must not invent.

For a case with no answer in the source, answerable: false is often more valuable than a reference response. It lets us check whether the product knows when to stop.

The same principle works for agents. I care more about the final state than one perfect sequence of tool calls. If the user asked to create a draft, the important outcome is that a draft exists and nothing was sent. The agent may reach that state in more than one valid way.

Overly rigid tests can punish a correct solution simply because it took a path the test author did not imagine.

Build a balanced set, not a flattering one

A test set can make a weak product look excellent if it asks only one side of the question.

If I test only whether the system answers when information is present, I may accidentally reward a system that always answers.

If I test only whether an agent calls a search tool when fresh information is needed, I may end up with an agent that searches for everything.

If I test only whether safety rules block dangerous actions, I may create a feature that refuses harmless work and feels impossible to use.

Good eval sets need both directions:

  • answer when the evidence is sufficient, and abstain when it is not;
  • use a tool when it helps, and avoid it when it adds no value;
  • ask for confirmation before a consequential action, and continue smoothly through a harmless one;
  • include the necessary detail, and avoid burying the user in everything the model knows.

This balance matters because a product can improve one metric by becoming worse in a different way.

A lower hallucination rate is not a clear win if the feature now refuses half the questions it could answer well.

Separate hard checks from human judgment

I prefer deterministic checks whenever the result can be verified directly.

They are cheap, fast, and easy to understand.

A normal test can verify that:

  • JSON matches the expected schema;
  • every citation refers to an existing source chunk;
  • a timestamp is inside the video's duration;
  • a tool name is allowed;
  • required parameters are present;
  • no write action occurred in a read-only scenario;
  • the response finished before a defined timeout;
  • token use stayed below a practical limit.

These checks do not need another model to judge them.

Then there are qualities that are harder to reduce to code: groundedness, completeness, usefulness, tone, or whether an answer quietly changes the meaning of the source.

Those are good candidates for a clear rubric and human review. A model-based grader can help scale the review, but I do not want it to become an invisible oracle.

An LLM judge is useful, but it is still a model

Using one model to grade another can feel circular.

It can still be useful.

A judge can compare an answer with the source, classify a failure, or score a specific property across hundreds of cases much faster than a person can. The quality improves when the question is narrow:

Does every factual claim in the answer have support in the supplied source?

That is better than:

Is this a good answer?

The second question hides too many decisions inside one word.

When I use a model-based grader, I want:

  • one criterion at a time;
  • a short rubric with concrete pass and fail conditions;
  • the source, user question, and answer available to the grader;
  • examples of difficult borderline cases;
  • periodic comparison with human decisions;
  • a way to inspect the grader's explanation when the score looks wrong.

For subjective comparisons, pairwise judgment can also be more useful than an absolute score. Asking which of two answers is better supported is often easier than deciding whether one answer deserves 73 or 81 points.

The point is not to automate judgment away.

The point is to use automation for coverage while keeping people involved in defining and checking the standard.

Read the failures, not only the score

A score is useful for noticing movement.

It is not an explanation.

Suppose a new retrieval strategy moves the pass rate from 78% to 84%. That sounds good. But I still need to look at what changed.

Maybe it fixed several common questions and broke one rare edge case. That may be a sensible trade-off.

Maybe it improved style scores while creating two unsupported claims in a high-risk workflow. That is not a sensible trade-off.

Maybe the grader changed its mind about nearly identical answers. Then the movement is noise, not product progress.

I want to read examples from four groups:

  • cases that improved;
  • cases that regressed;
  • failures with the highest user impact;
  • cases where automated and human judgments disagree.

This is where evals become engineering rather than reporting.

The score tells me where to look. The examples tell me what to change.

One average can hide the failure that matters

A single quality score is attractive because it makes decisions feel clean.

Version A scored 82. Version B scored 86. Ship version B.

Real products are rarely that simple.

I prefer a small scorecard with separate dimensions:

  • task completion;
  • groundedness;
  • citation correctness;
  • appropriate uncertainty;
  • latency;
  • cost per successful task;
  • critical safety or permission failures.

Some of these are metrics. Some are gates.

If a harmless formatting case gets worse, the overall decision may still be positive. If the system starts performing writes without confirmation, no average should wash that away.

The severity of failure matters.

This is also why model quality and product success should stay separate. A feature can achieve excellent eval results and still solve the wrong problem. Google's guidance on measuring ML project success makes the same distinction: model metrics do not automatically prove a business or user outcome.

The eval can tell me that answers are grounded and citations work. It cannot, by itself, tell me whether people find the feature worth returning to.

Latency and cost belong in the evaluation

Quality is not the only thing that can regress.

A prompt change may improve answer completeness while doubling response time. A larger model may pass three additional edge cases while making every common request much more expensive. A retrieval step may increase grounding but add so much delay that the workflow feels broken.

Those are product changes.

I like measuring cost and latency per completed task, not only per request.

A cheap answer that the user has to regenerate three times is not really cheap. A fast answer with the wrong timestamp is not really fast. The useful unit is the job the user came to finish.

This gives model selection a healthier shape.

The question stops being “which model is best?” and becomes:

Which configuration clears our quality bar for this job with acceptable speed and cost?

Sometimes the answer is a stronger model. Sometimes it is better retrieval, a clearer prompt, a smaller context, a deterministic check, or a simpler product promise.

The interface is part of the result

Backend evals can pass while the user experience still creates false trust.

Consider two interfaces showing the same uncertain answer.

One presents it as a confident paragraph under a large AI icon.

The other says that the source does not fully answer the question, shows the closest supporting passage, and gives the user a direct way to inspect it.

The model output may be identical. The product behavior is not.

For AI features, I want to test the surrounding experience too:

  • can the user see the source?
  • is uncertainty visible before they act on the answer?
  • is a generated draft clearly different from a sent message?
  • does a retry replace, duplicate, or append the result?
  • can the user recover when a tool call fails?
  • does the loading state explain what is happening?
  • can a consequential action be reviewed and cancelled?

Not every one of these checks belongs in an LLM eval. Some belong in normal integration tests, browser tests, or a manual product review.

That is fine.

The goal is to evaluate the feature, not to force every quality question into the same tool.

A practical evaluation loop from real tasks to release checks and production feedback

Production failures should become permanent tests

Before launch, the test set reflects what the team can imagine.

After launch, users contribute examples nobody imagined.

That is valuable evidence, as long as the product is designed to learn from it responsibly.

When a bad answer reaches production, I do not want the fix to end with a prompt edit.

I want to preserve the case:

  1. remove or protect any personal data;
  2. capture the relevant input, context, output, and system version;
  3. describe why the behavior was wrong;
  4. add the case to the regression set;
  5. confirm that the fix solves this case without breaking nearby ones.

Now the failure has changed the quality bar.

The next model upgrade, prompt rewrite, retrieval change, or tool refactor has to face it again.

This is one of the most useful habits in AI product work. Production stops being only the place where failures hurt. It becomes the place where the eval set becomes more honest.

The system should not log everything blindly. User data, private sources, and sensitive conversations need retention rules, access controls, and careful redaction. “We may want it for evals later” is not a reason to collect data the product does not need or does not have permission to keep.

A small-team workflow that is actually manageable

If I were adding evals to a small product today, I would keep the first version simple.

1. Name one critical job

Not “make the assistant better.”

Something concrete:

Answer a question from a video and link to the moment that supports the answer.

2. Write the quality bar in plain language

Define what must happen, what must never happen, and what the system should do when evidence is missing.

If two people on the team interpret success differently, the requirement is not ready yet.

3. Collect twenty to fifty cases

Use real manual checks, realistic examples, known failures, edge cases, and negative cases. Keep them understandable enough that a person can explain why each one passes or fails.

4. Add deterministic checks first

Validate schemas, citations, timestamps, tool arguments, permissions, latency, and cost with normal code wherever possible.

5. Add narrow judgment rubrics

Use human review for the first pass. Add a model grader only after the team can state the criterion clearly and compare its decisions with human judgment.

6. Establish a baseline

Run the current product before trying to improve it. Otherwise there is no reliable way to know whether the change helped.

7. Compare changes case by case

Look beyond the average. Read improvements and regressions, especially in high-impact scenarios.

8. Keep the release gate proportional to risk

A copy suggestion tool and an agent that can update customer records do not need the same threshold. The cost of being wrong should shape the strictness of the evaluation.

9. Feed real failures back into the set

Every meaningful production miss should leave behind a safer system, not only a closed bug report.

None of this requires a large AI research team.

It requires someone to own the question: what does “good enough to ship” mean for this feature?

What I would check before release

Before shipping a meaningful AI feature, I want clear answers to these questions:

  • What exact user job are we evaluating?
  • Does the test set resemble real inputs, including messy ones?
  • Do we test when the system should not answer or act?
  • Can we separate retrieval failures from generation failures?
  • Are factual claims connected to inspectable evidence?
  • Are deterministic checks doing the work that does not need an LLM judge?
  • Has a person reviewed a useful sample of passes and failures?
  • Do automated graders agree with human judgment often enough to be useful?
  • Are critical failures treated as gates instead of averaged away?
  • Are latency and cost measured alongside quality?
  • Does the interface communicate uncertainty and action state honestly?
  • Can production failures be added safely to the regression set?
  • Do we know how to roll back a model, prompt, or retrieval change?

If the team cannot answer these yet, the feature may still be ready for a small experiment.

It is probably not ready to be trusted quietly at scale.

Evals do not replace product judgment

There is a danger in becoming too comfortable with a test suite.

Once a number exists, it is easy to optimize it.

The team tunes prompts for the known cases. The model learns the shape of the rubric. The score rises. Everyone feels safer.

Meanwhile, user behavior changes. The content changes. The feature grows new tools. The old test set becomes easier and less representative.

Evals need maintenance for the same reason products do: the world around them does not stay still.

I see them as one layer in a larger quality system:

  • automated checks catch repeatable failures before release;
  • human review catches nuance and broken assumptions;
  • production monitoring shows what happens at scale;
  • user feedback reveals problems the team did not predict;
  • product metrics show whether the feature is useful at all.

No single layer is enough.

A good eval suite does not prove that an AI feature is safe, useful, or finished. It gives the team a much better way to see what it knows, what it does not know, and what changed.

That is already a major improvement over intuition.

Conclusion

A good demo answers the question you hoped someone would ask.

A good product also survives the questions you did not prepare for.

That is why I do not want to evaluate an AI feature by asking whether the model looks intelligent. I want to know whether the product completes a real job, uses the right evidence, behaves honestly when it is uncertain, and stays inside the boundaries the user expects.

The first eval does not need to be sophisticated.

It can be twenty cases in a file, a few deterministic checks, a small rubric, and an hour spent reading failures carefully.

What matters is the habit.

Define the job. Name the unacceptable failures. Test both sides of the behavior. Keep people involved in judgment. Measure cost and latency. Turn real mistakes into permanent regression cases.

Then, when a new prompt or model produces a beautiful answer, you can enjoy the demo without confusing it for proof.

FAQ

What is an LLM eval?

An LLM eval is a repeatable test that checks whether an AI system meets a defined expectation. Depending on the feature, it may use normal code, human review, a model-based grader, or a combination of all three.

How many evaluation cases do I need to start?

There is no universal number. For a small product, twenty to fifty carefully chosen cases can expose more useful problems than a large synthetic dataset. The set should grow as the team learns from user behavior and production failures.

Can an LLM reliably grade another LLM?

It can help when the criterion is narrow and the rubric is clear, but it should not be treated as unquestionable ground truth. Compare model grades with human judgment, inspect disagreements, and use deterministic checks whenever the result can be verified directly.

Should evals run before every release?

Critical regression cases should. Larger or more expensive evaluation suites can run on model changes, prompt changes, retrieval changes, scheduled checks, or before important releases. The frequency should match the risk and cost of the feature.

Are high eval scores enough to ship?

No. Evals measure the criteria represented in the test set. A release decision should also consider privacy, security, latency, cost, user experience, monitoring, rollback, and whether the feature solves a meaningful user problem.

Share this post

Send it to someone who might find it useful.

Discussion

Responses

0 responses

No approved responses yet.

Response

Join the discussion

Moderated
Guest responseGuestReviewed before publishing