Imagine a golf app tells you that you hit 64% of fairways last month. The figure sounds specific, so you believe it. But you never recorded a single fairway. The model saw an empty field, filled it with something that looked like golf data, and wrapped it in a friendly sentence.

One number like that can end a user's trust. The player checks it, finds it's wrong, and starts doubting every insight the product writes, including the correct ones.

This post explains how we prevent that in ScoreSmart.ai, a golf performance platform for players and coaches that we built and shipped in 2026. Its AI writes an insight after each round and a weekly report. Before any of that text is shown or saved, a check on the server compares every number in it against the figures the database supplied. The approach works for any product where a model writes about a user's own data, golf or not.

Key takeaways

  • Let the database own the numbers and the model own the words. Don't ask a model to calculate anything your code can calculate.
  • Check every number the model writes against the values you gave it, on the server, before the text is shown or saved.
  • If a check fails, ask once more and name the wrong numbers. If it fails again, show plain text built from the same data.
  • Decide carefully which numbers the model may use, and leave out placeholders and impossible counts.
  • Test the failure case. Feed in a wrong number on purpose and prove the product blocks it.

Why models invent numbers

A language model predicts likely text. It has no built-in way to check a figure against your database. Ask it about a golfer's round and it writes something that reads like a golf insight, and golf insights contain numbers. When the right number is in the prompt, the model usually copies it. When it's missing, or the model loses track of it, the model picks a number that fits the sentence.

OpenAI's own research says training and evaluation reward this. In a September 2025 paper on why models hallucinate, the authors argue that "models are encouraged to guess rather than say 'I don't know.'" A guess scores better than a blank, so models learn to guess.

Structured output doesn't solve it. Tools like the Vercel AI SDK can make a model return data in a fixed shape that you validate against a schema. That confirms the shape and the types. It can't confirm that a number is true, and the AI SDK documentation warns that structured output can still be incorrect or incomplete.

The cost of getting this wrong is real. In 2024 a Canadian tribunal ordered Air Canada to compensate a customer after its website chatbot gave him wrong information about a bereavement fare. The case wasn't about numbers, but it shows who answers for automated text. The airline argued that the chatbot was responsible for its own words, and the tribunal rejected that. If your product shows it, your product said it.

The principle: the database owns the numbers

The most important decision in ScoreSmart was a boundary, not a technique. Our code calculates every figure. The model only writes sentences about figures it has been given.

Users see the AI in two places:

  • a short insight after they finish a round;
  • the written commentary in their weekly report.

Everything else, including handicap tracking, shot analysis and the "scoring leaks" that show where a player loses strokes, is ordinary code that does the same thing every time. Most of a product's intelligence can live in plain arithmetic that you can test, and usually it should.

The weekly report goes a step further. Its layout, its data cards and its links are built by code. The model fills in the commentary. It can't add a card, change a link or decide the report should look different this week.

The model writes the sentences. It doesn't decide what's true or where a button goes.

AI

How a grounding check works

A grounding check sits between the model and the user, on the server. A typical loop looks like this.

A grounding loop, step by step

Step
What happens
Why
1. Build the facts
Code calculates the figures and lists the numbers the model may use
Anything not on the list can be rejected
2. Generate
The model writes its text, validated against a schema
This confirms the shape, not the truth
3. Extract
Every number in the text is pulled out and tidied up
Formatting can't hide a mismatch
4. Check
Each number must match a permitted value, allowing only for rounding
Anything else fails
5. Retry once
The model is asked again, told which numbers were wrong
One clear correction fixes most failures
6. Fall back
If it fails again, the product uses text assembled by code from the same facts
The user only ever sees supplied facts

We allow one retry. Every attempt costs time and money, and the first correction is where nearly all the benefit is. One retry and then the fallback keeps the worst case predictable.

In outline:

facts    = compute the figures in code
allowed  = the numbers those facts permit
repeat at most twice:
  copy   = ask the model, naming any numbers that failed last time
  failed = numbers in copy with no close match in allowed
  if none failed: return copy
return copy assembled by code from the same facts

Deciding which numbers are allowed

A check is only as good as its list of permitted numbers. Make the list too narrow and correct text keeps getting rejected. Make it too wide and invented numbers get through. Three questions help when you design yours.

What does "not recorded" look like in your data? Many products store a missing value as a placeholder. If that placeholder ends up on the permitted list, the model can quote it as if it were a measurement and the check will pass it. Leave placeholders off the list.

How do your users talk about numbers? People rarely write numbers the way a database stores them. A golfer says "3 over", not "+3". Your check has to accept the forms people use, and each extra form you accept is something the check can no longer tell apart. Make those choices on purpose and write them down.

What counts are even possible? Text often says things like "in four of your last five rounds". Limit counts to what the data could support. "Four of your last five" can pass. "Forty of your last fifty" can't, if there aren't fifty rounds to count.

Test the failure, not just the output

Most AI features are tested by reading the output and deciding it looks fine. That doesn't show what happens when the model gets something wrong, and that's the case that matters.

For ScoreSmart we did the opposite. The acceptance test fed in an insight with a deliberately wrong number and checked that the product blocked it. When the weekly report shipped, it got the same test, with the correction naming the wrong numbers back to the model.

2

places users see AI text: round insights and the weekly report

1

retry before a plain, accurate fallback

3

places the weekly report appears: app, email and PDF

These tests prove the check works. They don't prove the writing is any good. Whether an insight is useful and pitched right for a golfer is something people still judge by reading it. On our next AI product, a regular review of real output is in the plan from day one. Automated checks keep the numbers honest, and people keep the writing good.

Keep reports consistent

Two more decisions shape what users see.

Old reports stay as they were. If the analysis rules change, past reports keep the version they were made with, so a coach's view of last month doesn't change under them. This has a trade-off. If you recalculate old figures but keep the original AI text, the text can describe numbers that have since changed. Regenerating the text costs money, so decide which you'll do and document it.

One report, written once. The weekly report appears in the app, in an email and as a PDF, and all three use the same checked text. A player and their coach reading the same report see exactly the same words. The cost is that when the fallback text is used, it's used in all three places. We chose consistency over trying to get better wording in one channel.

What grounding costs

Grounding has real costs. Know them before you start:

  • The model can't do sums for the user. A number it works out correctly but wasn't given still gets rejected. If you want it to say "you three-putted twice as often as last month", your code has to calculate that ratio and pass it in.
  • The fallback text is plainer. Text built from templates is accurate but dull, so you want the fallback to be rare.
  • You design the facts first. You decide what the model may say before you write a prompt. It's more work at the start, and it pays off every time the product avoids an embarrassing mistake.
  • It checks values, not meaning. A permitted number can still be attached to the wrong metric, and a number written as a word ("three") gets past a check that looks for digits. Ask the model to use digits, and keep reading real output.
  • Numbers aren't the only facts. Names, dates, product codes and links need the same care. In ScoreSmart, the model never writes links at all.
Strong fit

Products where AI writes about a user's own data: reports, dashboards, summaries, coaching, finance, health and fitness, operations.

Strong fit

Any feature where a wrong figure leads to a support ticket, a refund or a legal problem.

Partial fit

Open-ended assistants and chat. You can ground the factual parts, but not the whole conversation.

Not needed

Creative drafting, where there's no source of truth to check against.

Our position

If an AI feature writes about a user's data, grounding is part of the feature, not an extra. Design the check before you write the first prompt, because the check decides what the model is allowed to say. Prompt wording matters less than people think.

It also means being clear about what the AI does. In ScoreSmart it writes round insights and a weekly report, both checked. A narrow AI feature that's always right is worth more than a broad one that's sometimes wrong.

A checklist for your own product

Before you ship an AI feature that writes about user data

  • List every number, name, date and link the AI text might mention, and decide which ones the model may use.

  • Calculate those values in code and pass them in. Don't ask the model to calculate.

  • Leave placeholders and "not recorded" values off the permitted list.

  • Check every number in the output on the server, allowing only for rounding.

  • Retry once with the wrong values named, then fall back to plain, accurate text.

  • Keep links and actions out of the model's hands.

  • Write a test that feeds in a wrong number and proves it's blocked.

  • Record which version of your rules produced each saved output.

  • Read real output regularly. Tests prove the check works; only people can judge the writing.

FAQ


It means tying what a model writes to a trusted source, such as your database, and checking it before anyone sees it. In practice the model may only quote values you gave it, and something verifies that it did.

If you're adding AI to a product and want people to trust it, talk to our AI engineering team. You can also read the ScoreSmart.ai case study for the full build.