Accuracy Over Automation: Why AI Hallucinations Are a Donor Trust Problem, and What to Demand Instead
Technology

Accuracy Over Automation: Why AI Hallucinations Are a Donor Trust Problem, and What to Demand Instead

Jan 28, 20269 min read

On this page

What an AI hallucination actually is

An AI hallucination is a plausible but false statement generated by a language model. That is OpenAI's definition, from its September 2025 research note on why language models hallucinate, and it is worth keeping the word "plausible" in view. A hallucination does not look like an error. It looks like an answer.

The same note explains where they come from, and the explanation is not what most people expect. Models are not malfunctioning when they hallucinate. They are doing what their training rewarded.

OpenAI's argument runs in two parts. First, the way models are evaluated rewards guessing. A model that says "I don't know" scores zero on a question, while a model that guesses has some chance of being right, so over thousands of test questions the guessing model looks better on the scoreboard. In OpenAI's words, standard training and evaluation procedures "reward guessing over acknowledging uncertainty."

Second, some facts cannot be learned from patterns at all. The note draws a distinction between things like spelling and grammar, which follow consistent rules and which large models get right almost every time, and what it calls "arbitrary low-frequency facts," which cannot be predicted from patterns no matter how large the model gets. Its example is a pet's birthday. You could show a model a million photographs labeled with the animal's birthday and it would still be guessing on the next one, because there is no pattern to learn.

Hold on to that example. It is the key to the rest of this page.

Why donor records are exactly what generic AI gets wrong

Ask what a general-purpose model is good at and the answer is patterns: tone, structure, the shape of a thank-you letter, the conventions of a grant narrative, the way a board summary is usually organized. It has seen millions of each. Ask it for the shape of a lapsed-donor email and it will give you a perfectly serviceable one.

Now ask it how much Margaret gave last year, or when your executive director last met the Whitfield Foundation, or whether the Chens have ever been asked for a planned gift. Those are not patterns. They are arbitrary, low-frequency facts, indistinguishable in kind from a pet's birthday. There is nothing in any training corpus that lets a model infer them. If the model has not been given the record, the only thing it can do is guess. And it has been trained to guess.

That is the structural problem with pasting donor questions into a generic assistant, and it is separate from the privacy problem, which is covered on our donor data security page. Even with privacy solved, a model working from patterns will produce confident, specific, wrong answers about individual donors, because individual donors are the one subject where patterns do not help.

The table below sorts common development tasks by which kind of work they are.

TaskMostly pattern, or mostly fact?What a generic model does
Draft a thank-you letter in your voicePatternUsually well
Suggest a structure for a board reportPatternUsually well
Summarize a donor's giving historyFactGuesses if not given the records, and sounds certain
State a donor's last gift amount and dateFactGuesses
Recall what was discussed at the last meetingFactGuesses, or invents a meeting
Say whether a foundation has funded you beforeFactGuesses
Pick which donors are at risk of lapsingFact, then judgmentCannot do it without the records; will produce names anyway

Everything in the lower half of that table is the job. The upper half is the formatting around the job. A tool that is excellent at the upper half and unreliable on the lower half is not "mostly right." It is wrong about the part that matters.

What it costs when it happens

There are three costs, and the sector tends to talk only about the first.

The relationship cost is the obvious one. A briefing that says a donor expressed interest in the capital campaign when no such conversation happened sends a gift officer into a meeting with false confidence. A thank-you that cites the wrong gift amount tells the donor you were not paying attention. These errors are rarely catastrophic individually and they compound, because each one is a small withdrawal from the trust that the entire relationship runs on.

The legal cost is newer, and the precedent is specific. In February 2024 the Civil Resolution Tribunal of British Columbia decided Moffatt v. Air Canada. A customer had asked the airline's website chatbot about bereavement fares, and the chatbot told them they could apply for the reduced fare retroactively within 90 days. The airline's actual policy said the opposite, and that policy was linked from the chatbot's own answer. When the customer claimed the difference, Air Canada argued, in the tribunal's summary, that "the chatbot is a separate legal entity that is responsible for its own actions." The tribunal member called that "a remarkable submission" and rejected it: "It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot." Air Canada was ordered to pay $812.02.

The amount is trivial. The principle is not. An organization owns what its AI tells people. If a tool on your website, in your inbox or in your donor communications states something false, the fact that a model generated it is not a defense. Nonprofits are not airlines, but the reasoning transfers without modification.

The donor-perception cost is the one with numbers attached. An August 2024 study of 1,006 donors to 501(c)(3) nonprofits, led by Nathan Chappell of Fundraising.AI and Cherian Koshy of Kindsight and published by Fidelity Charitable, found that 93% of donors rated transparency about how a nonprofit uses AI as very or somewhat important. Asked about AI-driven personalization, 39.8% expressed discomfort with how their data might be used, while 28.1% said they would accept it if the organization was transparent about its practices. Donors are not opposed to AI. They are opposed to not being told, and an organization that cannot explain where its AI's answers come from cannot be transparent about them.

Why "just fact-check it" does not scale

The standard advice, repeated on almost every page that ranks for this topic, is to review AI output before using it. That is correct and it is not sufficient, for two reasons.

The first is that hallucinations are designed, in effect, to pass review. A January 2024 study from Stanford's RegLab and Institute for Human-Centered AI tested more than 200,000 legal queries against GPT-3.5, Llama 2 and PaLM 2 and found hallucination rates of 69% to 88% on specific legal questions. The finding that matters here is not the rate. It is the authors' observation that "a common thread across all models is a tendency towards overconfidence, irrespective of their actual accuracy." The model does not flag the answers it is unsure about. The wrong answer and the right answer arrive in the same confident voice, so the reviewer cannot triage. They have to check everything or nothing.

The second reason is arithmetic. To check whether a model's summary of Margaret's giving is accurate, you open Margaret's record in the CRM. At that point you are looking at the source, and the model's summary has saved you nothing. Checking a generated answer against the record costs the same as reading the record, so the only way fact-checking "works" at scale is if people stop doing it. The 2026 Nonprofit AI Adoption Report from Virtuous, which surveyed 346 organizations in December 2025, found that 81% use AI on an ad hoc basis without documented workflows and 47% have no AI governance policy. That is the environment the advice is landing in. The review step exists on paper and is skipped in practice, because the tool makes skipping it feel safe.

There is a softer signal worth stating with its limits. A survey of 51 UK nonprofit and grassroots organizations, commissioned by the Joseph Rowntree Foundation and published in July 2024, found that 63% worried about the accuracy of generative AI. The sample is small and it is not US data, but it is the only published figure we could find that asks nonprofits the accuracy question directly, and the concern it records is rational.

Accuracy over automation, in practice

Accuracy over automation means that when a tool cannot be both fast and right, it must be right, and it must say so when it is neither. That sounds obvious. Most AI products are built the other way round, because automation is what demonstrates well and accuracy is what nobody notices until it fails.

Put concretely, it is four design requirements.

Retrieval before generation. The answer to a factual question about a donor is looked up in your records first, and the language model's job is to phrase what was found, not to supply it. If the fact is not in the records, there is nothing to phrase.

A citation on every claim. Each statement about a donor points back to the record, note, email or document it came from, so that checking it is a click rather than a search.

Abstention when the record is not there. The system says "there is no record of this" instead of producing a plausible sentence. This is the behavior OpenAI's research identifies as the fix and the one that standard training discourages, so it has to be built deliberately.

A person sends. Nothing reaches a donor, a board or a funder without a human approving it. The tool prepares; the team owns the relationship and the final word.

Deterministic, in this context, means the answer is determined by the record rather than by the model. Ask the same question of the same data and you get the same cited answer. It does not mean the system is incapable of error. It means the errors it can make are the errors in your data, which are visible, traceable and fixable, rather than errors invented at answer time, which are none of those things.

BehaviorPattern-only assistantRetrieval-grounded system
Asked for a donor's last giftProduces a figureReturns the figure from the record, with the record
Asked about a donor not in the dataProduces an answer anywayStates that no record exists
Same question asked twiceMay give two answersGives the same cited answer
Asked "where did that come from?"Cannot sayPoints to the source
Draft ready to goSends if you let itWaits for a person

Neither column is "AI" and the other "not AI." Both use a language model. The difference is what the model is allowed to be the source of.

Five accuracy questions to ask any AI tool

The privacy questions, about training data, tenant isolation and PII, are on our data security page and in the free Safe AI Policy Pack. These are the accuracy questions, and they are the ones vendors are less often asked.

  1. Where did that come from? Pick any factual statement the tool makes about a donor and ask for the source. The answer should be a specific record you can open. "The model generated it" is a no.
  2. What happens when the answer is not in our data? Ask about a donor who does not exist, or a meeting that never happened. The right response is a refusal. A fluent, specific answer is a failure, and a revealing one, because it tells you how the tool will behave on every question where your records are thin.
  3. If I ask twice, do I get the same answer? Ask the same factual question in two sessions. Two different answers mean the tool is generating rather than retrieving.
  4. Can it send anything without a person? Any path by which a draft reaches a donor without explicit approval is a path by which an error reaches a donor.
  5. What is logged? If something goes wrong, can you see what was asked, what was answered and what it was based on? A tool that cannot show its work cannot be audited, and an organization that cannot audit its AI cannot be transparent about it.

A ten-minute test you can run on your own data

You do not need a vendor's help to find out whether a tool hallucinates about your donors. You need five donors you already know well.

  1. Choose five donors whose history you can recite. Mix them: one major donor, one lapsed, one new, one recurring, one foundation contact. The point is that you already know the right answers.
  2. Ask three factual questions about each. Last gift amount and date. Total given over the relationship. The most recent substantive contact and what it was about. Write the answers down before you look at the tool's.
  3. Add one trap. Ask about a donor who does not exist, using a plausible name. Then ask about a meeting that never happened with a real donor.
  4. Score every answer as right, wrong or refused. Do not score "close." A gift of $8,200 reported as $12,000 is wrong. Count refusals separately, because on the trap questions a refusal is the correct answer.
  5. For every right answer, ask for the source. An answer that is correct but uncited is a correct guess, and you have no way to know that the next one will be.
  6. Re-ask two of the questions. Same wording, new session. Note whether the answers match.

Seventeen questions, ten minutes. A tool fit for donor work will be right and cited on the fifteen real questions and will refuse the two traps. Anything else tells you where the risk sits before you have put it in front of a donor.

Where Gratefully stands

Gratefully is built on the four requirements above, and we would rather describe the mechanism than make a promise. Grace answers donor questions by retrieving from your organization's own knowledge graph, the records, notes and documents you have connected, and every answer and draft is traceable to the record it came from. If the information is not in your records, Grace says so rather than filling the gap. Drafts are prepared in your voice with their sources attached, and nothing leaves your organization without a person approving it. All of it is logged and auditable. The full picture is on how the system works, and the comparison with a general-purpose assistant is on Gratefully vs ChatGPT.

We do not claim the underlying models have stopped hallucinating. OpenAI says plainly that they have not, and we see no reason to contradict them. We claim something narrower and more useful: that in donor work, the model should never be the source of a fact, and that a system built that way fails in ways you can see.

Last updated August 19, 2026.

Frequently asked questions

What is an AI hallucination?

An AI hallucination is a plausible but false statement generated by a language model. OpenAI's September 2025 research note on the subject argues that hallucinations persist because standard training and evaluation reward guessing over admitting uncertainty, and because some facts, which it calls arbitrary low-frequency facts, cannot be predicted from patterns at all.

Why do AI tools get donor information wrong?

Because individual donor facts, such as a last gift amount or the date of a meeting, are arbitrary and specific rather than patterned. A general-purpose model that has not been given the record cannot infer it, and it has been trained to guess rather than to say it does not know. The result is a confident, specific and wrong answer.

Is a nonprofit responsible for what its AI chatbot says?

Under the reasoning in Moffatt v. Air Canada, decided by the Civil Resolution Tribunal of British Columbia in February 2024, an organization is responsible for all the information on its website, whether it comes from a static page or a chatbot. The tribunal rejected the argument that a chatbot is a separate legal entity and ordered the airline to pay $812.02. The decision is Canadian, but the principle is widely cited.

How do you stop AI from hallucinating about donors?

You do not rely on the model for facts. Retrieve the answer from your own records first, have the model phrase only what was retrieved, require a citation on every claim, make the system refuse when the record is not there, and keep a person between the tool and the donor. Fact-checking output after the fact does not scale, because hallucinations sound as confident as correct answers.

What does "deterministic" mean for AI in fundraising?

It means the answer is determined by the record rather than by the model. The same question asked of the same data returns the same cited answer. It does not mean the system cannot be wrong. It means that the errors it can make are the errors in your data, which are visible and fixable, rather than errors invented at answer time.

Do donors care whether a nonprofit uses AI?

They care about being told. In an August 2024 study of 1,006 donors to 501(c)(3) nonprofits, led by Nathan Chappell and Cherian Koshy and published by Fidelity Charitable, 93% rated transparency about a nonprofit's AI use as very or somewhat important, 39.8% expressed discomfort with AI-driven personalization, and 28.1% said they would accept it if the organization was transparent about its practices.

Author

Muddsar Jamil, Founder, Gratefully

Muddsar Jamil is the founder of Gratefully and a 20-year Silicon Valley engineer (Adobe, Workday, SugarCRM) who spent nearly as long volunteering with Bay Area nonprofits. He built Gratefully so donor relationships survive spreadsheets, staff turnover, and guesswork. Connect on LinkedIn.

Ready to transform your donor relationships?

See how Gratefully can help you implement these strategies at scale with AI-powered donor intelligence.

Want more insights like this? Browse all articles or get in touch with our team.