Blog
AI AdoptionPortfolio OperationsPrivate Equity

Why Confident AI Answers Are Often Wrong

In finance the costly AI failure is a real figure sitting inside a frame too small to hold the actual cause. Here is how it survives review, and what changes the outcome.

ONE FILE · THE ANSWER STOPS AT ITS EDGEwhy?ledgerdebtor days +11disputemigrationtermsthree causes, none of them in the fileCONNECTED · THE QUESTION REACHES PAST ITwhy?One structured layerledgerdisputebillingtermsthe same question reaches the causes
One attached file yields a confident answer bounded by that file. A connected foundation lets the same question reach the causes sitting in other systems.

When Transport for London installed live arrival boards at bus stops, passengers reported shorter waits. Measured perceived waiting time fell from 11.9 minutes to 8.6, and 65 percent said they had waited less than before. Over the same period actual service reliability declined slightly. Sixty-four percent of passengers said they believed reliability had improved.

The boards changed nothing about the buses. They changed what passengers knew, and knowing turned out to feel much like being served better.

That gap, between the feeling of a closed question and the quality of the answer underneath it, is the central risk in how firms are now using AI. A model handed a file returns something clear, structured, and confident enough to forward. The confidence is a property of the prose rather than a measurement of anything.

Certainty is something people pay for

The preference is measurable. A stated-preference study of Dutch rail travelers found they would give up more than seven minutes of journey time in exchange for certainty about their wait. Research published in Nature Communications in 2016 found that a 50 percent chance of an electric shock produced more measurable stress, in skin conductance and pupil dilation, than a 100 percent chance. Uncertainty is aversive on its own terms, separately from the outcome it precedes.

Finance amplifies the effect. Work is expected to reconcile, and every conclusion eventually has to survive someone else’s review. An answer that closes the question is easier to present than an answer that says the data does not yet explain enough and someone needs to make a phone call.

The failure mode, at the operating level

Consider a portfolio company where cash collection slipped. Give a model the aging schedule and ask why, and it may report that debtor days stretched by eleven. That figure is real. It ties to the ledger and the arithmetic holds.

What the file cannot show is that one customer is withholding payment over an installation that went badly, that the billing team migrated systems in June and dated a batch of invoices incorrectly, and that a sales director conceded longer terms to close two contracts.

Three causes, sitting in three different functions, none of them in the document that was attached.

Hallucination describes a narrower problem: the model inventing a fact, a number, or a source. Nothing here was invented. The answer is defensible line by line and wrong about the company as a whole.

Why it survives review

A fabricated figure gives a reviewer something concrete to catch. A genuine figure inside an incomplete explanation passes several rounds, because the calculation is real and the narrative around it sounds finished. Once the explanation appears in a portfolio review it starts to acquire authority. It gets repeated in the board pack and the lender update, and becomes the accepted account of the quarter.

Larger context windows do not address this. Each new model ingests more tokens, which relieves a memory constraint. The questions that matter inside an operating business stay larger than the prompt and the handful of documents someone thought to attach: CRM records, customer conversations, pricing decisions, delivery problems, board discussions, and the judgment of people who know which exceptions matter.

Two things that change the outcome

The first is reach. The context that explains a number usually lives across systems that were never designed to be read together. Connecting them into one structured foundation, instead of querying five disconnected systems on every question, is what moves an answer from being bounded by a single file to being bounded by the business. We have written separately on why a foundation outperforms a set of connectors.

The second is visible workings. An answer that carries its sources, and that names the inputs it could not find or could not reconcile, can be interrogated. An answer delivered as clean prose cannot. Lineage is what turns a confident paragraph into something a reviewer can test rather than accept.

A lower-middle-market PE firm we worked with had five years of deals spread across drives, inboxes, and a CRM that nobody maintained. Answering “why did we pass?” meant half a day of folder archaeology, and the person doing it still could not be sure the answer was current. Connecting those deals into one queryable foundation, with every answer cited back to the document it came from, brought that down to seconds. The full case study sets out what was involved.

The person reading the answer

Until systems can reach that context on their own, the reviewer is the route it takes in. That role is usually described as checking whether the model got the answer right. Checking arithmetic is the smallest part of it, and it is generally already correct.

The useful work is taking the answer apart, pressing on the parts that look thin, and supplying what the model had no way of seeing: the outage, the disputed invoice, the concession agreed on a call in March. Context fed back that way is context the system holds the next time the question is asked, which is how a firm’s own priors end up inside the tooling rather than in someone’s head.

The trade

A system built this way will sometimes return less certainty than the process it replaced. It will name three plausible causes at different confidence levels where the old method produced one clean narrative. Presented to a board that wants the pack finished tonight, that reads as a downgrade.

Rory Sutherland describes the alternative to false precision as a “not-perfect-but-rarely-stupid conclusion”, reached by considering a wider range of factors than certainty allows. For a firm making capital allocation decisions, rarely stupid is the more valuable property.

Frequently asked questions

What is the difference between AI hallucination and an incomplete answer?
Hallucination means the model invented a fact, a figure, or a source. An incomplete answer contains only real, verifiable numbers, but the explanation is drawn from a frame too narrow to include the actual cause. The second failure is harder to catch because every component checks out.
Does a larger context window fix the problem?
No. A bigger window lets a model read more of what was attached. The causes of a business result usually sit in systems nobody attached: the CRM, the billing platform, a customer conversation, a decision recorded in a board pack. The constraint is reach, not capacity.
How do you tell whether an AI answer is missing context?
Ask what it could not see. A system with lineage will name the sources it used and flag the inputs that were missing or in conflict. An answer delivered as clean prose with no workings gives a reviewer nothing to test.
What should a reviewer actually be doing with an AI answer?
Checking the arithmetic is the smallest part of the job and it is usually already correct. The useful work is pressing on the parts that look thin and supplying the context the model had no route to: an outage, a disputed invoice, a concession agreed on a call.
Why would a better system return less certainty?
Because a wider frame often reveals several plausible causes with different confidence levels rather than one clean narrative. That reads as a downgrade against a process that always produced a single tidy explanation, and it is closer to how the business actually behaves.
Written by Harry Ratcliff

Co-founder of DealSage, the AI-native deal intelligence platform. He writes Acquisition Intelligence, a weekly read on AI in M&A for finance professionals.

Read the original issue

Get this thinking weekly.

Acquisition Intelligence is a weekly read on AI in M&A for deal-makers. No fluff, no hype.