The steering committee packet contained a two-page summary of an eighty-page technical due diligence report for a new core processing vendor. The summary stated the platform offered robust multi-region failover. The original report stated the failover was robust only if the legacy mainframe data was batched under a specific daily threshold, a limit the bank routinely exceeded on quarter-end days. The committee approved the vendor in February.

Ask a model to summarize a technical assessment and it gives you the document’s center of gravity. Risk lives at the edges. A language model is trained to produce fluent, authoritative text. Its attention layers weight definitive statements higher than qualifiers. Words like “might,” “assuming,” and “historically” reduce the fluency of the output. The model discards them to meet the implicit constraint for brevity. The summary is not wrong. It is confident.

An eighty-page vendor architecture blueprint usually contains ten pages of actual design conviction and seventy pages of conditional assumptions, regulatory caveats, and tail-risk disclosures. When you ask a model for a summary, it faithfully compresses the conviction and discards the conditions. The main failure is rarely an obviously false number. It is a change in modality. “Could,” “if,” and “under severe stress” become “will.” A steering committee can catch a wrong figure more easily than it can spot a condition that disappeared.

This happens structurally, not because of a poorly written prompt. To process a long document, the system breaks the text into chunks. The warning about the batch threshold sits in an appendix. The claim about robust failover is on page twelve. Chunking severs the link between the claim and the condition. The model evaluates them in isolation. When generating the synthesis, the positive claim wins the token-weighting battle because it contains more actionable technical terminology. If the assumption is in the appendix, the summary does not know it exists.

The problem compounds when evaluating multiple vendors. Procurement might ask for a consensus view of three competing proposals. The models naturally optimize for the median view. They identify overlapping claims and present them as the core truth. If all three vendors claim industry-standard encryption, the summary reports that all vendors meet encryption standards. It hides the fact that one vendor means encryption at rest, another means encryption in transit, and the third relies on a shared key management service. A summary that reads like consensus is a summary that has averaged away the disagreement.

The degradation continues as the document moves through the governance chain. The vendor report becomes the enterprise architect’s executive summary, which becomes the AI summary, which becomes the steering committee minute, which becomes the board risk minute. Each pass drops a hedge. By the time the language reaches the board, the tail risk is gone and the assessment is unanimous. You are trading institutional paranoia for grammatical fluency.

Banning the technology is unnecessary and impractical. The models are highly effective for triage, identifying which sections of a data room to read first, or translating a foreign-language compliance filing. The mistake is treating the narrative summary as the artifact of record. To use these tools safely, the prompt architecture must change.

Separate extraction from narrative. Instruct the model to pull structured fields first: failover thresholds, data residency constraints, SLA penalties, and named dependencies. A human checks the fields against the source. Only then summarize the narrative in a clearly labeled commentary section. Never let a paraphrase carry a technical specification into a decision document.

Require a verbatim quote budget. Paraphrase is where the conditions die. Mandate a minimum number of direct quotes in any summary of a risk document. Create a standing tail log in every output, listing the specific conditions named in the report and the interactions flagged between them. If a model cannot tag a sentence with the vendor’s own confidence language, it should not include it.

Run the same document through two different models. Where they diverge is where the report is doing something unusual or where the conditions are complex. Read that part yourself. Two models, diffed, beat one model trusted.

Build a test set of ten vendor reports you have already read. You know the conclusion and you know the caveat. Run them through your pipeline. Count how many caveats survive. That number is your error rate, and it is local to your documents.

When the language model removes the vendor’s doubt, the steering committee assumes the vendor’s liability.