What If Context Truncation Removes the Decisive Evidence?
Test whether an AI recommendation retains the clause that changes the decision. Separate missing sources, context selection, interpretation and action permission.
A shorter context can change the answer, not just its detail
An assistant reviewing a purchasing request sees a delivery proposal and recommends placing the order. The contract amendment that prohibits release without a capacity check was retrieved, but the application removed it while fitting documents into the model request. The answer can sound reasonable and cite the proposal accurately while missing the evidence that changes the decision.
The useful question is not whether the request fits the token limit. It is whether the evidence required for this decision survived retrieval, selection, trimming and interpretation. A larger window can help retain material, but capacity alone does not establish coverage. A complete packet still needs the correct revisions, an explicit question and a model that uses the relevant passages correctly.
The purchasing request, documents and fixture outcomes in this article are fabricated teaching examples. They are not legal advice, a customer engagement or measured model performance. The proposed controls support a bounded engineering review. They do not certify that an assistant understands every contract or may release an order. Recommendation quality and permission to act remain separate responsibilities.
Distinguish missing evidence from evidence the model overlooked
Lost in the Middle reports experiments on multi-document question answering and key-value retrieval in which changing the position of relevant information affected the evaluated models' performance. The paper was published in 2024. It is a reason to test position sensitivity, not a current benchmark for every model or proof that your particular assistant fails in the same way. It studies information present in an input, which is different from an application deleting that information before submission.
Claude's context-window documentation explains that input components and the response consume context capacity, and distinguishes compaction from ordinary context use. The engineering implication here is to budget the actual request, including tools and conversation state, rather than only the document text. This does not imply that every provider silently trims overflowing input. Identify whether your application, middleware or selected provider feature performs a transformation, rejects the request or returns a partial response.
Trace four different defects. The search stage may never retrieve an amendment. The packet builder may retrieve it and then drop it. A summary may keep its topic but lose its operative condition. Finally, the exact condition may reach the model and still be interpreted incorrectly. A generic label such as hallucination hides which control needs work. Record the defect at the first observed boundary without inferring later model behaviour from an earlier retrieval failure.
Define the decision before choosing what to retain
In the synthetic request P17, proposal D1 revision 3 offers delivery within six weeks. Amendment D2 revision 2 says the order must not be released until capacity review C7 is recorded. The task is to prepare an advisory release assessment, not to place an order. Under this declared fixture policy, D1 alone is insufficient. If D2 is applicable and C7 is not recorded, the expected assessment is hold for capacity review. If D2 cannot be obtained, the expected result is insufficient evidence, not an assumed permission to proceed.
Have the domain owner identify decision-changing evidence before tuning a retrieval score. That includes restrictions, exceptions, applicability, revision relationships and required evidence of satisfaction. A high-similarity proposal paragraph may be less decisive than a short amendment that uses different terminology. A ranking rule that consistently prefers the proposal is therefore not an acceptance rule for the release assessment.
Build an evidence register for the bounded task. Record source identity, revision, applicable scope, required passage and the observation supporting its status. A document identifier alone is not enough when only the first page was parsed or an attachment is missing. Keep the register outside the generated answer so a fluent response cannot retrospectively define which documents were necessary.
Do not claim that a register proves the unknown universe of documents is complete. A case may have an undiscovered amendment or incomplete intake. Record the reviewed collection and its known gaps. The owner must decide how the application behaves when required applicability checks or source retrieval are incomplete. For this example, it can explain the gap and request the missing review without making a release recommendation.
Inspect the final packet, not only the retrieval results
Capture the actual evidence packet delivered after deduplication, reranking, chunk selection, summarization and request assembly. Relate each retained span to its source revision and offsets or another supported locator. Preserve an omitted-source record and the reason for omission. A search log showing D2 at rank four does not prove the model received its restriction when the final builder admits only three chunks.
For the P17 fixture, verify that the release condition and its applicability survive together. Copying the phrase capacity review while deleting must not be released until changes the meaning. Splitting a paragraph at a chunk boundary can detach a condition from its exception. The acceptance check should compare the actual retained evidence against the independently specified condition, not merely search the request for a document title or keyword.
Claude's prompting guidance recommends clear document structure and extracting relevant quotations for long-document tasks. Quotation-first output can make this review easier, but a model-selected quotation is not proof that all decisive passages were included. Check that a cited passage exists in the stated revision and supports the conclusion. A correct quote from D1 does not resolve the missing D2 condition.
Restrict packet diagnostics as carefully as the underlying documents. Do not copy full contracts, personal data or secrets into unrestricted telemetry merely to investigate trimming. Use approved access, bounded retention and suitably minimized evidence. Sanitized records need enough detail to reproduce selection behaviour without implying that a redacted view contains every fact used in a real decision.
Use counterexamples that separate the failure stages
The worksheet changes one boundary at a time under the declared P17 policy. Expected outcomes come from the fixture contract, not from the model's preferred answer. These are tests to run in an inert advisory harness, not results obtained by Ampity or authorization to issue purchase instructions.
| Controlled packet condition | Expected advisory outcome | Evidence to inspect | | --- | --- | --- | | D1 and applicable D2 retained; C7 absent | Hold for capacity review | Final D2 condition and explicit unmet C7 | | Search retrieves D2 but packet builder removes it | Insufficient evidence, not release advice | Retrieved-to-retained difference and coverage failure | | D2 retained with the restriction moved into the middle | Same hold as the unchanged packet | Exact retained span, conclusion and position-sensitive comparison | | Summary keeps capacity review but loses the release prohibition | Reject the summary as adequate evidence | Original condition, summary difference and preserved uncertainty | | D2 is available but applicability cannot be established | Hold for applicability review | Missing applicability evidence and owned next step |
Move the same condition between early, middle and late positions while keeping its text and the intended task unchanged. Then test deliberate removal separately. A positional comparison asks whether the model uses available evidence consistently. A removal comparison asks whether the pipeline detects inadequate evidence. Combining both changes in one run makes the cause harder to interpret.
Use a deliberately weak negative control that recommends release whenever D1 contains an acceptable delivery date. The harness should catch its missing D2 and C7 reasoning. Inspect attempted recommendations as well as final records; a later reviewer correcting the decision does not erase an earlier unsafe recommendation. All adapters should remain non-sending and non-ordering during this exercise.
Fix the responsible boundary without hiding the trade-off
If retrieval misses D2, investigate collection coverage, revision linkage and queries before expanding the context window. If selection removes it, change the packet contract to retain required evidence or expose a coverage failure. If compression loses the condition, retain the authoritative span or reject the compressed packet. If interpretation fails despite complete evidence, revise the task design and evaluate the model with independent expectations. The same mitigation does not solve all four defects.
Avoid unconditional latest-revision rules where historical assessment is intentional. P17 might ask what was known when the proposal was reviewed, or what is permitted now. These are different questions. Bind the assessment to its chosen source snapshot and time context, and revalidate a recommendation before using it in a separately governed current action. A historical citation is not current authorization.
Compare fixes across the full evaluation set. A longer packet can preserve D2 yet add distracting, obsolete or inapplicable records. A required-evidence rule can improve one fixture while holding more incomplete cases for human review. Report coverage, supported conclusions, incorrect recommendations, justified holds, review workload and latency together. Keep unresolved and rejected packets in the denominator instead of reporting accuracy only on convenient accepted answers.
Decide what evidence permits wider use
Start with the bounded P17 counterexample and a recorded packet-building configuration. Pin the model, prompt, retrieval, parsing and compression revisions for the run. Check outcome differences after changing one stage, and repeat relevant cases where generation varies. Passing a finite set does not prove that every future document or model position will behave correctly. It does provide evidence for the particular task, source collection and configurations exercised.
Maintain a distinction between cannot answer and no restriction found. The first may result from a missing source or applicability gap. The second requires a defined search and review scope and still should not claim that no restriction exists anywhere. Show readers what was reviewed, what is missing and the next owned action instead of using a confident summary to fill the uncertainty.
Use the production AI evaluation workbook to define independent task expectations. Read retrieved instructions and action authority for the separate permission boundary. Ampity's AI evaluation services and RAG knowledge systems work can help scope a collection and acceptance review. The worksheet can be used without submitting an enquiry; contact is optional.