When an AWS AI Pilot Has Learned Its Acceptance Test

Identify evaluation contamination from prompt edits, case-family leakage and repeated winner selection. Preserve useful regression evidence and plan a genuinely...

Stop treating a case set as unseen acceptance evidence once its results influence the candidate you are judging. The cases remain useful for development and regression, but a stronger score on that exposed set cannot establish improvement on unfamiliar work. Record what was exposed, freeze the candidate and use independently administered new case families for the next comparison. Renaming a file or withholding the answer column after engineers have studied it does not restore independence.

This article helps an evaluation lead repair a contaminated comparison for a Bedrock-backed answer application. It focuses on the evidence claim, not a particular model or a training API. Prompt instructions, retrieval filters, examples, routing and human scoring decisions can all adapt to a test without changing model weights. The supplied counts below are fictional arithmetic. No model requests, customer evaluation or AWS jobs were executed.

Identify the decision that consumed the test

An engineer sees which cases fail and adds an instruction that fixes them. That is useful development. The evidence changes when the same cases later support a claim that the edited candidate generalizes. A test has supplied information to the decision that created the candidate. Keep the improvement as a known-case repair, with its original failure and revised output, rather than discarding useful engineering work or giving it a broader label.

Exposure can occur without opening the dataset. A scoreboard that says one configuration wins provides feedback for selecting that configuration. Repeatedly trying prompt variants against the same hidden set can therefore make the reported winner depend on that set. Hiding individual examples reduces one route of leakage, but does not remove selection through aggregate scores. Record how many candidates and decisions consumed those scores before interpreting the final result.

The scikit-learn maintainers' data leakage guidance explains why test information should not guide model choices and why splitting before fitted preprocessing matters. The recommendation here extends that principle to an application: any feedback that changes the assessed bundle belongs in its development history. It does not imply that a Python pipeline prevents prompt, retrieval or reviewer leakage.

Define the bundle broadly enough to find the leak. If the prompt stays fixed but a retrieval rule excludes the documents that caused test failures, the evaluated system changed. If the rubric changes after an inconvenient response, the outcome definition changed. Either can be defensible, but requires a new comparison record with the old result retained. A model identifier alone cannot explain what learned from the acceptance test.

Split by the relationship that carries the answer

Different filenames can contain the same underlying case. An original question, a paraphrase and a translated question about the same incident may share the facts needed to answer it. A document split can also place adjacent pages from one contract into separate sets. Random row assignment does not make those families independent. Choose a grouping unit from the task: incident, organization, policy revision, source package or another relationship that carries shared information.

Keep a family register beside the case register. Record the family basis, source revision and duplicates found by authorized inspection. Exact hashes find identical bytes; they do not prove that two differently worded cases are unrelated. An automated similarity screen can nominate candidates for review, but its threshold is not a certificate of independence. Preserve unresolved relationships rather than counting each paraphrase as new acceptance coverage.

For a retrieval application, legitimate source access needs a different interpretation. If the real task permits reading the current policy, the evaluation should permit it too. The source containing the answer is not automatically contamination. Injecting the evaluator's preferred response or hidden scoring notes into the retrieved corpus would change that boundary. Separate task evidence available at use time from labels and review reasoning intended only for evaluation.

Time also carries information. A forward-looking claim may require later policy versions or newly arriving case families, not a random split of one historical month. A chronological split can reveal a different challenge, but changing policy and input mix can confound a candidate comparison. State whether the study tests unfamiliar families, a later operating period or both. Each supports a narrower inference than universal robustness.

Preserve three dataset roles without turning them into a fixed ratio

Development cases support debugging and choosing prompts or thresholds. Acceptance cases support a defined decision for a frozen bundle under a declared task contract. Drift observations describe what later arrives and whether the earlier evidence still applies. Assign roles by purpose and access, not a universal percentage. Small source populations may not permit all three claims, and creating more near-duplicates does not solve that shortage.

When an acceptance failure is disclosed to implement a fix, retain it as a regression case. Preserve its former acceptance identity and exposure event; do not quietly mark it unseen again. Reserve fresh evidence for the claim affected by the repair. Some deterministic control tests can and should be run repeatedly against known cases. Their purpose is to detect a broken invariant, not estimate performance on a new population.

Later drift cases can enter development only after approved triage, privacy review and a recorded role change. Keep the original event reference and its outcome. A drift dataset made entirely from complaints cannot estimate the rate of failures among all arrivals. Conversely, withholding every observed incident from engineers to preserve a pristine test can leave users exposed to a known defect. Fix the defect and narrow the evidence claim while preparing a defensible next study.

Amazon Bedrock's evaluation overview describes automatic, human-worker and judge-model evaluation approaches. Selecting a job type does not administer these lifecycle boundaries for your team. Maintain candidate, case-set, rubric, exposure and access records independently of the service's scores.

Worked scenario: exposed success reverses on new families

Fictional candidate A is an evidence-summary application. Engineers reviewed failures from 40 case families and edited its instructions. A later presentation uses 80 cases: 40 from those exposed families and 40 from new families. Stipulate that all 40 exposed cases pass and 24 new cases pass. The displayed total is 64/80 = 80 percent. The new-family subset is 24/40 = 60 percent. Neither result establishes a causal benefit from the edits, but the split exposes how the headline hides the boundary.

Before a subsequent comparison, the evaluation lead freezes A and candidate B and supplies an independently administered set of 80 new families. The task, source access and rubric are shared. Stipulate 48 passes for A and 60 for B. Their observed fractions are 48/80 = 60 percent and 60/80 = 75 percent. Every count is invented for this example; no provider performance is implied. These margins also do not establish statistical significance or rare-harm tolerances.

Evidence packetSupplied passing casesClaim it can describe
A, exposed families40/40 = 100 percentBehavior on known repair families
A, new subset in the presentation24/40 = 60 percentThat separately identified subset
A, mixed presentation64/80 = 80 percentA mixture of exposed and new evidence
A, independently administered comparison48/80 = 60 percentFrozen A on the declared new-family set
B, same independently administered comparison60/80 = 75 percentFrozen B on that same set
A, exposed families
Supplied passing cases: 40/40 = 100 percent
Claim it can describe: Behavior on known repair families
A, new subset in the presentation
Supplied passing cases: 24/40 = 60 percent
Claim it can describe: That separately identified subset
A, mixed presentation
Supplied passing cases: 64/80 = 80 percent
Claim it can describe: A mixture of exposed and new evidence
A, independently administered comparison
Supplied passing cases: 48/80 = 60 percent
Claim it can describe: Frozen A on the declared new-family set
B, same independently administered comparison
Supplied passing cases: 60/80 = 75 percent
Claim it can describe: Frozen B on that same set

The mixed presentation cannot supply an unseen 80-case claim. The later shared comparison gives B the better stipulated pass fraction, reversing a preference based only on A's headline. However, if B has a prohibited disclosure among its 20 failures, the protected release decision remains held. A higher task-quality fraction cannot compensate for a critical failure. Investigate the causes and preserve both counts instead of rewriting the rubric to remove the disclosure.

An opposing outcome is also possible. If B passes 44 of the same 80 new cases, its fraction is 55 percent and A's is still 60 percent. Contamination does not prove that A is inferior; it prevents one particular claim about A's exposed evidence. The remedy is a valid comparison, not automatically selecting the incumbent or the competitor.

Repair the record when fresh cases are scarce

Hold the affected generalization claim and identify which result remains useful. Known-case regression, slice diagnosis and a documented fix may survive. Forecasted operating quality or a release recommendation dependent on unseen performance may not. Mark the next evidence task with an owner and deadline rather than deleting the previous report. If only aggregate scores were exposed, record that limitation and the candidate-selection history you can actually recover.

Use family-based resampling or an independently designed nested comparison where the population and study expertise support it. Do not call the outer cases untouched if engineers inspected them while selecting the inner procedure. Where the case population is too small, report that the proposed claim is not supported and narrow the operating scope. Repeatedly relabeling a scarce dataset as final does not manufacture additional information.

Review authorization before collecting replacement evidence. Bedrock evaluation data management documents a temporary evaluation copy and encryption choices. Deletion of that copy does not establish deletion of your originals, reviewer exports or incident records. A fresh dataset still needs permitted processing, evaluator access and retention. Never move sensitive case content into a public issue or enquiry form to explain a leak.

Next action: record exposure before the next candidate run

Use the following fields to decide whether the next score is development evidence, acceptance evidence or an unresolved claim. They are a proposed record, not proof that an actual dataset is independent.

Case family and purpose
Name the family rule, source revision, task and development, acceptance or drift role. Preserve aliases and related cases.
Exposure event
Record who saw cases, labels, explanations or aggregate scores, when and for which permitted purpose. Unknown history remains unknown.
Consumed decision
Identify prompt, retrieval, rubric, routing or candidate-selection changes that used the feedback. Bind the resulting bundle version.
Remaining claim
State known-case regression or the precise unseen comparison still supported. Identify invalidated claims and their recipients.
Replacement evidence
Assign independently administered families, frozen candidates, label authority and comparison design. Record rights and gaps.
Stop and recovery
Hold the affected decision, retain reports and notify the decision owner. Repair known defects through the normal change process without declaring the old set unseen.

The production evaluation workbook owns the wider execution procedure. The baseline selection article addresses whether the comparator and task mix are credible. Before the next run, ask the dataset custodian to reconstruct one case family's exposure history and the engineering lead to identify every change that consumed it. Resolve that record before presenting the next score as independent evidence.

Related services