Sample Rare, Consequential Errors in an AWS AI Pilot
Separate representative sampling from enriched challenge sets. Reconcile risk-slice denominators and interpret zero observed errors without claiming that rare...
Use a representative sample to describe the intended workload and a separately reported challenge set to examine consequential failure mechanisms. Deliberately adding risky cases can reveal a defect that an ordinary sample misses, but the enriched set's raw average is not the operating error rate. Keep slice counts, selection and uncertainty visible. Zero observed errors in a small or dependent sample cannot establish that a rare harmful outcome is absent.
This article is for an evaluation lead deciding what an AWS AI pilot's case budget can support. It covers sampling and the interpretation of risk-bearing cases, not a universal acceptable error rate or release permission. All populations, error counts and probabilities in the examples are stipulated. No AWS evaluation, customer data collection or provider response was observed. A qualified study design is needed before applying the probability model to consequential real decisions.
Name the error mechanism before looking for a sample size
Define the outcome that would matter: disclosure across a permission boundary, an unsupported high-impact recommendation, a missed material restriction or another task-specific failure. A generic accuracy label can merge errors with very different consequences. Keep correct denial, safe abstention, incomplete evidence, unavailable dependency and incorrect answer separately recognizable. If the rubric cannot distinguish them, increasing the case count produces more uncertain labels.
Identify where the mechanism can occur. A policy summary might fail on conflicting revisions or exceptions rather than routine retrieval. A tool proposal might fail when an authorization changes after approval rather than when the prompt contains an obvious malicious phrase. Document the input and control conditions, affected actor and evidence needed to adjudicate the outcome. Do not assign high-risk labels only after seeing a candidate fail.
Amazon Bedrock's evaluation overview describes computed scores and human ratings. Those outputs need a task-specific definition and denominator before they support an error claim. A model score cannot independently establish that a confidential answer reached an unauthorized recipient or that a later human decision was wrong. Evaluate the relevant application boundary separately.
Some mechanisms call for deterministic controls, not a statistical acceptance argument. Test permission enforcement, output validation and forbidden external actions under their own conditions. Repeated successful model answers cannot justify a missing authorization check. A challenge failure can invalidate a proposed boundary even when its real frequency remains unknown. The production evaluation workbook owns those broader system gates.
Give representative and enriched sets different jobs
A representative design needs a declared target population, selection process and observation window. Convenience samples from enthusiastic users or only completed easy tasks need narrower labels. If some task groups are inaccessible, record that gap. A sample cannot represent work excluded by the collection process merely because its final size is large.
An enriched challenge set deliberately spends more cases on a mechanism than its apparent operating frequency warrants. This is useful for debugging and testing boundary behavior. It may contain synthetic faults that should never be recreated with real sensitive records. Report what each case tests and which conditions were supplied. Passing an engineered case shows behavior for that case, not the prevalence of that mechanism in production.
If population weights and a valid sampling design support an operating estimate, retain the weights, their source and the conditional slice estimates. Do not copy a familiar market percentage into the record as a task weight. Missing slice evidence remains missing. Weighting a tiny high-risk slice can lower an overall number without improving the evidence for a prohibited outcome within that slice.
Keep outputs from the two sets distinguishable even when one evaluation job handles them. A category label helps filtering but cannot prove selection integrity. The business population register, family relationships and adjudication records live outside the score. If the service cannot preserve the distinctions needed for the claim, report separately or hold the affected inference.
Worked scenario: the average changes when sampling changes
Stipulate a fully adjudicated fictional operating ledger with 1,000 eligible policy-summary tasks: 900 routine and 100 requiring a material exception. There are nine routine errors and eight exception errors. Overall error is 17/1,000 = 1.7 percent. Within slices it is 9/900 = 1 percent and 8/100 = 8 percent. The overall fraction does not make the exception path acceptable; the task owner must assess its consequences and defined limits.
Separately stipulate a challenge ledger containing 100 routine and 100 exception tasks, with one and eight errors respectively. Its conditional fractions happen to match the first ledger by construction. Its raw error fraction is 9/200 = 4.5 percent. That higher aggregate reflects the supplied 50/50 mix, not evidence that the candidate became worse. These are two teaching ledgers, not independent observations establishing stable rates.
| Ledger or calculation | Routine slice | Exception slice | Aggregate |
|---|---|---|---|
| Fictional operating ledger | 9/900 = 1 percent | 8/100 = 8 percent | 17/1,000 = 1.7 percent |
| Fictional enriched challenge ledger | 1/100 = 1 percent | 8/100 = 8 percent | 9/200 = 4.5 percent |
| Reweight the supplied conditional fractions to the stipulated 90/10 mix | 0.90 × 1 percent | 0.10 × 8 percent | 0.9 + 0.8 = 1.7 percent |
- Fictional operating ledger
- Routine slice: 9/900 = 1 percent
- Exception slice: 8/100 = 8 percent
- Aggregate: 17/1,000 = 1.7 percent
- Fictional enriched challenge ledger
- Routine slice: 1/100 = 1 percent
- Exception slice: 8/100 = 8 percent
- Aggregate: 9/200 = 4.5 percent
- Reweight the supplied conditional fractions to the stipulated 90/10 mix
- Routine slice: 0.90 × 1 percent
- Exception slice: 0.10 × 8 percent
- Aggregate: 0.9 + 0.8 = 1.7 percent
The last row is arithmetic under supplied weights, not a validated estimate from an arbitrary challenge pack. Its transfer would require appropriate within-slice sampling, trustworthy labels and applicability to the target population. If the exception cases were chosen because they already exposed failures, their conditional fraction cannot be substituted into an ordinary population estimate without addressing that selection.
Now consider an opposing change: the exception share in the target workload becomes 30 percent while the stipulated conditional fractions stay unchanged. The weighted fraction becomes 0.70 × 1 + 0.30 × 8 = 3.1 percent. A preserved 1.7 percent headline would describe the old mix. Alternatively, a new candidate could eliminate one known exception defect while introducing a prohibited disclosure. Averaging the two changes into a better score would hide the release-blocking event.
Read a zero-error result at the denominator that produced it
Suppose a separate fictional representative sample contains 200 tasks, including only four exception tasks, and no errors are observed. The exception result is zero out of four, not zero out of 200. The other 196 cases cannot supply missing evidence about that exception mechanism. Report the actual slice count before applying any interval or planning rule.
NIST's binomial distribution reference defines a fixed probability model for binary trials. Under independent trials with one fixed error probability, the probability of zero errors is (1 - p)^n. Inverting a fixed-sample zero-event test gives the exact one-sided 95 percent upper bound 1 - 0.05^(1/n). NIST's proportion interval guidance discusses exact binomial intervals for small failure counts.
For n = 200, that derived upper bound is about 1.49 percent. For n = 4, it is about 52.71 percent. These are conditional statistical calculations, not a statement that the true rate has a 95 percent probability of lying below the bound. The design must justify independent, comparable trials and accurate binary judgments. Repeated paraphrases, shared source defects or one reviewer applying a flawed rule can break those assumptions.
A useful planning calculation is also limited. Under the same fixed model, 299 independent zero-error trials yield an upper bound about 0.997 percent. This does not mean “299 cases certify safety.” It concerns one defined binary endpoint and sampled population, excludes model selection on the sample and says nothing about an unseen subgroup. Multiple endpoints and repeated decisions need their own analysis. Stopping early when the result looks favorable is not the fixed-sample procedure just derived.
A mechanism test can matter more than a prevalence estimate
Suppose an invented rare failure occurs with fixed probability 0.008 per eligible task. Under independent trials, a 200-task sample has a 1 - 0.992^200 probability, about 79.94 percent, of observing at least one. It can still miss the failure. This probability uses a stipulated rate that a real pilot does not know in advance. It illustrates why an attractive sample size alone cannot guarantee discovery.
Enriching the relevant mechanism can improve diagnostic opportunity without supplying the missing operating rate. Test a conflicting-policy input, a stale permission and a deliberately malformed source through the permitted isolated path. Keep expected behavior and independent adjudication visible. If the model appears to handle the prompt but the application releases a cached unprotected result, the mechanism test failed at the release boundary.
Do not turn every suspected harm into a synthetic production experiment. Use approved fixtures where real disclosure, financial effects or irreversible actions would be unacceptable. Confirm that injected dependencies and sandbox outputs cannot reach normal destinations. If appropriate rights or isolation cannot be established, hold the test and assign the missing prerequisite. A sample-design problem is not permission to expose a person to risk.
Preserve disputed, missing and delayed labels
A sample is not fully adjudicated while consequential outputs await qualified review. Report inspected, resolved, disputed and pending counts. Investigating only flagged failures while calling every unflagged case correct biases the result. When an outcome matures later, retain its task identity and update the observation under the original rule. Do not classify pending cases as successful to finish a report.
The labeler must have the relevant source evidence and authority. An LLM grader can assist triage, but agreement with the candidate does not establish independent correctness. Resolve unclear rubrics through a recorded revision and re-score affected candidates when the revision changes their comparison. Keep known failures in regression and protect genuinely new acceptance families from tuning feedback.
Next action: prepare a risk-sample claim record
- Endpoint and consequence
- Specify the binary event or richer disposition, affected actor, evidence required and separate critical exclusion.
- Population and selection
- Record eligible tasks, window, slice/family definition, representative or enriched purpose and exclusions.
- Counts and weights
- Keep selected, adjudicated, errors, disputed and pending by slice. Attach the weight basis before reporting a population estimate.
- Uncertainty assumptions
- Declare fixed sample size, independence, stable probability, labeling limits and any repeated selection or testing. Do not force a binomial interval onto dependent cases.
- Mechanism and stop
- Name the isolated adverse case, expected containment, test owner and condition that holds exposure regardless of the average.
- Supported action
- Separate debugging evidence, a bounded estimated rate and release authority. Assign the unrepresented slice or unresolved label before expanding the claim.
Use the pilot decision paper to choose whether another experiment could alter the investment choice. For the next review, the evaluation lead should select one consequential mechanism, count how many independent adjudicated cases actually cover it and write the claim that denominator can support. If the claim is narrower than the proposed operating scope, retain the narrower claim and commission the missing evidence.
Related services
AI Observability, LLM Monitoring & Governance
LLMOps consulting for AI observability, LLM monitoring, evaluation and guardrails. Review production answer quality, operating failures and cost evidence.
AI Readiness Assessment Services
AI readiness assessment covering workflow needs, feasibility and operating constraints. Identify the evidence, team capability and controls needed for adoption.