Which AI Pilot Is Worth Running? Decision Value, Exposure and Adoption
Compare no pilot, shadow replay, human assistance and limited live use using decision-reversal evidence, separate review capacity and honest adoption denominators.
Executive summary: select an experiment that changes a decision
An AI pilot is worth running when its result can change a decision the organization has to make. A fluent demonstration is insufficient if the same funding, operating or release choice will be made regardless of the result. Define the decision, the unresolved uncertainty and the action under opposing results before building the pilot. Then compare no pilot, a shadow replay, human assistance and limited live use. These options answer different questions and expose different people and systems. More exposure is not automatically better evidence.
This paper argues for a constrained selection rather than a weighted maturity score. Data authority, unacceptable consequences and the ability to stop are prerequisites, not points that a large potential benefit can offset. Review capacity is part of feasibility. User adoption requires observation under a usable operating arrangement, not inference from model accuracy. An experiment that measures output quality but cannot reveal whether anyone can use its result may justify another experiment, not an investment in a complete product.
The worked example concerns fictional internal change-request evidence summaries. The model may draft a source-linked summary, but cannot approve a deployment, alter a ticket or send a notification. All quantities are stipulated for inspection, not observed outcomes. A shadow replay is initially preferred because review effort and correctness remain unresolved. A later supplied packet can justify recommending a small human-assist experiment. Changed staffing and discovered accepted errors reverse that recommendation. Neither conclusion authorizes a real pilot or establishes return on investment.
1. Write the decision that the experiment can reverse
Begin with a statement such as: “Should we fund a reviewer-assisted evidence-summary path for this change-request family, retain the current manual path, or stop this use case?” Name the person accountable for that choice and the deadline by which the evidence must matter. Distinguish it from the later decision to enable a production capability. A sponsor can approve money for an experiment while a data owner withholds source access and an application owner withholds release authority.
Specify an opposing result. In the fictional example, a supported draft that takes less active checking than manual summarization could justify investigating human assistance. A draft that repeatedly omits a material dependency or takes longer to verify would support retaining the manual path or redesigning the task. If both results lead to the same full build, the activity is implementation with a demonstration attached, not a decision experiment. Be honest about that purpose rather than pretending it is exploratory research.
Also ask whether the information arrives soon enough. A lengthy pilot may finish after the procurement decision or staffing change it was meant to support. Some uncertainty is cheaper to resolve through a rights review, an interface walkthrough or a deterministic rule comparison. The AWS Generative AI lifecycle treats scoping as a business and feasibility activity. The proposed reversal record makes that work explicit without assuming that generation is the preferred intervention.
2. Separate uncertainty from a missing prerequisite
A learnable uncertainty might be whether reviewers can identify unsupported claims from the proposed evidence interface. A missing prerequisite might be that the organization has no permission to copy the source records into the intended evaluation path. A pilot cannot ethically resolve the second by doing it anyway. Classify the item before estimating learning value. Record who can supply the prerequisite, which evidence closes it and whether an isolated synthetic substitute answers the same decision or a narrower one.
Use four categories: factual unknown, design hypothesis, operating prerequisite and non-negotiable exclusion. A factual unknown needs an observation. A design hypothesis needs a discriminating comparison. An operating prerequisite needs an owner and acceptance record. An exclusion stays excluded even when an output looks excellent. The categories can change after investigation, but their history should remain visible. Do not turn “unknown” into “pass” because a deadline approaches or because a vendor feature appears in a console.
The AWS Responsible AI design principles connect narrow use cases to applicable risks and release criteria. Here the task excludes deployment approval and customer communication. That restriction makes a safe evidence-summary experiment more plausible, but does not eliminate confidentiality, misleading content or reviewer over-trust. A read-only output can still affect a consequential judgment. The experiment contract must say who sees it, what they may infer and which decision it cannot make.
3. Compare four experiment modes without implying a ladder
No pilot is a legitimate choice when an unresolved prerequisite dominates the decision, the uncertainty is immaterial or an ordinary workflow change addresses the problem. Retain the current process and spend only on the missing evidence when that is the defensible option. It is not a failed innovation score. For example, a well-maintained template and an evidence checklist may solve inconsistent summaries without a model. Compare that possibility before treating the current process as an intentionally weak opponent.
Shadow replay uses approved historical or synthetic inputs and produces outputs that do not influence operational decisions. It can test supported claims, evidence coverage, malformed inputs and independent review effort. It cannot establish live adoption, interruption cost or how people behave when a recommendation appears beside real work. The replay path must not secretly write to business systems. Calling a production endpoint in “test mode” does not establish isolation; its downstream effects require inspection.
Human assistance presents a proposal to a bounded user group while the person retains the defined decision. It can reveal active review time, correction, nonuse and interface friction. It also creates exposure to persuasion and inappropriate reliance. Limited live use permits a separately authorized operational effect for a narrow cohort with enforced controls. It can observe delivery and recovery, but should not be selected merely because it seems closer to a finished product. There is no automatic progression between these modes.
- No pilot
- Useful when rights, decision relevance or the non-AI alternative dominate. Evidence sought: prerequisite closure or workflow comparison. Exposure: existing process, with its own known risks retained. Cannot establish candidate output quality.
- Shadow replay
- Useful for supported-claim quality and independent review feasibility. Exposure: approved dataset and evaluators only, no operational effect. Cannot establish actual adoption or live recovery.
- Human assistance
- Useful for real review effort, correction and nonuse in a bounded interface. Exposure: participating reviewers can be influenced. Cannot authorize automatic completion or prove causal benefit without a suitable comparison.
- Limited live use
- Useful for observing a separately permitted effect and its recovery. Exposure: declared operational cohort and consequence. Requires rights, capacity, safety controls and release authority before entry; evidence is not a universal rollout verdict.
Figure 1. A proposed option comparison, not a maturity ladder or provider architecture. The adjacent records contain the equivalent decision boundaries. The document composition preserves the same option boundaries at print width.
4. Define decision value without invented probabilities
Decision value means the importance of resolving an uncertainty for the choice at hand. It is not identical to predicted revenue, model score or the excitement generated by a demo. Record which option would be chosen under each plausible finding, what commitment that choice creates and whether the finding is likely to be distinguishable with the proposed evidence. If the study cannot separate the alternatives within the available time and permissions, narrow the question or choose another method.
Do not attach convenient probabilities to unknown outcomes simply to calculate an attractive expected monetary value. A sponsor may supply a probability model, consequence valuations and a basis for them. Those assumptions need their own review and sensitivity, including how the preferred experiment changes when they change. This paper instead uses a decision-reversal table expressed as records: favorable supported summaries and feasible review lead to a human-assist recommendation; excessive checking or unsupported dependencies lead to redesign or no pilot.
Include the value of a negative result. An experiment that prevents a difficult-to-reverse commitment may be useful even when it produces no deployable feature. Conversely, learning that a model can summarize familiar examples is low-value when that was never the disputed issue. Ask the sponsor to identify the smallest piece of new evidence that could alter the next tranche. Keep learning claims narrow: rejecting an experiment design is not proof that all AI approaches to the business problem lack value.
5. Set an exposure budget before selecting the attractive option
State who and what can be affected, not just the number of requests. Exposure includes data copied, outputs shown, decisions influenced, actions reachable, duration of unresolved effects and effort needed to contain them. A ten-user experiment can be more consequential than a thousand isolated replays. Use a positive scope definition and named exclusions. The fictional experiment includes internal, non-emergency change summaries and excludes privileged incident records, customer secrets, emergency changes and approval recommendations.
Define stop conditions before the first output. An unauthorized source, absent material dependency, failed access check, inability to find the original evidence or unavailable pause control requires the owner to hold further exposure. The threshold comes from the use case, not a universal percentage. Preserve in-flight outputs and explain their disposition. A switch that hides the interface without accounting for already relied-on summaries is not a complete stop mechanism. The experiment owner needs a way to withdraw and correct those records.
Configured safeguards remain components of the system, not business authorization. Amazon Bedrock Guardrails documents several filter types, including probabilistic sensitive-information detection. A successful filtering response does not establish that a summary is factually complete or that its reader may approve a deployment. Test the actual protected boundary separately. The limited scope in this paper avoids proposing operational commands, but the same separation matters when an experiment later includes a real action.
6. Budget review and adjudication as separate constrained lanes
Operational review and independent adjudication serve different purposes. The operational reviewer checks a proposal before using it. The adjudicator checks whether a reviewer-accepted result was actually correct under an independently defined rubric. The same person may contribute in different roles only if the design addresses independence and time explicitly. Do not count the same available minutes in both lanes, or assume that a sampled pass review can replace required review of every consequential proposal.
For the fictional ten-working-day human-assist design, cap intake at 20 distinct tasks per day. Assume six active review minutes per task, covering correction and all visits, for 120 minutes daily. A review lead assigns 150 minutes to this lane, leaving 30 modeled minutes of margin. Independently inspect four accepted tasks per day at eight minutes each: 32 minutes against a separate 40-minute adjudication allocation, leaving eight minutes. These stipulated averages are neither observed staffing evidence nor waiting-time guarantees.
The adverse case matters. At eight review minutes, the same intake needs 160 minutes and exceeds the first lane by ten. At twelve adjudication minutes, four checks need 48 and exceed the second by eight. A favorable sum across both lanes cannot transfer authority or capacity between them. Reduce intake, obtain appropriate capacity or hold the experiment. The review-capacity article covers arrivals, backlog and recovery in detail; this selection decision uses it to reject an infeasible experiment before commissioning.
7. Keep the comparator credible and the sample inspectable
The comparator is the current workflow as it operates, not a caricature of people making every summary slowly and flawlessly. Preserve comparable task difficulty, evidence access and reviewer experience. In a replay, paired cases can show how each method handles the same evidence, but reviewers recognizing the same item can contaminate effort and quality observations. Counterbalance ordering or use a defensible assignment design and explain what it controls. Do not imply randomization or causal inference when none occurred.
Freeze the rubric before looking at favorable outputs. A correct summary must preserve the relevant change, supporting evidence and material unresolved dependency, not merely read well. Independent labels require qualified review. Difficult cases with genuine uncertainty should retain an unresolved disposition rather than a guessed gold answer. Keep task identifiers separate from attempts and preserve failures, timeouts, withdrawals and exclusions. Excluding inconvenient records after observing results changes the experiment, so retain the original and revised population definitions.
Amazon Bedrock's evaluation overview distinguishes automatic, human-worker and judge-model approaches. Selecting a service mode does not supply the task's acceptance contract or make its graders independent and accurate. Check selected-model evaluation support for the intended path rather than inferring support from a model listing. No particular model, Region or evaluation job is executed or recommended by the fictional comparison here.
8. Measure accepted errors, not only approvals
A reviewer acceptance is an observed workflow state, not ground truth. Record independently confirmed correctness after acceptance, including the reason and severity of any disagreement. An accepted summary that invents a missing dependency can be more dangerous than a conspicuously incomplete one because it invites reliance. Define false acceptance as a reviewer-accepted result that independent adjudication finds unacceptable under the frozen task contract. Keep model error, reviewer error and unclear rubric distinct during investigation, even when they combine into one unsafe result.
Report both the denominator among accepted results and the eligible task population. In a supplied fully adjudicated synthetic ledger, 120 of 200 offered tasks are reviewer-accepted, but four accepted summaries are incorrect. The observed accepted-error fraction is 4/120, approximately 3.33 percent. Independently correct accepted work is 116/200, or 58 percent of offered tasks. Neither percentage is a deployment tolerance. In this scenario one of the errors is a prohibited material omission, so widening is held regardless of the useful-task proportion.
Partial adjudication cannot support those full-ledger claims. If only 40 accepted records were independently inspected, describe the sample and its selection, not “all 120 are correct.” Even zero errors in 40 independent representative trials does not prove a rare-error rate of zero. Under a fixed independent binomial model, the exact one-sided 95 percent zero-event upper bound is 1 - 0.05^(1/40), about 7.22 percent. NIST's proportion interval guidance explains the need for an appropriate interval method. Clustered users, changing inputs and flawed labels can invalidate that simple illustration.
9. Track adoption without changing the denominator
Define adoption at the unit the business needs. Opening a proposal is not starting the task; starting is not completing; completion is not correct acceptance; correct acceptance is not reuse in the intended workflow. In the fictional ledger, 200 eligible offered tasks lead to 160 started, 140 completed, 120 reviewer-accepted, 116 independently correct and 90 actually reused. The supplied counts are nested. Reuse is 90/200, or 45 percent of offered tasks, and 90/116, approximately 77.59 percent of correct accepted tasks. Show both because they answer different questions.
Record why the other 110 offered tasks did not yield reuse. Some users may not need a summary; others may lack time, evidence access or confidence in the interface. Do not label all nonuse model failure, and do not delete nonusers from the headline denominator. Identify whether tasks were genuinely offered, whether the user had a reasonable opportunity and whether observation had matured. A task with an outcome due next week should remain pending, not become a failure or a conveniently omitted record today.
Shadow replay supplies no real adoption observation. Human assistance can observe behavior in its recruited cohort, but participants may be unusually motivated or supported. Limited live use adds operating context without automatically eliminating selection bias. Compare the intended population with the observed one and record what remains unrepresented. If the intervention replaces part of the user's process, check whether it adds double entry or shifts effort elsewhere. Apparent uptake can coexist with an unusable operating product after the pilot team withdraws.
10. Reconcile the worked budget and avoid a false ROI claim
The ten-day human-assist budget uses hypothetical USD amounts, not vendor prices or Ampity rates: setup 2,400, platform and evaluation 200, review 20 hours at 40 per hour for 800, and adjudication 320 minutes divided by 60 at 60 per hour for 320. Total modeled cost is 3,720. The categories are disjoint; setup does not include the separately counted review or adjudication work. Replace this basis with actual procurement and staffing records before using it in an investment decision.
A stipulated baseline of ten active minutes for each of 200 comparable summaries uses 2,000 minutes. The proposed operational review uses 1,200, giving an 800-minute gross difference. Subtract the additional 320 adjudication minutes to show a 480-minute net recorded effort difference on this narrow boundary. Setup, engineering, support, waiting and downstream correction are not included in that effort difference. It is not a cash saving and must not be multiplied into a return claim while the actual adoption and correctness observations are absent.
Do not divide a future cost budget by a preferred future adoption count and label it observed unit economics. Use the AI workflow unit-economics paper and AI accepted-result calculator when equivalent task populations and supplied costs can be reconciled. Here the appropriate output is an experiment commitment and sensitivity, including the adverse staffing case and an option that stops. Funding, credits or a partner programme cannot substitute for task value or be assumed as delivery cash.
11. Keep data, evaluator and trace authority explicit
Document who permits source access, model processing, independent evaluation and retention. Permission for an application to read a record does not necessarily permit copying it into a separate dataset, granting evaluators access or retaining its output indefinitely. A synthetic pack can test defined failure behavior without containing real sensitive material, but it cannot establish how frequent that behavior is in the real population. Minimize the copied fields and preserve the reason each is needed for the decision.
An evaluation service can introduce copies outside the application's familiar stores. Bedrock evaluation data management describes temporary AWS-owned evaluation copies and encryption choices. Do not extrapolate the documented job-copy deletion to original datasets, exported reports or customer logs. The owner needs a complete retention and access inventory. This paper does not configure an evaluation job, choose a key policy or establish legal rights to any customer's records.
Tracing also needs an explicit boundary. Supported Bedrock runtime invocation logging is disabled by default and can contain input and output content when configured; endpoint coverage is not universal. A trace field saying “summary complete” cannot authenticate the sources or the later human decision. Keep the smallest approved references sufficient to join task, candidate, reviewer disposition, independent correctness and use. If that join is unavailable, hold the affected claim rather than supplying optimistic missing values.
12. Evidence does not become production authority automatically
An experiment report should separate four evidence streams: task correctness, reviewer capacity, data and security boundaries, and adoption/operating feasibility. Each has a qualified owner and a specific version. A strong quality result with unresolved staffing is not a balanced average; it is an incomplete proposal. A populated adoption ledger with unknown data rights also remains held. The named release owner receives the packet and makes a separately authorized bounded decision under the organization's actual controls.
Bind the proposed action to the accepted task family, user cohort, system version, permitted effects and expiry triggers. A model, prompt, evidence source, policy, interface or staffing change can invalidate some evidence while leaving other findings useful. Explain which part must be recollected. Do not erase historical failures after a favorable rerun. The production evaluation workbook provides the execution procedure; this paper supplies the upstream option-selection and adoption record that it should serve.
Figure 2. Solid arrows mean evidence submission or a separately owned decision, not traffic or an automatic permission. HOLD retains an owner and next check. The diagram is conceptual and does not assert an AWS deployment.
The NIST AI Risk Management Framework provides voluntary lifecycle risk context, not a compliance certificate. Treat documentation as a way to support scrutiny, not a substitute for scrutiny. If the release owner cannot identify the input evidence, its limits and the action being permitted, the right next step is to improve the packet. A sign-off called “AI approved” is too broad to bind the decision this example needs.
13. A filled record shows why the selected experiment changes
The initial fictional packet has approved synthetic replay inputs and a fixed summary rubric, but no observed review effort and no approved live operating cohort. The sponsor chooses shadow replay to test material omissions and active checking effort. The result that changes the next decision is not “a good summary.” It is independently supported task quality within the named scope together with review/adjudication feasibility. If either is absent, keep the human-assist proposal held and assign the missing observation.
The following filled record describes a later hypothetical recommendation using stipulated favorable capacity values. It is not a completed experiment. The companion arithmetic can reconcile inputs, but cannot authenticate the evidence references, evaluate the task or authorize data use. Its favorable state is explicitly a model review result. An actual release owner must still verify the packet and permissions. The opposing records immediately below show how a superficially similar proposal reaches a different recommendation.
The executable record separates prospective plan P40, with 40 budgeted adjudications, from illustrative ledger L120, with 120 adjudications in another population and observation window. Each has a supplied identity, scope, basis and window. A separate prior cohort is also permitted but requires external applicability review; its outcomes do not establish this plan's completion, correctness or adoption. Missing or unknown relationships hold. A same-plan record must match the declared plan identity, scope, population and window. Its 120 adjudications cannot borrow the 40-check budget: the model retains both and holds the overrun. Rebudgeting all 120 checks gives 960 adjudication minutes and USD 4,360, still exceeding the separate daily capacity. Prospective cost is never reported as realized cost or divided by another cohort's correct outcomes.
- Decision and owner
- Recommend the next bounded human-assist experiment or retain shadow; fictional product owner Mira owns the next tranche, not deployment approval.
- Task and exclusions
- Internal non-emergency change evidence summaries; no approvals, ticket writes, messages, privileged incident sources or customer secrets.
- Uncertainty and reversal
- Can an evidence-linked proposal reduce active checking without a prohibited omission? Favorable evidence supports limited human assistance; excessive checking or a material omission supports redesign or no pilot.
- Alternative and comparator
- No pilot, shadow, human-assist, limited live. Compare current manual summaries on the same task contract; limited live stays held for missing operating prerequisites.
- Input and rights reference
- Fictional SYN-20 and rights-review-R1 are supplied aliases, not authenticated permissions or customer data.
- Exposure and stop
- 20 tasks/day for ten working days; named review lead stops further display on a material unsupported dependency, unauthorized evidence or unaccounted output.
- Capacity by lane
- Review 120/150 minutes daily; independent adjudication 32/40. Margins 30/8. Neither lane borrows the other's authority or minutes.
- Outcome and adoption denominators
- Budget tasks 200. No outcome observed. Separate illustrative fully adjudicated ledger 200/160/140/120/116/90 is not evidence for this proposed run.
- Commitment and unknowns
- Hypothetical USD 3,720. No realized saving, actual reviewer distribution, supported model placement, live adoption or financial return established.
- Prerequisites and next authority
- Data/security, task evaluation, capacity and adoption/operations each require separate owners. Release owner receives the packet; no automatic production grant.
- Version and invalidation
- SUMMARY-R1, RUBRIC-R1, SYN-20. Changed evidence, task policy, interface, model path or staffing triggers affected recheck.
- Next action and retained record
- Mira asks evaluation and review leads to challenge the packet and adverse case before spending; retain both recommendations and hold reasons.
Two opposing recommendations using explicit assumptions
Case A: capacity-supported recommendation. All required synthetic packet references are supplied, critical accepted errors are stipulated zero, review needs 120 of 150 daily minutes and adjudication 32 of 40. The arithmetic permits a model-only human-assist recommendation. It does not establish that those conditions exist or that zero observed errors justify a rare-harm bound. Limited live is still not an automatic next step.
Case B: adverse work and false acceptance. Keep the same scope, task count and budget, but increase review to 160 of 150 and adjudication 48 of 40; additionally supply one critical accepted error. The recommendation is HOLD, with three distinct reasons. Cheap platform use or a higher reuse proportion cannot offset them. If the error instead comes from an unclear rubric, resolve that ambiguity before relabeling the case, preserving both versions and the independent judgment.
14. Use the blank record to challenge the preferred option
Complete the record with the people who will carry the work, not only the people proposing the model. Bring the existing workflow owner, evaluator, reviewer lead, data/security owner and sponsor together. Require an evidence reference and an owner for each asserted prerequisite. Empty fields are open items, not zeros or approvals. A useful first meeting can end with no pilot and one owned evidence task. That outcome is better than approving a broad demonstration that cannot inform the later commitment.
- Decision and owner
- State the next actual investment choice, accountable person and date; distinguish it from release authority.
- Task and exclusions
- Name the user, supported job, effect and positively excluded sources/actions/cohorts.
- Uncertainty and reversal
- State the opposing findings and what action each would change. Explain why the evidence discriminates.
- Alternative and comparator
- Keep no pilot and a plausible non-AI option. Specify what shadow, assistance or live use can and cannot observe.
- Input and rights reference
- Record dataset/corpus version, provenance, permitted processing, evaluator access and retention. Unverified aliases remain unverified.
- Exposure and stop
- Record intake, duration, reachable effects, affected people, stop owner and disposition of already exposed work.
- Capacity by lane
- Record distinct review/adjudication allocations, full active visits, arrival variation and adverse capacity case.
- Outcome and adoption denominators
- Define offered, started, completed, accepted, independently correct and reused, including pending, rejected and unsampled work.
- Commitment and unknowns
- State cost basis/currency/categories and unknowns. Keep capacity, cash and expected value separate.
- Prerequisites and next authority
- Name independent evidence owners, missing conditions and the separate person who may permit the bounded next effect.
- Version and invalidation
- Bind task, rubric, system, evidence and cohort versions; name changes that require affected re-evaluation.
- Next action and retained record
- Choose one owned observation or hold remediation, deadline and retained record of previous failures.
15. Limits, acceptance checklist and next useful action
This method does not estimate a universal probability of AI success, certify safety, settle data rights or prove investment return. The fictional packet uses supplied aggregates and trusted opaque references. Changing their underlying meaning without changing the reference can conceal an evidence change; immutable-reference discipline and validation are external responsibilities. The offline companion can expose inconsistent nested counts and capacity arithmetic, but cannot detect false labels, selection bias, hidden costs, actual processing or user intent. A favorable calculation therefore remains non-authorizing.
Before commissioning, ask whether opposing results truly change the choice, whether the selected mode can observe that uncertainty, whether data and exposure are permitted, whether the two staffing lanes are feasible under adverse conditions, and whether accepted work is independently checked. Require a denominator for nonuse and pending work, a separately owned release boundary and a record of what invalidates the conclusion. Do not turn these checks into a single score that hides a critical failure.
The next useful action is for the product owner to complete the reversal, exposure and capacity fields with an evaluation lead and reviewer lead before selecting a provider. Ask the data/security owner to resolve one blocking source or retention assumption. Then use the production evaluation workbook to execute only the permitted study. If the next decision cannot be stated or no study result would alter it, retain the current process and revise the proposal. Optional Ampity contact can address that specific evidence gap; it is not required to read this paper or inspect its records.
Reproduce the bounded arithmetic offline
Download the experiment decision companion. It includes a blank decision record, fictional packet, supplied cases, the evaluator and its tests. With Node installed, extract the archive and run node --test experiment.test.mjs, then node experiment.mjs packet.json. These commands run locally without credentials, provider calls or external storage. The output separates the prospective budget from illustrative, prior or same-plan supplied outcomes. It never authorizes an experiment or production use, verifies rights, authenticates evidence or reports realized savings. Keep sensitive source contents out of the packet.
Related services
AI Readiness Assessment Services
AI readiness assessment covering workflow needs, feasibility and operating constraints. Identify the evidence, team capability and controls needed for adoption.
AI Observability, LLM Monitoring & Governance
LLMOps consulting for AI observability, LLM monitoring, evaluation and guardrails. Review production answer quality, operating failures and cost evidence.