Keep Shadow Traffic Away from Real Business Effects
Trace shadow requests through workers and external destinations, restrict their authority and test denied effects before mirroring customer traffic.
Follow the request beyond the shadow endpoint
Before copying a customer request to a replacement service, identify every downstream effect that request can cause. A candidate that cannot update the primary database may still enqueue work, call a partner, emit a webhook or write an object that another worker interprets as a new instruction. The comparison boundary must cover that entire path.
Consider a hypothetical delivery-quote replacement. Its handler calculates an amount and writes a diagnostic object. A shared worker reacts to that object, submits a supplier request and sends a notification. The handler's response may look harmless while the worker has already started a real business operation. Labelling the endpoint as shadow does not remove the worker's authority or change the destination.
This article proposes an isolation review for platform and migration owners. The quote scenario is illustrative, not an Ampity incident or a verified deployment. The goal is to collect candidate behavior without creating a second customer instruction. Keep any intentionally permitted diagnostic writes explicit, separated from business execution and subject to the same data-handling review.
Treat mirroring as delivery, then build isolation
Envoy's request-mirroring documentation describes shadow requests that do not delay the primary response while waiting for the shadow result. It also describes modifying the host or authority header, with configuration to change that behavior. Our recommendation is to inspect the receiving application's actual path. A shadow-marked hostname is diagnostic information, not proof that the receiver cannot execute a payment, write or notification.
Record the route, candidate identity, effective configuration and every destination used by the candidate. Include secrets, inherited environment variables, connection strings, queue producers, storage paths and feature settings. A deployment copied from production can retain a production endpoint even when its public address differs. Check what the running workload uses rather than relying on the intended configuration file.
Inspect asynchronous paths separately. A request may publish an event through a shared library whose consumer runs under another identity. That consumer can regain permissions the shadow handler lacks. Trace the event's destination and the authority used when it is handled. A read-only role on the first service cannot constrain a different role on a later service.
Assign a named owner to each boundary. The migration owner controls admission; the identity owner controls permissions; the destination owner verifies permitted behavior. Keep the observed boundary and any untested path in the review record. If the team cannot identify a consequential destination, leave the corresponding route out of the comparison cohort.
Restrict credentials, destinations and dependent workers
Use a dedicated comparison identity with only the permissions required to read approved fixtures and write approved diagnostics. Remove partner submission credentials from that workload. Give diagnostic data an isolated destination that business consumers do not watch. Verify both the candidate's identity and the identity of any worker it can cause to run.
Network controls can provide another boundary, but their scope must match the protocol and destination. A broad internet allowance can still reach the production partner API. A blocked database port does not stop an HTTPS business action. Inspect permitted paths with the network owner and avoid claiming total isolation from one denied connection.
AWS's IAM policy-simulator guidance explains that simulations evaluate policies without making a real service request. It warns that simulated results can differ from the live environment. Our recommendation is to use simulation as preliminary policy evidence, then verify the intended denial in an authorized isolated exercise. Do not send an actual charge or supplier change to production to prove that a control should reject it.
Read access and diagnostic writes also need limits. Copy only the fields approved for the comparison, restrict who can inspect discrepancies and set an owned retention period. A production customer payload stored in a shadow account remains customer data. Use synthetic records for initial boundary tests, then obtain the appropriate approval before any scoped production-data comparison.
Define what each effect is allowed to do
Record each effect's destination, permitted comparison behavior and evidence. The worksheet below is a proposed review artifact for the hypothetical quote service. Fill in the actual identity and observed result for each path. It is not an implemented security policy.
| Path from the candidate | Permitted comparison behavior | Required boundary evidence | | --- | --- | --- | | Production business database | Read only the approved fixture scope | Running identity cannot perform the prohibited transition | | Diagnostic object and event consumer | Store isolated diagnostics without business processing | Destination is separate and no consequential consumer receives the event | | Supplier or payment integration | Use an approved test adapter without production submission authority | Credentials, destination and receiving result identify the test path | | Notifications and webhooks | Capture intended output in an isolated sink | No customer destination is invoked by the handler or its workers |
Separate unavoidable operational telemetry from intentional business effects. A diagnostic request may legitimately create access logs or metering records. Define what those records contain and where they go. Do not promise that the comparison writes nothing anywhere when the approved design includes logs, but do not use that exception to permit a hidden business action.
Preserve the original operation identity for diagnostic correlation without allowing it to authorize production execution. A production destination might accept a repeated identifier as a duplicate, but duplicate prevention can expire or have narrower coverage than expected. It should not be the only safeguard against shadow requests reaching that destination.
Exercise the prohibited paths before using customer traffic
Build synthetic fixtures that reach every consequential branch, including errors and delayed work. For the delivery quote, test a missing price, a retry after a lost response, a diagnostic-object write and an event handled after the request has ended. Confirm that each branch remains within its approved destinations and identities.
Observe at the receiving boundary. A handler log saying that submission was disabled does not establish whether a shared helper or worker made another call. Capture permitted test-adapter receipts, denied operation evidence and event-sink records. Verify which identity performed the attempt. Keep credentials and unnecessary payload values out of the report.
Test a deliberately misconfigured destination in an isolated environment. The expected result should be rejection before an unauthorized business effect, with an observation that identifies the enforcement layer. A DNS error can demonstrate a failed lookup while leaving production authority unchanged. Record that narrower result rather than presenting every failed request as a successful permission test.
Include delayed effects in the observation window. A scheduled retry or queue consumer can run after the traffic copy stops. Choose the window from the configured execution and retry paths, and record any branch whose delay is not bounded. If the team cannot observe the receiving result, classify that path as unverified rather than inferring safety from a quiet dashboard.
Stop admission without assuming earlier work vanished
Set limits for copied request rate, queue depth, dependency pressure and diagnostic storage. Shadow work can compete with the serving workload even when it cannot make business changes. Measure the customer's path while the comparison runs. A candidate with a healthy error rate can still place enough load on a shared service to affect production.
Rehearse an independently available stop control. Stopping new copies should not depend solely on the candidate service whose behavior is under investigation. Record the last admitted comparison identity and inspect outstanding requests, diagnostic events and retry state. Distinguish work that never started from work already accepted downstream.
If an unexpected effect occurs, preserve its original identity and obtain destination evidence before another attempt. Disabling the mirror contains future admission; it does not undo a supplier change or retract a delivered message. Assign the business owner to any corrective action and record its result separately from technical containment.
Keep the candidate isolated during investigation. Restoring production credentials to make a failing shadow test pass defeats the acceptance boundary. Repair the test adapter, required diagnostic path or intended read permission through the appropriate review, then rerun the same fixture with its history intact.
Extend the same boundary to AI-assisted candidates
An AI component may select a tool or propose a follow-up action that the deterministic handler did not anticipate. Inventory the tools available in the comparison environment and bind them to the same test destinations and limited identities. Instructions telling the model to avoid writes cannot replace receiving-side restrictions.
Check retries and background tool execution as part of the fixture. A model that replans after a timeout can invoke another path, and an orchestration worker may use a separate credential set. Record what the tool actually attempted and the destination's observed result. If those records are missing, retain an unresolved path in the acceptance report.
A denied effect can be a successful isolation test while still showing incomplete application behavior. Record both conclusions. The migration owner may need further testing before accepting the candidate's business result, even after the security owner accepts its containment. Keep those approvals separate so isolation does not get mistaken for functional readiness.
Start the next review with one copied request and a complete downstream-effect inventory. For output acceptance, use the old-and-new comparison contract. For phased replacement, use the capability migration playbook. Bring the tested boundaries and open paths to a platform modernization review or cloud security review. Readers can use these resources without providing an email address.