Test the Client Retry Contract Before Moving an AI Request to Bedrock

Inspect SDK-owned attempts, authentication, endpoint paths and timeout handling before changing an AI client's base URL. Includes an offline probe and a worked...

Before changing an AI client's base URL, test the request contract at the transport boundary. Count the attempts the SDK actually produces, check the assembled path and authentication mode, and observe how injected errors reach the application. A request that parses under the same method name can still violate your application's retry, deadline or recovery contract.

This article is for the engineer adapting an existing synchronous, text-only Chat Completions call to an Amazon Bedrock endpoint. Its output is a client-behavior record, not a model-quality evaluation or permission to send data. The worked probe uses OpenAI Python SDK 3.26.1 and HTTPX2 2.13.1 with an in-memory mock transport. Every response, token placeholder and error is synthetic. No AWS request, model invocation, credential lookup, customer traffic or account inspection was executed.

The result is deliberately narrow. It shows what this pinned client does when its transport supplies particular outcomes. It cannot establish which outcome Bedrock will return for a real request, how long inference takes, whether the account has access, or whether generated text satisfies the product's task. Keep those questions in separate evidence records.

1. Write the route as a complete request contract

The candidate documentary route is https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1, the Chat Completions operation and model identifier openai.gpt-oss-120b-1:0. AWS's Chat Completions guide documents that Runtime path and bearer-key or SigV4 authentication. It also warns that Runtime does not provide OpenAI-compatible model listing. A successful client.models.list() check on another endpoint therefore cannot be your acceptance test for this route.

The gpt-oss-120b model card identifies the Runtime model ID and Chat Completions support, while excluding Responses on that endpoint. This is documentary support for the chosen shape, not observed account eligibility. Do not combine the Mantle model ID with the Runtime path or turn a Chat Completions response into proof that a stored Responses workflow migrated.

The probe checks that the SDK appends /chat/completions exactly once. Its resulting URL is the complete Runtime operation path, not a base URL already ending in that operation. The mock also checks POST, the model field and the synthetic message. These assertions catch mistakes before any transport that could reach a provider is supplied. They do not inspect DNS, TLS, proxy routing or actual destination processing.

Keep the current application capability explicit. The capability-split article owns decisions about retaining hosted search or conversations. This article assumes the request has already been narrowed to a self-contained text operation. If the operation still depends on provider-held objects, stop this client adaptation and resolve that dependency first.

2. Keep authentication evidence out of the prompt

The fixture supplies an unmistakably invalid placeholder token. Its transport asserts that the configured bearer header arrives, then records only a boolean indicating the expected synthetic header was present. It never records a credential value. The hostname in the mocked request does not turn this into authenticated AWS traffic: the in-memory transport supplies every result without a socket or service call.

AWS's API-key usage guide explains bearer headers and supplying a Bedrock key as an SDK API-key parameter. That does not mean an existing direct-provider key can be reused. AWS's API-key overview distinguishes short-term keys tied to their generating principal from long-term exploration keys. Actual key permissions, expiry and refresh remain outside this offline test.

A SigV4 configuration is a different authentication path. Merely renaming an environment variable or changing a URL does not demonstrate signing, credential refresh or permission enforcement. The current SDK also has a Bedrock provider configuration; its configuration must be checked separately rather than assumed equivalent to this explicit bearer-mode fixture.

For an authorized integration test, the security owner should provide a bounded credential mechanism through secret settings, not a copied key in a worksheet. Capture its permitted identity and authentication mode by reference. Test denied access and expiry without publishing secrets, and record whether the failure was observed at the provider, signer, transport or application boundary. A displayed “authentication error” alone does not answer that provenance question.

3. Observe SDK attempts before adding application retries

The pinned OpenAI Python client documentation describes two default retries for selected connection/status failures and configurable retry and timeout settings. The versioned base-client implementation is the implementation reference for the exact release. Those are client policies, not a statement that every retried inference is safe or free.

In the offline probe, an always-successful injected response produces one transport attempt. Injected 401, 403 and 404 outcomes produce one attempt and their corresponding typed SDK errors. Injected 408, 409, 429 and 500 outcomes produce three attempts with the configured two retries. A synthetic 429 followed by a synthetic 200 produces two attempts and a parsed fixture response. Disabling SDK retries produces one attempt for the same synthetic 429.

The comparison below is an observed local client result, not a Bedrock error-rate measurement. Every response is generated inside the test. The fixture supplies a small retry-delay hint so status-failure tests finish promptly; it is not a claim about a provider's actual delay headers or production backoff distribution.

Injected transport sequenceSDK retry settingObserved attemptsLocal result
20021Parsed synthetic response
401, repeated21AuthenticationError
403, repeated21PermissionDeniedError
404, repeated21NotFoundError
408, 409, 429 or 500, repeated23 per separate caseSDK status error after budget exhaustion
429 then 20022Parsed synthetic response
429, repeated01RateLimitError
Synthetic read timeout, repeated23APITimeoutError
Synthetic application ValueError21Original application exception

Retain the complete attempt sequence instead of recording only the final success. Two requests can have the same final answer field while differing in attempted work and operational exposure. Conversely, three local mock attempts prove no remote charges, accepted outcomes or repeated side effects. The counting unit is transport attempts under one SDK call, not users, logical tasks or accepted answers.

4. Separate a transport timeout from the task deadline

The fixture raises a ReadTimeout immediately. It does not wait for a real two-second timeout or simulate partial provider execution. The observed three attempts show how this SDK handles that injected exception. They cannot measure wall-clock delay, establish when a request crossed the provider boundary, or prove that cancellation stopped remote work.

Your product's deadline may start before the SDK call, including queueing, retrieval and approval. Its timeout configuration may apply to connection, write, read or pool behavior rather than one global end-to-end stopwatch. Record those scopes. A setting of two seconds cannot be multiplied into a guaranteed total runtime unless the relevant implementation actually enforces that complete bound.

The HTTPX2 migration guide documents the current transport family and changes affecting custom clients and certificate trust. The mock exercises that transport interface, not actual certificate verification. A deployment with a corporate proxy or container trust store needs its own permitted integration checks. Disabling TLS verification to obtain a green request would defeat the purpose of the review.

Give the application an independent monotonic deadline and a disposition for late or uncertain work. Test whether an expired task suppresses subsequent attempts and rejects late results. Do not claim that from this probe, which intentionally has no production task scheduler. Its timeout evidence stops at typed client behavior. A changed transport, hook or async runtime should reopen that evidence rather than borrowing this synchronous observation.

5. Work the nested-retry exposure before choosing an owner

Consider a synthetic planning case with one application retry and two SDK retries per application call. If every inner call exhausts its retry allowance, there are 2 × 3 = 6 transport attempts for one logical task. That is a structural maximum for this stipulated nesting, not the measured result of a load test or a recommended retry policy.

To illustrate the deadline conflict, stipulate that each attempt consumes two seconds, the two internal waits sum to one second per SDK call, and the application's wait between calls is one second. The planning total is 6 × 2 + 2 × 1 + 1 = 15 seconds. A ten-second task requirement is incompatible with that particular worst-case plan. All durations are assumptions, not measured provider timing or verified timeout caps.

With the same one outer retry but zero SDK retries, stipulate two attempts and the same one-second application wait. The planning total becomes 2 × 2 + 1 = 5 seconds. This fits the teaching budget arithmetically, but does not establish acceptable failure handling. A task may still need to stop after an uncertain first attempt rather than dispatch a second one. Fewer attempts alone do not grant retry authority.

Stipulated policyMaximum attempts in the planAssumed wait totalAssumed elapsed totalTeaching conclusion
One outer retry, two SDK retries63 seconds15 secondsExceeds the ten-second planning requirement
One outer retry, zero SDK retries21 second5 secondsArithmetic fits; retry safety and real timing remain unproved
No outer retry, zero SDK retries10 seconds2 secondsOne planned attempt; no automatic recovery

The companion's arithmetic accepts explicit bounded integer retry counts and finite nonnegative time assumptions. Missing values, booleans, strings and nonfinite times are rejected rather than guessed. Zero time is permitted as an explicit mathematical assumption, not proof that work is instantaneous. Every result keeps execution authorization, provider observation and deadline enforcement false.

Choose the retry owner based on the application's evidence and control needs. A deliberately configured SDK budget can be appropriate for a bounded operation. An application-owned policy can provide better task-level deadline and attempt tracking when the SDK's behavior is insufficient. Avoid enabling both independently because each layer appeared conservative in isolation. Preserve the rejected nested plan so a future configuration change cannot quietly reintroduce it.

6. Translate errors into owned dispositions

An SDK exception is an input to recovery, not the recovery decision itself. A denied identity or wrong route should send work to configuration or security review rather than a loop that keeps trying. A throttled task may wait only if its remaining deadline, access and budget permit that action. An uncertain dispatched attempt needs an explicit status and evidence owner even when the operation generates text rather than writing a business record.

Keep two views of the outcome. The client view records status or exception, attempt count and sanitized receipt metadata. The application view records whether the logical task remains pending, is held, was rejected, or produced an output that still needs acceptance. A parsed response should not mark a draft accepted or attach it to a different task because the original requester already timed out.

The mock's custom ValueError case checks that an application-origin exception is not silently reclassified as a transient provider failure in this fixture. Extend the actual application's tests for its own hooks and transports. If a middleware layer changes that behavior, retain the old and new observations with their dependency versions. Do not use a broad catch-and-retry handler that hides configuration defects as outages.

Useful diagnostic signals include more transport attempts than the declared task budget, successful late responses after task expiry, repeated denied-credential outcomes, requests to an unintended path, and a final error with no retained attempt trail. A favorable metric is attributable completed work inside the policy, not merely fewer visible exceptions. The existing agent architecture paper owns the broader approval and reconciliation model.

7. Record the actual client evidence and what is still unknown

The filled record describes only the local fixture. The matching blank record helps an application owner replace it with permitted evidence without pasting raw prompts, bearer keys or customer data into a public review.

Operation and route
Synchronous text Chat Completions; Runtime us-east-1 base path and openai.gpt-oss-120b-1:0. Documentary route; no provider dispatch.
Client and transport
OpenAI Python 3.26.1, HTTPX2 2.13.1, in-memory MockTransport with trust_env disabled. Exact local dependency identity, not production deployment proof.
Authentication boundary
Explicit invalid fixture token; expected bearer-header shape asserted locally and token omitted from observations. No SigV4, credential refresh or IAM check.
Attempt and timeout evidence
Two configured SDK retries; injected 429 gives three attempts, 429 then 200 gives two. ReadTimeout is injected immediately, not a measured wait.
Application policy and disposition
No outer scheduler or production recovery controller in the probe. Parsed fixture response and exception type are observations; task acceptance and replay authority remain unestablished.
Budget and next action
Six-attempt nested planning example totals 15 stipulated seconds against a ten-second requirement. Application owner must choose a reviewed retry owner and establish actual timing and safe recovery separately.
Operation and route
Record exact capability, API, assembled path, model/version, source Region and account evidence. Name excluded state, tools, streams or modalities.
Client and transport
Record installed versions, configuration and hooks, proxy/TLS scope, fixture or actual environment and immutable observation references.
Authentication boundary
Reference permitted identity, bearer or signing mode, credential source, expiry/refresh and denied-access evidence. Exclude secret values.
Attempt and timeout evidence
Record inner/outer retry owners, per-attempt observations, delay assumptions, timeout scopes and which facts remain uncertain.
Application policy and disposition
Record task/attempt identities, pending or held work, late-result rejection, safe retry prerequisites and owner of unresolved effects.
Budget and next action
State task deadline, time/cost/attempt bounds, observed versus assumed timing, violated gates and the smallest separately authorized follow-up test.

8. Change one boundary, then preserve the opposing results

The offline SDK probe and saved observations contains the pinned client probe and the inspectable exposure calculation. Run it without real credentials. A fresh environment needs the two declared package versions and their resolved dependencies; the probe refuses changed SDK or transport versions. It dispatches no provider request and uses no HTTP socket, so its success establishes local client behavior only.

An authorized integration review should start with denied access, exact path assembly and one bounded successful operation, then add the selected failure and deadline cases. Do not use production customer traffic merely because the code passed a mock. Model task acceptance, permitted data processing, current account prerequisites, quotas and provider lifecycle remain independent gates. Do not infer absence of remote execution from a timeout or turn a mock status into an AWS guarantee.

Start with one application's existing retry configuration. Count its inner and outer attempt allowances, write the task deadline and preserve one injected-failure trace. Ask a second engineer to challenge the attempt count, provenance and late-result disposition before any paid trial. An AI integration review can scope those specific gaps while keeping the currently accepted path unchanged. Keep mock-client evidence separate from any decision to invoke a provider.

Related services