Bedrock Inference Placement: Choose a Route Your Workload Can Admit
Compare in-Region, geographic and global Bedrock inference against an explicit processing boundary. Separate route availability, authorization, retained copies and...
Abstract: choose the permitted route before optimizing it
An application calling a model from one AWS Region does not thereby establish that inference occurs only there. An inference profile can select compute elsewhere while familiar logs remain near the originating call. A permitted endpoint, a working policy and a successful answer are different acceptance facts. The engineering decision is whether the exact route satisfies the workload's approved processing boundary, and what remains unknown before exposure begins.
This paper recommends a source-specific placement contract. Compare route availability, possible processing locations, authorization context, retained copies and task-capacity evidence separately. Admit a candidate only within its supported and approved scope. Reject an incompatible route even if it is cheaper, and hold a route whose evidence is incomplete even if it is reachable. A deterministic or delayed task can be preferable to quietly widening the permitted boundary.
The reader is an engineering or data owner selecting a Bedrock route for a bounded task. Examples and diagrams are educational references, not a customer design or executed AWS test. Documentation was checked October 7, 2026. No inference request, profile read against an account, IAM/SCP change, logging enablement or spend occurred. This is not legal advice, a compliance certificate, guaranteed regional availability or a model recommendation. Applying this framework requires resource-specific review by the responsible workload owners.
1. State the task constraint without inventing a country rule
Write the useful function before the model choice: explain one approved document, classify a fabricated ticket or draft a source-linked internal summary without external actions. Name its users, permissible inputs, required answer evidence, deadline and owner. A route decision cannot establish retrieval authority or answer correctness. Those prerequisites belong to the existing enterprise RAG evaluation framework.
The placement contract must state what the organization actually permits. A fictional owner may require inference within a named set of Regions. Another may accept an expanding commercial-Region boundary for public, non-sensitive input. Neither choice follows automatically from the customer's office address. India and US workloads need the same disciplined record, with different constraints only when their approved purpose, agreement or data decision requires them. Do not manufacture an India-only legal mandate or treat an APAC label as India-only processing.
Separate processing permission from storage and access permission. A rule permitting execution in selected locations does not automatically permit prompt retention, operational exports or unrestricted support access. Conversely, an approved source-Region log store does not authorize worldwide inference. The data owner records these decisions distinctly, including unresolved obligations and who may accept a bounded exception. Engineering supplies the factual route inventory, not a jurisdictional verdict.
2. Record the exact route, not just the model family
The decision unit includes model/version, endpoint, API, source Region, profile identity where used, requester and task configuration. A family name alone hides differences in available routing paths and supported features. Record the actual configuration reference that would select the route, including indirect selection through application configuration. A diagram labeled Bedrock cannot identify those choices.
For a dated discovery example, AWS's Claude Sonnet 4.5 detail page identifies anthropic.claude-sonnet-4-5-20250929-v1:0, runtime Converse/Invoke support and the associated geographic/global profiles. Its October 7 snapshot marks in-Region unavailable, and Mumbai/Hyderabad as global rather than geographic source options. That is a documentation observation, not an account entitlement, present execution result or India-only path. Recheck the exact card and account prerequisites before adopting any candidate.
Do not assume that every desired model has all three alternatives. The correct output can be a different evaluated model, a different task deadline or a held function. An unavailable in-Region option is not made available by assigning it a favorable comparison score. A supported profile also does not establish that its output meets the existing task acceptance criteria. Keep model suitability and route admissibility as independent columns.
3. Discover destinations from every planned source
AWS's profile support guidance directs readers to check the selected profile from each intended source Region. Destination sets can differ by source. Geographic profiles have defined destinations; global routing can expand as commercial Regions are added. The account's opt-in list is not a complete proof of excluded destinations. A copied profile prefix or another team's manifest is insufficient evidence.
Use the GetInferenceProfile schema to define a readback record: profile ID/ARN, type, status, models[].modelArn, update time and the source endpoint used for collection. Where returned model ARNs specify Regions, retain those Regions alongside the model card. Where a global resource is Region-agnostic, do not invent a closed destination list from an empty ARN field. Its documented routing boundary needs separate evidence.
These are proposed collection fields, not an executed API response. No customer identifier or fabricated successful JSON appears here. A future authorized reviewer should retain the actual permitted response and source references, then reconcile contradictions. A profile's ACTIVE status is configuration readiness, not proof that the owner permits its destinations or that useful workload capacity exists. Missing read access leaves a gap; it does not justify asking for general administrator privileges.
4. Treat global expansion as a different acceptance choice
A geographic route with a documented bounded set and a global route with expansion potential require different owner decisions. A finite owner allowlist cannot be proven compatible with an unknown future expansion merely because today's selected destination is allowed. Do not freeze the global contract to one observed request or to a copied table. An organization can explicitly accept the documented expanding boundary, but that acceptance must name its data scope and review conditions.
The global routing guide separates the originating profile and source-model evaluations from a Region-agnostic global foundation-model evaluation. For that evaluation aws:RequestedRegion is unspecified, not an asserted physical destination. A Region-name condition is therefore not automatically a destination fence for this route. This paper does not supply a policy that purports to force global inference into a chosen country.
If that distinction cannot be resolved within the approved boundary, reject global for this task or hold the decision pending specialist review. Do not broaden an SCP until a request works and retrospectively call the result an approved placement design. Configuration permission and business permission remain different. Global may be appropriate for an explicitly accepted scope, but mere API success cannot supply that acceptance.
5. Review authorization at the resource evaluation it actually uses
For geographic inference, AWS's route-specific guidance describes profile, source-model and destination-model authorization. bedrock:InferenceProfileArn is available in foundation-model evaluations, not the originating profile-resource evaluation. A scoped exception is not equivalent to adding every destination to an organization's unrestricted Region allowlist, and does not bypass the originating-call restriction. Review the exact request and policy context rather than copying an example across route classes.
The Bedrock service authorization reference documents actions, resource types and bedrock:InferenceProfileArn. Map the selected API to its authorization actions, distinguishing readback from invocation and streaming. The security owner includes applicable identity policies, boundaries and organization restrictions. No permission review or simulator output is promoted here into observed inference behavior.
There is a source-quality issue worth keeping visible. The prerequisites page names aws:InferenceProfileArn in prose while its example uses the Bedrock-prefixed key. The current service-authorization and geographic/global references support the Bedrock-prefixed name. Resolve the exact evaluated context through those references and authorized validation; do not copy the inconsistent phrase into a deployable policy. This paper provides a handoff, not an IAM/SCP configuration or mutation instruction.
6. Do not confuse the logging Region with inference placement
The global guide describes source-Region monitoring even when compute is selected elsewhere. Treat a source-region log as an observation of the call's originating context unless its particular fields explicitly support another conclusion. It is not a universal receipt of physical processing location. A retained log and a documented possible-destination boundary answer different questions.
The invocation-logging reference describes same-account/Region destinations and supported runtime calls, with separate coverage for other endpoints. Logging can retain full input/output material, not merely counters. Do not enable it casually to fill the placement worksheet, and do not assume deleting its configuration disposes of every prior copy. An approved evidence plan can use minimal configuration and request references where sufficient.
Inventory applicable prompt/result retention, abuse-review handling and diagnostic copies separately. Geographic guidance explicitly discusses destination storage for applicable abuse detection. Resolve the actual model/path's handling before asserting a storage boundary. The existing observability data-governance paper owns collection minimization, exports and disposal; this paper does not repeat its retention policy. If permitted evidence cannot support the desired statement, shorten the statement instead of collecting private payloads everywhere.
Make selected features explicit rather than treating optional invocation logs as the only retained representation. The selected Sonnet 4.5 card supports explicit prompt caching. Record caching enabled, disabled or UNKNOWN, the exact request/configuration reference and feature-specific handling evidence. A documented cache lifetime is not a complete location or disposal guarantee. Exclude caching from the proposed candidate until that handling is resolved, or hold the candidate if its required behavior depends on caching. Other selected features need their own applicable evidence.
7. Use the placement map as a question register
The first figure separates three locations: the originating application/runtime call, possible inference compute and configured invocation-data storage. It is a deliberately hypothetical route view. Region labels are symbolic, not a claim that a named profile uses the illustrated destinations. The dashed retained-copy branch is separate from the processing path. Global expansion is an accepted-boundary question, not another static box pretending to list the world.
Blue arrows represent the proposed inference request. Dashed gray is an optional configured retained-copy relationship. Brown marks unresolved or expanding scope. No route, log or invocation was executed.
For each illustrated boundary, attach a worksheet reference. What selected configuration permits that relationship? What source-specific documentation or readback supports the possible destination set? What owner permits retained data? Where is the evidence incomplete? This is more useful than adding a country flag to the application frame. A nearby source endpoint is a starting condition, not a substitute for the destination contract.
8. Compare alternatives without offsetting a hard constraint
An in-Region candidate is useful when available for the exact model/API and compatible with the task. It can simplify the processing-location question, but does not eliminate retention, authorization, quota, quality or regional-dependency review. A different supported model can make that option possible, with its own evaluation cost. Do not describe the different model as semantically interchangeable without task evidence.
A geographic profile can expand available capacity within its documented destination boundary. Its operating burden includes source-specific inventory, applicable policy review and evidence that every candidate destination fits the contract. A global profile can be considered only when the owner accepts its broader and potentially changing boundary. A fixed-country restriction is not satisfied by selecting the source endpoint in that country. These alternatives are conditional, not ranked recommendations.
The fourth option is to hold generation or preserve a useful deterministic workflow. It avoids unapproved routing while reducing capability or increasing delay. Record the customer consequence, task validity window and owner; do not present held work as completed. The existing provider continuity paper owns task fallback and exit policy. This route decision can feed that policy without duplicating its effect-reconciliation framework.
Compare these options with four separate fields: admissible scope, prerequisites still missing, useful task capacity/cost assumptions and operating burden. A prohibited destination cannot be compensated for by a lower price. An allowed route with unknown capacity is not ready merely because permission is settled. A candidate may pass placement review while remaining held for task evaluation or account authorization.
9. Build one reusable admission worksheet
Complete one record for each model/API/source/profile combination. Blank fields remain unknown. Broadly shared copies should use sanitized references, while authoritative identifiers belong in the permitted repository. Never place credentials or complete sensitive prompts into this artifact.
Decision ID / revision / task / accountable owner / relying reader:
Input class / permitted purpose / prohibited consequence:
Owner-approved processing boundary / fixed set or accepted expansion:
Separate retained-copy / abuse-review / diagnostic boundary:
Candidate model/version / exact endpoint / API / configuration reference:
Source Region / account / requester / profile ID and ARN:
Profile type/status/update time / documentation check date:
Returned models[].modelArn / source-specific destination evidence:
Global Region-agnostic resource / unresolved scope or contradiction:
Profile/source/destination/global authorization evaluation references:
Relevant actions, policies, conditions / security reviewer / gaps:
Task evaluation reference / output checks / excluded features:
Prompt caching: enabled / disabled / UNKNOWN / explicitly excluded:
Selected feature configuration / handling evidence / unresolved retained scope:
Route quota / workload mix / retry budget / deadline / budget owner:
Permitted evidence collector / storage / access / independent stop owner:
Expected fixture / actual result: NOT EXECUTED unless observed:
Candidate outcome: rejected / held / placement-compatible only:
Remaining exposure prerequisites / acceptance owner / review trigger:
Exception consequence, expiry and closure evidence / next action:Keep placement compatibility distinct from production readiness. The first says the documented route boundary fits the declared constraint under recorded assumptions. The second needs actual authorized implementation, task evaluation, capacity, observation and release evidence. An owner may commission a synthetic validation from a compatible record without authorizing customer traffic or a wider data class. Record exactly which next action the decision permits.
10. Work a fictional comparison without fictional execution
Suppose task T explains fabricated documents and its owner permits processing only in symbolic Regions A and B. Candidate G has documented, source-specific destinations A/B under the stipulated example; candidate X has A/B/C. These labels are invented, not actual profile output. Set inclusion makes G placement-compatible on those assumptions and X incompatible because C is outside the allowed set. Neither becomes production-ready: model/API availability, authorization and useful task evidence are still pending.
A second fictional task permits an explicitly accepted expanding commercial-Region boundary for public content. A global candidate may be considered, but unresolved model/path documentation, retention or task evidence still holds it. Conversely, task T with its fixed A/B rule rejects that expansion rather than freezing a present observation as a future promise. The mathematical check is simply whether the possible boundary fits the approved boundary, not a workload benchmark or compliance algorithm.
Illustrative record T / proposed revision 1
Approved processing set: {A, B}; fixed, owner-stipulated
Candidate G destinations: {A, B}; fictional source-specific premise
Placement comparison: compatible under premise
Account/profile readback: NOT EXECUTED
Authorization, task quality and capacity results: UNKNOWN
Overall exposure decision: HELD
Next action: obtain dated exact-route evidence in authorized scope
Non-claims: no real model/profile availability or customer readinessA checker that accepts G as deployed-ready has failed the example. A checker that changes T's boundary to include C just to pass X has changed the business decision, not repaired evidence. A checker that sees no logged C request and declares C impossible has confused observed sampling with the routing contract. These counterexamples make the worksheet useful even before a live environment exists.
11. Admit, reject or hold the candidate explicitly
The second figure asks which disposition follows from the evidence, not where packets go. It starts with exact-route evidence, then distinguishes incompatibility from unknown scope. Only a supported boundary matching the owner contract reaches the remaining exposure prerequisites. Missing evidence at either stage holds the affected decision. A compatible location contract is an intermediate conclusion, not automatic release approval.
Proposed decision relationships, not AWS API order. Brown holds missing evidence; red rejects an incompatible candidate. Blue advances supported evidence. No branch changes the permitted boundary or authorizes a policy mutation.
Attach the reason to every disposition. Reject means this candidate cannot satisfy the current contract; it does not mean Bedrock is universally unsuitable. Hold means the needed evidence or authority is incomplete; it does not assert an incident. Compatible means only that this part of the review supports its narrow statement. This vocabulary prevents a green intermediate row from becoming an unsupported public deployment claim.
12. Evaluate safely, then budget useful work
Begin with offline manifest checks and fabricated constraints. Verify that the checker rejects an extra destination, unknown global scope, changed source, unsupported endpoint and source-log-as-location substitution. The expected answer must be independent of the candidate manifest. These educational expectations are NOT EXECUTED. They require no live prompt, customer payload or policy broadening. A local test would establish only that checker's logic, not AWS's account behavior.
For a separately authorized account validation, define exact read actions, synthetic invocation scope if permitted, data class, owner, spend limit, observer and stop path before execution. Stop on changed requester, unexplained profile selection, inconsistent destination evidence, unapproved retained content or incomplete observations. A successful synthetic answer can support only its actual request and task contract. It does not prove that future global routing remains inside a fixed set.
The quota reference supplies model/route-specific discovery, not a throughput guarantee. Record actual account quota and token accounting separately from measured useful task completion. Long inputs, retries, evaluation calls and concurrency can change consumption. Compare costs only after rejecting inadmissible options, using dated rates and the same accepted workload. This paper provides no dollar figure, claimed discount, latency ranking or universally sufficient quota.
13. Maintain the route and its exceptions
Separate task classes before making one application-wide decision
Consider a fictional support application with two queues. One produces explanations of a public product manual. The other summarizes internal incident records. Both use the same interface, but the owners have approved different input handling. The public-manual queue permits an explicitly accepted expanding processing boundary. The incident queue permits only symbolic Regions A and B and requires a separately reviewed retained-copy boundary. These are stipulated organizational choices, not legal conclusions about either data class or any real AWS destination.
An application-wide switch to a global profile would erase that distinction. A profile selected for one queue is not evidence that the other queue is admitted. Record two candidates with separate configuration references, input selectors, owners and release decisions. If the software cannot reliably distinguish the queues before constructing the request, neither candidate is ready for mixed traffic. The missing control is the routing selector and its evidence, not another destination label on the diagram.
The admission test must examine the input assembled for the model, not only the document's original classification. A public question can acquire private context through retrieval, account metadata, conversation history or a tool result. That transformed request belongs to the more restrictive review until its owner supports another conclusion. A user's choice of a public queue cannot override the actual payload boundary. This paper does not design retrieval controls; it requires their accepted output contract to reach the route selector intact.
For this fictional application, compare three implementation choices. Separate workers with separate configuration can make route selection easier to inspect, but require two maintained release and observation paths. A shared worker with an explicit selector can reduce duplication, but adds a critical classification and configuration dependency. One restricted route for both queues simplifies selection when such a route is actually available and task-compatible, while potentially changing model choice, capacity or useful behavior. None is universally superior. Choose the smallest arrangement that the team can verify under the approved inputs.
Write the decisive failure fixture before choosing that arrangement. Give the selector a public question carrying a synthetic restricted attachment and expect it to hold or choose only the separately admitted restricted route. Give it missing classification and expect a held request, not a convenient default to the broadest profile. Give it a stale route revision and expect the admission boundary to reject the mismatch. These are proposed offline checks using fabricated material. They have not been executed and do not establish AWS enforcement.
Make configuration changes invalidate the right evidence
Treat the candidate manifest as a versioned input to the release, not an informational document beside it. The application configuration should identify which reviewed model, API, source and profile it intends to use. The release record should point back to that manifest revision. This is a proposed application control, not an AWS feature or a supplied implementation. Without the linkage, a complete worksheet can coexist with an unreviewed environment-variable change.
Classify changes by the conclusion they invalidate. A new source Region reopens source-specific destination and authorization evidence. Enabling caching reopens the selected-feature handling record. Expanding the input class reopens the owner's permission and task evaluation. Replacing the model reopens support, behavior and potentially lifecycle evidence. A documentation update that affects one field need not erase unrelated observed facts, but it must not inherit the old conclusion automatically. Keep the earlier revision as history and identify the exact rows requiring a new decision.
For example, suppose the fictional public queue passes its scoped synthetic evaluation under revision 3, while the incident queue remains held. A release that changes only the incident queue's profile cannot use the public queue's result as acceptance. Conversely, withdrawing the incident candidate need not disable a still-admitted public function if isolation and configuration evidence support that separation. The application owner must define which functions remain useful and which status users see. A generic service-healthy signal cannot communicate that distinction.
Use the same discipline for emergency substitution. If the admitted route becomes unavailable, a fallback that changes processing scope is a new candidate, not merely a retry. Hold the affected function or use an already admitted alternative under its own task and input contract. An operator responding to an outage should not have to invent the data-owner decision under pressure. Retain an explicit list of permissible alternatives and their current evidence, including the option that no generation can proceed for that class.
Measure useful completion without hiding withheld work
After placement compatibility, the capacity review needs a declared task population. In a fictional observation window, suppose 100 synthetic tasks enter the selector. Twenty are held because the proposed route lacks evidence for their input class. Eighty are admitted to the separately authorized synthetic evaluation, of which 72 meet the predeclared output and deadline checks. The conditional completion ratio is 72/80, or 90 percent. The end-to-end useful completion ratio is 72/100, or 72 percent. Both denominators matter. These numbers are arithmetic examples, not a benchmark or a passing result.
Reporting only the first ratio hides the capability withheld by the placement decision. Reporting only the second can wrongly blame model quality for deliberately held traffic. Preserve three populations: held before invocation, attempted under the admitted candidate, and accepted useful outputs. Keep task-quality failures, timeouts and unknown outcomes separately inspectable within the attempted population. A request that returns text is not automatically accepted useful work, and an unresolved timeout must not become a new business task simply to improve the success count.
Compare financial assumptions on the same population and horizon. Include applicable attempts, retries, selected caching behavior, evaluation consumption and other explicitly scoped service costs using dated rates and actual usage evidence. Divide by accepted useful tasks only when explaining what that ratio includes and excludes. If no task passes, the unit-cost ratio is undefined, not zero. A cheaper admissible route can still be a worse operating choice if it fails the task deadline or requires a costly manual queue. No numeric AWS rate or forecast is supplied here.
The business owner then makes a capability decision with engineering: accept the limited public function, fund evidence for the restricted queue, choose a different evaluated model, or retain a deterministic process. The worksheet must record the customer consequence of that choice. The aim is not to make the largest routing boundary win a comparison. It is to make the permitted useful service and the work still withheld understandable to the person funding and operating it.
Review checklist for a bounded route release
- Bind the task and assembled input class to the owner's processing and retained-copy permissions. Missing classification holds the affected request.
- Identify the exact model, API, source, profile and configuration revision. Preserve source-specific destination evidence and unresolved contradictions.
- Distinguish placement compatibility, authorization, selected-feature handling and observed task capacity. No single green row clears the others.
- Specify adverse selector and manifest-change fixtures before evaluation. Record their expected and actual results separately, with an independent observer and stop owner.
- Report held, attempted and accepted task populations without dropping failures or unknown outcomes. Cost comparisons use the same admitted scope and declared exclusions.
- Record permitted alternatives, user-visible held behavior, expiry and invalidating changes. A fallback with wider processing scope needs its own admission decision.
Review when model/version, profile, source endpoint, API, requester, policy, input class or owner boundary changes. Global expansion needs its declared review treatment; a periodic meeting cannot automatically freeze provider behavior between reviews. A geography-specific profile replacement is a new configuration identity even if its marketing label is unchanged. Keep prior manifests as history, not apparent evidence for the new route.
Current source summaries are not fully aligned. The profile overview's global-model/source example is narrower than the dedicated global guide and selected model card. Preserve that discrepancy and prefer exact model/path references plus permitted readback for a scoped decision. If they cannot be reconciled, mark the affected availability conclusion unknown. Do not silently substitute a newer model, or turn an old EOL lower-bound date into proof of retirement. The lifecycle operating model owns that separate clock.
An exception should identify the changed constraint, approving authority, permitted input scope, expiry, monitoring limits and closure evidence. It is not a technical workaround hidden in a role policy. On withdrawal, stop new admissions to the candidate through the approved control, preserve in-flight and retained-copy obligations, and verify the relevant resulting state. Reverting the model ID does not delete old logs, retract sent data or establish that every previously issued session lost permission.
14. Conclusion and limitations
The engineering deliverable is a candidate-specific placement record with supported facts, unknowns and an owned next action. It cannot certify legal compliance, every provider-controlled copy, universal authorization correctness or future model availability. A documentation inventory is not an AWS execution test; an execution sample is not an exhaustive destination guarantee. Maintain those differences when explaining the decision to procurement, security and application teams.
Choose one task and one exact source/model/API/profile combination. Write the owner's processing and retained-copy boundaries, reconcile current route evidence, compare admissible alternatives and commission only the missing bounded validation. If the desired model cannot fit the constraint, make the capability tradeoff visible. Ampity's AI evaluation work or cloud security scope can support that review; using this worksheet does not require an enquiry. The goal is a route the workload can legitimately use, not the broadest routing option the account can technically invoke.
Related services
AI Observability, LLM Monitoring & Governance
LLMOps consulting for AI observability, LLM monitoring, evaluation and guardrails. Review production answer quality, operating failures and cost evidence.
Cloud Security Architecture Consulting
Cloud security architecture consulting for AWS and GCP. Review IAM, network boundaries and compliance requirements, then implement agreed security controls.