Check the Tool Contract Before Moving a Request to Bedrock

Compare one completed Responses function proposal with a Bedrock Converse tool proposal, preserving argument meaning, native identities and independent execution...

A model request can succeed after an integration change while its tool proposal no longer means what the application expects. Before moving a tool-enabled request to Bedrock, compare the exact declaration, completion envelope, parsed arguments and application acceptance rule. Keep native call correlation separate from business-operation identity. A matching proposal is evidence for an adapter review, never permission to run the tool.

This article works one narrow comparison: OpenAI Responses with gpt-4.1-2025-04-14 and Bedrock Runtime Converse with the US inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0, originating in us-east-1. Both declarations explicitly request strict tool arguments. The fictional tool reads a fictional inventory page. There are no model invocations, AWS requests, customer records or executed tools in the example. All numbers and identifiers are stipulated test inputs.

The useful output is a compatibility record that an application engineer can inspect: what each representation carries, what the adapter preserves, what it deliberately holds and which evidence is still missing. Streams, tool-result submission, server-side tools, reasoning blocks, stored conversation transfer, retries, pricing and performance comparisons are excluded. This is not an endpoint-replacement recipe.

Download the offline tool-contract fixture to inspect the five-file Python package and run its 25 standalone test methods without an account or contact form. These are synthetic, non-authorizing parser and comparison checks, not model trials or an SDK certification. Read the included README before adapting it.

Two synthetic proposal branches preserve their own native IDs. Source arguments require a second duplicate-safe JSON parse. Target input needs original decoding evidence because a later object cannot recover lost duplicates. Compare only the permitted tool, alias and count; equal actions reach review only, never execution authority.

Solid arrows carry a supplied representation or an extracted action into local review, not a network request. Native IDs stay separate. A raw duplicate, unsupported block or missing completion evidence holds the specimen. A prior permissive target decode can erase that evidence and still yield a favorable local result; it remains an unresolved production capture gap. Exact declaration review is separate from action comparison. No provider call or tool execution occurred.

1. Pin the two contracts before comparing them

OpenAI's GPT-4.1 model page lists the selected snapshot, Responses and function-calling support. The function-calling guide documents the Responses function declaration. Write strict: true explicitly. The guide describes different omitted-strict behavior across Responses and Chat Completions; a migration record should not depend on a remembered default from the other API.

On the target side, the Haiku 4.5 card documents Runtime Converse, client-side tools and structured outputs. It also identifies the US profile and excludes structured outputs through this model's India profile. The selected Runtime path is not a Bedrock Responses path. Keep the model, endpoint family, API, profile, source Region and strict setting together in the record.

These documentation observations justify investigating the representation pair. They do not establish entitlement in an account, actual invocation success, permitted processing locations or continued availability. A US source endpoint does not establish single-Region inference. Resolve those questions through the separate Bedrock placement decision. The model card contains lifecycle wording that this article does not use to establish a retirement date or availability commitment.

Pin the application contract too. This specimen permits one proposed inventory_lookup action for fixture-catalog-a, requesting an integer from one through ten. No real catalogue is connected. Changing the allowed alias or accepted count changes application policy even if both providers continue returning valid JSON. Record the parser, projection and semantic-rule revisions alongside the request configuration.

2. Use a closed schema without inventing shared schema support

The specimen declares an object with two required properties: resourceAlias as a string and itemLimit as an integer. Additional properties are prohibited. OpenAI places that schema in parameters; Converse places it in toolSpec.inputSchema.json. The fixture retains the complete declarations in adapter.py rather than pretending those different wrappers are interchangeable request objects.

{
  "type": "object",
  "properties": {
    "resourceAlias": {"type": "string"},
    "itemLimit": {"type": "integer"}
  },
  "required": ["resourceAlias", "itemLimit"],
  "additionalProperties": false
}

This is the shared schema fragment, not a complete provider request. OpenAI's strict requirements include closed objects and required properties. AWS's structured-output reference describes a supported JSON Schema subset and excludes numeric minimum and maximum. Consequently the example keeps its one-through-ten rule in application validation. Removing an unsupported keyword from a declaration must not remove the business limit from the executor's controls.

The fixture checks equality with this exact declaration profile. It rejects an added enum as a changed profile, not as a claim that AWS rejects enums. It also rejects an added description or alternate required-field ordering rather than attempting semantic schema equivalence. That conservatism is useful for identifying a changed comparison input, but it is not a general schema validator or proof that another declaration is unsupported.

Review an application's actual keywords individually before adapting a larger schema. Optional values, unions, nesting, references and constrained strings can change the decision. If a source schema needs a feature the target cannot express, the options include keeping that check in application code, redesigning the tool contract under a new revision, choosing another evaluated route or retaining the existing integration. Silently dropping the requirement is not a compatibility fix.

3. Interpret the native completed proposal before normalizing it

The official Responses function-call type declares string arguments and a call_id, with a separate optional item id. The target ToolUseBlock carries a JSON-valued input, tool name and toolUseId. Those differences require explicit interpretation rather than a field rename followed by automatic execution.

Here is the fictional source projection. It is intentionally smaller than a complete Responses response. Request metadata, usage and other unrelated envelope fields are not represented. Its strict item allowlist is described in the companion README.

{
  "model": "gpt-4.1-2025-04-14",
  "status": "completed",
  "output": [{
    "type": "function_call",
    "name": "inventory_lookup",
    "arguments": "{\"resourceAlias\":\"fixture-catalog-a\",\"itemLimit\":3}",
    "call_id": "call-source-17",
    "id": "item-source-71",
    "status": "completed"
  }]
}

The corresponding fictional target projection uses a completed tool proposal, not a completed inventory operation:

{
  "stopReason": "tool_use",
  "output": {"message": {
    "role": "assistant",
    "content": [{"toolUse": {
      "toolUseId": "tool-target-29",
      "name": "inventory_lookup",
      "input": {"resourceAlias": "fixture-catalog-a", "itemLimit": 3}
    }}]
  }}
}

The Converse reference documents stopReason; tool_use identifies the tool-use ending used here. The fixture holds other endings except its reviewed text-only end_turn path. It does not extract a tool fragment from a token-limit or malformed response and pretend that the fragment is a completed supported proposal.

Projection itself is a review boundary. Keep every output or content item when deciding whether the response fits the specimen. Filtering out an unfamiliar block just to leave one familiar tool call can conceal a refusal, another proposal or unsupported content. A production projection needs its own tested field mapping from the selected raw response or SDK object. No such live projection is delivered or certified by this offline pack.

The source function-call status is optional in the official type. The specimen requires it to be present and completed. It can therefore hold a legal provider response. Likewise, both APIs can return more than one item, while this specimen holds multiple items even if each is individually well formed. An intentional narrow hold is different from claiming that a legal provider response is invalid.

4. Reject duplicate members before they disappear

Suppose the argument text contains itemLimit twice, once with three and once with nine. A parser that keeps only the last occurrence has already selected a meaning before application validation sees the object. Both values are within the allowed range. A range check on the surviving value cannot establish that the original representation was unambiguous.

The offline parser uses a duplicate-member hook recursively. It rejects duplicate keys in the outer projection and performs a second duplicate-safe parse of the source's argument string. It also rejects non-JSON numeric constants rather than treating them as business values. These are chosen specimen rules, not claims about either provider's decoder.

Converse input introduces a practical capture problem: an SDK may already expose a decoded object. Serializing that object again cannot recover duplicate-member history that an earlier parser discarded. The raw capture location, permitted retention and decoding behavior must be resolved before making a production ambiguity claim. Keep missing capture evidence in the record instead of describing the local parser as protection for a path that never uses it.

This parser has no production resource budget for hostile sizes, nesting or computation. Use only the small synthetic inputs supplied here. An operational implementation needs bounded input size, parsing limits, exception handling, safe diagnostic retention and review of the actual library behavior. A local test passing on a dozen short strings does not establish a production security boundary.

5. Work the same action and the cases that change the answer

In the fictional ordinary case, each representation proposes inventory_lookup with resourceAlias equal to fixture-catalog-a and itemLimit equal to three. The source adapter parses the argument string. The target adapter checks the already represented object. Both preserve the exact tool and values. Their native call identifiers remain different and namespaced by provider.

The comparison returns EQUIVALENT_PROPOSAL_REVIEW_ONLY. This means only that the supplied projections produce the same allowed action under the local rule. It does not assert that either model produced them, that a user requested them or that the catalogue exists. The specimen always returns false for execution authorization and provider validation; individual results also explicitly return false for tool execution.

Now change only the target count to four. Each proposal remains individually inside the application limit, but they no longer represent the same action. The comparison holds. That distinction prevents a shape-only test from declaring equivalence while the intended query changed. It also explains why different native IDs alone do not make otherwise equal actions incompatible.

Change the target alias to another catalogue, or make the count eleven, and the individual proposal holds on application scope. A strict argument shape does not establish a permitted resource or a suitable amount of work. Change three to the string "3", boolean true or decimal lexical form 3.0, and the specimen holds instead of coercing it. Its Python integer rule deliberately excludes integral floats, which is narrower than some general JSON Schema validators. Preserve that limitation in any interpretation of the result.

A legitimate text-only answer, such as a request for clarification, becomes NO_TOOL, not an invented zero-argument action. An explicit source refusal becomes REFUSED. Even matching NO_TOOL results do not produce action equivalence in this checker. If the intended task permits clarification, a separate task-level rubric should assess that behavior. Do not force a tool merely to improve the number of matching actions.

6. Retain a filled compatibility record

The following record is fictional and corresponds to the ordinary supplied projections. Named roles are example responsibilities, not staffed customer reviewers. Application acceptance, real capture and deployment remain unverified.

Record and owner
B021-C17, revision 1; fictional application integration owner.
Source configuration
OpenAI Responses; gpt-4.1-2025-04-14; strict true; single completed function proposal.
Target configuration
Bedrock Runtime Converse; us.anthropic.claude-haiku-4-5-20251001-v1:0; source us-east-1; strict true.
Declaration and policy
Exact two-property closed schema in adapter.py; inventory_lookup; fixture-catalog-a; integer one through ten.
Representation inputs
Fictional raw JSON projections; source arguments additionally parsed as JSON. No provider capture.
Native correlation
Source call-source-17 and item-source-71; target tool-target-29. Provider namespace retained, no business-operation identity inferred.
Expected local outcome
Equal permitted action with count three; EQUIVALENT_PROPOSAL_REVIEW_ONLY. No authorization, execution or provider validation.
Opposing case
Target count four is individually permitted but not equivalent. Duplicate itemLimit, missing completion evidence and out-of-scope aliases hold.
Capture and eligibility gaps
Actual SDK projection and raw parsing behavior, account access, placement permission and real task behavior remain UNKNOWN.
Next action and invalidation
The integration owner requests independent specimen review before separately authorized no-write integration tests. API, model, profile, schema, parser or policy changes reopen the affected evidence.

Keep the filled record beside the exact artifacts and result, not beside an unrelated successful endpoint screenshot. A reviewer should be able to reproduce the changed-count hold and distinguish it from a malformed JSON failure. Preserve unfavorable cases when changing an adapter. Replacing the evidence with only the newest successful sample hides what the previous revision failed to preserve.

7. Run the offline specimen without confusing it with a model evaluation

The companion offline fixture uses Python 3.10 or later and the standard library. From its directory, run python3 -m unittest -v test_adapter.py test_expectations.py. The ordinary source and target packets are fictional projections; the tests construct them without contacting any provider. The predeclared expectation file supplies additional raw malformed and opposing inputs with named expected dispositions.

Read the tests before adapting the checker. They cover changed declarations and routes, wrong tool names, incomplete states, absent call identities, unexpected properties, numeric coercion attempts, duplicate members, multiple proposals, text-only answers, refusal and unequal allowed actions. Test-method totals are not a count of model trials or representative customer tasks. No benchmark percentage is meaningful here.

The artifact does not implement replay protection, authorization, catalogue lookup, result submission or the complete provider API. Native ID checks require nonempty trim-stable strings, not every provider pattern and length constraint. It does not authenticate a supplied route manifest. A fabricated but structurally matching response can compare favorably because the exercise evaluates supplied representations. This is a declared trust boundary, not proof of authentic origin. Real invocation and request binding require separate evidence outside the model's text.

Use failure reasons to assign work. A declaration-contract hold belongs to the schema/configuration owner. A JSON ambiguity belongs to the capture/parser investigation. Application-scope holds belong to the task and policy owners. Unknown native blocks require a documented adapter decision, not an unconditional ignore rule. An owner can narrow supported behavior or add independently reviewed handling, but the old favorable cases and adverse cases must remain inspectable.

8. Choose the next integration step, not an automatic migration

If an actual task needs multiple tool proposals or mixed text and tools, this one-item specimen is insufficient. Extend the contract with explicit ordering, completion, correlation and allowed combinations before extending code. If the only useful function is deterministic inventory pagination, a direct application capability may be simpler than model-directed tool choice. If the required schema cannot transfer without losing a rule, retain the existing route or redesign the tool under a separately accepted contract.

Keep the executor independent whichever option is chosen. A read-only lookup still needs current caller and resource permission because the returned data can be sensitive. A native tool ID identifies provider interaction; it does not establish user authority, an approved task, an idempotency key or a successful effect. AWS client-side tooling leaves tool implementation in application code. This specimen deliberately stops before that implementation.

For a real candidate, the integration owner should first complete the blank record below using sanitized evidence references. Then obtain independent review of the projection and application rules. Only a separately authorized no-write integration evaluation can examine actual native outputs for representative permitted tasks. Its scope, data handling, budget, observer and stop conditions belong in that separate approval. Passing this local pack grants none of them.

Record, revision and accountable owner
Enter the task family, comparison revision and integration owner.
Exact source configuration
Record API, model snapshot, explicit strict setting and dated supporting sources.
Exact target configuration
Record endpoint family, API, model/profile, source Region and model-specific feature evidence.
Declaration and independent task rule
Reference schema/adapter revisions, allowed tool, resource scope, types and limits. Mark unsupported requirements.
Capture and projection
Identify raw versus decoded inputs, duplicate-member boundary, preserved blocks and deliberate exclusions.
Correlation and operation identity
Record native call/item identities separately from trusted application task, attempt and effect identities.
Predeclared expectations
Specify equal-action, different-action, no-tool, refusal, incomplete, malformed and unauthorized-resource cases.
Actual bounded evidence
Record exact artifact hashes and observed local results, or NOT EXECUTED. Never infer a provider trial.
Remaining authority and eligibility
Reference current caller/resource checks, data-placement/access decisions and unverified deployment prerequisites.
Owned next action
Name the reviewer, missing evidence, permitted next test, stop conditions and changes that invalidate this record.

Use the existing fallback output-contract article for the broader task-equivalence decision and the guardrail playbook for enforcement at protected destinations. The concrete next action here is smaller: have the integration owner reproduce one equal proposal and one independently specified adverse case, retain their exact projections and ask a separate reviewer whether the mapping preserves the declared action contract.

Related services