# Bedrock text-stream trace-review contract Offline, synthetic. These are decoded-event PROJECTIONS, not complete API payloads, framing fixtures, recorded traffic or runtime proof. Omitted fields are not assertions that those fields are optional. `local`, `attempt` and `drain` are application/test labels, not vendor fields. Never send this pack to a provider. ## Pinned comparison and release policy Source: synchronous Responses/SSE, `stream:true`, `background:false`, `gpt-4.1-2025-04-14`. Target: Runtime ConverseStream, `https://bedrock-runtime.us-east-1.amazonaws.com`, `us.anthropic.claude-haiku-4-5-20251001-v1:0`, source Region `us-east-1`. Actual client/SDK, account access and runtime outcomes are UNKNOWN/NOT EXECUTED. One assistant text answer, one text block for this specimen, no tools/reasoning/stop sequences. Complete-answer release. Release only for the active, locally eligible attempt after a normal native ending, clean drain and semantic acceptance. Any exceptional event before clean drain holds release. Metadata is accounting evidence, not content acceptance. Native unknown/unexpected forms hold for review. Retrying creates a fresh attempt, never a concatenated continuation. Supplied task: answer from “The public desk opens Wednesday from 09:00 to 17:00 UTC. It is closed Thursday.” Accept only preservation of both days, Wednesday hours and UTC without invented additional hours. Do not require one exact string. No tools are authorized by this task. ## S1: normal source projection All entries belong to local attempt `s1`. Sequence numbers are supplied ordered observations, not a claim about a complete stream's numbering. ```json [ {"type":"response.created","sequence_number":0,"response":{"id":"resp_s1","status":"in_progress"}}, {"type":"response.output_text.delta","sequence_number":4,"item_id":"msg_s1","output_index":0,"content_index":0,"delta":"The public desk opens Wednesday from 09:00 to 17:00 UTC. "}, {"type":"response.output_text.delta","sequence_number":5,"item_id":"msg_s1","output_index":0,"content_index":0,"delta":"It is closed Thursday."}, {"type":"response.output_text.done","sequence_number":6,"item_id":"msg_s1","output_index":0,"content_index":0,"text":"The public desk opens Wednesday from 09:00 to 17:00 UTC. It is closed Thursday."}, {"type":"response.completed","sequence_number":9,"response":{"id":"resp_s1","status":"completed","output":[{"type":"message","id":"msg_s1","role":"assistant","status":"completed","content":[{"type":"output_text","text":"The public desk opens Wednesday from 09:00 to 17:00 UTC. It is closed Thursday."}]}]}}, {"local":"drain","outcome":"clean"} ] ``` Expected: before the final three release conditions are satisfied, fixed progress only. After clean drain and semantic acceptance, one answer and one matching history entry. No duplicated text from appending the final response to deltas. These are partial envelope projections and intentionally omit other valid events/fields; no wire parser may treat this as a complete source transcript. ## S2: source text closure followed by failure Use S1's created event and first delta. Then supply: ```json [ {"type":"response.output_text.done","sequence_number":6,"item_id":"msg_s1","output_index":0,"content_index":0,"text":"The public desk opens Wednesday from 09:00 to 17:00 UTC. "}, {"type":"response.failed","sequence_number":9,"response":{"id":"resp_s1","status":"failed","error":{"code":"server_error","message":"Synthetic failure"}}}, {"local":"drain","outcome":"clean"} ] ``` Expected: failed attempt; no answer/history/export/copy candidate. A clean transport end does not reverse native failure. Use a separate equivalent source case for `response.incomplete` with `incomplete_details.reason:max_output_tokens`, and a separate stream `error` case when testing the deployed source adapter. Their shapes must be verified against that client's decoder before being described as framing fixtures. ## T1: normal target projection All entries belong to attempt `t1`. There is no text `contentBlockStart`. Metadata numbers are stipulated synthetic accounting inputs, not model measurements or comparable source token units. ```json [ {"messageStart":{"role":"assistant"}}, {"contentBlockDelta":{"contentBlockIndex":0,"delta":{"text":"The public desk opens Wednesday from 09:00 to 17:00 UTC. "}}}, {"contentBlockDelta":{"contentBlockIndex":0,"delta":{"text":"It is closed Thursday."}}}, {"contentBlockStop":{"contentBlockIndex":0}}, {"messageStop":{"stopReason":"end_turn"}}, {"metadata":{"usage":{"inputTokens":30,"outputTokens":20,"totalTokens":50},"metrics":{"latencyMs":900}}}, {"local":"drain","outcome":"clean"} ] ``` Expected: same accepted semantic answer as S1, once, after the release gate. This supplies no performance or usage comparison. At block stop alone, candidate remains private. `messageStop` alone does not establish the semantic result or clean drain. ## Negative target projections and predeclared outcomes Each case starts a new attempt. A substitution means replace the named suffix of T1, never append a contradictory second normal ending. Local labels remain outside native event objects in a real harness. | ID | Supplied trace construction | Native/local classification | Expected UI/history/export/copy result | | --- | --- | --- | --- | | T2 | `messageStart`; first text delta; block stop; `{"modelStreamErrorException":{"message":"Synthetic stream failure"}}`; clean drain | Exceptional, partial candidate | Interrupted status; no candidate in accepted destinations. | | T3 | T1 through block stop, with only the first sentence; replace ending with `{"messageStop":{"stopReason":"max_tokens"}}`, then metadata and clean drain | Limited, not accepted in this profile | Limit status; no answer/history candidate, despite fluent partial. | | T4 | T1 through first delta, then local EOF without `messageStop` | Native outcome unknown, transport ended | Interrupted/unknown; no successful answer. | | T5 | T1 event structure, but one delta containing `The desk opens Thursday.`; `end_turn`, metadata, clean drain | Normal native ending, rejected meaning | Rejected-content status; no accepted history entry. | | T6 | `messageStart`; `{"contentBlockStart":{"contentBlockIndex":0,"start":{"toolUse":{"toolUseId":"tool_synthetic","name":"unexpected_lookup"}}}}` | Profile mismatch | Hold; no tool dispatch and no stringified tool data in answer. | | T7 | T1 first delta; local Stop for this attempt; remaining T1 events including `end_turn` and clean drain | Locally stopped, native normal ending later | Stopped status; late events do not create answer/history. | | T8 | Old attempt first delta; local Stop old; start new attempt; old second delta; full new T1 trace; old late terminal | Old ineligible, new independently eligible | Only new attempt's complete answer accepted once; no old text appended. | T8 requires the harness to attach attempts at reader creation. If it merely rewrites every incoming callback with the current attempt, it defeats the test. Save old and new reader callbacks separately, deliver interleaved events and inspect the new buffer as well as final destinations. Additional boundary variants: unexpected second text block, non-text delta, duplicate/mismatched source final text, an unknown future stop reason, exception after normal ending but before drain, and a Stop ordered after an already accepted answer. Define expected results before implementation. This specimen holds the first five for review; Stop after acceptance does not undo the already observed answer. Metadata absence is an accounting anomaly, never invented zero usage. Whether that anomaly separately blocks release must be recorded in the deployed policy. ## Filled review record: T7 | Field | Supplied decision | | --- | --- | | Comparison revision | Text-stream comparison 2026-10-09; exact pair above | | Reader/parser revision | NOT PROVIDED; no integration executed | | Attempt correlation | Local `t7`; native target has no Responses response ID | | Native observation | `end_turn` after local Stop; clean drain supplied | | Candidate | Complete supplied answer after late deltas; still private | | Local ordering | Stop recorded before normal terminal and acceptance | | Semantic rule | Opening hours and Thursday closure, not exact-string equality | | Release decision | Ineligible due to Stop, regardless of correct late answer | | Accepted destinations | Expected absent; actual UI/history/export/copy NOT EXECUTED | | Authority | No tool configured, proposed or authorized | | Remote cancellation | UNKNOWN; local action is not a provider receipt | | Evidence gap | Actual client abort propagation, reader lifecycle and destination tests | | Independent reviewer | UNASSIGNED; not approved | ## Blank application record Copy for each real trace. A blank is UNKNOWN, not PASS. | Field | Fill with application evidence | | --- | --- | | Source API/model/request settings | | | Target API/endpoint/profile/source Region | | | Client/SDK/parser and release-rule revisions | | | Supported output profile and excluded forms | | | Trace origin: synthetic projection, decoded capture or framed capture | | | Attempt ID and native identities/indexes retained | | | Event order, including Stop/retry/drain observations | | | Candidate assembly and source final-text agreement | | | Native outcome and native reason/exception | | | Local release eligibility at decision time | | | Semantic acceptance rule and content result | | | Transport/drain and accounting observations | | | Expected answer/status/history/export/copy destinations | | | Actual destination readback with evidence paths | | | Tool/effect authority, if outside this text-only profile | | | Remote cancellation evidence or explicit UNKNOWN | | | Decision: retain integration, adapt, expand profile, hold | | | Remaining gaps, named owner and independent review | | This worksheet is reusable precisely because it permits rejection and UNKNOWN. It does not contain an executable classifier that accepts only its own literal examples. A real acceptance test must drive the application reader/release path and inspect destinations. Wire/client/provider checks are separate and remain unperformed.