Redact Sensitive Telemetry Before It Crosses a Storage Boundary

Decide where telemetry redaction must happen by tracing the actual collection, buffering and export path. Keep useful operational evidence without copying request...

A clean dashboard can conceal an earlier raw copy

Redaction must happen before the first boundary where a prohibited value could be stored or exposed. Filtering data before the final monitoring backend is useful, but it does not establish what the application logger, local agent, retry buffer or diagnostic exporter already received. Trace those boundaries before deciding where to put the control. Prevent collection of fields the team does not need, then use downstream filtering to contain mistakes and third-party instrumentation you cannot fully control.

Imagine an AI-assisted document-intake service that records an exception with the request body attached. The monitoring dashboard hides the body through a display rule. Meanwhile, an application log file still contains the uploaded text, and an agent may have forwarded the raw exception before that display rule ran. The operator sees a clean screen but has not established a clean collection path. This is an illustrative design review, not a claim about Ampity's production telemetry or a customer incident.

The operational objective is specific: diagnose a failed intake without duplicating the source document into the observability estate. A task identifier, stage, failure category and timing may be enough for the first investigation. If an engineer genuinely needs source content, the application can provide a separately authorized path with its own access and retention controls. A log search accessible to every operator should not become an accidental document repository.

Give each collected field a diagnostic purpose

OpenTelemetry's sensitive-data guidance assigns responsibility to the implementer, recommends data minimization and describes ways to remove or transform data. It also warns that hashing predictable identifiers may not provide the anonymization a use case needs. See OpenTelemetry guidance on handling sensitive data. Use that guidance to review what your instrumentation emits; the framework cannot decide the sensitivity of your business data for you.

For the illustrative intake service, define a small field contract. The task identifier correlates stages. A bounded error category distinguishes unavailable dependencies from invalid input. Duration supports latency investigation. A source-system category helps isolate an integration failure. The document text, access token and user's free-text message are excluded from ordinary telemetry. These are proposed choices that need review against the actual diagnostic requirements, not a universal attribute schema.

An allowlist is easier to review when fields have stable meaning and type. A permissive field such as details can absorb an entire request or model response as the code evolves. Rejecting every unknown field may also remove evidence needed for an incident, so define how a proposed addition is reviewed and deployed. Record who can approve it, why it is needed, what values it can contain and whether it changes storage or access requirements.

Avoid calling a transformed identifier anonymous merely because it is no longer readable at a glance. Ask whether it can be linked back to a person, joined across systems or guessed from a small set of possible inputs. Use the least identifying correlation that supports the investigation. The appropriate data classification and legal treatment require the organization's privacy and security review; a processor setting does not establish compliance.

Draw the collection path through queues and extra exporters

Follow one event from instrumentation to the destinations actually configured. Include the application's own logging path as well as SDK exporters, agents, gateways and the final backend. At each step, identify what is held in memory, what can be written to disk, who can access it and what happens when the downstream service is unavailable. A diagram containing only application and dashboard misses the places where raw data can persist.

OpenTelemetry documents exporter sending queues and optional persistent storage using the file-storage extension. With persistent storage enabled, queued telemetry can survive a Collector restart and be sent later. See OpenTelemetry Collector resiliency guidance. That durability is useful for delivery, but it also makes the content entering the queue part of the storage review. Inspect the configured version and component behavior rather than assuming every Collector uses persistent buffering.

Place redaction before a queue or destination that must never receive the value. A filter at a later gateway cannot retroactively remove data already written by an earlier agent. If the application can avoid constructing the sensitive attribute at all, that removes one source of accidental copies. Where a component necessarily receives raw data, document that boundary, restrict access and decide whether the design is acceptable before describing the path as sanitized.

Check parallel outputs. A production exporter may apply a filter while a diagnostic exporter, local file output or alternate telemetry pipeline receives the unfiltered event. Include error output and collection diagnostics in the review. Do not enable verbose payload logging in production to prove that payload logging is safe. Test with synthetic markers in an isolated setup that represents the relevant components and configuration.

Review the value inside the field, not only its name

An attribute named error.message can contain a customer identifier copied from a dependency response. A URL field can contain a query parameter. A model error can include part of the submitted prompt. Removing an attribute named email leaves those other routes untouched. The field contract should constrain values as well as keys, especially for free text and nested data.

The following matrix is a proposed test set for the intake example. It is deliberately small enough for an engineer to exercise, with one prohibited synthetic marker placed in each surface. Decide the marker format before the test so it does not collide with normal identifiers. Never use real documents, credentials or personal information as redaction probes.

| Test surface | Allowed diagnostic evidence | Prohibited synthetic content | | --- | --- | --- | | Exception message | Bounded error category and task identifier | A document excerpt embedded in free text | | Request URL | Route template and response status | A sensitive query parameter value | | Model call observation | Provider category, duration and failure class | Prompt or retrieved document text | | Metric attributes | Reviewed bounded operational dimensions | A user email used as a dimension | | Retry and diagnostic output | Queue state and sanitized event metadata | The raw event copied into debug output |

Keep evidence of useful fields arriving as well as prohibited fields being absent. Dropping the entire event can make a leak test pass while making operational investigation impossible. The acceptance result should show that the task can still be correlated with its failure category and timing. If the design intentionally drops an event, count and explain that behavior without logging the discarded payload.

Exercise failure paths that bypass normal export

An unavailable backend changes how long telemetry is buffered and which diagnostics are produced. Test the configured retry behavior with synthetic data, then inspect the permitted evidence surfaces before and after recovery. Include a restart if the path uses persistent buffering. An exporter result after recovery does not establish what was present in the queue before recovery, so inspect that boundary through approved tooling and restricted test access.

Test malformed input and processing errors. Decide what the pipeline does when redaction cannot parse a body or encounters an unsupported value. Continuing with the raw event preserves delivery but violates the prohibited-data contract. Dropping it protects that contract but loses evidence. A proposed policy can keep a bounded failure counter and sanitized event identifier while refusing to forward the unsafe payload. Validate that the failure counter itself does not include the sensitive value.

A regex rule has limitations. Encoded values, nested structures and new message formats can escape a pattern designed for a known sample. Use structured field removal where the data permits it, then test remaining free-text surfaces. A passing marker test covers the selected paths and values, not every future instrumentation library. Repeat relevant tests when instrumentation, processors, exporter configuration or log formatting changes.

Sampling should not be treated as redaction. Reducing the number of retained events still leaves the selected events exposed if they contain prohibited values. Check sensitive content before relying on sampling for cost control. Review whether a component buffers unsanitized data to make its sampling decision; the exact behavior depends on the implementation and configuration under test.

Keep emergency diagnostics controlled and temporary

During an incident, engineers may want more detail than the normal field contract provides. Define that path before the incident: who can approve it, which workload and interval it covers, where evidence goes and how access expires. Prefer a narrowly scoped diagnostic change to enabling raw request or prompt logging across all tenants. Record the configuration version so the team can identify events collected during the exception.

The closure task must include the additional data already collected. Turning off verbose logging stops future collection but leaves earlier files, buffered events and exported records. Follow the applicable retention and incident process for those copies. Do not promise deletion across every destination unless each destination and relevant retained copy has been reconciled. Where deletion cannot be completed immediately, record the restricted-access and expiry plan.

Distinguish operational observability from a business audit record. A privacy-minimized error stream may intentionally omit content that a regulated workflow must preserve elsewhere. Keep that evidence in the appropriately governed system, with separate access and retention. Broadly accessible telemetry should not carry the full business record just because engineers find it convenient to search.

Prove one event path before expanding the rule

The next action is to choose one sensitive workflow and follow one synthetic event through the actual collection path. Agree the allowed field contract, mark every storage or exposure boundary, then run the matrix during normal operation, backend failure and recovery. Record both useful evidence arriving and prohibited markers being absent at the inspected boundaries. Leave uninspected destinations visible as gaps rather than declaring the whole estate clean.

For AI workflows, the companion article on sensitive data in tool logs focuses on what a tool result may contain before instrumentation sees it. Use this collection-boundary review with that input review. Bring the observed path and unresolved storage boundaries to an observability review before broadening collection. Neither this article nor a redaction configuration is a compliance certification.