An Incident Timestamp Is Not Proof of Causality
Separate event time, collection time and clock uncertainty in incident reviews. Establish dependency evidence before turning a sorted timeline into a cause.
The first line in a timeline may not be the first event
Two services report errors eighty milliseconds apart. A dashboard sorts their logs and presents a clean sequence. It is tempting to conclude that the earlier error caused the later one. That conclusion requires more than a timestamp: the records must refer to comparable event boundaries, their clocks must support the claimed ordering, and evidence must connect the events through a plausible mechanism.
Consider a fictional reservation workflow. Service A records an accepted state transition; service B reports a downstream failure. A's source timestamp is lower, but its log reaches the collection system after B's log. The incident reviewer now has two different orderings, neither of which alone explains the failure. A buffered log can arrive late without its underlying operation happening late.
This guide proposes an evidence worksheet for that problem. All identifiers, times and uncertainty bounds below are synthetic. They are not an Ampity incident, measured clock error or executed telemetry experiment. The goal is to make the strongest supported statement without turning missing observations into an invented history.
Identify what each timestamp actually measures
OpenTelemetry's logs data model distinguishes source Timestamp from ObservedTimestamp, measured when the collection system observes an event. Our operational recommendation is to retain those meanings in the investigation export. If a source time is absent, label the available value as collection time rather than silently treating it as when the business operation happened.
Even a field called event_time needs an instrumentation definition. Was it captured before sending a request, after receiving a response, at transaction start, after a commit acknowledgment, or while formatting a buffered message? Those boundaries are not interchangeable. Ask for the code or documented emitter behavior that establishes the meaning of the field, together with the application version that produced it.
The PostgreSQL date/time reference explains that CURRENT_TIMESTAMP and now() represent transaction start, while other functions have statement or current-clock semantics. Our operational recommendation is not to read a row's default timestamp as its commit time without checking the actual expression. Replacing the expression with a clock-reading function would still not, by itself, prove when the transaction committed.
Retain the original timestamp, units, time-zone representation, source process and collection path. Normalize a copy for comparison, preserving the original evidence. Record whether the platform rounded values or displayed only seconds. Multiple records with the same displayed second are not evidence of simultaneity, and extra fractional digits do not establish that the source clock was accurate to that precision.
Use uncertainty bounds before asserting chronological order
RFC 3339 describes timestamp representation and numeric offsets. Our operational recommendation is to normalize valid offsets consistently but treat synchronization accuracy as a separate question. Converting a timestamp to UTC does not correct a clock that was wrong or a record emitted after the event it claims to describe.
For this worksheet, both source values are expressed in milliseconds from the same arbitrary reference instant. A reports 1,000 and B reports 1,080. Assume each value has a defensible total error bound of plus or minus 100 milliseconds around the event boundary being compared. The bound includes the relevant clock and capture uncertainty for that boundary. A's possible interval is 900 through 1,100; B's is 980 through 1,180. They overlap, so these values cannot establish a strict order.
A different hypothetical bound of twenty milliseconds gives A the interval 980 through 1,020 and B the interval 1,060 through 1,100. Those intervals support A preceding B under the assumptions. They still do not prove that A caused B. Another event might cause both, or their chronological sequence could be incidental to the failure being investigated.
units: milliseconds from one reference instant
A source value: 1000
B source value: 1080
broad total bound: 100
tight total bound: 20
touching-case B value: 1040
A collection value: 1500
B collection value: 1300
strict order: latest possible A < earliest B
causal claim: needs independent mechanism evidenceThese bounds are assumed, not measured. In a real review, retain the evidence supporting the bound over the incident period, including clock changes and capture behavior. A later synchronization check cannot certify an earlier outage window. If the bound is unknown, do not substitute a convenient number. Keep order unresolved unless other reliable evidence establishes it.
For strict ordering, equality is not enough: with A at 1,000, B at 1,040 and bounds of twenty, the intervals touch at 1,020. This rule is a conservative worksheet comparison, not a synchronization protocol. It does not apply unchanged to arbitrary duration measurements or to timestamps whose event meanings have not been aligned.
Connect records through identity and a verified dependency
The W3C Trace Context specification defines identifiers used to propagate request context and discusses security considerations. Our operational recommendation is to use trace information to find related records, then verify the relevant operation and dependency. A shared trace identifier is not proof that one record describes the cause of another, and an incoming header is not an authenticated business event.
Distinguish the business operation from its execution attempts. A reservation command can have one stable operation identity and several retries, each with its own attempt and transport records. Joining only on customer identity or trace ID can combine unrelated work. Conversely, creating a new trace at a queue boundary can hide a valid relationship unless an explicit link or message identity is preserved.
For the fictional workflow, stronger evidence would include a retained message M17 associated with the accepted transition and a worker record identifying consumption of that same message. The reviewer also needs the verified path that publishes only committed transitions, and evidence that the worker record belongs to the actual consumer attempt. Under those conditions, publication precedes consumption even if unsound wall clocks display the opposite order.
That establishes a workflow dependency, not the root cause of B's failure. The failure may depend on B's configuration, another dependency or a defective interpretation of the message. Inspect the mechanism and competing hypotheses. Do not infer successful business completion from message consumption or from a span labeled success without knowing what its instrumentation considers successful.
Record supported conclusions and missing evidence separately
Write the observation before the explanation. An observed collection ordering, a bounded event ordering and a verified message dependency are different findings. Keep the specific evidence references beside each statement so another reviewer can check whether the conclusion exceeds what the record says.
The five cases below are independent expected review decisions. They are not output from a trace backend or evidence that the example's clocks and messaging path were validated.
| Available evidence | Supported statement | Not yet established | | --- | --- | --- | | A 1000 and B 1080, each with total bound 100 | Event intervals overlap; strict order is unresolved | Which event preceded the other or caused the failure | | A 1000 and B 1080, each with total bound 20 | A precedes B under the stated timestamp assumptions | A caused B or explains the observed failure | | One collector observes B at 1300 and A at 1500 | B was observed before A by this collector | Order of the underlying business events | | Two records carry the same trace identifier | Records are candidates for correlation | Authenticated identity, direct dependency or root cause | | Committed-transition publication and matching M17 consumption are verified | Publication precedes that consumption | Why the consumer failed or whether business work completed |
Do not turn absent logs into a negative finding unless collection coverage supports that interpretation. Sampling, retention, exporter failures, filtering and an uninstrumented path can each remove records. The defensible statement may be that no record was found in a named source during a specified export, not that the operation never occurred.
Keep recovery actions distinct from causal conclusions. An operator may need to stop unsafe writes while the root cause remains unresolved. That action can be justified by the current risk and approved incident procedure without claiming that the timeline has already identified the responsible component.
Let AI summarize evidence without manufacturing the missing edges
An AI assistant can group records, identify inconsistent timestamp meanings and draft hypotheses for review. It should receive evidence with source identities and uncertainty labels, not only a sorted list stripped of provenance. Ask it to distinguish observations, assumptions and proposed tests. A fluent narrative can make an unsupported sequence harder to notice.
Do not let the assistant infer clock accuracy from timestamp precision, invent a message dependency from adjacent lines, or collapse retries into one successful operation. Its proposed relationship needs a supporting record and a checked instrumentation meaning. Where evidence is missing, the output should retain the gap and suggest the narrowest next observation rather than manufacture an explanation.
Protect the evidence boundary as well. Incident records can contain credentials, personal data, customer payloads and confidential topology. Use approved access and redaction before moving them into an AI or external analysis system. Read the telemetry redaction guide for that separate control. Correlation convenience is not a reason to copy every raw payload into a model prompt.
Review deliberately broken controls in an isolated evidence fixture: delayed collection, an unknown clock bound, a transaction-start time mislabeled as commit time, and unrelated records sharing a supplied trace ID. The reviewer should reject each unsupported inference. Such a fixture tests the interpretation rule; it does not certify the accuracy of a production collector or establish the cause of an actual outage.
Turn the timeline into a testable investigation
Start with one consequential question, such as whether a worker consumed a command before its associated state was durable. Identify the source records and event meanings needed to answer it. Preserve raw identifiers, versions, timestamp semantics, uncertainty evidence and collection coverage. Write the alternative explanations before choosing the most convenient one.
The next useful step is a permitted test that can distinguish those explanations. That might inspect retained message identity, compare the deployed emitter code with the recorded field, or reproduce a buffered export in an isolated environment. Do not run failure injection against customer traffic simply because the timeline contains a gap. Agree the test's safety boundary and expected evidence with the incident owner.
Use the production debugging guide to organize hypotheses and the observability review to examine instrumentation and collection boundaries. A cloud reliability assessment can connect those findings to recovery decisions and accepted service behavior. Define the actual scope and evidence required rather than treating this worksheet as a complete incident process.
The local example checks interval arithmetic and review wording only. No clock calibration, collector experiment, distributed trace or incident reproduction was executed for this article. Reading does not require an email address. If you choose to contact Ampity, share the disputed event boundary and the evidence you have permission to disclose, rather than an unsupported root-cause statement.