The First 30 Days of a Production Platform Reset
A day-by-day operating playbook to establish system truth, contain urgent risk, baseline delivery and reliability, prove one production improvement, and leave an...
trigger="A new technology leader, recovery sponsor, or platform owner needs to regain control of a production system whose reliability, delivery flow, cost, security, or ownership is not understood well enough to plan responsibly." owner="The executive or senior engineering leader accountable for the platform reset and empowered to resolve priorities, access, ownership, and cross-team dependencies." timebox="30 calendar days. The goal is one proven operating loop and a defensible next horizon, not a complete transformation." participants={["Product owner", "Application engineering", "Platform or SRE", "Security", "Data owner", "Cloud or FinOps", "Support or operations", "Business sponsor"]} prerequisites={[ "Access to source, delivery pipelines, cloud accounts, production telemetry, incident records, support signals, architecture material, and cost data.", "Named owners for at least one critical user journey and authority to pause unsafe changes or escalate unresolved dependencies.", "Agreement that the first month will preserve evidence, avoid speculative rewrites, and expose uncomfortable facts without blame." ]} outputs={[ "A current-state journey, service, dependency, data, ownership, and risk map.", "Baselines for delivery flow, user-centered reliability, cloud cost, security exposure, and operating burden.", "One bounded production improvement with release, rollback, recovery, and acceptance evidence.", "A sequenced 90-day decision backlog with owners, value, evidence, dependencies, stop conditions, and residual risk." ]} doneWhen={[ "Leaders and operators use the same system map, definitions, and evidence sources.", "Urgent risks are contained or explicitly accepted by an accountable owner.", "One user journey has moved through the improved delivery and operating loop in production.", "The next 90 days are ordered by consequence and evidence, not by the loudest backlog request." ]} />
Use the first month to regain control, not announce a transformation
A platform reset is needed when leaders cannot answer basic operating questions with confidence. Which user journeys matter most? What fails most often? How does a change reach production? Which team owns a dependency? What does the system cost under real demand? Which recovery claims have been rehearsed? Which risks are accepted, and by whom?
The first 30 days should make those questions answerable and prove that the organization can turn evidence into one safe production improvement. It should not produce a long target architecture, a generic maturity score, or a rewrite promise before the current system is understood.
The working principle is simple: stabilize the decision system before scaling the change program. Create one shared map. Establish a small set of user-centered measures. Contain urgent risk. Select one thin journey. Improve and release it through the intended controls. Rehearse failure. Accept the evidence. Then sequence the next horizon.
Day 0: establish mandate and guardrails
Before the clock starts, the accountable sponsor writes a one-page reset mandate. It names the reason for the reset, affected product boundary, decision authority, operating constraints, required participants, evidence access, and the conditions that justify pausing a production change.
State what the reset will not do. Typical exclusions include a whole-platform rewrite, vendor selection without workload evidence, organization redesign, cloud migration without a bounded journey, and arbitrary cost-reduction targets. Exclusions protect the first month from becoming a container for every unresolved technology concern.
Create three guardrails:
- No destructive change without a tested recovery path and named approval.
- No platform-wide recommendation without evidence from a real user or operator journey.
- No metric is treated as success unless it changes a decision or explains a user consequence.
Open a decision log and an evidence register. Every important statement should reference a source, owner, collection time, and confidence. Mark unknowns explicitly. A visible unknown is safer than a confident diagram based on stale assumptions.
Days 1 to 3: map journeys, authority, and access
Select three to five journeys that represent product value and operating risk. Examples include account creation, checkout, claim submission, order fulfillment, data publication, scheduled export, partner onboarding, or an AI-assisted decision that can affect a customer.
For each journey, record:
- user or operator and desired outcome;
- entry point and completion evidence;
- services, data stores, queues, third parties, and human steps;
- identity, tenant, region, and policy context;
- owner for product behavior, code, data, runtime, support, and vendor relationship;
- known failure and recovery path;
- current telemetry and blind spots;
- commercial or regulatory consequence of failure.
Do not begin with a complete application inventory. Trace real journeys and add components as they appear. This exposes dependencies that architecture repositories often miss, including manual exports, scheduled jobs, shared credentials, spreadsheet reconciliation, and vendor callbacks.
Validate access in parallel. The team should be able to read source and configuration history, inspect pipeline behavior, query telemetry, view recent incidents, access cost reports, and identify production authority. Record gaps and the person who can resolve them. Do not accept screenshots as a substitute for reproducible access when decisions depend on the underlying data.
Exit gate: the team can trace at least one critical journey from user action to durable outcome and name the authority for each changing fact.
Days 4 to 5: baseline delivery flow
Choose a representative change from request to production and map every wait, handoff, approval, build, test, environment transition, release, and validation step. Distinguish elapsed time from active work. Count rework and the percentage of items that arrive complete enough for the next step.
The DORA continuous delivery guidance recommends value-stream mapping to expose long elapsed time, low value-add time, and poor complete-and-accurate flow. Use the map as an investigation tool, not a performance score for individuals.
Capture a small baseline:
| Signal | Definition | Decision it supports | | --- | --- | --- | | Change lead time | Accepted work start to production evidence | Where work waits and whether smaller batches help | | Deployment frequency | Production releases for the bounded system | Whether the release path is routinely usable | | Change failure | Releases requiring rollback, fix-forward, or user-impacting intervention | Whether feedback and release controls are sufficient | | Recovery time | User-impact start to restored journey | Whether detection, ownership, and recovery work | | Rework | Work returned because criteria or quality were incomplete | Whether readiness and acceptance are clear | | Acceptance latency | Candidate ready to recorded decision | Whether stakeholder delay is part of delivery risk |
Inspect the pipeline itself. Identify the canonical artifact, versioning, environment differences, secrets, manual steps, test duration, flaky checks, database changes, infrastructure changes, approval boundaries, and rollback behavior. A green pipeline does not prove that the release path is safe if operators bypass it for urgent changes.
Days 6 to 7: baseline reliability and operating burden
Define one or two service-level indicators for the selected user journeys. Begin with what the user needs, then work backward to measurement. Google’s service-level objective guidance recommends this direction because easy-to-measure component health often misses the complete experience.
For an interactive journey, measure eligible attempts, successful outcomes, latency, and unknown results. For a scheduled job, measure correct completion by the promised time. For a data product, measure freshness, validity, and availability to intended consumers. For a marketplace transaction, measure reconciled terminal outcomes rather than only API responses.
Review incidents and support signals from at least the previous quarter if available. Classify them by user journey, detection source, time to ownership, cause class, recovery method, recurrence, and follow-up status. Look for repeated manual work and near misses, not only severe incidents.
Estimate operating burden:
- pages and alerts per on-call shift;
- percentage of alerts that lead to action;
- recurring manual repair or reconciliation;
- deployment outside normal working hours;
- privileged production intervention;
- support cases that require engineering diagnosis;
- time spent on unplanned work;
- undocumented dependencies known to only one person.
Exit gate: the sponsor can explain how the selected journeys fail from the user’s perspective and how the team currently detects, owns, and recovers them.
Days 6 to 7: baseline cloud cost and exposure
Cost is an architectural signal when it is tied to demand, product behavior, and ownership. Separate fixed platform cost, variable workload cost, shared services, data transfer, observability, storage growth, committed-use coverage, and idle or orphaned resources.
Choose a useful denominator such as active tenant, fulfilled order, processed document, API request class, model inference, or data volume. Avoid a denominator that makes a product look efficient while excluding a material shared cost.
Review cost together with reliability and demand. A resource may look oversized because it absorbs an unbounded batch job or compensates for a dependency failure. A commitment may lower unit price while increasing lock-in to an unhealthy architecture. A log-retention bill may be high because the system emits uncontrolled payloads or lacks a tiered evidence policy.
Also record immediate exposure: public assets, stale credentials, broad roles, missing backups, untested restores, unsupported dependencies, expiring certificates, unencrypted sensitive data, unbounded queues, unowned domains, and third-party integrations without operational contacts. The AWS Operational Excellence pillar connects operational success to business outcomes, baselines, observability, owned alerts, runbooks, and playbooks. Use that relationship rather than treating cost and risk as separate audits.
Days 8 to 10: contain urgent risk
Contain risks that can create immediate customer, security, data, financial, or recovery harm. Containment is not the same as permanent remediation. It creates a safer window for evidence-based improvement.
Examples include:
- rotate or revoke exposed and unowned credentials;
- protect or disable an unsafe administrative path;
- add a queue limit or workload budget to prevent resource exhaustion;
- preserve backups and prove that the intended restore material exists;
- freeze an unrepeatable production migration path;
- add an idempotency control around a duplicate external effect;
- create a manual reconciliation queue for unknown transaction outcomes;
- route a critical unowned alert to a named operator;
- pause a rollout whose failure or rollback behavior is unknown.
Every containment action needs an owner, expiry or follow-up decision, verification, and reversal path. Temporary controls that silently become permanent are a common source of fragility.
Do not use the reset to blame the team that inherited the system. Fragility often reflects accumulated constraints, incentives, deadlines, and incomplete ownership. The goal is to change the operating conditions that reproduce it.
Days 11 to 14: choose one thin-slice improvement
Select a journey where a bounded change can improve both user outcome and the delivery system. A good thin slice crosses real production boundaries but remains reversible.
Score candidates on:
| Criterion | Strong candidate | | --- | --- | | User consequence | Meaningful and understood | | Evidence | Baseline and acceptance can be measured | | Boundary | Small enough for one controlled release | | Dependencies | Known and owned within the timebox | | Recovery | Rollback or safe recovery is feasible | | Learning | Tests an important target operating capability | | Reuse | Produces a pattern other teams can adopt |
Possible slices include moving one service to the canonical pipeline, adding end-to-end trace context for one journey, separating an unbounded batch workload, introducing a reconciliation record for one external effect, proving one database migration pattern, or implementing one user-centered SLO with an owned response.
Write an acceptance brief before implementation. Include representative load and data, failure cases, release cohort, security and privacy constraints, cost guardrail, rollback, reviewer, and decision time. Keep a list of adjacent improvements explicitly out of scope.
Days 15 to 18: improve the delivery path
Build the smallest controls required to release the selected slice responsibly. Typical work includes a reproducible build, canonical artifact, automated unit and integration tests, database compatibility checks, infrastructure review, vulnerability checks, environment configuration, feature exposure, telemetry, and a smoke test that verifies the user journey.
Do not rebuild the entire platform toolchain. Prefer one paved path that can be repeated. Keep service-specific exceptions visible and owned. If the current system has no canonical release record, create one that joins source revision, artifact, configuration, migration version, approvals, environment, time, and result.
Use production-like conditions where feasible. Test data, integration behavior, identity, and workload often differ materially from a developer environment. Record those differences instead of interpreting a staging success as production evidence.
The DORA deployment automation guidance describes an automated path from canonical packages through environment preparation, deployment, and smoke testing. The exact tooling can vary. The important result is that the process is repeatable, reviewable, and usable under normal and urgent conditions.
Days 19 to 21: release with observation and recovery
Define a release plan with cohort, time window, observers, stop conditions, rollback authority, in-flight work behavior, support communication, and data reconciliation. Confirm telemetry before exposing the change.
Release to the smallest representative cohort that can produce useful evidence. Watch the user journey, not only component health. Compare the result with the baseline and monitor guardrails for reliability, security, cost, and downstream effects.
If a stop condition is reached, stop new admission or exposure first. Preserve logs, traces, state versions, and external references. Roll back or recover according to the pre-agreed path. Diagnose after the user consequence is contained. Do not erase failed evidence.
The release record should answer:
- which version and configuration ran;
- which users, tenants, regions, or workloads were exposed;
- which migrations and external effects occurred;
- which evidence was observed and by whom;
- whether stop conditions or alerts fired;
- what happened to in-flight work;
- whether rollback or recovery was used;
- what decision allowed expansion, correction, or closure.
Days 22 to 24: rehearse failure
Select one credible failure that crosses the journey. Examples include a dependency timeout after a local commit, queue backlog, stale identity policy, partial database migration, lost provider response, failed regional dependency, or a release that must be reversed while work is in flight.
Run a tabletop first. Ask participants what they would see, who owns the incident, how they contain impact, which source is authoritative, which actions are safe to repeat, and how they verify recovery. Then run the safest practical technical exercise in a controlled environment or bounded production cohort.
Capture detection time, ownership time, decision time, recovery time, evidence gaps, unsafe manual steps, and follow-up. A runbook is accepted only when a person other than its author can use it under realistic conditions.
Do not confuse rollback with recovery. Rolling back application code may not undo a schema change, external payment, emitted event, user notification, or data mutation. Name compensating actions and their authority separately.
Days 25 to 27: reconcile outcomes and operating cost
Recompute the journey outcome from authoritative records. Compare requests, durable intent, external provider state, internal transaction state, data changes, user-visible result, and any financial or operational ledger. Find missing records as well as mismatched values.
Compare the new slice with the baseline:
- Did user success or operator task success improve?
- Did lead time, rework, deployment pain, or acceptance latency change?
- Did failure become easier to detect, contain, and recover?
- Did the change move cost, security exposure, or support burden elsewhere?
- Can another team repeat the path without the original authors?
- Which assumptions were invalidated?
Record residual risks and their expiry. Do not represent one successful release as proof of a long-term trend. State the sample, limitations, and what further evidence is required.
Days 28 to 29: build the 90-day decision backlog
Convert findings into decisions and bounded work, not a list of observations. Every backlog item should include the user or operating consequence, evidence, proposed boundary, accountable owner, dependencies, acceptance method, stop condition, and reason for its order.
Group work into five horizons:
- Contain: remove immediate unsafe exposure.
- Stabilize: make critical journeys measurable, recoverable, and owned.
- Simplify: remove duplicate paths, unowned components, manual state changes, and unnecessary variation.
- Enable: create repeatable delivery, platform, data, and operating capabilities.
- Transform: undertake larger architectural or organizational change only when evidence supports it.
Limit work in progress. One team cannot simultaneously rewrite the platform, migrate clouds, introduce an internal developer platform, adopt a new observability stack, change the data architecture, and deploy production AI responsibly. Sequence based on consequence and dependency.
Include decisions to stop, retain, or defer. A roadmap that only adds work is not a strategy.
Day 30: run the evidence review
The final review is not a presentation contest. Give leaders the evidence pack in advance. Walk through the selected journey, baseline, urgent containment, production slice, failure rehearsal, measured result, limitations, residual risk, and next decisions.
Ask the accountable owners to record acceptance or specific disagreement for:
- current-state map and known blind spots;
- urgent risk treatment;
- baseline definitions and evidence quality;
- thin-slice acceptance;
- operating ownership and support readiness;
- 90-day priorities, dependencies, and capacity;
- risks accepted, deferred, transferred, or requiring further investigation.
Close temporary privileged access that is no longer required. Assign every retained artifact and operational control to a durable owner. Schedule the first 30-day follow-up on the new roadmap.
Reset artifacts to retain
'Reset mandate, system boundary, guardrails, decision rights, and exclusions', 'User journey, service, dependency, data authority, vendor, and ownership maps', 'Evidence register with source, owner, collection time, confidence, and known gaps', 'Delivery-flow, reliability, operating-burden, cost, and security baselines', 'Urgent-risk register with containment, verification, expiry, and permanent owner', 'Thin-slice acceptance brief and release record', 'Rollback, recovery, and failure-rehearsal evidence', 'Reconciliation result and residual-risk decisions', 'Ninety-day backlog with consequence, evidence, owner, dependencies, and stop conditions', 'Access closure and operating handoff record', ]} />
Failure modes and stop conditions
Stop and escalate when production authority is unclear, access would breach policy, a critical backup or restore claim cannot be verified, the proposed change has irreversible effects without approval, or the organization cannot name an owner for the resulting service.
Pause the thin slice when representative evidence is unavailable, a required dependency is late, scope expands beyond the recovery plan, or urgent risk consumes the team’s operating capacity. A pause is not failure. It is evidence that the original boundary or readiness decision was wrong.
Do not declare the reset complete because the documents exist. Completion requires one accepted production journey, one functioning ownership loop, and a next horizon that leaders are prepared to resource.