Run a PostgreSQL Point-in-Time Restore Drill

Rehearse an isolated PostgreSQL restore, verify the recovery target and business invariants, reconcile external effects and measure time to an accepted service.

trigger="A PostgreSQL backup reports success, but the team has not measured whether a restored application can safely resume its agreed business function." owner="The service owner accountable for accepting the recovered application." participants={['Database operator', 'Application engineer', 'Security reviewer', 'External integration owner', 'Independent exercise observer']} prerequisites={['Approved isolated recovery environment and cost allowance', 'Identified backup mechanism, artifacts and target', 'Synthetic markers and business acceptance fixtures', 'Blocked production integrations and a tested stop control']} outputs={['A recovery-target evidence record', 'A milestone timing log', 'Application invariant and permission results', 'An external-effect reconciliation register', 'An owned recovery decision and follow-up defects']} doneWhen={['The actual recovery target is independently verified', 'Required application operations pass with authorized identities', 'Unauthorized controls remain denied', 'External effects are reconciled or explicitly held', 'The complete measured duration and limitations are accepted by the service owner']} />

Measure recovery until the agreed service is usable

Use this playbook to rehearse a PostgreSQL point-in-time restore without connecting the recovered copy to production traffic. The result should identify the recovered database state, prove agreed application behavior, account for external side effects and measure elapsed time through the service owner's acceptance. A listening database port is an intermediate milestone.

Choose one bounded business function for the first drill. An order lookup and an isolated test-order update make the exercise more specific than “the application works.” Identify the engine version, backup mechanism, deployment configuration and dependencies. A logical export, a physical archive and a managed provider's recovery operation need different execution runbooks.

This is proposed engineering guidance, not a command sequence for your production database. The order example and 105-minute timing illustration are hypothetical. They do not describe an Ampity customer, a provider guarantee or a recovery promise for your system. A database operator must supply the version-appropriate restore instructions and obtain approval for resources, data handling and costs.

The companion PostgreSQL restore-proof article explains why backup status cannot establish application recovery. The recovery-window article covers choosing a defensible target. This playbook turns those questions into assigned actions and a repeatable acceptance record.

1. Define the exercise contract and abort criteria

Owner: service owner with incident commander. Output: approved drill contract. State the selected business function, target environment, permitted data, participating identities and acceptance conditions. Specify who can hold the drill, who can accept its result and who owns each dependency. Approval for this rehearsal does not permit switching the production endpoint.

Define the exercise's clock. For example, start at the simulated recovery declaration and stop when the owner accepts the agreed function after all required checks. Also record detection and decision time if the wider objective begins earlier. Preserve the individual milestones so an operator cannot hide artifact retrieval or identity restoration outside an unexplained stopwatch boundary.

Agree stop conditions before allocating resources. Stop if the restored application reaches a production payment or notification destination, if the target identity is ambiguous, if protected data appears in uncontrolled evidence, or if a required backup artifact is missing. Hold any cutover recommendation until those conditions have an owned disposition.

State what a failed exercise means. A held result can still provide useful measurements, but it cannot establish recovery acceptance. If the team needs another restore attempt, record it as another attempt with its own target and timestamps. Do not delete the failed attempt from the report or count its elapsed time as zero.

2. Identify the backup mechanism and dependency chain

Owner: database operator. Output: artifact and mechanism register. Record the engine version, backup identity, start and completion evidence, archive location, required credentials, encryption access and dependency chain. Include relevant configuration, extensions, tablespaces, application schema version and any tooling needed to reconstruct the selected backup. Keep secrets out of the register.

PostgreSQL 18's continuous-archiving documentation describes recovery from a base backup with a continuous archived WAL sequence. It distinguishes that method from logical dumps and requires the stopping point to follow the base backup's completion. Use an appropriate earlier backup when the desired target predates that boundary; do not relabel an unsuitable artifact as a viable restore source.

For an incremental chain, identify all required parents and the supported reconstruction procedure. For a managed service, identify the selected source instance and currently reported restorable interval instead of assuming that retention settings prove every timestamp is available. Record the mechanism-specific evidence the operator actually observed.

Test retrieval with the exercise identity. Confirm that it can locate and read the needed artifacts and obtain the required decryption access in the isolated environment. An artifact accessible only through one absent administrator's account is an operational dependency. Record retrieval duration and failures separately from database replay time.

3. Prove the restore target and comparison fixtures

Owner: application engineer with database operator. Output: target evidence packet. Select a target that exercises a real recovery question. In a synthetic order scenario, the team might need the last valid state before an incorrect bulk status update. Identify the event and transaction evidence that distinguish the accepted state from the unwanted change.

Use a timestamp with an explicit timezone and document its relation to transaction completion. Where the selected mechanism supports another target type, record its semantics and version-specific settings. PostgreSQL 18's recovery-target configuration describes target selection, inclusive behavior, timelines and the action at the target. Verify the chosen settings rather than assuming that a displayed time represents the exact intended transaction boundary.

Prepare synthetic markers before and after the proposed stopping point in the test source. Record the expected presence or absence of each marker and the expected business records. The acceptance evidence should establish both that necessary committed work survived and that the unwanted change did not. A single recovered row is too weak for that conclusion.

Keep a reference dataset or approved evidence query independent of the restore target. Record how the observer obtains it and what its limitations are. If the bad event's timing is uncertain, label the target as a candidate and plan another isolated attempt. Do not call a guessed timestamp the last known good state.

4. Establish isolation before restoring data

Owner: platform operator with security reviewer. Output: isolation checks. Verify the target account or project, network, storage and identifiers before running the restore. Keep ordinary users away from the copy. Grant the exercise team only the access required to perform and observe the rehearsal, and document how temporary access will be removed.

Block production outbound effects before application startup. Cover payment providers, email, SMS, webhooks, object writes, scheduled jobs and queue consumers relevant to your application. An environment-variable label such as “staging” is not proof that every destination is isolated. Test the destination controls with approved synthetic requests and inspect the actual observed paths.

Restore sensitive data only under an approved handling plan. A production-shaped rehearsal may need volume and schema fidelity without exposing complete personal records to every participant. Define masking or restricted access, evidence redaction, retention and deletion ownership before copying data. Never use a public screenshot or unrestricted PDF as a place to store recovery payloads.

For Amazon RDS, AWS's point-in-time restoration guidance states that the operation creates a new instance without modifying the source. It also describes security-group and parameter/option-group choices. Inspect the resulting settings explicitly. A new instance and endpoint provide separation, but do not by themselves establish your application's network or outbound isolation.

5. Validate artifacts, then run the isolated restore

Owner: database operator with independent observer. Output: restore execution log. Run the integrity check appropriate to the backup mechanism and version. Record its configuration, exit result and checks excluded by format or option. Preserve a bounded log and the identity of the artifacts examined. Treat a missing or inconsistent artifact as a hold, not an invitation to skip verification.

The PostgreSQL 18 pg_verifybackup documentation describes checking base backups against a manifest and notes that its checks cannot cover everything a running server will check. It recommends test restores and checking the resulting data. A passing manifest verification is evidence about those artifacts, not the acceptance result for the application.

Execute the separately reviewed restore procedure in the confirmed target. Record when provisioning begins, when artifacts are available, when replay or provider recovery starts, and when the database reaches the expected state. Record retries, manual interventions and waiting for credentials or capacity. Preserve the original backup and source; the drill should not overwrite evidence needed for another attempt.

Check the actual recovery result against the selected target and markers. Verify the expected database identity, engine configuration and recovery status using supported observations. If the database cannot reach the target, record the reason and hold acceptance. Do not promote a different stopping point and present it as the requested one merely because it starts successfully.

6. Verify application invariants and permission boundaries

Owner: application engineer with business record owner. Output: application acceptance results. Connect a separately configured exercise application to the confirmed recovered endpoint. Verify its schema compatibility, extensions, roles and required dependencies. Record the application build used for the checks. A restored database paired with an incompatible application release may need a different recovery plan.

Run checks for the selected function through its supported interface. In the hypothetical order example, verify order identity, line-item relationships, currency, permitted status transitions and referenced documents. Compare expected totals using the application's business rules and correct precision. A row count alone cannot establish that orders agree with their payments or fulfillments.

Use authorized and unauthorized control identities. Confirm that permitted users can complete the agreed test operation and that another tenant or restricted role cannot read or change the fixture. Check restored grants and application authorization together. A recovery that removes permission boundaries is not acceptable simply because the legitimate user can log in.

Record individual results with fixture identity, expected state, actual observation and owner. Mark unknown or untested checks explicitly. Avoid a single green screenshot covering authentication, data integrity and business behavior. Preserve enough evidence for a reviewer to reproduce the acceptance decision without exposing unnecessary customer data.

7. Reconcile external effects and unsettled jobs

Owner: external integration owner with application engineer. Output: reconciliation register. List effects outside the database that may have happened after the chosen recovery point. Include payments, emails, object uploads, webhook deliveries and jobs relevant to the function under test. The recovered database cannot rewind those systems merely by losing their local receipts.

For each effect, identify an independent lookup, event or provider receipt that can establish what happened. Bind the evidence to the intended business operation and target, not just a matching amount or timestamp. Where reliable absence cannot be proved, hold the item for reconciliation. Do not resend a payment or message simply because the restored local row says “pending.”

Test the dangerous case with synthetic systems: an external provider commits an effect, while the selected database recovery point predates its receipt. Observe whether the recovered worker tries to issue the effect again. Keep that worker isolated and require the designed duplicate-prevention or reconciliation behavior before treating the workflow as resumable.

The regional unsettled-write article discusses this mismatch across failure boundaries. Retain the actual register from your drill, with resolved, disputed and unknown states. A manual recovery team also needs a procedure for deciding which queued work remains blocked and who can authorize its eventual continuation.

8. Measure the full duration and state the coverage

Owner: independent exercise observer. Output: milestone worksheet. Use recorded timestamps to calculate declaration-to-acceptance duration. Separately report database-ready time, application-check completion and reconciliation completion. Record whether the exercise used pre-provisioned capacity, cached artifacts, a healthy identity service or other conditions that would differ during an incident.

In the hypothetical sequential illustration, artifact access takes 15 minutes, restore and target verification take another 35, application checks take 28, external reconciliation takes 22, and the final acceptance decision takes 5. Database validation completes at minute 50. Service acceptance occurs at minute 105, so describing this as a 50-minute service recovery would omit 55 minutes of required work.

Do not assume those durations will repeat. Report the attempt's data size, replay volume, target capacity, dependency conditions and number of operators. If phases overlap, calculate elapsed time from the clock rather than adding every participant's work duration. Operator effort and service outage duration answer different planning questions.

Compare the measured result with the approved recovery objective using the same start and stop boundaries. If the objective was 90 minutes, this hypothetical 105-minute exercise misses it by 15 minutes. Identify the measured constraint and propose a specific follow-up, such as testing artifact retrieval under the intended recovery identity. Do not claim improvement before repeating the changed drill.

9. Rehearse the cutover decision without switching production

Owner: incident commander with service owner. Output: cutover decision rehearsal. Walk through who would authorize a real endpoint change and which evidence they would require. Include current-source state, remaining writes, client reconnection, routing, monitoring and the accepted disposition of external effects. This drill's acceptance record does not itself authorize production cutover.

Define how the team prevents two writable authorities during a real transition. The mechanism depends on the system, and its behavior must be tested under a separately approved plan. A DNS change or a new connection string alone does not establish that old workers and existing connections can no longer write to the original database.

Describe the reversal limit. Returning traffic to the old system after new writes occur may require data and effect reconciliation. Do not call that action a simple rollback without checking which system now owns accepted work. Record the point at which the incident commander must stop and obtain another decision instead of executing a prewritten reversal blindly.

Verify the monitoring that would detect a failed transition. Use synthetic reads and writes, permission controls and queue observations appropriate to the accepted service. Assign an observation window and an owner. A database health check cannot cover broken customer authentication, repeated notifications or application sessions that remain connected to the wrong endpoint.

10. Accept the evidence, retain findings and remove test access

Owner: service owner with security reviewer. Output: signed report and cleanup record. Record accepted, held and failed checks separately. Include target identity, backup mechanism, complete timing boundaries, fixture coverage, application build, external-effect findings and environment differences. Keep any unsupported recovery promise out of the report and public service claims.

Use these acceptance criteria before recording a successful rehearsal:

  • The backup and required dependency chain are identified and retrievable.
  • Target markers establish the intended boundary and exclude the unwanted change.
  • The restored application uses the isolated endpoint and approved configuration.
  • Required business invariants pass through supported application operations.
  • Unauthorized controls remain denied and evidence contains no exposed secrets.
  • External effects and unsettled jobs have resolved or explicitly held dispositions.
  • Measured service duration uses the agreed declaration and acceptance boundaries.
  • A real cutover would require a separate owner decision with defined reversal limits.

Assign every defect a remediation owner and a repeat-test condition. Preserve the failed evidence alongside the successful attempt. Remove temporary access and retire exercise resources according to the approved retention and deletion plan, after confirming that the report no longer depends on an unretained copy. Record completion rather than assuming a resource request means deletion finished.

Start the next drill with the slowest or least-supported recovery dependency from this report. Ampity's reliability review can help define the evidence and rehearsal scope around your application, identity and integrations. Share sanitized milestone logs, fixture results and unresolved boundaries; production credentials are not needed to discuss the recovery decision.