Sync Your Cloud
Sync Your Cloud

๐ŸชWe value your privacy

We use cookies to enhance your browsing experience, analyse site traffic, and personalise content. Choose your preferences below.

Essential CookiesRequired

Required for the website to function properly. These cannot be disabled.

Functional CookiesOptional

Enable personalised features like remembering your preferences and settings.

Home/Insights/Payment Failure Runbook: QSA Expectations
PCI DSS 4.0.1 ยท Incident Response

Payment Failure Runbook: What a QSA Expects to See

A payment failure runbook is written for engineers who need to restore service at 2am. A QSA reads it differently โ€” looking for evidence that incident response procedures exist, are tested, and satisfy the specific sub-requirements of PCI DSS Requirement 12.10. Most runbooks satisfy neither purpose fully. This article shows what needs to change.

13 min read PCI DSS 4.0.1 Incident Response

Why Payment Runbooks and PCI DSS Diverge

Engineering teams write runbooks to solve operational problems: a payment processor is returning 5xx errors, a Lambda timeout is causing transaction failures, a database connection pool is exhausted. The runbook documents the diagnostic steps and the fix. The audience is the on-call engineer.

PCI DSS Requirement 12.10 addresses incident response from a different angle: it requires evidence that the organisation can detect, contain, investigate, and recover from security incidents affecting cardholder data โ€” and that these capabilities are documented, assigned, and tested. A payment failure is potentially a security incident. A runbook that treats it only as an operational event fails the compliance requirement.

The distinction that matters: An operational runbook asks "how do we restore service?" A PCI DSS-compliant incident response procedure asks "was cardholder data exposed during this failure, and what do we do if it was?" Most runbooks answer the first question. Very few answer the second.

PCI DSS Requirements a Runbook Must Satisfy

Sub-reqRequirementWhat the runbook must containQSA verification method
12.10.1Written incident response plan covering all system componentsNamed roles and responsibilities for payment failure scenarios; escalation path to CISO/DPO; acquirer notification procedureReview: role names match current org chart; notification contacts are current; plan version is dated within 12 months
12.10.2Incident response plan reviewed and tested at least annuallyEvidence of tabletop exercise or simulated payment failure scenario within 12 months; test outcome documentedReview: test date confirmed; participants named; gaps identified in test documented with remediation status
12.10.3Specific personnel available 24/7 to respond to suspected security breachesOn-call rota or escalation matrix showing coverage for payment failure/breach scenarios; contact details currentReview: on-call schedule covers 24/7; contacts verified as current employees with correct roles
12.10.4Personnel trained on incident response proceduresTraining completion evidence for all personnel named in the runbook; training dated within 12 monthsReview: training records match personnel list in runbook; content covers payment breach scenarios
12.10.5Monitoring and alerting for all critical system components in CDECloudWatch alarms or SIEM alerts referenced in runbook; alert thresholds documented; alert-to-runbook linkage confirmedReview: alert names in runbook match deployed alarms; thresholds are quantified not qualitative
12.10.7 โ˜…Incident response procedures for suspected PAN exposureExplicit procedure for suspected/confirmed CHD exposure: contain, preserve evidence, notify acquirer, engage PFI if requiredReview: CHD exposure scenario is a named scenario in the runbook, not implied; PFI contact is listed

โ˜… New or materially changed in PCI DSS 4.0.1

The Anatomy of a QSA-Ready Payment Failure Runbook

A runbook that satisfies both operational and compliance requirements has a specific structure. The following sections are required โ€” sections that are standard in engineering runbooks are noted, sections that are compliance-specific are marked.

01
Incident classification matrix [Compliance-specific]
Before any diagnostic step, the runbook must classify the incident. Payment failures fall into at least three categories: (1) operational failure โ€” service unavailable, no CHD exposure; (2) suspected security incident โ€” anomalous behaviour that may indicate CHD exposure; (3) confirmed breach โ€” CHD confirmed as exposed. The classification determines the response path. Most runbooks have only one path โ€” the operational path.
02
Immediate containment steps [Compliance-specific]
If the failure is classified as a suspected security incident, the runbook must specify containment actions before diagnosis. Containment typically means: isolate affected Lambda functions or EC2 instances from the CDE network, revoke active IAM sessions, preserve CloudTrail logs and VPC Flow Logs for forensic use, and suspend automated scaling that could spin up new compromised instances. These steps are distinct from โ€” and must precede โ€” service restoration.
03
Evidence preservation checklist [Compliance-specific]
PCI DSS and acquirer contracts require forensic investigation of suspected breaches. Investigation requires logs. The runbook must include a checklist of evidence to preserve before any remediation action that could destroy it: CloudTrail log export to immutable S3, VPC Flow Log archive, application log preservation, Lambda function version snapshot, RDS point-in-time snapshot. Remediation steps that follow must note which evidence-preservation steps must precede them.
04
Diagnostic procedures [Standard engineering content]
The standard engineering runbook content: check AWS Service Health Dashboard, review CloudWatch metrics for the payment processor integration, inspect Lambda error logs, verify RDS connection pool, check downstream API response codes. This section is usually well-written in engineering runbooks โ€” the compliance gaps are elsewhere.
05
Escalation and notification matrix [Compliance-specific]
The runbook must specify: who is notified at each classification level, within what timeframe, and with what information. For a suspected breach: CISO within 1 hour, DPO within 4 hours (GDPR obligation), acquirer within 24 hours (acquirer contract obligation), card brands within 24 hours of confirmed or suspected CHD exposure (Visa CDRP / Mastercard ADCR obligations). Contact details โ€” names, direct lines, email addresses โ€” must be current. A runbook that lists "CISO" without a name or contact is not compliant.
06
Remediation and recovery steps [Standard engineering content]
Service restoration steps. These are standard engineering runbook content. The compliance requirement here is that remediation steps do not destroy forensic evidence โ€” covered by the evidence preservation checklist. Additionally, the runbook should note that service restoration does not close the incident โ€” the incident remains open until root cause is confirmed and breach determination is made.
07
Post-incident review procedure [Compliance-specific]
Req 12.10.6 requires the incident response plan to be updated following any security incident. The runbook must specify: when the post-incident review occurs (within 72 hours), who attends, what outputs are required (timeline, root cause, control gap identification, runbook update), and how updates are approved and versioned. Without this, the runbook itself becomes stale evidence.

AWS-Specific Evidence Requirements During a Payment Incident

When a payment failure escalates to a suspected breach, a PFI (PCI Forensic Investigator) will request the following AWS evidence. The runbook should specify how each is preserved:

Evidence typeAWS sourcePreservation actionRetention required
API call historyCloudTrail (management + data events)Copy trail to immutable S3 with Object Lock; disable log rotation for incident period12 months minimum; 3 months immediately accessible
Network traffic recordsVPC Flow LogsArchive to S3; note: Flow Logs have a delay of up to 15 minutes โ€” archive immediately to prevent gapDuration of incident window plus 30 days
Application logsCloudWatch LogsExport log groups covering incident window to S3; set retention to max during incidentDuration of incident window plus 90 days
Database query logsRDS audit logs (if enabled), CloudTrail data eventsConfirm RDS audit logging was enabled before incident โ€” if not, this evidence does not existDuration of incident window
Lambda function versionsLambda version history, ECR image digestsRecord function version ARNs deployed during incident window; do not delete versionsUntil investigation closed
IAM activityCloudTrail: IAM eventsExport CloudTrail filtered to IAM events for incident period and 30 days priorUntil investigation closed
WAF logsCloudFront/WAF logs in S3Confirm WAF logging was enabled; archive logs for incident windowDuration of incident window plus 30 days

Critical gap: RDS audit logging and CloudTrail data events are not enabled by default. If they were not enabled before the incident, the database query history and S3 access history for the incident window do not exist. A PFI investigation without this evidence significantly increases the difficulty of determining whether CHD was accessed. Enable these before an incident, not during.

Common QSA Findings When Reviewing Payment Runbooks

  • No incident classification: The runbook has one response path โ€” operational. There is no branch for "suspected security incident." A QSA cannot confirm that the entity would escalate a payment anomaly to a security incident response.
  • Contact details are role-based not named: "Notify the CISO" does not satisfy Req 12.10.3. The CISO must be named, with a direct contact number and backup contact. Role-based references become useless when the role is vacant.
  • Acquirer notification absent: Most runbooks reference notifying "the payment processor." The acquirer is a separate entity and has a separate โ€” and stricter โ€” notification timeline. QSAs will flag the absence of acquirer contact details.
  • Evidence preservation not mentioned: Runbooks that proceed directly from "incident detected" to "restore service" give no instruction on evidence preservation. A QSA treating this as a security incident finding will note that evidence destruction is possible under the current runbook.
  • No annual test evidence: The runbook exists but there is no evidence it was tested in the last 12 months. A document that has never been exercised does not satisfy Req 12.10.2. Tabletop exercise records must be produced.
  • Version not controlled: A runbook without a version number, author, and last-reviewed date cannot be confirmed as current. QSAs will request the version history and flag runbooks that have not been reviewed within 12 months.

Interactive assessment

Apply this in your assessment

This assessment maps directly to the controls in this article. Use it to generate gap evidence, cardholder data flow diagrams, or scoping documentation for a PCI DSS assessment โ€” free with a Sync Your Cloud account.

Failure Playbook Generator
Generate a QSA-ready payment failure runbook โ€” structured to satisfy PCI DSS 4.0.1 Requirement 12.10, with escalation matrices, rollback procedures, and evidence fields formatted for audit review.