Operational playbook · Operations & review

What automation should stop during an ad-platform incident

Platform incidents, delayed reporting, and traffic anomalies can make automated rules amplify bad data. Data freshness, anomaly thresholds, and degraded modes matter more than continued scaling.

Editorial synthesisReviewed 2026-08-01Verify before action
DECISION BRIEF

Three points to take away

  1. 01

    Pause irreversible actions when data is stale.

  2. 02

    Retain both platform status and internal logs.

  3. 03

    Restart in stages rather than all at once.

01

Incident controls

When reporting delay, error rate, or key metrics exceed thresholds, stop budget increases and bulk changes, switch to read-only monitoring, and notify the owner.

02

Recovery checks

Confirm data backfill, payment, and landing pages before restoring small-budget units. Record all incident-period suggestions so stale actions are not executed later.

VERIFY BEFORE ACTION

Verification checklist before action

  • Simulate stale data and API failures to confirm graceful degradation instead of continued execution.

  • Run a bounded trial with non-sensitive samples and retain successes, failures, and human corrections.

  • Before wider use, name an owner, data boundary, stop condition, and review date.

RELATED NOTES

More notes on this topic