Webhook Ingestion Service
Retries happen. Duplicate side effects should not.
A Go reliability case study fixing duplicate call events, drifting account totals, lost recording jobs, and unsafe cache writes with PostgreSQL transactions and a durable outbox.
The problem
At-least-once webhook delivery exposed a SELECT-then-INSERT race, partial database writes, concurrent map writes, and recording tasks tied to a canceled HTTP request context.
How it comes together
A unique event ID and a single transaction protect the event insert, call upsert, aggregate increment, and recording-job enqueue. A worker claims durable jobs with expiring leases and SKIP LOCKED. Startup restores cached statistics from the database.
Follow the flow.
Treat duplicate event IDs as successful no-ops, with all durable changes committed atomically.
Persist recording work in a PostgreSQL outbox so pending jobs and abandoned leases survive restarts.
Retry recording failures with structured logs, and atomically complete the call update and job deletion.
Protect cache writes with a mutex and restore durable account totals at startup.
Why this approach?
Keep the correctness guarantee in PostgreSQL, where the affected state already lives. Redis deduplication would introduce a second source of truth and a crash window between the two systems.
An incident-repair engineering exercise. The solution discusses partitioning, batch processing, and sharded counters for a future 10,000-webhooks/second design; that throughput is not a measured result.