SignalOps · Critical IMS
From noisy infrastructure signals to actionable incidents.
An asynchronous incident management system with durable ingestion, failure debouncing, audit storage, and root-cause-gated incident closure.
The problem
Repeated infrastructure failures can flood operators with duplicate signals, while slow downstream databases can turn an ingestion endpoint into a bottleneck.
How it comes together
The API validates and rate-limits signals, publishes to Redpanda, and returns before persistence completes. Workers separate raw audit data, incident state, hot-path cache, and time-series aggregates.
Follow the flow.
Kafka-compatible ingestion and asynchronous workers.
Redis-backed rate limiting and debounce windows.
Separate raw audit payloads and incident workflow records.
Root-cause requirements for closure, plus MTTR calculation.
Why this approach?
Commit message offsets after persistence succeeds. Transient retries and broker lag provide a defined backpressure path when downstream storage slows.