On March 24, 2026, at 19:27 UTC, one network shard experienced intermittent connectivity affecting a subset of customers. The affected users may have experienced elevated latency and temporary error responses related to their subscription requests. The instability was caused by an atypical surge in message volume within a shared processing environment that had improperly configured resource limits. This led to high resource utilization and triggered automated system restarts. PubNub Engineering resolved the issue by implementing the proper limits after expanding infrastructure capacity to accommodate the increased load. Service was fully stabilized once the environment was tuned to the new traffic profile.
Infrastructure Tuning: Adjusted automated scaling parameters to provide greater headroom for rapid traffic fluctuations.
Enhanced Traffic Management: Deployed refined monitoring heuristics to better isolate and manage high-volume traffic patterns without impacting shared resources.
Dynamic Resource Allocation: Accelerating the rollout of enhanced vertical scaling technology to allow individual processing nodes to adapt more fluidly to demand spikes.
Operational Coordination: Strengthening internal protocols for high-capacity events to ensure large-scale traffic shifts are proactively transitioned to dedicated environments.