Яндекс.Метрика
Postmortem

Service availability incident on October 16, 2024

Root cause analysis and timeline of the service outage on October 16, 2024.

Incident summary

Date and timeOct 16, 2024, 05:24 (UTC+3)
CauseLoss of connectivity between our cloud provider's data centers
SeverityHigh
Duration4 hours 29 minutes of full downtime
Affected systemsapi.pachca.com, all platforms
Affected usersAll users

What happened

On October 16, 2024, at 05:24 Moscow time, Pachca went completely down because our cloud provider's data centers lost connectivity with each other. The problem was on the provider's side and outside our control. The service was restored at 09:53 once the provider fixed the issue.

No data was lost.

We apologize for this incident. Below are the timeline, root cause analysis and the steps we're taking.


Timeline

All times are UTC+3.

  • 05:24 – First signals that the service was unavailable.
  • 05:41 – We started investigating the cause.
  • 06:24 – We got in touch with the provider for details.
  • 07:11 – The provider isolated the problem in its cloud infrastructure.
  • 07:41 – The provider declared a public incident, acknowledging a possible large-scale outage.
  • 07:51 – The provider started a configuration sync on its side (expected to take 40–60 minutes).
  • 08:30 – With the provider's recovery dragging on, we started failing over to a database replica.
  • 09:53 – Full access to the service was restored after the provider fixed the issue.

Root cause chain

  1. Why did the service go down? The API servers lost their connection to the database and other infrastructure components.
  2. Why was the connection lost? Our cloud provider's data centers lost connectivity with each other.
  3. Why did recovery take more than 4 hours? The problem was entirely on the provider's side, so we depended on how fast they responded. In parallel, we started failing over to a database replica, but the provider fixed the issue first.

What we're doing to prevent this from happening again

Faster recovery

  • Database replica failover runbook. We're finishing and rolling out a runbook that will cut maximum downtime in similar incidents to 15 minutes.

Long-term resilience

  • Diversified infrastructure. We're bringing in capacity from a second, independent provider to make the service more fault-tolerant.