Яндекс.Метрика
Postmortem

Service availability incident on August 16, 2024

Root cause analysis and timeline of the service outage on August 16, 2024.

Incident summary

Date and timeAug 16, 2024, 12:06 (UTC+3)
CauseHuman error: a mistake during deployment and in the container orchestration settings
SeverityHigh
Duration47 minutes of full downtime + 12 minutes of slow performance
Affected systemsapi.pachca.com, all platforms
Affected usersAll users

What happened

On August 16, 2024, at 12:06 Moscow time, Pachca went down because of a mistake during deployment and in the container orchestration settings. Our primary database became unavailable, which brought the API to a complete stop. The service was restored 47 minutes later, followed by 12 more minutes of slowdowns.

No data was lost.

We apologize for this incident. Below are the timeline, root cause analysis and the steps we're taking.


Timeline

All times are UTC+3.

  • 12:06 – First signals that the service was unavailable.
  • 12:07 – The team started investigating and working on a fix.
  • 12:08 – The problem was traced to the API. Recovery work began.
  • 12:20 – Root cause identified: a mistake during deployment and in the container orchestration settings.
  • 12:53 – The service started to recover.
  • 12:57 – Access to the service was fully restored, though some slowdowns remained.
  • 13:05 – Performance was fully restored.

Root cause chain

  1. Why did the service go down? The primary database became unavailable to the API servers.
  2. Why did the database become unavailable? A mistake during deployment and in the container orchestration settings disrupted the cluster.
  3. Why did the incident affect all replicas? Still under investigation. We'll update this section as we learn more.

What we're doing to prevent this from happening again

Immediate actions

  • Cluster rebuild. We rebuilt the machine cluster for the API services to restore stable operation.

Long-term actions

  • Server cluster redundancy. We're looking into adding redundancy to make deployments more reliable and prevent similar failures in the future.