Яндекс.Метрика
Postmortem

Incident postmortem

Root cause analysis and timeline of the technical outage on December 22, 2025.

Incident summary

Date and timeDec 22, 2025, 13:23
CauseDelayed failover to backup infrastructure
SeverityHigh
Duration9 min of downtime + 52 min of slow performance
Affected systemsAPI server (api.pachca.com)
Affected usersAll users

Incident details

Resolution timeline:

  • 13:23 – Monitoring systems detected that the service was down.
  • 13:25 – The engineering team started diagnosing and fixing the problem.
  • 13:26 – Cause identified: the primary infrastructure had been shut down. We started a restart on backup capacity.
  • 13:32 – The service was back online. We saw performance issues, most noticeably in the iOS app.
  • 13:34 – We started going through incoming user reports.
  • 14:19 – We found an error in the network configuration.
  • 14:24 – Performance was fully restored.

Root causes:

  • Our cloud provider carried out scheduled maintenance, which shut down our primary infrastructure. The backup capacity we'd prepared in advance wasn't activated automatically when the shutdown happened.
  • After access was restored, we found an error in the network configuration that was slowing the service down.

Actions taken:

  • We improved how we prepare for scheduled maintenance that requires failing over to backup infrastructure.
  • We added extra checks to make sure the network configuration is correct after a failover.

Lessons learned and recommendations:

  1. Run regular failover drills to backup infrastructure to make sure it's ready.
  2. Automate the activation of backup capacity during scheduled maintenance.
  3. Strengthen configuration validation after infrastructure failovers.