Яндекс.Метрика

Incident history

A full record of past Pachca incidents and outages.

2026

The database slipped into a degraded state with no visible increase in load. Pachca was down for 3 hours 15 minutes on all platforms. Service was restored after a database restart. Full postmortem.

A mass drop of WebSocket connections set off a flood of reconnects that overloaded the service. Downtime and degraded performance lasted about 4 minutes. Full postmortem.

Our cloud provider shut down our infrastructure because of a billing error on their side. 11 minutes of full downtime plus 58 minutes of partial downtime for the desktop apps. Full postmortem.

2025

After an SSL certificate renewal and a web server restart, connections were reset. The resulting surge of requests exhausted our resources and took the service down for 5 minutes.

An uncontrolled spike in load on one of our Redis servers brought it down and degraded service performance. Recovery was automatic. To prevent similar incidents, we changed our rate-limiting policy. Full postmortem.

An update that optimized @mentions, released on August 11, 2025, put heavy load on an old query and caused 43 minutes of slowdowns. Full postmortem.

A wave of internet blocking in Russia temporarily hit the infrastructure of our cloud provider, Selectel. Uploading, opening and downloading files and images was affected.

A DBMS-level error made the service unavailable for 15 minutes. We'll fix the underlying DBMS bug by upgrading to a newer version.

2024

The WSS server went down, which triggered a flood of API requests. The API was overloaded, causing 44 minutes of downtime.

Our provider's data centers lost connectivity with each other, leaving Pachca unavailable for 4.5 hours in the morning, Moscow time. Full postmortem.

Our primary database became unavailable due to human error: a mistake during deployment and in the container orchestration settings. The API was down for an hour. Full postmortem.

From roughly 12:00 to 14:00, Pachca ran slower than usual because a large number of users suddenly changed how they used the app. We made changes to search, which fixed the API slowdown.

A bug in an update caused slow backend response times for 2 minutes.

A backend configuration change caused an error, which we fixed quickly. The service was down for 3 minutes, and search took a few more minutes to recover.

A faulty release caused the backend to return errors for 13 minutes.

2023

An SSL certificate expired after the auto-renewal system failed. Downtime: 25 minutes.