Postmortem
Service availability incident on August 27, 2025
Root cause analysis and timeline of the service performance degradation on August 27, 2025.
Incident summary
Date and timeAug 27, 2025, 11:00 (UTC+3)
CauseAn uncontrolled spike in load on one of our Redis servers brought it down and degraded service performance
SeverityHigh
Duration31 minutes in total (two episodes)
Affected systemsapp.pachca.com: web, mobile and desktop apps
Affected usersAll users
What happened
On August 27, 2025, Pachca's performance degraded twice after one of our Redis servers went down. The first episode lasted about 10 minutes, the second about 21 minutes. No data was lost.
We apologize for this incident. Below are the timeline, root cause analysis and the steps we're taking.
Timeline
All times are UTC+3.
- 11:00 – We started getting a large number of reports that the service was slow.
- 11:01 – The team started investigating and narrowing down the problem.
- 11:10 – Full access to the service was restored.
- 11:53 – We started getting a large number of reports about problems sending messages, receiving login codes and creating chats.
- 11:54 – The team started investigating and narrowing down the problem.
- 12:14 – Full access to the service was restored.
Root cause chain
- Why did performance degrade? One of our Redis servers went down under load, which slowed down request processing.
- Why did Redis go down? An uncontrolled spike in load exceeded the server's capacity.
- Why were there two episodes? The root cause wasn't fully fixed after the first recovery, so performance degraded again.
What we're doing to prevent this from happening again
Immediate actions
- Redis recovery. We made changes to the Redis state to stabilize it.
- RateLimiter policy change. We updated our limits to prevent uncontrolled load spikes.
Long-term actions
- Focus on stability. We decided to focus on improving service stability in fall 2025.
- Performance optimization. We're continuing to optimize infrastructure performance.