History
Aug 2026 - Oct 2026
Outage
OpenMail API timing out
Affected services:
OpenMail API
Sending emails
Console
Postmortem

Summary

At 23:35 EEST on 24 September the OpenMail API stopped responding. It recovered at 07:58, froze again 20 seconds later, and was fully stable from 08:24. Total downtime: about 8 hours 50 minutes. A single inbound email with a malformed spreadsheet attachment caused it.

Impact

Every API request timed out for the whole window. Sending mail, the console, and WebSocket connections all failed.

Inbound mail was delayed, not lost, for nearly everyone: messages queued upstream and were delivered within minutes of recovery. Two inbound messages from late in the window exceeded the retry period and could not be recovered. If you're missing an email from between 23:35 and 08:24 EEST, contact us and we'll trace it.

Root cause

An email arrived with a 6.9 MB spreadsheet. One sheet declared 968,194 rows and 871,688 cells; only 5,152 held data. The rest were empty cells created by formatting whole columns.

Our attachment text extractor loads spreadsheets fully into memory before reading them. On this file it never finished. It ran on the API's main thread, and Node.js is single-threaded, so while it spun, nothing else ran: no requests, no health checks, no timers. Our 15-second timeout depended on one of those timers, so it never fired.

When we restarted the API, the queue handed the same email straight back and froze the new process too.

Resolution

  1. 07:57: restarted the API.
  2. 08:23: removed the stuck email from the queue; stable from 08:24.
  3. 10:03: attachment parsing now runs in an isolated thread with hard time and memory limits. Over-limit attachments are skipped; the email still delivers.
  4. 10:06: redelivered the triggering email.
  5. 10:20: deployed a watchdog that restarts the API automatically if it stops answering.

Lessons learned

What went well

Once diagnosed, the fix, redelivery, and watchdog shipped within three hours. Inbound mail queued rather than bounced.


What could be improved

A timeout that shares a thread with the work it guards is not a timeout. Health checks only ran at deploy time. Detection took eight hours.

Action items

  1. Isolate attachment parsing with hard limits
  2. Automatic restart on unresponsive API
  3. Extend inbound retention so delayed mail is never dropped
  4. Alert on a stuck job before it stalls the queue


Sep 26, 8:51 PM
Resolved
We've investigated the issue and posted a post-mortem for a read.
Sep 26, 8:49 PM