A customer clicked "Export to CSV" on 380,000 records. The server loaded them all into memory, OOMed, and took the API offline for every tenant. The fix was straightforward. The real question is why nobody caught it sooner.
A client's API had been running Node 16 for two years past end-of-life. When a critical OpenSSL vulnerability dropped, the "we'll upgrade next quarter" plan collapsed into a three-week fire drill. The upgrade itself wasn't hard. Undoing two years of drift was.
A Node.js service kept getting OOMKilled, but only on Tuesdays. The batch job was a red herring. The real problem was a lazy cache eviction strategy that nobody thought to question.