The story
A 504 From Behind the Proxy
A SaaS client's slow report page kept timing out with a 504, and we spent two days blaming the wrong layer. The real answer was a query nobody had looked at.
By Marcus Feld, Infrastructure Editor · 6 min read
·
Disclosure: Some links on this page are affiliate links — if you sign up through one, we may earn a commission at no extra cost to you. It never changes our ratings, rankings or verdicts: we don't sell hosting and take no pay-for-placement.
One of my longer-running clients is a small SaaS startup. They run a dashboard for their customers, and I look after the servers underneath it. Their reports page pulls together a lot of numbers, and it had worked well for months. Then the founder messaged me on a Tuesday evening to say a handful of customers were seeing 504 Gateway Timeout, always on the biggest accounts.
A 504 is a funny error. It doesn’t mean the page is broken. It means one server, sitting in front, asked another server behind it for an answer and gave up waiting. In their setup, a proxy sits in front of the application, and that proxy has a patience limit.
The wrong layer
Because the error said gateway, I decided the proxy was the problem. I’ve spent years on call watching proxies misbehave, so my working theory came easily: the timeout was set too short. I raised it from sixty seconds to ninety.
The report still timed out for the biggest customer, just later. So I raised it to a hundred and twenty. Same thing. Two days went by and I was turning a dial that wasn’t connected to anything.
I’ll be honest, it was habit and a little pride that kept me turning it. The proxy was mine, so it felt like the place to look. It was comfortable, and it was wrong. I’ve been doing this a long time, and I still skipped the first rule.
The scramble
On the third morning I did what I should have done on the first. I skipped the proxy and asked the application directly. I connected to the app on its own port from inside the server and requested the report, with a stopwatch running. It took a hundred and forty seconds to answer. It wasn’t a network problem at all. The app itself was crawling.
So the proxy had been doing exactly its job. It waited, ran out of patience, and reported a timeout.
I turned on the database’s slow query log, and there it was. One query behind the report joined three large tables and filtered on a column that had no index. For small accounts it scanned a few thousand rows. For their biggest customer it scanned millions, every single time the page loaded.
The fix
We added an index on that column, and the query went from two minutes to a little under a second. The client’s developer and I also stopped the page from running the report live every time. Now it builds the summary in the background and shows the stored result, with a small note saying when it was last refreshed.
I put the proxy timeout back to its original value. A long timeout is a bit like leaving a tap running. It feels like you’ve solved the flooding, but really you’ve just let the bath overflow more slowly.
The lesson that stung most was how long I’d trusted the error message’s vocabulary. Gateway sounded like the gateway’s fault. In reality, a 504 just tells you where the waiting stopped, not where the delay began.
What I do on day one now
If I replayed those three days, I’d skip the first two. The moment a gateway error shows up, I run a direct request to the application and time it. It takes two minutes. It tells you straight away whether the proxy is fine, and it saves you from rewriting a setting that was never the problem.
I also keep a rule from my years on the support desk, from the other side of the ticket. The person who wrote to us in a hurry, certain the fault was ours, was usually looking at the wrong layer. Being that person for a change was a useful reminder to check my own assumptions first.
There’s a wider point here that I give every client now. When something is slow, there are usually three suspects in a row: the front door, the application, and the database behind it. Most people only ever blame the first one, because it’s the one that speaks. The error message comes from the proxy, so the proxy takes the blame. I’ve seen the same thing with load balancers, CDNs and firewalls. The component that reports the problem is rarely the component that caused it.
We’ve since added a simple alert that warns us when any page takes longer than ten seconds to respond. The founder would rather hear about it from a graph than from their biggest customer, and so would I.
A few things I’d tell you
- Test the application directly. If it’s slow on its own, the proxy is innocent.
- Don’t keep raising timeouts. It’s a painkiller, not a cure.
- Check for slow database queries. Missing indexes are the most common cause of slow reports.
- Look at timing details, such as how long the proxy waited and whether the app responded at all.
- Move heavy reports to the background and show users saved results.
If you’re not an engineer, the same idea applies. When something times out, ask which step was slow, rather than which step complained. They are often not the same one.
To sum it up
- A 504 Gateway Timeout means a front server gave up waiting for the application behind it.
- Raising timeouts hides a slow request; it doesn't fix one.
- Test the app directly, bypassing the proxy, to learn which layer is slow.
- Slow reports are very often slow database queries.