← All posts

Your monitor says the site is up. Your server is dying.

October 7, 2026 · 7 min read · The CleverQA Team

Disaster always happens while you're busy with something else.

That's not even the worst part. The worst part is that everything looks like it's working, right up until it isn't. The dashboard is green. The checks are passing. You have no reason to go and look, so you don't, and the problem gets a seven-hour head start on you.

What still bothers me about this particular one isn't the outage. Outages happen, you fix them, you move on. It's the twenty minutes I spent arguing with a man on the other end of a support email, politely, professionally, while I had a dashboard open in the other window telling me everything was fine.

All green. Every check passing. And he was right and I was wrong.

What I didn't know while it was happening

Here's the timeline, reconstructed afterwards, which is the only way you ever get a clean timeline.

14:02. I deployed. Dependency bump, one new endpoint. Maybe forty lines of diff. I remember thinking it was a boring release, which is a thought I now treat as a warning sign.

14:09. Memory on the API box starts climbing. About 40MB a minute. Steady. If you'd graphed it you could have put a ruler against it.

I didn't graph it. Nobody graphs memory on a Tuesday afternoon.

16:30. It crosses 70%.

Seventy percent is fine. Seventy percent is a number you'd see and not think about. I want to be clear that there was nothing to notice here even if I'd been looking, which I wasn't, because I was doing something else entirely by then and the deploy had gone out two and a half hours earlier and in my head it was finished.

19:45. The box starts swapping. Response times go from 180ms to about 400.

Still fine, as far as the monitoring was concerned. The timeout was three seconds. Four hundred milliseconds is not three seconds. Green.

21:12. The kernel runs out of patience and kills something.

21:19. A customer emails me.

Every single uptime check in those seven hours passed. The monitoring did exactly what it was built to do. It just wasn't built to tell me what I needed to know.

The thing about green dashboards

I think what got under my skin was how confident it looked.

A red dashboard is honest. A red dashboard says: something is wrong, go look. But a green dashboard that's wrong doesn't just fail to help you, it actively argues against you. It takes the side of the problem. You sit there with a customer telling you the checkout is broken and a screen telling you it isn't, and for a few minutes you genuinely consider that maybe the customer is confused.

They weren't confused. They were seven hours ahead of me.

So I did the obvious thing, and it also didn't work

Alert at 85% memory. Alert at 90% CPU. Done. I felt quite good about this for a while.

Go back to the timeline and run the numbers, though. A leak climbing 40MB a minute crosses 85% somewhere around 20:40. Thirty minutes of warning.

Thirty minutes sounds generous until you're actually in it. It's twenty to nine at night. Your phone buzzes. memory > 85% on web-01. You're not at your desk — nobody is at their desk at 20:40 — so you find a laptop, you VPN in, you SSH to the box, you run top, and you see:

node   2.1GB
node   1.4GB
node   890MB
node   340MB

Four Node processes. Which one is the API? Which one is the worker? You genuinely cannot tell, because you named none of them, because why would you, and now it's 20:51 and you've burned eleven of your thirty minutes finding out something you could have known in advance.

There's a worse version of this that took me much longer to learn. Two services on one host. One of them goes greedy. The other one, which is innocent, starts timing out because it's being starved. Your alerts fire on the innocent one, because that's the one that's visibly failing, and you spend twenty minutes reading its logs looking for a bug that isn't there.

The victim is always louder than the culprit. I've lost whole evenings to that.

Three questions

What I actually needed, that night and every night like it since, comes down to three things:

Which service. Not which box. Boxes don't have bugs. checkout-api is holding 3.2GB is something you can act on. web-01 is at 88% memory is the beginning of an investigation you're conducting in the dark.

Since when. This one is quietly the most valuable and almost nobody collects it. A number tells you where you are. A curve with a start time tells you what happened. Memory started climbing at 14:09; I deployed at 14:02. That's not a server problem anymore. That's a diff, and diffs are small and readable and I know how to deal with them.

How fast. Forty megabytes a minute and two megabytes a minute are completely different situations. A threshold alert reports them identically.

None of this is exotic. All of it is sitting in /proc the entire time. It's just that an HTTP check is standing outside the building, knocking on the door, reporting back that someone answered.

What we built, and what we refused to build

So that's the gap CleverQA's agent fills. I'll be specific, because vague is how this category usually talks.

It's a stock OpenTelemetry Collector running our config. Not something exotic we wrote ourselves. It samples every five seconds and maps processes onto the service names you actually use, so you get checkout-api instead of PID 8814, and four Node workers belonging to one service get added together instead of showing up as four mysteries.

The rule isn't a threshold. Memory above 80% and climbing at least 3% a minute, fitted over a three-minute window. Both conditions, which means the service that sits at 87% forever never pages you, and neither does a spike at 40%.

On that Tuesday, the alert would have looked roughly like this, around two in the afternoon:

🔴 checkout-api — high memory on web-01

checkout-api is using 512MB memory and climbing (+3.5%/min) on web-01.
Next: worker (310MB), api-gateway (120MB).
At risk of saturation — trending up, not a certainty.
Rule: memory > 80% and rising.

Now look at what that doesn't say.

It doesn't tell you the box dies at 21:12. We could print a number like that. We decided not to, because a slope fitted over three minutes cannot honestly support a countdown and we'd rather be less impressive than wrong. There's a longer post in that decision somewhere.

What it does tell you is which service, how fast, and what else is on the box. At two in the afternoon, that's enough.

And then it becomes a defect — with the evidence, an owner, a lifecycle, the AI's read on what's probably happening, pushed into Jira or Linear or GitHub or Azure DevOps. Not a message in a channel that scrolls past while you're asleep.

You're going to object to the agent

Yes. It means running something on your server. I'd hesitate too.

So, honestly: it reads process-level resource usage and sends it out. It doesn't read your application data, your environment variables or your database. The process scraper is configured so that command lines, PIDs and process owners never leave the host, and there's a whitelist on our side that drops them again if they somehow did. Only the executable name travels.

But it's still a process. It needs updating. It's one more thing with the capacity to break at an inconvenient moment, and I'm not going to pretend otherwise.

If you're entirely serverless, none of this applies to you and you should go read something else.

If you're not, the real choice isn't between an agent and no agent. It's between finding out at 14:09 and finding out at 21:19, from a customer, with a green dashboard open in the other window, arguing.

Dead is expensive. Dying is cheap.

Same leak. Same service. Same fix. The only variable is when you hear about it.

Caught at 21:12: an incident, a rollback under pressure, a status page nobody updated in time, an apology, and some number of people who tried your product for the first time that evening and will not be trying it again.

Caught at 14:20: you restart something during office hours and look at a diff from twenty minutes earlier. Nobody notices. There is no story. There's nothing to write up, which is the whole point — the best version of this ends with no post at all.

I still think about that support email though.


CleverQA watches from inside the box as well as outside it. Free plan, no card: Start free