← All posts

Why we never deploy on Fridays

October 7, 2026 · 7 min read · The CleverQA Team

We have a rule. Nothing goes out on a Friday.

People laugh at that rule. It sounds superstitious, like not walking under ladders. And it is a bit superstitious, honestly. But every team that has one got it the same way, which is by learning it the hard way, once, and never needing to be told twice.

Here's where ours came from.

The Friday

The release was ready. Genuinely ready — this wasn't a case of pushing something half-finished out the door because the sprint was ending. It had been tested. It had been reviewed. Someone had clicked through the main flows and everything did what it was supposed to do.

I remember feeling good about it. That's the part I'd forgotten until I started writing this down. It wasn't a nervous deploy. It was the kind where you watch the pipeline go green, close the laptop, and feel like you've earned the weekend.

So I closed the laptop and went to have a weekend.

The question I didn't ask

Before I shut everything down, there was one question I didn't ask myself, and it's the entire reason this story exists.

Is anything watching this?

Not the homepage. The homepage was fine, the homepage is always fine. I mean: if one of the endpoints behind it stopped answering at nine o'clock on a Friday night, with nobody at a keyboard, would anything anywhere make a noise about it?

No. Nothing was watching. There were no monitors on the individual services. If something broke quietly, it would stay broken, quietly, for as long as it took a human being to notice.

I knew this. I'd known it for months. It was on a list somewhere, below things that felt more urgent.

The weekend, in two parts

Part one was lovely. I'd recommend it.

Part two started with my phone ringing on Sunday morning, which is not when people call you with good news.

It was the CEO. He was not calm. He'd had customers contacting him directly, which is the worst version of finding out about a problem, because by the time it reaches the person at the top it has already been through several other people who are also now annoyed.

Something had been broken since Friday evening. Not everything — if everything had gone down, somebody would have spotted it within the hour. One flow. One of those paths where things fail politely, return something plausible, and nobody notices immediately.

It had been broken for roughly thirty-six hours.

I'll skip the conversation. You can imagine the conversation.

Sunday, starting from nothing

Here's the part I want to describe properly, because it's the part the rule is really about.

I didn't start by fixing the bug. I started by trying to work out what the bug was, with essentially no information.

No alerts, because nothing was configured to alert. No timeline, because nothing had been recording. I couldn't tell you when it started. I couldn't tell you whether it had been failing constantly or intermittently. I didn't know whether it was the deploy at all — I assumed it was, because the timing fit, but assuming is not the same as knowing and I spent real hours chasing that difference.

So it was logs. Scrolling, grepping, trying to reconstruct Friday evening from fragments on a Sunday afternoon. Trying to work out whether a particular error had been happening before the release or only after, which is the kind of question that takes four minutes to answer if you have the data and most of a day if you don't.

I found it eventually. The actual fix was small. It usually is.

Debugging first, fixing later — and the first part took about six times longer than the second.

The lesson is not about Fridays

The rule we made afterwards was "no Friday deploys," and we've kept it, and I think it's a reasonable rule.

But if I'm honest, it's the wrong lesson. Or rather, it's a lesson that treats the symptom.

Friday wasn't the problem. The deploy wasn't really the problem either — releases break things, that's just true, and no amount of testing gets you to zero. The problem was that something went wrong and nobody and nothing noticed for thirty-six hours, and then when we finally did notice, we were starting from a blank page.

Deploying on a Tuesday would have meant somebody caught it within a few hours, probably by accident. That's not a monitoring strategy. That's just having more people awake.

The rule exists because we didn't trust our ability to find out. That's the actual problem, and "don't deploy on Fridays" is what you do instead of solving it.

So I went and solved it

That weekend is a reasonable chunk of why CleverQA exists.

Not in a dramatic origin-story way. It's more that I'd had versions of that Sunday several times across different projects, and at some point I got tired enough of it to build the thing I kept wishing I had at 11am on a Sunday with no data and a furious phone call behind me.

Here's what it does, told as the same Friday, with it running.

Something is watching every service, not just the front door. Each endpoint, each API, the background jobs, the certificates, the database queries that matter. Not "is the site up" but "is each of the things the site depends on still doing its job." The flow that broke on that Friday had a specific endpoint behind it. That endpoint would have had a check on it.

The alert finds you, wherever you are. Email, SMS, or a webhook into whatever you already use, with escalation if nobody acknowledges it — so it doesn't sit unread somewhere until Monday morning. Friday 21:04, not Sunday 10:30.

It also watches the server itself. There's an optional agent that sits on your machine and watches resources per service, so when something starts eating memory or pinning the CPU, you get told which service — by name — rather than being told the box is unhappy and left to work out why. A lot of weekend failures aren't code at all. They're a process slowly consuming everything while nobody's looking.

If you have an app, the crashes come in too. Crashlytics and Sentry feed in, and crashes get grouped into one signature per real bug instead of a wall of near-identical stack traces.

Everything ties back to the release that caused it. This is the one I'd have paid for on its own that Sunday. Crashes, errors, incidents — all attached to the version that shipped them. A release whose score drops sharply against the previous one gets flagged as a regression, rather than left for you to infer at 11am from timestamps.

You get a guess at the cause, in plain language. Not a dashboard to interpret. An actual sentence: this failed, here's what it looks like, and here's how confident we are. When the evidence is thin it says so rather than inventing something authoritative — which matters, because a confident wrong answer at the start of a debugging session can cost you hours.

Every release gets a number. 0 to 100, built from the uptime, crash and defect signals, with a plain ship-or-hold recommendation. You can gate your pipeline on it through the API. Friday afternoon, you look at the score before you close the laptop.

And it becomes a defect, not a notification. Incidents, fatal crashes and release regressions open defects automatically: something with an owner and a status and the evidence attached, which you can send to Jira or Linear or GitHub or Azure DevOps. Not a message that scrolls away over the weekend. When you sit down on Monday there's a thing in your tracker that already knows what happened, when it started, and which release it came from.

That last point is really the whole product. Alerts evaporate. Work persists.

There's a status page too, which sounds minor and isn't. If the CEO had been able to look at a page and see "yes, we know, we're on it," that Sunday call would have been an entirely different conversation.

We still don't deploy on Fridays

Old habits. And there's something to be said for not pushing changes right before everyone disappears, regardless of how good your tooling is.

But the rule means something different now. It used to mean we won't find out until Monday. Now it's just a preference about when we'd rather be interrupted.

That's a much better reason to have a rule.


CleverQA watches every service, ties what breaks back to the release that broke it, and turns it into work instead of a notification. Free plan, no credit card: Start free