Skip to main content

Slow Is the New Down — Why Performance Is Availability

Users abandon slow sites long before a timeout, and degradation usually precedes outages. Why response-time trends belong in your uptime monitoring.

The CompleteStatus Team 6 min read

Slow Is the New Down — Why Performance Is Availability

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

Your uptime monitor says 100%. Your users say the site is down. You're both right, and that's the problem. A page that takes twelve seconds to respond will pass a conventional uptime check, because the check asks "did the server answer before my 30-second timeout?" Your visitor asks something else: "did this load before I gave up?" Their timeout is measured in single-digit seconds, enforced by a thumb hovering over the back button.

The binary up/down model made sense when the alternative to a working site was a connection error. In practice, most availability loss happens in the gray zone between fast and dead, and if your monitoring only alerts on dead, you're measuring the smallest slice of the problem.

Users time out long before your server does

The research has been consistent for years. Exact figures vary by study and industry, but the shape doesn't:

  • Abandonment starts early. Widely cited findings, including Google's own mobile research, show the odds of a visitor bouncing rise sharply once load time passes about three seconds; Google's figure was that 53% of mobile visits are abandoned at that point.
  • Seconds cost conversions. Retail studies have long reported that an extra second or two of load time can cut conversions by ten percent or more. The famous old Amazon figure (every 100ms costs about 1% of sales) may be dated, but retests land in the same neighborhood: the curve is steep exactly where most degraded sites live.
  • Search engines have made it official. Since Google's "Speed Update" rolled out in July 2018, page speed is a ranking factor for mobile searches. A slow site loses visitors it never even had; they were filtered out on the results page.

Put together, a site responding in 8–15 seconds has, for business purposes, most of the availability loss of a hard outage, while generating none of the alerts.

Degradation is how outages introduce themselves

There's an operational reason to watch response times even if you don't care about conversion curves: most outages aren't born as outages. They start as slowness, minutes or hours earlier, because the classic failure mechanisms are gradual:

  • Connection-pool exhaustion. The database pool has 100 slots. A slow query makes each connection linger a bit longer; utilization creeps up; requests start queueing for a free slot. Response times climb (300ms, 900ms, 3 seconds) until the pool saturates and everything fails at once. The up/down monitor fires at the last step. The response-time trend had been climbing for half an hour.
  • The swap death spiral. A memory leak nudges the server into swap. Everything slows; requests pile up; piled-up requests hold more memory; the machine swaps harder. From outside, response times double every few minutes, a curve distinctive enough to diagnose from the graph, ending in a load average in the hundreds and a server that might as well be off.
  • Slow resource leaks: file handles, worker threads, a disk filling with logs. Each produces the same signature: a healthy baseline, then a drift, then a cliff.

The cliff is the outage; the drift is the warning. Response-time monitoring gets you paged at the drift, while the site is still up and the fix is a calm restart instead of a 2 AM incident.

Alert on latency, not just liveness

In practice, that means giving your monitor a second threshold:

  • Set a response-time alert at "degraded," well below "dead." If your homepage normally answers in 400ms, an alert at 2–3 seconds sustained across several consecutive checks catches every mechanism above with time to act. Require the condition to persist: one slow sample is internet weather; five in a row is a trend.
  • Base thresholds on your baseline, not a universal number. An API that normally answers in 80ms is badly wrong at 1.5s; a report-generation endpoint might be fine at 4s. A week of monitoring history tells you what normal looks like. Alert on departures from it.
  • Watch trends between alerts. A response-time graph drifting from 300ms to 700ms over a month never crosses an alert line, but it's telling you something about data growth, a missing index, or traffic outgrowing a server. Slow drift is capacity planning; fast drift is an incident.

Measure TTFB, from the outside

For monitoring, the most useful number is time to first byte: how long from request until the server starts answering. It isolates your stack (network in, TLS, application logic, database, rendering the response) from page weight and frontend concerns, so it's the number that moves when the backend is sick. You can break it down yourself:

curl -s -o /dev/null -w \
'dns        %{time_namelookup}\nconnect    %{time_connect}\ntls        %{time_appconnect}\nttfb       %{time_starttransfer}\ntotal      %{time_total}\n' \
https://yoursite.com/

A healthy site might show TTFB a few hundred milliseconds above the TLS handshake time. A site entering the connection-pool spiral shows dns/connect/tls unchanged and TTFB climbing, which tells you it's the application, not the network, before you've logged into anything.

Two rules make the measurement meaningful:

  • Measure from outside your network. Localhost TTFB skips DNS, real network paths and TLS setup, and the visitor's experience includes all three.
  • Measure continuously, not when curious. A single number is trivia; the same number every minute for six months is a baseline and an early-warning system.

Redefining "up"

The useful mental shift is to stop treating availability as a boolean. A site that's up but slow is partially down: losing some fraction of its visitors, revenue and search standing every minute, and quite possibly heading for a full outage. Your monitoring should see it that way too.

CompleteStatus records response time on every check, charts the trend for every monitor, and lets you alert on response-time thresholds the same way as downtime, with the same confirmation logic and the same channels, measured from outside, where your users are. See what it tracks, or create a free account and look at your own response-time graph after a week. Most teams find at least one surprise in it, and it's better to find it on a graph than in an incident.

Written with AI assistance and reviewed by the CompleteStatus team.