Skip to main content

Your VPN Is Production Now: Monitoring the Infrastructure Remote Work Depends On

A month into working from home, the VPN, SSO and internal tools are business-critical. What to monitor, and what not to expose in a hurry.

The CompleteStatus Team 7 min read

Your VPN Is Production Now: Monitoring the Infrastructure Remote Work Depends On

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

A month ago, most of us were commuting. Now a large share of the working world is at a kitchen table, and a lot of infrastructure that used to be a convenience has quietly become the most important thing the company runs.

The scale of the shift is hard to overstate. Zoom's CEO wrote at the start of April that daily meeting participants had gone from around 10 million in December to more than 200 million in March. In Europe, Netflix and YouTube cut their default streaming quality to take pressure off networks. Inside individual companies the numbers are smaller but the change is just as sharp: the VPN concentrator that was sized for the sales team on the road is now carrying everyone, every day.

This post is about the parts of your stack that were never treated as production and now are, and about the shortcuts that are tempting right now and worth resisting.

The new critical path

Before March, "is the business working?" mostly meant "is the website up?". For a remote company, a normal working day now depends on a longer chain:

  • The VPN or remote access gateway. If it's down, nobody can reach internal systems.
  • The identity provider. SSO, directory services, MFA. If sign-in fails, every SaaS tool the company uses fails with it.
  • Internal web apps. The admin panel, the wiki, the ticketing system, the build server. These were often reachable only from the office network, with monitoring to match (i.e. none, because someone would notice).
  • Chat and video. Mostly someone else's infrastructure, but still a dependency worth knowing the status of.
  • Home internet. Yours and your on-call engineer's.

The shift in thinking is simple: anything in that list that isn't monitored with alerting is a place where the first alert is a message from a colleague saying "is it just me?". That was tolerable when the colleague was at the next desk. It's a lot less tolerable when it's forty people across four time zones.

What to monitor on a VPN

A VPN can fail in more ways than "down". Worth covering:

  • Reachability of the endpoint. A TCP check on the right port (443 for most SSL VPNs, 1194 for OpenVPN over TCP) from outside your network. For UDP-only setups such as WireGuard or IPsec, a TCP check won't tell you much, so check the management interface or use a synthetic client if you can.
  • The certificate on the portal. SSL VPN portals have TLS certificates, and they expire like any other. They're also frequently forgotten, because nobody visits them in a browser. An expired VPN certificate on a Monday morning locks out the whole company at once.
  • Capacity. Concurrent sessions against the licence limit, CPU on the appliance, and bandwidth on the uplink. Many appliances are licensed per concurrent user, and plenty of teams have discovered in the last month that their limit was set years ago for a very different workforce.
  • The thing behind it. A VPN that connects but can't route to the internal network is down from the user's point of view. If you can, run a check from a machine that connects through the VPN and fetches an internal page.

If you only do one thing from this list, check the certificate expiry dates on every remote-access endpoint today, and put them somewhere that will warn you a few weeks in advance.

Watch sign-in like you watch checkout

For many companies the identity provider has become the single most important dependency, and it's usually the least watched. If you host your own (Active Directory Federation Services, Keycloak, a self-hosted LDAP behind a login page), give it the same treatment as your public website:

  • An HTTP check on the login page, with a check that the expected content is there rather than just a 200.
  • Certificate expiry checks on every hostname involved, including the ones used for SAML metadata and token signing where you can.
  • Response time monitoring. A sign-in page that takes 15 seconds to load during the 9am rush is an early capacity warning, and slow is the new down for internal tools too.

If your identity provider is a SaaS product, subscribe to its status page and route its notifications somewhere your team will actually see them.

Don't expose things just because it's quicker

The most common shortcut right now is to make an internal tool reachable from the internet because the VPN is overloaded or someone can't get it working at home. Opening port 3389 so staff can use Remote Desktop directly, publishing the admin panel on a public hostname, or allowing the database port "just for a week". Security researchers have reported noticeable increases in the number of RDP servers exposed on the internet since March, and attackers scan for them constantly.

If you have to make something reachable quickly:

  • Put authentication in front of it that isn't the application's own login: a reverse proxy with SSO, or at minimum basic auth over HTTPS plus the app's login.
  • Restrict by IP where you can, even if the list is long.
  • Never expose RDP, SSH with passwords, or database ports directly. Use the VPN, a bastion host with keys, or a gateway product.
  • Write it down. Keep a list of every emergency exception with an owner and a date to revisit it. A temporary firewall rule without an expiry date is a permanent one.

Then monitor what you've exposed from the outside, so you know it's still up and still behind the protection you put in front of it. A useful check: request the page without credentials and alert if it returns a 200 instead of a 401 or a redirect to sign-in.

# Should NOT be 200 without credentials
curl -s -o /dev/null -w "%{http_code}\n" https://admin.example.com/

On-call from a spare bedroom

On-call has changed too. A few practical points that teams have been learning the hard way:

  • Your home connection is now a dependency of incident response. Make sure on-call engineers have a fallback, even if that's tethering to a phone, and that the tethered connection can actually reach the VPN.
  • Alerts need to reach phones. If alerts used to be noticed on a big screen in the office, that screen now shows them to nobody. Move alerting to channels that wake a person: push notifications, SMS or phone calls.
  • Escalation needs names, not rooms. "Someone in the office will see it" no longer works. Write down who is primary, who is secondary, and how long before it escalates.
  • People are tired. Many of your engineers are also looking after children or relatives right now. Keep on-call rotations humane: shorter shifts, clear handovers, and fewer noisy alerts that wake someone up for nothing.

A short checklist for this week

  1. List every system a remote employee needs to do a normal day's work.
  2. For each one, confirm there's a monitor, where it runs from, and who it alerts.
  3. Check certificate expiry on every VPN, SSO and remote-access hostname.
  4. Compare current concurrent VPN sessions with the licence and hardware limits.
  5. Review every emergency firewall change made since March and give each one an owner and an end date.

None of this is glamorous work, and none of it was on anyone's roadmap for 2020. But the infrastructure that keeps a remote company running deserves the same attention as the infrastructure your customers see, because right now, for your own staff, it is the product. If you're adding checks for these systems, CompleteStatus can watch the external ones from outside your network; whatever you use, make sure something is watching.

Written with AI assistance and reviewed by the CompleteStatus team.