System Monitoring and Alerts in Plain English: Who Finds Out First When the Site Goes Down, and How to Tier Alerts So They Do Not Cry Wolf
“I think the website is down?” If you hear that from a customer on the phone, the system has already been down for a while, and you are the last to know. System monitoring and alerts exist to solve exactly that: to make sure the responsible person spots a problem before it affects customers, or at least as soon as it starts.
Many small and medium businesses have their systems built by an outside vendor, and once they go live nobody actively watches them. Whether the server is healthy, whether forms are being sent, whether payments are going through — it all relies on “if nobody complains, assume it’s fine.” Most of the time that looks harmless, but when something does go wrong, it leaves you in the most reactive position possible.
What follows explains in plain terms what monitoring watches, how uptime monitoring differs from error monitoring, how to tier alerts and arrange on-call, what a status page is for, and how to run a review after an incident.
What monitoring actually watches
Think of monitoring as fitting the system with a dashboard and an alarm. The dashboard lets you see its condition at any time, and the alarm notifies someone when a number looks wrong. There are four common areas to observe:
- Availability: Can the system be reached? An outside service “knocks on the door” at regular intervals, and no answer means a problem.
- Errors: The system may be reachable, but is something failing behind the scenes? For example, a button that does nothing when clicked, or a piece of code that keeps throwing errors.
- Performance: How fast does it respond? Pages getting steadily slower is usually an early warning sign.
- Resources: Does the server still have enough processing power, memory, and disk space? A full disk is a very common and very preventable cause of outages.
Beyond these four technical measures, there is one area that often gets overlooked: business metrics, such as today’s orders, form submissions, or sign-ups. The system can be technically fine while the contact form has not sent a single email all day. That kind of “silent failure” only shows up when you watch business metrics.
Uptime monitoring: knocking on the door from the user’s side
The most basic monitoring, and the first thing to set up, is uptime monitoring. The idea is simple: an external monitoring service connects to your website or system at regular intervals and checks that it responds normally.
Monitoring only the homepage is not enough
Many sites monitor only the homepage, but a homepage that loads does not mean everything is working. A more complete approach:
- Monitor key pages. Beyond the homepage, add the product pages, login page, checkout page, and other pages that actually bring in revenue or deliver the service.
- Monitor key flows. Simulate a user walking through “log in → add to cart → go to checkout” to confirm the whole flow completes. This is called synthetic monitoring.
- Check the response content. “Did it respond?” is not enough, because a server returning an error page still counts as a response. You can configure a check for specific text on the page.
- Monitor certificate and domain expiry. If the site’s encryption certificate or domain is not renewed, the whole site suddenly becomes unreachable. This is entirely something you can be warned about in advance.
- Check from several locations. If one location cannot connect, it may just be a network issue. Failures from several locations at once are far more likely to mean the site is really down, which cuts false alarms.
Error monitoring: catching problems that are “up but broken”
Uptime monitoring can only tell you the system is alive. Error monitoring tells you whether it is failing behind the scenes. These tools collect errors that occur while the code runs, organize them into a list, and show how often each happened, which users were affected, and where it occurred.
The value of error monitoring:
- Finding problems before customers complain. If the submit button does not work in a particular browser version, most customers will not report it; they will just quietly leave.
- Shortening the time to diagnose. Engineers get a specific error location and the context at the time, rather than “a customer says it’s acting weird.”
- Seeing whether a new release introduced problems. Watching how the error count changes after each update makes it quick to decide whether to roll back to the previous version.
Error monitoring works hand in hand with the deployment process. For how new versions go live and how to roll back when something goes wrong, see Environments and Deployment Explained.
Alert tiers: not everything should wake someone up at night
Once monitoring is in place, the most common problem is not “no alerts” but “too many alerts.” When dozens of notifications arrive every day, people stop reading them, and when the one that really matters arrives, it gets ignored along with the rest. This is known as alert fatigue.
A practical three-tier scheme
| Tier | What it means | Examples | How to notify |
|---|---|---|---|
| Urgent | Customers are already affected; handle immediately | Site completely down, payments failing, login failing | Phone call or instant message straight to the on-call person |
| Important | Some functions are failing or about to; handle the same day | Error count clearly rising, disk nearly full, certificate about to expire | Instant message or team group chat |
| Notice | Worth watching but not urgent | Responses slightly slow, occasional one-off errors | Daily summary report |
Principles for tiering
- Judge by customer impact, not technical severity. If the database restarts once and customers notice nothing, it does not need an urgent alert.
- Every alert needs a matching action. An alert that leaves the recipient unsure what to do should be removed or downgraded.
- Fire only after a problem persists. Momentary blips are common. Setting “notify only after several consecutive failures” greatly reduces false alarms.
- Review regularly. Every so often, look at which alerts were never acted on. Those are usually noise.
On-call and escalation: send alerts where someone is sure to see them
The biggest blind spot with alerts is “they were sent, but nobody received them.” Sending them to a shared inbox or a group chat nobody watches is the same as having no alerts at all.
The basic on-call questions
- Who receives alerts first? With an in-house engineering team, set up a rotation. Without one, it is usually the maintenance vendor, with a copy to one internal contact at the company.
- How long without a response before escalating? For example, if nobody acknowledges an urgent alert within a set time, the next person in line is notified automatically.
- What about outside business hours? If customers use the system at night and on weekends, agree in advance how problems during those hours are handled and what is covered.
- Who communicates with others? Let the technical people concentrate on the fix while a separate person handles communication with customers and colleagues, so the person fixing it is not constantly interrupted.
If maintenance is outsourced, these arrangements must be written into the contract, including response times, scope, and contact channels. The relevant terms are covered in the Website Maintenance Cost and Contract Guide.
Status pages: saying it first beats being asked
A status page is a separate web page that announces whether the system is currently working, what problem is happening, and when it is expected to recover. It is useful for:
- Reducing repeated questions. Customers and colleagues can check the situation themselves, so customer service does not have to answer them one by one.
- Showing you are in control. Announcing “we know about it and we’re working on it” maintains more trust than letting customers discover the problem on their own.
- Announcing maintenance windows. Posting planned maintenance in advance lets customers work around it.
A status page must be hosted somewhere separate from the main system; otherwise, when the main system goes down, the status page goes down with it and loses its point. The announcement should also be in language ordinary people understand:
Some members are reporting that they cannot log in. We have identified the problem and are working on it. Orders and payments are not affected. Our next update is expected within an hour.
There is no need for technical details, but state clearly the scope of the impact, the current status, and when the next update will come.
Post-incident review: the goal is to keep it from happening again
Once the system is back, the work is not finished. Every incident that affected customers deserves a review, and the focus is on finding weaknesses in the process and the system, not on deciding whose fault it was. As soon as a review turns into blame, people tend to hide things, and the next problem gets discovered even later.
The skeleton of a review
Follow this structure; one page is enough:
- What happened: a timeline from when the problem began, when it was discovered, when handling started, and when service recovered.
- Impact: which functions, which customers, and for how long.
- How it was discovered: by a monitoring alert or by a customer report? If a customer noticed first, monitoring has a gap.
- Root cause: do not stop at “the server crashed.” Keep asking why it crashed and why there was no early warning.
- What went well: which responses worked and should be kept next time.
- Improvement actions: each one with an owner and a completion date, for example adding a particular monitor, writing an operating guide, or adjusting an alert threshold.
Improvement actions often point to shortfalls in backups, capacity, or security. For what to do about those, see Website Backup and Disaster Recovery and Traffic Spikes and Load Testing.
A monitoring self-check owners can start with
You do not need to build a complete monitoring setup all at once. Start by checking your current situation against this list:
- When the website or system goes down, is there an automatic notification rather than relying on customers to report it?
- Does monitoring cover key flows such as checkout, login, and form submission, not just the homepage?
- Do you get a reminder before the domain and encryption certificate expire?
- Is someone definitely watching the person or group that receives alerts?
- Is it written down who handles problems outside business hours and how to reach them?
- Is there somewhere you can announce system status to customers?
- After the last system problem, was there a record and a set of improvement actions?
Wherever you cannot tick a box, start there.
If you are not sure who would find out first when your system has a problem, NETVANA can help identify monitoring gaps, design alert tiers and escalation, and fold monitoring into routine maintenance. All software work is quoted after a consultation, so get in touch and tell us about the size of your system and how it is maintained today. What the relevant services include is listed in the software services overview.
Further reading: For getting your data back after something goes wrong, see Website Backup and Disaster Recovery. To confirm your system can hold up before a campaign, read Traffic Spikes and Load Testing. For how new versions go live and get rolled back, see Environments and Deployment Explained. For the basic protections you need beyond monitoring, read Website Security Basics for Business. And for choosing a hosting plan, see Cloud Hosting Options for SMBs.