Difference between revisions of "Information Systems:Site24x7 Monitoring Service"
m (→Overview) |
|||
| Line 13: | Line 13: | ||
* Network; ping: Our WAN interfaces (WAN IPs) are pinged. |
* Network; ping: Our WAN interfaces (WAN IPs) are pinged. |
||
====On-premise poller and reporting agents==== |
====On-premise poller and reporting agents==== |
||
| + | [[Information Systems: Site24x7 On-premises Poller|The on-premise poller is installed on WDS.]] |
||
| − | This service also allows you to install agents on servers and/or an on-premise poller for detailed monitoring/reporting of onsite hardware. The agents are custom software that install on individual servers with supported OSes, while the on-premise poller operates similarly to Spiceworks, polling the network via SNMP. |
||
| + | |||
| − | I've tried this on a personal basis. It's incredibly powerful and provides a wealth of insight (vCenter, VMs, Linux VPS, AWS instances), but after a while, you realize that your servers are pumping out gobs of resource monitor data to some..erm.."cloud service"... We may want to analyze this closely for feasibility/desirability. As an example, I got a CPU-usage alert for one of my servers at home the second some started playing a movie that it had to transcode. IoT is scary... [[User:Norwinu|norwizzle]] ([[User talk:Norwinu|talk]]) |
||
===Probe locations=== |
===Probe locations=== |
||
There are many locations from which monitors can be run. They represent continental regions and can consist of up to 6 servers. For the most part, North America is chosen as the source, and is made up of 6 locations spanning North America that I chose randomly. Some monitors like DNS are set to be tested from other parts of the globe just for variety. |
There are many locations from which monitors can be run. They represent continental regions and can consist of up to 6 servers. For the most part, North America is chosen as the source, and is made up of 6 locations spanning North America that I chose randomly. Some monitors like DNS are set to be tested from other parts of the globe just for variety. |
||
Latest revision as of 11:19, 26 April 2021
Overview
A external monitoring service helps IS manage its infrastructure by probing for and alerting on potential issues. In April 2017, a subscription to Site24x7 was acquired on a trial basis and is proposed to replace Keynote Red Alert to provide such a monitoring service. Whereas Keynote Red Alert provided basic ping and HTTP-request monitors, Site24x7 offers more advanced monitors and a more advanced overall configuration. Its setup will thus be detailed in this article.
Configuration
Monitors
Monitor types
One of the most appealing factors of this service is the many types of monitors available. The ones currently in use are:
- Website; basic HTTP requests: An HTTP GET/POST is performed and either the HTTP return code or HTML result is analyzed to determine uptime. This is set up for Exware sites, Web Orders etc.
- Website performance; MTFB,response times: Page load is done and the individual elements (CSS files, images, included javascript files) are analyzed for response time.
- Email; SMTP connect: A basic SMTP connection to the mail server is tried.
- DNS; record lookup: CNAME, A, and MX record lookups can be tried against a specified DNS server (e.g. Dyn, uni3sys etc.)
- Network; ping: Our WAN interfaces (WAN IPs) are pinged.
On-premise poller and reporting agents
The on-premise poller is installed on WDS.
Probe locations
There are many locations from which monitors can be run. They represent continental regions and can consist of up to 6 servers. For the most part, North America is chosen as the source, and is made up of 6 locations spanning North America that I chose randomly. Some monitors like DNS are set to be tested from other parts of the globe just for variety.
Monitor groups
Due to the many individual monitors set up under this service, monitor groups may be the more convenient way to peruse the interface or analyze for uptime (the main list of all individual monitors can be an overwhelming site otherwise). Monitor groups are meant to represent a 'whole' (i.e business services or technical areas), therefore comprising individual monitors that test parts of that whole. Monitor groups (and their constituent individual monitors) include:
- Exware sites (Web monitor for medicinecentre.com, Web monitor for unipharm.com)
- Apache instances on Bart (Web monitor for orders.unipharm.com, Web monitor for infonet.unipharm.com)
- Email (SMTP check for Barracuda servers, SMTP check for mail.unipharm.com, DNS MX record check for unipharm.com)
- Web Orders (Web monitor for orders.unipharm.com, Web monitor for unipharm.com (required for login to Web Orders), DNS record for orders.unipharm.com)
- Telus WAN connectivity (Network ping test for superman.unipharm.com, network ping test for mail.unipharm.com)
Scheduled maintenance
The service allows for scheduling of time windows where monitoring (and thus alerts) are suppressed to account for planned downtime. The web servers on Bart are brought down every night during the overnight process, so a scheduled maintenance entry is set for Bart-associated monitors from 1:00AM to 5:30AM. Scheduled maintenance can also be activated ad-hoc for the next 5-60 minutes through the interface.
Alerting
- Site24x7 allows for many alerting methods: email, call, SMS, Twitter, IM, and phone app. Probably, the primary alerting methods will be email, SMS and phone app.
- Alert targets (user groups) can be set per monitor or per monitor group.
- Thresholds can also be specified to quell monitor oversensitivity e.g. monitor has to be down from 'x' number of locations for it to be considered down.
Administration details
- Each user in the IS team has their own credentials that they can use to access the web console or app.
- The current subscription is the middle-tier Business plan, allowing for 40 basic/advanced monitors and 200 SMS/call credits per month.