- Pure stdlib Python (3.12), runs on python:3.12-alpine as non-root - Checks: ping, HTTP (expected codes), TCP ports, Netdata alarms + metrics - Netdata v2 compatible: system.cpu (no idle dim), system.ram, disk_space.* discovery - AlertEngine dedup: first alert, reminders, recovery (only after real alert) - Baseline first-run (no alert storm on deploy), atomic state file, --once mode - 20 unit tests passing; verified live against 192.168.0.5
11 KiB
ServerMonitor — Dockerized server monitoring service
Build a complete, production-quality server monitoring service in this repository. The service runs inside a Docker container on a REMOTE VPS and monitors a home server (Unraid at 192.168.0.5) over the network. It periodically runs health checks and sends Telegram alerts on failures, recovery messages when services come back, and periodic reminders while a service stays down.
Hard constraints
- Python 3.12, stdlib ONLY (urllib, socket, json, subprocess, time, signal, logging, threading, unittest). NO third-party dependencies, no pip installs, no requests/pyyaml.
- Must run on
python:3.12-alpine(busybox, no bash — onlysh). - The main loop must NEVER crash because of a single failed check. Every check is wrapped in try/except, errors are logged, loop continues.
- Every network operation has a timeout.
- Graceful shutdown on SIGTERM/SIGINT: save state, log "shutting down", exit 0.
- Non-root user in the Docker image.
- Commit NOTHING and do not run git. Just write the files.
Files to create
1. checks.py
Check functions, each returns a tuple (ok: bool, message: str) and never raises:
ping_check(host, timeout)—subprocess.run(['ping', '-c', '1', '-W', str(timeout), host], ...). Busybox ping:-Wis seconds. If ping is unavailable or permission denied (no ICMP in container), log a warning once and return ok=True with message "ping unavailable" (degrade gracefully — do not fail the check).http_check(name, url, expected_codes, timeout)— urllib.request with timeout; returns ok if HTTP status in expected_codes; message includes status, latency ms. Distinguish connection error / timeout / bad status. Follow no redirects.tcp_check(name, host, port, timeout)— socket.create_connection; ok if connects; message includes latency ms.netdata_alarms(base_url, timeout)— GET{base_url}/api/v1/alarms?active. Response JSON:{"hostname":..., "status": true, "alarms": {name: {status, info, value, ...}}}. Ifstatusis false → not ok. Ifalarmsdict non-empty → not ok; message lists each alarm asNAME [STATUS]: value — info. If alarms empty → ok.status: truemeans Netdata itself is fine (it'sfalsewhen Netdata cannot evaluate).netdata_metrics(base_url, timeout)— collect host metrics via Netdata data API to evaluate thresholds. UseGET {base_url}/api/v1/data?chart=system.cpu&after=-1&points=1&format=json(dimensions includeidle, plus user/system/iowait/etc.; values are percentages; utilization = 100 - idle) andGET {base_url}/api/v1/data?chart=system.ram&after=-1&points=1&format=json(dims: used, free, cached, buffers in MB). AlsoGET {base_url}/api/v1/data?chart=disk.space._&after=-1&points=1&format=json— actually usechart=disk.space(it aggregates or errors; handle both, and if the chart is missing just skip disk) for per-mount used percent (dim name = mount path, values are percent used; dimensionavailmay exist — prefer the dims whose names are paths). Return dict with keyscpu_percent,ram_used_mb,ram_total_mb(fromsystem.ramfirst value ofused+free+cached+buffersOR parse/proc/meminfo-style — no, keep it to the API; ram_total_mb = sum of used/free/cached/buffers),disk_used_percent(max over mounts). Wrap EVERYTHING in try/except; on any error return None values and log a warning (metrics are best-effort, never alert on inability to fetch — only on exceeding thresholds when data IS available).- Also
netdata_info(base_url, timeout)— GET{base_url}/api/v1/info, ok if JSON withversionpresent. (This is effectively covered by the HTTP check on :19999 root; you may skip if redundant.)
2. notify.py
send_telegram(bot_token, chat_id, text, timeout)— POSThttps://api.telegram.org/bot<token>/sendMessagewith form-encodedchat_id,text,parse_mode=HTML,disable_web_page_preview=true. Return bool. If bot_token empty → return True without sending (log-only mode) and log the would-be message at INFO.escape_html(text)— escape & < > for parse_mode=HTML.class StateStore— JSON file persistence (default path/data/state.json): dict keyed by check name:{"state": "OK"|"CRIT", "last_alert_ts": float, "alert_count": int}. Methods:load(),save(),get(name),set(name, state, last_alert_ts, alert_count). Atomic write (write temp file + os.replace). Never crash if file corrupt — start fresh.class AlertEngine— dedup logic:should_alert(name, now, state_store, reminder_interval)→ (bool, is_recovery):- CRIT and prev state != CRIT → alert (first failure).
- CRIT and prev state == CRIT and now - last_alert_ts >= reminder_interval → reminder (re-alert).
- OK and prev state == CRIT → recovery message.
- All alert decisions go through here; tests target this class.
3. monitor.py
Main entry point:
- Load
config.jsonfrom the same directory as the script (allow override via envCONFIG_FILE). - Env overrides applied after config load:
HOST,CHECK_INTERVAL_SEC,TELEGRAM_BOT_TOKEN,TELEGRAM_CHAT_ID,STATE_FILE. - Setup logging: stdout, format
%(asctime)s %(levelname)s %(message)s, INFO level. - Register SIGTERM/SIGINT handler → set a threading.Event, main loop exits cleanly, state saved.
- Main loop (every
check_interval_sec, default 60):- Run all checks: ping, each http_check, each tcp_check, netdata_alarms, netdata_metrics thresholds (cpu_percent > threshold → CRIT "CPU 95%"; ram_avail_mb < threshold → CRIT; available = ram_total - ram_used; disk_used_percent > threshold → CRIT).
- For each check that is CRIT or recovered: decide alert via AlertEngine, send Telegram (sequential is fine, but keep total loop time bounded; timeouts are short).
- Log every check result at INFO:
Gitea OK (200, 15ms)/Gitea FAIL (502, expected [200], 12ms). - On alert: log at WARNING. Message format (HTML):
- Failure:
⚠️ <b>Server alert</b>\n\n🔴 <b>Gitea</b>: HTTP 502 (expected 200), 12ms(include host name from confighostin header line:<b>Server alert</b> · 192.168.0.5). - Reminder:
⏰ <b>Still down</b> · 192.168.0.5\n\n🔴 <b>Gitea</b>: HTTP 502 (expected 200), 12ms+\n(already 5 min)— track downtime duration if easy; otherwise skip duration. - Recovery:
✅ <b>Recovered</b> · 192.168.0.5\n\n🟢 <b>Gitea</b>: HTTP 200, 10ms. - Netdata alarms:
🔴 <b>Netdata alarms</b>: CPU_USAGE [CRITICAL]: 95 — ...
- Failure:
- Save state each iteration.
- On first run with no state file: do NOT alert immediately for pre-existing failures — establish baseline: first iteration only logs and records state without sending alerts (prevents alert storm on deploy/restart). Implement via
state_filemissing + abaseline_doneflag in state (or: treat first iteration as baseline when state file absent). --onceCLI flag: run a single check cycle and exit (useful for cron/testing).
4. config.json
Default config (services discovered on the target server):
{
"host": "192.168.0.5",
"check_interval_sec": 60,
"reminder_interval_sec": 1800,
"timeout_sec": 5,
"state_file": "/data/state.json",
"telegram": {"bot_token": "", "chat_id": ""},
"thresholds": {"cpu_percent": 90, "ram_avail_mb": 2048, "disk_percent": 90},
"ping": {"enabled": true},
"http_checks": [
{"name": "Gitea", "url": "http://192.168.0.5:3000/", "expected": [200]},
{"name": "Unraid WebUI", "url": "http://192.168.0.5/", "expected": [200, 301, 302, 307]},
{"name": "MetaCubeXD", "url": "http://192.168.0.5:9090/", "expected": [200]},
{"name": "Netdata", "url": "http://192.168.0.5:19999/", "expected": [200]},
{"name": "Plex", "url": "http://192.168.0.5:32400/", "expected": [200, 401]},
{"name": "Nginx 8080", "url": "http://192.168.0.5:8080/", "expected": [200]}
],
"tcp_checks": [
{"name": "MySQL", "host": "192.168.0.5", "port": 3306},
{"name": "SMB", "host": "192.168.0.5", "port": 445},
{"name": "Proxy 1984", "host": "192.168.0.5", "port": 1984}
],
"netdata": {"base_url": "http://192.168.0.5:19999", "check_alarms": true, "check_metrics": true}
}
5. Dockerfile
FROM python:3.12-alpine- create user
app(adduser -D -u 10001 app), WORKDIR /app, copy files,USER app - ensure
/dataexists and is writable by app (mkdir -p /data && chown) — note: /data will be a volume; make it owned by app in image. ENV PYTHONUNBUFFERED=1CMD ["python", "-u", "/app/monitor.py"]- Healthcheck: use python one-liner checking state.json mtime freshness:
python -c "import os,time; p='/data/state.json'; raise SystemExit(0 if os.path.exists(p) and time.time()-os.path.getmtime(p)<300 else 1)"— interval 60s, retries 3, start_period 30s.
6. docker-compose.yml
- service
server-monitor: build.,restart: unless-stopped,env_file: .env, volumes:monitor-state:/data, logging:{driver: json-file, options: {max-size: "10m", max-file: "3"}}.
7. .env.example
# Telegram bot for alerts (get from @BotFather). Leave empty for log-only mode.
TELEGRAM_BOT_TOKEN=
# Your chat id (get from @userinfobot)
TELEGRAM_CHAT_ID=
# Check interval in seconds (default 60)
CHECK_INTERVAL_SEC=60
8. README.md — in Russian
Sections: что это; как это работает (схема: VPS → чеки → Telegram); быстрый старт (скопировать .env.example → .env, вписать токен, docker compose up -d --build); как получить токен бота (BotFather) и chat_id (@userinfobot); настройка чеков в config.json (таблица полей); формат алертов (примеры); лог-режим без токена; локальный запуск без Docker (python3 monitor.py --once); запуск тестов (python3 -m unittest discover -s tests -v); troubleshooting (контейнер не видит хост: сеть, firewall; ping в контейнере).
9. tests/test_checks.py — unittest, stdlib only
http_checkreturns ok for 200 expected 200: spin uphttp.serveron 127.0.0.1:0 in a thread, hit it.http_checkfails on wrong status (e.g. expected [200], server returns 404): use a handler that returns 404.http_checkconnection refused → not ok (closed port).tcp_checkok on listening socket; fail on closed port.netdata_alarmsparsing: monkeypatchurllib.request.urlopento return a fake response JSON (active alarms present → not ok, message contains alarm name; empty alarms → ok; status false → not ok).AlertEnginededup: first CRIT alerts; second CRIT within reminder window does NOT alert; CRIT after reminder window alerts again; OK after CRIT sends recovery exactly once.StateStoreround-trip and corrupt-file resilience (write garbage, load → empty dict).escape_htmlescapes&<>. All tests must pass:python3 -m unittest discover -s tests -v
Quality bar
- Clean, readable, commented code (docstrings on functions).
- No print() in library modules — use logging.
- No global mutable state in checks.py (pass config around).
- File count: exactly the files above (SPEC.md already exists, don't modify it).
- Total code should be compact but complete — prefer clarity over cleverness.