What we were handed.
A growing stack of hosts, tunnels, and services means outages get noticed by users before operators. Babysitting dashboards doesn't scale, and 3 a.m. incidents don't wait for business hours.
Worse, simple uptime checks lie: a wedged machine can still answer pings while every service on it is silent. We learned that one in production.
The system.
A fleet of self-hosted agents that patrol the infrastructure on timers: watchdogs that check real service responses (not just pings), a sentinel that investigates anomalies and writes up what it found, and scheduled sweeps for the routine chores.
A hybrid local/cloud AI brain does the reasoning, and everything reports into Discord — the same place the ops assistant lives — where the team can talk to the fleet and issue commands.
Real HTTP-level checks against every service — because pings lie.
An agent that investigates anomalies and arrives with a write-up.
Scheduled agents for the routine chores nobody should do by hand.
Findings, alerts, and commands all flow through Discord.
What changed.
The infrastructure notices its own problems and often arrives with a diagnosis before a human has looked. Operations happen in chat, with a paper trail for free.
The next study could be yours.
Bring us the problem — the messy, manual, “there has to be a better way” kind. We'll design the system, build it, and run it.
