Edge
A Cloudflare Worker in front of this site and both demos. Visitors reach nothing else.
It adds a secret the services check, so nobody can skip it and call them directly, and it passes on the visitor's address in a header that can't be forged from outside.
Code deploy/cloudrun/edge/worker.js deploy/cloudrun/edge.tf
Public pages
The front page, the stories and the status page, drawn on the server from the database.
Every figure on the site is read from the checks Lighthouse has recorded. The pages are complete without JavaScript; scripts only add to them.
Code internal/web/site.go internal/web/pages.go internal/status/status.go
More about this in the story →
Monitors and checks
What to watch, when it is next due, and every result, each row belonging to one tenant.
Row-level security keeps each visitor's sandbox apart from every other and from the public monitors. The uptime figures on the site are sums over these rows.
Code internal/store/monitors.go internal/store/checks.go internal/store/tenants.go
More about this in the story →
Incidents
Each outage, with a timeline of when it opened and resolved.
Written in the same transaction as the result that caused it, under a lock on the monitor's row, so two results arriving together can't open two incidents.
Code internal/store/incidents.go
Scheduler
Finds the monitors that are due and claims each one before checking it.
There is no queue service. Claiming is one update in the database that only one worker can win, so two instances running at once never check the same monitor twice.
Code internal/monitor/scheduler.go internal/store/monitors.go
More about this in the story →
Probe
Makes the request and times it, holding no database transaction while it waits.
A check can take 30 seconds. Holding a transaction open that long would tie up a database connection for every slow site, so the probe runs between two short transactions.
Code internal/monitor/probe.go
More about this in the story →
State machine
Decides whether a result changes anything. One failure isn't an outage; several in a row are.
A pure function with no database and no clock: the monitor's state and a result go in, the new state and what changed come out. That makes every edge testable, and property tests run it over thousands of random sequences of results.
Code internal/monitor/state.go
More about this in the story →
Alerts
Emails the owner when an incident opens or resolves.
Sent after the change is saved, never before, so an email is never about something the database doesn't know.
Code internal/alert/alert.go
Clock
Cloud Scheduler calls Lighthouse every 15 minutes with a signed request, and that is the only thing that starts a round of checks.
Lighthouse and its database both sleep when nothing is happening, which is why this runs for free. The request carries a token that proves it came from the scheduler's own account.
Code deploy/cloudrun/scheduler.tf internal/web/tick.go internal/oidc/oidc.go
Sites being watched
The demos, and whatever a visitor adds in a sandbox.
A visitor's monitor can point anywhere on the public internet and nowhere else: addresses inside the network are refused at the moment of connecting, after the name has been looked up.
Code internal/monitor/probe.go
More about this in the story →
Walk-through, step by step
From the page you are reading to the check behind it
Start with a page like this one and follow it back: where its figures are kept, what puts them there, and what happens when a site being watched stops answering.
-
You open a pageEdge → Public pages
Your request comes through the edge, the only way in. Lighthouse draws the page on the server and sends it whole.
-
Every figure is a rowPublic pages → Monitors and checks
The uptime and response times on the page are read from checks recorded in the database. So where do those rows come from?
-
A clock wakes itClock → Scheduler
Every 15 minutes Cloud Scheduler calls with a signed request. Lighthouse starts if it was asleep.
-
ClaimScheduler → Monitors and checks
One update claims each monitor that is due. If another instance is running, only one of them gets each monitor.
-
Hand overScheduler → Probe
The claim's transaction has ended. The probe starts with no database connection held.
-
ProbeProbe → Sites being watched
One request, through a guard that refuses internal addresses. This time the site doesn't answer.
-
What does it mean?Probe → State machine
A single failure might be a blip, so the monitor is checked again straight away. Enough failures in a row and the state machine says it is down.
-
Record and openState machine → Incidents
One transaction saves the result, the monitor's new state and the incident, under a lock on the monitor's row.
-
Tell someoneState machine → Alerts
Only after that transaction commits is the email sent.
-
Back on the pagePublic pages → Incidents
The next person to open the status page sees the incident, read from the row that was just written.
How we know: Tests run two schedulers against one database and check each monitor is checked once (internal/monitor/scheduler_test.go); property tests check the state machine changes state only after the set number of results in a row (internal/monitor/state_test.go).