Skip to content
A live portfolio. Every project on this page is running now, and every figure is measured by the site itself. Wednesday 30 September 2026 · Nairobi
The LighthouseThe engineering portfolio of Martin Omwenga
Projects / Lighthouse · Infrastructure · The system behind this page

The website that monitors itself, logs its own failures and publishes the evidence

Every 15 minutes it checks each project on this page. When one goes down it opens an incident, writes up what happened, and closes the case when service returns. Nobody has to press a button.

Lighthouse serves these pages, checks every project on a schedule, and keeps its results and incidents in PostgreSQL.

The problem

A portfolio of live demos has two problems. Demos go down, and a recruiter who clicks a dead link doesn't come back. And a demo dropped on someone without context is half wasted: they don't know what to try, or why it was hard to build.

Lighthouse turns both into features. It watches every project and publishes what it finds, including its failures, and every launch starts with a short technical introduction while it confirms the demo is up, so nobody lands on a broken page.

How it works

A scheduler finds monitors that are due and checks them with a bounded pool of workers: HTTP checks for the real projects, simulated ones for the public sandbox. Each result goes through a small state machine that opens an incident after several failures in a row and closes it after several successes, so one blip never pages anyone and a flapping site doesn't open an incident every minute.

It also counts its readers without cookies: a visitor is a hash of their IP address keyed with a salt that is deleted every day, browsers asking not to be tracked are skipped, and public numbers under five are hidden.

Everything is one Go binary: server-rendered public pages, plus a small Vue console, built into the binary, where the owner and sandbox visitors manage monitors and incidents. The owner signs in with GitHub (one account is admitted); visitors can create a throwaway sandbox with its own monitors, isolated from everything else by PostgreSQL row-level security. When a visitor launches a demo, the launch page polls the demo's health address, which wakes it (idle demos sleep, to cost nothing), and opens the demo when it answers.

“Every figure on these pages is measured by the site itself, and so is every outage.”
1ClaimAn atomic update claims a due monitor, so no two workers or instances check it twice.
2ProbeThe check runs outside any database transaction, through a guard that blocks internal addresses.
3RecordThe result and the monitor's new state are saved together, under a row lock.
4Open or resolveCrossing a threshold opens or resolves the incident, with a public timeline entry.
How a check becomes an incident, and how it closes again.

Key decisions

What was chosen, why, and what it costs.

The database is the queue

The choice. Due monitors are claimed with one conditional UPDATE, instead of a separate queue or lock service.

Why. Nothing extra to run on a small cluster, and instances can be added freely. A test runs four schedulers against the same monitors and checks each is checked exactly once.

The trade-off. Scheduling precision is bounded by the polling tick (one second), which is plenty for uptime checks.

Probes never hold a transaction

The choice. Claim in one short transaction, probe with no database connection held, record in another.

Why. A check can take up to 30 seconds. Holding a connection that long would cap concurrency at the pool size.

The trade-off. A monitor deleted while its probe runs has to be handled on the way back, and is.

Guard against SSRF after DNS resolution

The choice. The dialer checks the address actually being connected to, rejecting private, loopback, link-local and shared ranges.

Why. Checking the URL alone misses hostnames that resolve inward, DNS rebinding and redirects. Checking at dial time catches all three.

The trade-off. The owner has to opt in per monitor to watch services inside the cluster.

Hysteresis in a pure state machine

The choice. The rules for opening and resolving incidents live in one pure function, separate from the database.

Why. It can be tested exhaustively. Property tests run random check sequences and assert an incident is never opened twice, never resolved without one, and always matches the monitor's health.

The trade-off. None worth noting; it made the scheduler simpler.

Server-rendered, and complete without JavaScript

The choice. Go templates for every public page; the only script is the launch page's slideshow.

Why. These pages are read far more than anything else, and a status page has to work when things are broken. Go's templates escape by context.

The trade-off. The console is the exception: a small Vue app, because editing live data every few seconds wants client-side state. It is built into the same binary and held to the same Content Security Policy.

Visibility fixed at creation

The choice. An incident's public flag is copied from its monitor when it opens.

Why. Deriving it at read time would have made a private monitor's incidents public the moment the monitor was deleted.

The trade-off. Changing a monitor's visibility doesn't change its past incidents.

How it is tested

Integration tests run against real PostgreSQL in containers, each test with its own database cloned from a migrated template in milliseconds. They try to read, change and link to another tenant's rows, and to switch row-level security off as the application's role. The HTTP tests sign in against a fake GitHub (owner, stranger, forged state) and check the public pages leak nothing private.

Everything runs under Go's race detector, with property tests (rapid) and fuzzing for the input validation and the charts. The key tests were checked to fail with their protection removed: the RLS policy, atomic claiming, the origin check, the owner check and private-incident hiding. CI adds staticcheck, govulncheck, gitleaks, CodeQL and a Trivy scan of the image.

What it doesn't do yet

Checks run every 15 minutes, not every few seconds: the databases are on a free plan with a monthly compute allowance, and a database woken every minute would use it up by mid-month. So an outage can take up to about 45 minutes to become an incident. Aggregations are computed from raw checks on each request, which is fine at this scale; daily rollups would be the first optimization. The Kubernetes deployment (k3s, Flux, Cloudflare Tunnels) is written and proven on a local cluster, but isn't the live one.

Try it yourself Start a sandbox, break a site, and watch the incident open by itself.
← All projects