phare-nix: Declarative Uptime Monitoring for NixOS

A NixOS module that keeps the monitors on phare.io in sync with the services on your server

In my post about a systemd service failure notification system I mentioned in a PS that, in addition to the notifications at the systemd level, I use an external monitoring service to make sure that all public-facing services are reachable via the internet. This post is about that second part and how it became declarative as well.

The services on my server are declared in my NixOS configuration. When I add a new website, I add a virtual host to services.nginx.virtualHosts, set forceSSL and enableACME, deploy, and NixOS takes care of the rest, including the SSL certificate. The monitoring, however, lived outside this world. For every new service I had to log into a web interface and create a monitor by hand. When I removed a service, I had to remember to remove the monitor as well. Usually I did not, and the monitoring configuration and the server configuration slowly drifted apart.

That is exactly the kind of problem NixOS solves for everything else on the server. So I wanted monitoring to be just another option: flick a switch on the virtual host and the monitor exists. As there is (to my knowledge) no standardized API for monitoring, I had to pick a concrete service. I chose phare.io, which has a well-documented API for everything I needed. The result is phare-nix, a NixOS module for phare.io monitor declarations.

What it looks like

After adding the flake to your inputs and importing phare-nix.nixosModules.phare, you enable the module and point it to a file containing a phare.io API token:

services.phare = {
  enable = true;
  tokenFile = "/run/secrets/phare-token";
  regions = [ "eu-deu-fra" "eu-swe-arn" ];
};

The most common case is an nginx virtual host. For those, a single option is enough:

services.nginx.virtualHosts."pascal-wittmann.de" = {
  forceSSL = true;
  enableACME = true;
  enablePhare = true;
};

This creates an HTTP monitor named after the virtual host. Because forceSSL is set, the monitor checks https://pascal-wittmann.de. If the defaults do not fit, every attribute of the monitor can be overridden via the phare option of the virtual host.

Everything that is not an nginx virtual host is declared directly. phare.io supports HTTP, TCP and ICMP monitors:

services.phare.monitors = {
  ssh = {
    protocol = "tcp";
    request = {
      host = "pascal-wittmann.de";
      port = 22;
      connection = "plain";
    };
  };

  ping = {
    protocol = "icmp";
    request.host = "pascal-wittmann.de";
  };
};

Run nixos-rebuild switch and the monitors appear on phare.io. Remove the enablePhare line, rebuild, and the monitor is paused.

How it works

The module consists of two parts: the Nix side, which turns the declarations into JSON, and a small Python tool called sync-with-phare, which talks to the phare.io API.

From Nix options to a JSON file

The options of a monitor are modeled after the phare.io API. The type of the request depends on the chosen protocol, so a TCP monitor with a url or an HTTP monitor without one is rejected. The regions, intervals and timeouts are enums of the values phare.io accepts, and length limits of the API (e.g. at most 45 characters for a monitor name) are checked with assertions. I find this important: a typo should fail the evaluation on my machine and not the sync on the server at 3 o'clock in the morning.

The enablePhare option for virtual hosts uses the same trick as my failure notification system. NixOS allows you to re-declare an option that already exists, and the module system merges the submodule types:

services.nginx.virtualHosts = mkOption {
  type = types.attrsOf (
    types.submodule {
      options.enablePhare = mkEnableOption "phare.io management for the virtual host";
      options.phare = mkOption {
        type = types.submodule monitorModule;
        default = { };
      };
    }
  );
};

All virtual hosts with enablePhare set are then turned into monitors and merged with the ones from services.phare.monitors. The result is written to the Nix store with builtins.toJSON. The Nix options are in camel case while phare.io expects snake case, so the keys are converted by the Python tool before they are sent.

Syncing on every rebuild

The JSON file is passed to a systemd service:

systemd.services.sync-phare-monitors = {
  wantedBy = [ "multi-user.target" ];
  wants = [ "network-online.target" ];
  after = [ "network-online.target" ];
  serviceConfig = {
    Type = "oneshot";
    RemainAfterExit = "yes";
    ExecStart = "${syncWithPhare}/bin/sync-with-phare sync-monitors --monitorfile ${monitors-json}";
    # ...
  };
};

There is no explicit "run this after a rebuild" hook here, and none is needed. The path of the JSON file is a store path that changes whenever the monitor declarations change. This changes the unit, and nixos-rebuild switch restarts every unit that changed. If nothing changed, nothing happens.

The sync itself is a reconciliation of the declared state with the state on phare.io. Monitors are identified by their name:

  • A monitor in the configuration but not on phare.io is created.
  • A monitor in both is updated if it differs.
  • A monitor in the configuration but paused on phare.io is resumed.
  • A monitor on phare.io but not in the configuration is paused.

I deliberately pause absent monitors instead of deleting them, as deleting a monitor also deletes its history. If you do not care about the history, you can set deleteAfterDays, and monitors that have been paused for longer are deleted.

Deciding whether a monitor "differs" turned out to be trickier than expected. phare.io returns more fields than it accepts: runtime state such as last_checked_at, and the defaults of fields that were never declared. A naive comparison considers every monitor as changed on every run. The tool therefore restricts the response of phare.io to the fields that are declared locally before comparing them.

The service runs with DynamicUser, gets the token via LoadCredential and is sandboxed with the usual systemd hardening options, so the script that talks to the internet does not see much of the system.

Maintenance windows

A rebuild restarts services. If one of them is down for a few seconds while a check runs, phare.io opens an incident, and I get a notification for something that is not a problem. With enableMaintenanceModeOnRebuild, an activation script pauses all monitors during a nixos-rebuild switch, and the sync service resumes them afterwards.

For planned work, maintenance windows can be declared next to the monitors:

services.phare.maintenanceWindows.database-upgrade = {
  monitoringMode = "pause";
  startsAt = "2026-10-15T14:30:00Z";
  endsAt = "2026-10-15T16:30:00Z";
  monitors = [ "pascal-wittmann.de" ];
};

The monitors are referenced by name and resolved to their phare.io ids when the window is synced, so the module checks at evaluation time that they exist.

My server upgrades itself every night with system.autoUpgrade. phare.io only knows maintenance windows with absolute timestamps; there is no "every night at 3". So phare-nix has rolling maintenance windows:

services.phare.rollingMaintenanceWindows.nixos-rebuild = {
  startAt = "03:00"; # same schedule as system.autoUpgrade.dates
  duration = "1h";
  monitors = [ "pascal-wittmann.de" ];
  cleanupAfterDays = 30;
};

startAt is a systemd OnCalendar expression. For every rolling window the module creates a systemd timer that fires on this schedule and a service that creates the concrete windows on phare.io: the one that just started and the next one. I did not want to reimplement systemd's calendar semantics in Python, so the tool simply asks systemd:

$ systemd-analyze calendar --iterations=1 03:00
  Original form: 03:00
Normalized form: *-*-* 03:00:00
    Next elapse: Tue 2026-10-06 03:00:00 CEST
       (in UTC): Tue 2026-10-06 01:00:00 UTC
       From now: 5h 12min left

The (in UTC) line is exactly what phare.io needs. As a bonus, the syntax and the timezone handling are identical to the timer of system.autoUpgrade. phare.io keeps completed windows as history, so cleanupAfterDays removes old ones after a while.

Incidents for failed systemd services

This brings me back to where I started. The email notification from my old post attaches an email@%n.service to the onFailure of every systemd service. phare-nix does the same, but instead of an email, it opens an incident on phare.io:

services.phare.failureNotify = {
  enable = true;
  impact = "partial_outage";
  monitors = [ "ssh" ];
};

phare.io requires an incident to affect at least one monitor. For services without a monitor of their own, a monitor that stands for the host itself (like the TCP check of the SSH port above) is a good fallback. Services that have a monitor can override this, e.g. systemd.services.nginx.phareIncident.monitors = [ "pascal-wittmann.de" ];.

The incident service records the id of the incident in /var/lib/phare-notify. While an incident for a service is open, further failures of that service do not open another one, which solves the "stop notifying after many failures in a row" problem I mentioned in my old post. A timer checks every minute whether the service has recovered and, if so, marks the incident as recovered on phare.io.

How incidents are communicated is configured on phare.io with alert rules, which can be declared as well. Incidents opened by phare-nix are not detected by a monitor, so they are selected with the event setting type = "manual":

services.phare.alertRules.service-failed = {
  event = "uptime.incident.created";
  eventSettings.type = "manual";
  integrationId = 1234; # e.g. an email or ntfy integration
};

This takes care of the other improvement I mentioned back then: channels other than email. phare.io handles email, ntfy, Slack, webhooks and so on, and I only need to pick one.

Testing

A tool that pauses and deletes monitors in an account somewhere on the internet should be tested. The flake contains a NixOS VM test that runs the module against a small mock of the phare.io API written in Python, and a faster test that runs without KVM. The latter is part of nix flake check and runs in CI on every push.

Conclusion

My monitoring is now part of my server configuration. A new service gets a monitor in the same commit that introduces it, removing a service pauses its monitor, and nightly upgrades no longer wake anybody up. The first proof of concept was written in March 2025, and version 1.2.0 was released today.

You can find the source code at codeberg.org/pSub/phare-nix and the documentation of all options at phare-nix.quine.de. If you use it and something does not work as expected, feel free to open an issue.

Comments

Kommentare für diesen Eintrag als RSS Feed

No comments

Leave a Reply