Infrastructure monitoring

See insideevery server you run.

The Capture agent reports CPU, memory, disk, network and temperature from your machines through a small REST API. The data lands in your database and nowhere else.

prod-web-01Online
Capture agent · reporting every 15 s
CPU
Memory
Disk
Temp
Network throughput
InOut
Per-core usage
C1C2C3C4C5C6C7C8
Host
Uptime42 days
Load avg0.61 · 0.54 · 0.48
Swap0.2 / 4 GB
Compute

CPU and memory, with history.

Per-core CPU usage, load averages and memory pressure, charted over time. A box that creeps from 60 to 95 percent over a month is a capacity decision you get to make early, instead of an outage that makes it for you.

The dashboard shows current gauges with the history behind them, so a spike and a trend look different at a glance.

  • Per-core usage and load averages
  • Memory pressure over time
  • Instant gauges backed by history charts
web-01 · CPU8 cores
Usage
46%
all cores
Load 1m
1.84
5m 1.62 · 15m 1.41
Memory
9.2 GB
of 16 GB
Per core
cpu0
34%
cpu1
61%
cpu2
28%
cpu3
47%
cpu4
82%
cpu5
39%
cpu6
55%
cpu7
44%
Storage

Disks that warn you before they fail.

Capacity and I/O tell you when a disk is filling or struggling. S.M.A.R.T. health indicators tell you when it is dying. Checkmate reports both, for spinning drives and solid-state alike.

Full disks remain one of the most preventable outages in self-hosted setups. This is the prevention.

  • Capacity and I/O per disk
  • S.M.A.R.T. health indicators
  • Works with HDDs and SSDs
storage-01 · DisksS.M.A.R.T. on
/dev/nvme0n1
38 MB/sHealthy
62% used
/dev/sda
6 MB/sHealthy
87% used
/dev/sdb
2 MB/sWarning
41% used
Network and heat

Throughput and temperature, per machine.

Inbound and outbound bandwidth per interface, with historical charts that make a saturated uplink or a chatty backup job easy to spot.

Temperature readings round out the picture, which matters when your servers live in a closet instead of a climate-controlled datacenter.

  • Bandwidth per interface, in and out
  • Historical network charts
  • Temperature readings per machine
web-01 · Networketh0
Inbound
184 Mb/s
Outbound
97 Mb/s
Temperature
54°C
package
inout
Thresholds

Alerts when a line gets crossed.

Set thresholds on hardware metrics and get paged when a machine crosses one. The alert arrives on the same 12 channels as your uptime alerts, so infrastructure and availability page the same way.

No second alerting stack to configure, no second place to check.

  • Hardware thresholds you define
  • Same 12 notification channels
  • One alerting pipeline for everything
Hardware thresholdsweb-01
CPU usage above 90%
Memory above 85%
Disk usage above 80%
Temperature above 80°C
Alerts to:EmailSlack
Who runs this

From one closet server to a rack of them.

Homelabs

Watch temperatures, disks and load on the hardware in your house without sending any of it to someone else's cloud.

Ops teams

One dashboard for a fleet's health, with thresholds that page the on-call channel before capacity problems become incidents.

Cost-conscious teams

Per-host pricing on hosted monitoring adds up fast. Capture is open source, so another server is another container, not another line item.

Scope

What it covers, and what it doesn't.

Covered

  • Per-core CPU usage and load averages
  • Memory pressure with history
  • Disk capacity, I/O and S.M.A.R.T. health
  • Network throughput per interface
  • Temperature readings
  • Docker container state alongside host metrics
  • Hardware thresholds with alerts

Out of scope

  • Application performance monitoring: no code-level tracing or profiling
  • Log collection and search for your applications
  • Custom application metrics: Capture reports hardware, not business counters
Under the hood

For the technically curious.

A small, open agent

Capture is its own open-source project (bluewave-labs/capture on GitHub). It runs on the machines you monitor and exposes readings over a small REST API that your Checkmate instance polls.

Nothing leaves your perimeter

The agent talks to your Checkmate server, results land in your database and no telemetry goes anywhere else. Air-gapped and compliance-sensitive environments stay that way.

Auditable, like the rest

Agents are a trust decision: this one ships its source. You can read what Capture collects before you put it on a production box.

Uptime and hardware, one view

Infrastructure monitors sit in the same dashboard as uptime, Docker and page speed monitors. When a service goes down, the host's CPU, memory and disk history is one click away.

FAQ

Frequently askedquestions.

CPU per core with load averages, memory pressure, disk capacity, I/O and S.M.A.R.T. health, network throughput per interface and temperature. Docker container state is reported alongside the host metrics.

Yes, for hardware metrics. The Capture agent runs on each machine you want to watch and reports through a small REST API. It is open source, so you can audit exactly what it collects.

No. Capture reports to your own Checkmate instance and the data stays in your database. There is no SaaS backend and no telemetry pipeline.

Yes. Set thresholds on hardware metrics and Checkmate alerts on the same 12 channels used for uptime monitoring, so your whole alerting story lives in one place.

Yes. Container status, health checks and per-container resource usage are covered, through Docker monitors as well as the host metrics Capture reports.

Yes. Checkmate and Capture are both open source under AGPL-3.0. Monitor as many machines as you like with no per-host pricing.

Get started

Every feature,no paywall.

Checkmate is open source under AGPL-3.0. Self-host it and this feature ships free, on your servers, with your data.