Incident management

Every outage,on the record.

Checkmate opens an incident on the first failed check, closes it on recovery and keeps the history. What broke, when, for how long and who resolved it: answered from a dashboard instead of a spreadsheet.

Incidents1 active
api.example.com4 m agoongoingactive
postgres-primaryyesterday12 mauto
mail.example.com3 days ago41 mmanual
cdn-edgelast week3 mauto
Opened on the first failed check, closed on recovery
Automatic

Opened and closed by the checks themselves.

The first failed check opens an incident and fires alerts on the monitor's channels. When a check passes again, the incident resolves automatically and a recovery notification follows, so the timeline writes itself while you work the problem.

Start time, end time and duration come from the checks, not from whoever remembered to update a ticket at 3am.

  • Incident opened on the first failed check
  • Resolved automatically on recovery
  • Recovery notifications on the same channels
api.example.com7 m 21 s
14:02:11Check failed · connection timeout
14:02:14Incident opened · alerts sent
14:09:32Check passed · 200 in 210 ms
14:09:32Incident resolved automatically
Start, duration and recovery recorded without anyone typing
Manual resolution

Close it yourself, comment included.

Fixed the problem out of band and don't want to wait for the next scheduled check? Resolve the incident manually from the incidents page, with an optional comment on what you did.

The resolution records who closed it and what they wrote. 6 months later, when the same alert fires, the comment from last time is the head start.

  • Manual resolve with an optional comment
  • Resolver's identity stored with the incident
  • Automatic and manual closes both labelled
Resolve incidentmail.example.com
Comment (optional)
Failed over to the backup SMTP relay. Root cause was a full disk on mail-01, cleaned up and re-queued.
Resolve manuallyRecorded as manual · by ops@example.com
History

The history your SLA report is missing.

Filter incidents by monitor, date range and status, with durations and an average resolution time computed for you. That is the raw material for SLA reporting and the fastest way to spot patterns.

Incidents that land at the same hour usually trace to a cron job. Durations that keep growing point at a slowing recovery process rather than new failures. The history is where those patterns become visible.

  • Filter by monitor, date range and status
  • Average resolution time across the selection
  • Every incident keeps its full timeline
Last 30 daysall monitors
Incidents
7
down from 11
Avg resolution
18 m
across all incidents
Auto-resolved
5
recovered on their own
Manual
2
closed with a comment
Filter by monitor, date range or resolution type
Who runs this

For teams that outgrew "was it down yesterday?"

Small ops teams

No dedicated incident commander, no war room. The incident record assembles itself from the checks, which is the only process a 3-person team will consistently follow.

Anyone reporting against an SLA

Agencies and MSPs with uptime commitments need defensible numbers. Incident durations and resolution times come straight from check data, on your own servers.

Teams that do postmortems

The timeline answers when it started, when it ended and how long it took before the meeting starts, and the resolution comment holds the one-line summary of the fix.

Scope

What it covers, and what it doesn't.

Covered

  • Incidents opened automatically on the first failed check
  • Automatic close on recovery, with recovery notifications
  • Manual resolution with an optional comment
  • Resolver identity recorded on manual closes
  • Duration and average resolution time
  • Filters by monitor, date range and status

Out of scope

  • On-call scheduling and escalation policies: pair Checkmate with its native PagerDuty channel for paging workflows
  • Public incident communication: that lives on your status page, which reads from the same monitors
  • Long-form postmortem documents: the comment holds the summary, the full writeup belongs in your wiki
Under the hood

For the technically curious.

A span, not a log line

An incident is the stretch between the first failed check and the recovery, one record per outage rather than a pile of alerts. Active incidents show elapsed time as they run.

Resolution type on every incident

Every incident closes as automatic or manual, and the label is stored. You can see at a glance how much of your recovery is self-healing and how much needed hands.

Who closed it, and why

Manual resolutions store the resolving user and the comment they left. The audit trail exists without anyone maintaining it, which is how audit trails survive.

Resolution time, aggregated

Average resolution time is computed across whatever filter you are looking at, per monitor, per month or fleet-wide. The number your SLA conversation starts from.

FAQ

Frequently askedquestions.

Automatically. The first failed check on a monitor opens an incident and fires alerts, and the incident stays active until a check passes again or someone resolves it manually. No manual declaration step.

Yes. The incidents page offers a manual resolve on every active incident, with an optional comment. The incident records the resolver and the comment, labelled as a manual resolution.

Yes. Every incident keeps its start time, end time, duration, resolution type and, for manual closes, who resolved it and their comment. The history is filterable by monitor, date range and status.

That is what it is for. Durations and average resolution time are computed from check data on your own instance, so the numbers behind an SLA report are yours to verify and export.

Through its notification channels, yes: PagerDuty is a native channel, alongside Slack, email and 9 others. Escalation policies and rotations stay in your paging tool, Checkmate feeds it.

No. Like everything in Checkmate it ships in the open-source core under AGPL-3.0, self-hosted on your servers with no tiers to unlock.

Get started

Every feature,no paywall.

Checkmate is open source under AGPL-3.0. Self-host it and this feature ships free, on your servers, with your data.