openstatus logoPricingDashboard

Incident management for small teams

Declare an incident from Slack or the dashboard, run it in its own channel, keep one timeline and leave with a postmortem. Next to the monitors and the status page it belongs to.

Free to start, with a 14-day Starter trial and no card needed. Paid plans from $30/mo.

Checkout API 503s in EUIncident #42
Severity
Major
Status
Resolved
Commander
Gilfoyle
Duration
55m
Timeline
11:20Postmortem approvedGilfoyle
11:02Postmortem draftedAgent
10:36ResolvedGilfoyle
10:14MitigatedGilfoyleRollback complete in eu-west-1 and eu-west-2.
09:49NoteGilfoyle · via SlackEdge config deploy at 09:38 changed EU routing. Rolling back.
09:44Status report linkedGilfoyleElevated errors on Checkout API
09:43DeclaredGilfoyleMajor · Checkout API 503s in EU
Started 09:41 UTC, when the monitor alerted7 events

The incident page for "Checkout API 503s in EU" at the fictional Pied Piper: severity Major, status Resolved, commander Gilfoyle, duration 55m, and a timeline from Declared at 09:43 through a note pinned in Slack, the linked status report, Mitigated at 10:14 and Resolved at 10:36 to the approved postmortem at 11:20.

Trusted by teams who ship transparency

Cal.comTwentyDocumensoTraefikPassboltHankoWhiteBITSuperwallOpenPanelProboRoundtableSmplrspace

Why run incidents here

It starts where the alert lands

The monitor that caught the failure, the incident and the status page live in one workspace. Nothing is copied between tools while the Checkout API is down.

Internal and public stay apart

The team works with severities, notes and a commander. Customers read a status report in your words. The two are linked, never mixed.

The record writes itself

Every status change, note and public update is timestamped as it happens. The postmortem and the audit trail are built from it, not from memory.

Declare from Slack or the dashboard

Run /openstatus incident declare with a title and a severity, approve the card, and the incident exists. The dashboard has the same form, with a commander, a start time you can set in the past and an open status report to link.

#incidentsslash command
G
Gilfoyle09:43
/openstatus incident declare Checkout API 503s in EU --sev major
os
openstatusAPP09:43
Declare incident
Title
Checkout API 503s in EU
Severity
Major
os
openstatusAPP09:43
Incident Checkout API 503s in EU declared (major). Open in openstatus
Channel #inc-2026-10-05-checkout-api-503s-in-eu openedGilfoyle invited

Gilfoyle runs /openstatus incident declare Checkout API 503s in EU --sev major in #incidents at 09:43. The openstatus app posts an approval card with the title and severity, then confirms the incident is declared and that a dedicated channel was opened with Gilfoyle invited.

A Slack channel per incident

#inc-2026-10-05-checkout-api-503s-in-euMAJOR · open
os
openstatusAPP09:43 · pinned
Checkout API 503s in EU
Severity
Major
Status
Open
Open in openstatus · Pin a message with 📌 to add it to the timeline.
G
Gilfoyle09:49
Edge config deploy at 09:38 changed EU routing. Rolling back.
📌 1✅ 1
Pinned messages land on the timelineReminder after 4h without an update

The incident's own Slack channel, named inc-, the date and checkout-api-503s-in-eu. The pinned card shows severity Major and status Open. Gilfoyle's 09:49 message, "Edge config deploy at 09:38 changed EU routing. Rolling back.", carries a pin reaction and a check mark from the app: it is now a note on the timeline. A reminder follows after 4 hours without an update.

Every incident gets a channel with the severity and status in its topic and the incident card pinned on top. React to any message with 📌 and it becomes a timeline note, with a link back to Slack.

If the incident goes quiet, the channel gets a reminder: after one hour for critical, four for major and a day for minor.

From the monitor that caught it

When a monitor goes down, declare the incident from its downtime row and the start time is already filled in. The duration counts from when the impact began, not from when someone got around to declaring it.

#incidentsopenstatus · alert
os
openstatusAPP09:41
Checkout API is failing
POST https://api.piedpiper.dev/v1/checkout
Status
503
Regions
lhr, ams, cdg, koyeb_fra
Latency
4,388 ms
Cron Timestamp
2026-10-05T09:41:12.000Z
Error
Expected status code 200, received 503
Also PagerDuty paged · Email sent · Webhook 2004 of 6 regions failed the assertion

The Slack alert that started it at 09:41: Checkout API returned 503 from lhr, ams, cdg and koyeb_fra at 4,388 ms, with 4 of 6 regions failing the status-code assertion.

Internal status, public update

#inc-2026-10-05-checkout-api-503s-in-euslash command
G
Gilfoyle10:14
/openstatus incident mitigate Rollback complete in eu-west-1 and eu-west-2.
os
openstatusAPP10:14
Incident Checkout API 503s in EU is now mitigated.
Linked status report
Elevated errors on Checkout APIMonitoring
The rollback has completed in eu-west-1 and eu-west-2. Error rates are back to baseline; we are monitoring for the next hour.
Internal status and public update move separately1,337 subscribers notified

At 10:14 Gilfoyle runs /openstatus incident mitigate Rollback complete in eu-west-1 and eu-west-2. in the incident channel and the app confirms the incident is now mitigated. Below it, the linked status report "Elevated errors on Checkout API" reads Monitoring, with its own public message and 1,337 subscribers notified.

Mitigate, resolve, cancel or reopen from the channel or the dashboard, with a note that lands on the timeline. The linked status report keeps its own status and wording, and you post the public update from the same page.

A postmortem with a first draft

Once the incident is resolved, the agent drafts a blameless postmortem from the timeline, the public updates and the channel history. Where it has no facts it says so instead of guessing.

Edit it in markdown, approve it, and close the incident.

PostmortemDraft · drafted by the agent
Summary

A configuration change in the European edge made the Checkout API return 503 for requests routed through Europe.

Impact

Checkout requests routed through Europe failed for 55m, from 09:41 to 10:36 UTC. Payments outside Europe were not affected.

Timeline
  • 09:43 UTC Declared
  • 09:44 UTC Status report linked
  • 09:49 UTC Note
  • 10:14 UTC Mitigated
  • 10:36 UTC Resolved
Root cause

The 09:38 edge config deploy changed EU routing and was not validated against a European region before rollout.

What went well

The monitor alerted within a minute and the first public update went out three minutes later.

What went wrong

The deploy had no canary, so every European region failed at once.

Action items
  • [ ] Canary edge config changes in one EU region first
  • [ ] Add a pre-deploy check from lhr and koyeb_fra
Built from the timeline, public updates and the Slack channelApprove, then close the incident

The postmortem draft, written by the agent, in seven sections: summary, impact (55m, from 09:41 to 10:36 UTC, Europe only), a UTC timeline from Declared to Resolved, root cause (the 09:38 edge config deploy), what went well, what went wrong, and two action items as a checklist. Buttons to approve it or draft again with the agent.

For agents, with a paper trail

Audit log13 events
11:20:48incident_postmortem.updategilfoyle@piedpiper.dev · slack→ approved
11:02:14incident_postmortem.creategilfoyle@piedpiper.dev · slackdraft · agent
10:36:20incident.updategilfoyle@piedpiper.dev · slack→ resolved
10:36:00status_report.updategilfoyle@piedpiper.dev · slack→ resolved
10:14:25incident.updategilfoyle@piedpiper.dev · slack→ mitigated
10:14:03status_report.updategilfoyle@piedpiper.dev · slack→ monitoring
09:52:41notification.sendsystememail 1,337 · rss · slack-connect
09:52:40status_report.updategilfoyle@piedpiper.dev · slack→ identified
09:49:17incident_event.creategilfoyle@piedpiper.dev · slacknote · pinned in slack
09:44:31notification.sendsystememail 1,337 · rss · slack-connect
09:44:30status_report.creategilfoyle@piedpiper.dev · slackinvestigating · Checkout API
09:43:08incident.creategilfoyle@piedpiper.dev · slackmajor · Checkout API 503s in EU
09:41:12monitor.alertprobeCheckout API · lhr, ams, cdg, koyeb_fra
Every mutation, with actor and sourceexport CSV / JSON

The audit log for the same incident, with the newest entry at the top. From oldest to newest: monitor.alert at 09:41:12, incident.create at 09:43:08 by gilfoyle@piedpiper.dev via Slack, the status report and its updates, incident.update to mitigated and resolved, and the postmortem created by the agent and approved at 11:20:48.

The MCP server and the Slack agent share the incident tools: declare, update, change status, add a note, draft the postmortem. Anything that changes state asks for approval first.

Every one of those steps is in the audit log with the person behind it, whichever surface they used.

Get started

Free to start, with a 14-day Starter trial and no card needed. Paid plans from $30/mo.

Frequently asked questions

What is incident management?

Incident management is how a team responds when something breaks: someone declares the incident, one person takes command, the team works the problem in one place, customers are kept informed, and afterwards a postmortem records what happened and what changes. Openstatus keeps the declaration, the timeline, the public status report and the postmortem in one record.

What is the difference between an incident and a status report?

An incident is internal: severity, commander, timeline notes and the postmortem are for your team. A status report is public: it is what customers read on your status page and receive as a subscriber. You link the two, and each keeps its own status, so the team can mark an incident mitigated while the status report still says monitoring.

Do I need Slack to manage incidents?

No. You can declare, update, resolve and write the postmortem from the dashboard. Connecting Slack adds the channel per incident, slash commands, pinning messages to the timeline, stale-incident reminders and agent drafting.

How does the postmortem draft work?

Once an incident is resolved, the agent drafts a blameless postmortem from the incident's fields, its timeline, the public status report updates and the history of the incident's Slack channel. The draft has seven sections: summary, impact, timeline, root cause, what went well, what went wrong and action items. Where it has no facts it says what is unknown. You edit the draft, approve it and close the incident.

Which severities and statuses are there?

Three severities: critical for a major outage or data loss, major for significant degradation, and minor for limited impact. An incident is open, mitigated, resolved or canceled. Closing an incident freezes it: status, severity, commander and notes can no longer change.

Can an AI agent run an incident?

Yes. The MCP server and the Slack agent share the same incident tools: declare, update, change status, add a note, draft and approve the postmortem. Every action that changes something asks for approval first and is written to the audit log with the person behind it.