Incident management for small teams
Declare an incident from Slack or the dashboard, run it in its own channel, keep one timeline and leave with a postmortem. Next to the monitors and the status page it belongs to.
Free to start, with a 14-day Starter trial and no card needed. Paid plans from $30/mo.
The incident page for "Checkout API 503s in EU" at the fictional Pied Piper: severity Major, status Resolved, commander Gilfoyle, duration 55m, and a timeline from Declared at 09:43 through a note pinned in Slack, the linked status report, Mitigated at 10:14 and Resolved at 10:36 to the approved postmortem at 11:20.
Trusted by teams who ship transparency
Why run incidents here
It starts where the alert lands
The monitor that caught the failure, the incident and the status page live in one workspace. Nothing is copied between tools while the Checkout API is down.
Internal and public stay apart
The team works with severities, notes and a commander. Customers read a status report in your words. The two are linked, never mixed.
The record writes itself
Every status change, note and public update is timestamped as it happens. The postmortem and the audit trail are built from it, not from memory.
Declare from Slack or the dashboard
Run /openstatus incident declare with a title and a severity, approve the card, and the incident exists. The dashboard has the same form, with a commander, a start time you can set in the past and an open status report to link.
/openstatus incident declare Checkout API 503s in EU --sev major- Title
- Checkout API 503s in EU
- Severity
- Major
Gilfoyle runs /openstatus incident declare Checkout API 503s in EU --sev major in #incidents at 09:43. The openstatus app posts an approval card with the title and severity, then confirms the incident is declared and that a dedicated channel was opened with Gilfoyle invited.
A Slack channel per incident
The incident's own Slack channel, named inc-, the date and checkout-api-503s-in-eu. The pinned card shows severity Major and status Open. Gilfoyle's 09:49 message, "Edge config deploy at 09:38 changed EU routing. Rolling back.", carries a pin reaction and a check mark from the app: it is now a note on the timeline. A reminder follows after 4 hours without an update.
Every incident gets a channel with the severity and status in its topic and the incident card pinned on top. React to any message with 📌 and it becomes a timeline note, with a link back to Slack.
If the incident goes quiet, the channel gets a reminder: after one hour for critical, four for major and a day for minor.
From the monitor that caught it
When a monitor goes down, declare the incident from its downtime row and the start time is already filled in. The duration counts from when the impact began, not from when someone got around to declaring it.
POST https://api.piedpiper.dev/v1/checkoutExpected status code 200, received 503
The Slack alert that started it at 09:41: Checkout API returned 503 from lhr, ams, cdg and koyeb_fra at 4,388 ms, with 4 of 6 regions failing the status-code assertion.
Internal status, public update
/openstatus incident mitigate Rollback complete in eu-west-1 and eu-west-2.At 10:14 Gilfoyle runs /openstatus incident mitigate Rollback complete in eu-west-1 and eu-west-2. in the incident channel and the app confirms the incident is now mitigated. Below it, the linked status report "Elevated errors on Checkout API" reads Monitoring, with its own public message and 1,337 subscribers notified.
Mitigate, resolve, cancel or reopen from the channel or the dashboard, with a note that lands on the timeline. The linked status report keeps its own status and wording, and you post the public update from the same page.
A postmortem with a first draft
Once the incident is resolved, the agent drafts a blameless postmortem from the timeline, the public updates and the channel history. Where it has no facts it says so instead of guessing.
Edit it in markdown, approve it, and close the incident.
A configuration change in the European edge made the Checkout API return 503 for requests routed through Europe.
Checkout requests routed through Europe failed for 55m, from 09:41 to 10:36 UTC. Payments outside Europe were not affected.
- 09:43 UTC Declared
- 09:44 UTC Status report linked
- 09:49 UTC Note
- 10:14 UTC Mitigated
- 10:36 UTC Resolved
The 09:38 edge config deploy changed EU routing and was not validated against a European region before rollout.
The monitor alerted within a minute and the first public update went out three minutes later.
The deploy had no canary, so every European region failed at once.
- [ ] Canary edge config changes in one EU region first
- [ ] Add a pre-deploy check from lhr and koyeb_fra
The postmortem draft, written by the agent, in seven sections: summary, impact (55m, from 09:41 to 10:36 UTC, Europe only), a UTC timeline from Declared to Resolved, root cause (the 09:38 edge config deploy), what went well, what went wrong, and two action items as a checklist. Buttons to approve it or draft again with the agent.
For agents, with a paper trail
The audit log for the same incident, with the newest entry at the top. From oldest to newest: monitor.alert at 09:41:12, incident.create at 09:43:08 by gilfoyle@piedpiper.dev via Slack, the status report and its updates, incident.update to mitigated and resolved, and the postmortem created by the agent and approved at 11:20:48.
The MCP server and the Slack agent share the incident tools: declare, update, change status, add a note, draft the postmortem. Anything that changes state asks for approval first.
Every one of those steps is in the audit log with the person behind it, whichever surface they used.
Get started
Free to start, with a 14-day Starter trial and no card needed. Paid plans from $30/mo.Frequently asked questions
What is incident management?
Incident management is how a team responds when something breaks: someone declares the incident, one person takes command, the team works the problem in one place, customers are kept informed, and afterwards a postmortem records what happened and what changes. Openstatus keeps the declaration, the timeline, the public status report and the postmortem in one record.
What is the difference between an incident and a status report?
An incident is internal: severity, commander, timeline notes and the postmortem are for your team. A status report is public: it is what customers read on your status page and receive as a subscriber. You link the two, and each keeps its own status, so the team can mark an incident mitigated while the status report still says monitoring.
Do I need Slack to manage incidents?
No. You can declare, update, resolve and write the postmortem from the dashboard. Connecting Slack adds the channel per incident, slash commands, pinning messages to the timeline, stale-incident reminders and agent drafting.
How does the postmortem draft work?
Once an incident is resolved, the agent drafts a blameless postmortem from the incident's fields, its timeline, the public status report updates and the history of the incident's Slack channel. The draft has seven sections: summary, impact, timeline, root cause, what went well, what went wrong and action items. Where it has no facts it says what is unknown. You edit the draft, approve it and close the incident.
Which severities and statuses are there?
Three severities: critical for a major outage or data loss, major for significant degradation, and minor for limited impact. An incident is open, mitigated, resolved or canceled. Closing an incident freezes it: status, severity, commander and notes can no longer change.
Can an AI agent run an incident?
Yes. The MCP server and the Slack agent share the same incident tools: declare, update, change status, add a note, draft and approve the postmortem. Every action that changes something asks for approval first and is written to the audit log with the person behind it.

