Staff Incident Manager

July 9, 2026
Application ends: October 6, 2026

Job Description

REQUIREMENTS

  • Proven experience running major incident response in a production environment, ideally as an incident commander.
  • Strong working knowledge of modern distributed systems and cloud-native operations.
  • Experience defining or maturing incident management practices, including severity frameworks and runbooks.
  • Proficiency with observability and incident tooling (metrics, logging, tracing, alerting).
  • Hands-on experience with Datadog is strongly preferred.
  • Familiarity with paging platforms such as PagerDuty or OpsGenie.
  • Excellent written and verbal communication skills, with the ability to remain calm and decisive under pressure.
  • Availability for on-call responsibilities, including critical incidents outside of standard business hours.
  • Advanced proficiency in English.

Preferred

  • Familiarity with PCI-DSS and the operational obligations of handling payment flows at scale.
  • Business-level proficiency in Spanish.
  • Prior experience in payments, fintech, or other high-availability, regulated domains.
  • Exposure to SRE practices, error budgets, and SLO-driven prioritization.
  • Experience managing incidents across multi-timezone, remote-first engineering organizations.

RESPONSIBILITIES

  • Act as Incident Commander for major and critical incidents, owning coordination, decision-making, and escalation from detection through resolution.
  • Drive down MTTR (Mean Time to Recover) and own the operational discipline required to meet 99.99% uptime targets.
  • Manage the on-call program, including rotations, escalation policies, paging hygiene, and alert quality.
  • Lead internal stakeholder alignment and drive clear, accurate merchant-facing updates via status pages.
  • Conduct blameless postmortems and ensure action items are concrete, owned, and tracked to closure.
  • Partner with engineering teams to translate recurring incident patterns into reliability roadmap items.
  • Define and maintain incident severity levels, response runbooks, and the overall incident operating model.
  • Report on incident trends, reliability posture, and SLA/SLO performance to engineering leadership.

Are you interested in this position?


Apply by clicking on the “Apply Now” button below!

#CrossChannelJobs #JobSearch
#CareerOpportunities #HiringNow
#Employment #JobOpenings
#JobSeekers
#FacebookLinkedIn

Related Jobs