Lead Infrastructure & Site Reliability Engineering

September 9, 2026
Application ends: December 8, 2026

Job Description

REQUIREMENTS

  • Bachelor’s degree in Computer Science, or related field—or equivalent experience.
  • 8+ years in infrastructure, site reliability, platform, or DevOps engineering, with recent hands-on delivery.
  • Deep, hands-on expertise operating production workloads on Microsoft Azure.
  • Proven experience with Infrastructure as Code using Bicep and/or Terraform.
  • Hands-on experience with observability tooling across metrics, logs, and traces (e.g., Grafana, Prometheus, Azure Application Insights).
  • Proven experience defining and operating SLIs/SLOs and error budgets.
  • Hands-on experience securing Azure cloud environments—identity/access, network security, posture management (e.g., Microsoft Defender for Cloud, Azure Policy), and vulnerability management.
  • Experience scaling Microsoft Fabric and Microsoft Purview.
  • Experience owning incident management, on-call, and blameless post-incident reviews (e.g., Better Stack, PagerDuty, or comparable).
  • Strong scripting/automation ability (e.g., PowerShell, Python, Bash) and CI/CD experience (Azure DevOps preferred).
  • Experience supporting business continuity and disaster recovery with defined RTO/RPO.
  • Experience building internal developer platforms, golden paths, and self-service infrastructure.
  • Cloud cost optimization / FinOps experience.
  • Exposure to AI-assisted operations (AIOps) and modern reliability automation.
  • Microsoft Azure certifications (e.g., Azure Solutions Architect Expert, Azure DevOps Engineer Expert, Azure Security Engineer Associate) preferred.
  • Experience in insurance, travel, fintech, or SaaS (regulated environments) preferred.
  • Prior leadership or mentoring of distributed/offshore teams desired. Formal people-leadership experience a plus

RESPONSIBILITES

Infrastructure Strategy

  • Own and evolve TII’s Azure infrastructure and reliability roadmap, aligned to the Azure Well-Architected Framework.
  • With the AVP, Infrastructure to co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains.
  • Define standards for compute, network, and platform services that scale with business growth and new channels.
  • Drive cloud cost optimization (FinOps)—balancing performance, resilience, and spend.
  • Partner with engineering architecture to ensure infrastructure enables design-time resilience and delivery velocity.

Site Reliability Engineering

  • Own run-time reliability across availability, performance, scalability, and capacity for TII’s platform.
  • Mature and expand the SLO practice—defining SLIs, refining the 99.9% (and higher, where warranted) SLOs, and operating error budgets to balance reliability and delivery speed.
  • Lead capacity planning and performance engineering to support the platform’s growth.
  • Drive operational readiness reviews for new services and major releases.

Observability

  • Own and mature the observability platform across the three pillars—metrics, logs, and traces—enhancing Grafana/Prometheus and Azure Application Insights.
  • Implement distributed tracing across the GraphQL/REST services to accelerate diagnosis and reduce time-to-detect and time-to-resolve.
  • Establish meaningful alerting and telemetry that reduce noise and surface real signals.
  • Build reliability dashboards that give teams and leadership clear visibility into service health.

Platform Engineering

  • Own the internal developer platform and self-service infrastructure capabilities that enable engineering teams to provision and operate safely.
  • Define golden paths / paved-road templates that make the reliable, secure, and compliant way the easy way.
  • Establish and champion Infrastructure as Code standards using Bicep and Terraform.
  • Improve developer experience and engineering enablement through automation and reusable platform services.

Operational Excellence & Automation

  • Drive automation across provisioning, configuration, deployment, and remediation to eliminate toil.
  • Establish operational runbooks, self-healing patterns, and proactive reliability practices.
  • Continuously improve deployment safety and rollback capability in partnership with CI/CD owners.

Cloud Security & Compliance

  • Own the engineering and operational security of the Azure cloud platform—including identity and access management, network security, configuration hardening, and secrets/key management.
  • Manage and improve the cloud security posture (e.g., Microsoft Defender for Cloud, Azure Policy), including continuous vulnerability management and remediation.
  • Implement security monitoring and alerting as part of the observability platform to detect and respond to threats.
  • Embed DevSecOps and secure-by-design practices into platform and IaC workflows, enforcing guardrails and policy-as-code within golden paths and self-service tooling.
  • Partner with the Security function on policy, governance, and compliance in a regulated insurance (PII) environment.

Incident & Problem Management

  • Own the major-incident process and incident command, matured on the on-call platform (Better Stack).
  • Lead blameless post-incident reviews and drive systemic problem management to prevent recurrence.
  • Improve on-call health, escalation paths, and mean-time-to-detect / mean-time-to-resolve.
  • Maintain and enhance the mature Business Continuity and Disaster Recovery capabilities, including RTO/RPO targets and periodic testing.

Leadership

  • Will co-lead a small team of infrastructure, cloud, & system engineers.
  • Uplift the reliability and platform capability across engineering—raising standards and building a reliability culture.
  • Mentor engineers, demonstrating the leadership behaviors that support growth into a formal infrastructure leadership role.
  • Establish standards, documentation, and ways of working that scale across teams.
  • Other duties as assigned

Are you interested in this position?


Apply by clicking on the “Apply Now” button below!

#CrossChannelJobs #JobSearch
#CareerOpportunities #HiringNow
#Employment #JobOpenings
#JobSeekers