Lead Infrastructure & Site Reliability Engineering
Job Description
REQUIREMENTS
- Bachelor’s degree in Computer Science, or related field—or equivalent experience.
- 8+ years in infrastructure, site reliability, platform, or DevOps engineering, with recent hands-on delivery.
- Deep, hands-on expertise operating production workloads on Microsoft Azure.
- Proven experience with Infrastructure as Code using Bicep and/or Terraform.
- Hands-on experience with observability tooling across metrics, logs, and traces (e.g., Grafana, Prometheus, Azure Application Insights).
- Proven experience defining and operating SLIs/SLOs and error budgets.
- Hands-on experience securing Azure cloud environments—identity/access, network security, posture management (e.g., Microsoft Defender for Cloud, Azure Policy), and vulnerability management.
- Experience scaling Microsoft Fabric and Microsoft Purview.
- Experience owning incident management, on-call, and blameless post-incident reviews (e.g., Better Stack, PagerDuty, or comparable).
- Strong scripting/automation ability (e.g., PowerShell, Python, Bash) and CI/CD experience (Azure DevOps preferred).
- Experience supporting business continuity and disaster recovery with defined RTO/RPO.
- Experience building internal developer platforms, golden paths, and self-service infrastructure.
- Cloud cost optimization / FinOps experience.
- Exposure to AI-assisted operations (AIOps) and modern reliability automation.
- Microsoft Azure certifications (e.g., Azure Solutions Architect Expert, Azure DevOps Engineer Expert, Azure Security Engineer Associate) preferred.
- Experience in insurance, travel, fintech, or SaaS (regulated environments) preferred.
- Prior leadership or mentoring of distributed/offshore teams desired. Formal people-leadership experience a plus
RESPONSIBILITES
Infrastructure Strategy
- Own and evolve TII’s Azure infrastructure and reliability roadmap, aligned to the Azure Well-Architected Framework.
- With the AVP, Infrastructure to co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains.
- Define standards for compute, network, and platform services that scale with business growth and new channels.
- Drive cloud cost optimization (FinOps)—balancing performance, resilience, and spend.
- Partner with engineering architecture to ensure infrastructure enables design-time resilience and delivery velocity.
Site Reliability Engineering
- Own run-time reliability across availability, performance, scalability, and capacity for TII’s platform.
- Mature and expand the SLO practice—defining SLIs, refining the 99.9% (and higher, where warranted) SLOs, and operating error budgets to balance reliability and delivery speed.
- Lead capacity planning and performance engineering to support the platform’s growth.
- Drive operational readiness reviews for new services and major releases.
Observability
- Own and mature the observability platform across the three pillars—metrics, logs, and traces—enhancing Grafana/Prometheus and Azure Application Insights.
- Implement distributed tracing across the GraphQL/REST services to accelerate diagnosis and reduce time-to-detect and time-to-resolve.
- Establish meaningful alerting and telemetry that reduce noise and surface real signals.
- Build reliability dashboards that give teams and leadership clear visibility into service health.
Platform Engineering
- Own the internal developer platform and self-service infrastructure capabilities that enable engineering teams to provision and operate safely.
- Define golden paths / paved-road templates that make the reliable, secure, and compliant way the easy way.
- Establish and champion Infrastructure as Code standards using Bicep and Terraform.
- Improve developer experience and engineering enablement through automation and reusable platform services.
Operational Excellence & Automation
- Drive automation across provisioning, configuration, deployment, and remediation to eliminate toil.
- Establish operational runbooks, self-healing patterns, and proactive reliability practices.
- Continuously improve deployment safety and rollback capability in partnership with CI/CD owners.
Cloud Security & Compliance
- Own the engineering and operational security of the Azure cloud platform—including identity and access management, network security, configuration hardening, and secrets/key management.
- Manage and improve the cloud security posture (e.g., Microsoft Defender for Cloud, Azure Policy), including continuous vulnerability management and remediation.
- Implement security monitoring and alerting as part of the observability platform to detect and respond to threats.
- Embed DevSecOps and secure-by-design practices into platform and IaC workflows, enforcing guardrails and policy-as-code within golden paths and self-service tooling.
- Partner with the Security function on policy, governance, and compliance in a regulated insurance (PII) environment.
Incident & Problem Management
- Own the major-incident process and incident command, matured on the on-call platform (Better Stack).
- Lead blameless post-incident reviews and drive systemic problem management to prevent recurrence.
- Improve on-call health, escalation paths, and mean-time-to-detect / mean-time-to-resolve.
- Maintain and enhance the mature Business Continuity and Disaster Recovery capabilities, including RTO/RPO targets and periodic testing.
Leadership
- Will co-lead a small team of infrastructure, cloud, & system engineers.
- Uplift the reliability and platform capability across engineering—raising standards and building a reliability culture.
- Mentor engineers, demonstrating the leadership behaviors that support growth into a formal infrastructure leadership role.
- Establish standards, documentation, and ways of working that scale across teams.
- Other duties as assigned
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#CrossChannelJobs #JobSearch
#CareerOpportunities #HiringNow
#Employment #JobOpenings
#JobSeekers