Senior Site Reliability Engineer (SRE)
Job Description
REQUIREMENTS
- 6+ years of progressive experience in Site Reliability Engineering (SRE).
- Deep expertise in multi-cloud environments (AWS, GCP, or Azure) including networking, compute, and storage.
- Extensive experience with container orchestration (Kubernetes, EKS, or ECS).
- Proficiency with Infrastructure as Code (Terraform, Pulumi, or CloudFormation).
- Experience with GitOps practices (ArgoCD, Flux) and modern deployment patterns (Blue/Green, Canary).
- Strong background in platform security, secrets management, and IAM.
- Familiarity with AI/ML infrastructure requirements (model inference, vector databases, or GPU workloads).
- Programming proficiency in Node.js, NestJS, or Python.
Preferred
- Experience building self-service developer platforms to increase engineering velocity.
- Experience managing Multi-Cloud API Gateways and Edge Routing (Kong, Traefik, or Cloudflare).
- Hands-on experience with security hardening tools like Falco or eBPF.
RESPONSIBILITIES
- Architect, build, and scale next-generation infrastructure and AI-powered capabilities to support modern applications and advanced AI workloads.
- Design resilient cloud architectures and ensure platforms remain secure, scalable, and high-performing.
- Implement robust automation across CI/CD, infrastructure provisioning, and operations to reduce manual overhead.
- Integrate AI systems into production environments, including pipelines for model hosting, embeddings, and vector search.
- Strengthen observability through improved logging, monitoring, tracing, and alerting (Prometheus, OpenTelemetry).
- Lead incident response processes, define SLIs/SLOs, and improve on-call readiness.
- Mentor team members to elevate platform engineering and DevOps maturity across the organization.
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#CrossChannelJobs #JobSearch
#CareerOpportunities #HiringNow
#Employment #JobOpenings
#JobSeekers
#FacebookLinkedIn