Technical Lead – GPU Infrastructure
Job Description
REQUIREMENTS
- Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
- Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
- GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
- High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
- Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
- Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
- HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
- Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
- Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
- A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
- Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
- Excellent written and spoken English. Most partner and leadership work happens in writing.
RESPONSIBILITES
- Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
- Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
- Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
- Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
- Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
- Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
- Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
- Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#CrossChannelJobs #JobSearch
#CareerOpportunities #HiringNow
#Employment #JobOpenings
#JobSeekers