Role Opportunity Score
Expected applicants: ~80–150
39 AI's Senior GPU Infrastructure SRE role is advertised on a custom career page, with an estimated 80–150 strong applicants across the EU. Experience with Kubernetes and PyCUDA can shift your odds, as many will lack this specialized knowledge.
Evidence Dossier
Remote Details
Scope: EU Remote
Role Summary
Build and operate the infrastructure that supports large-model training and online inference for GPU clusters at scale. Responsibilities include maintaining GPU clusters of 100+ cards, responding to production incidents, and working closely with training and inference teams to ensure stability, performance, and resource efficiency. The role involves hands-on infrastructure work with Linux, Slurm, Kubernetes, Ceph, RDMA networking, GPU drivers, observability, and automation.
Stack Required
Compensation
Requirements
AI Tools Culture
AI mentionedLast seen on career page: 28d ago
Posted: 28d ago