Featured Job

Staff GPU Infrastructure

San Francisco, California Contract On-site 10/02/2026 Job ID: 000371
Apply Now
Linux Bare-metal provisioning Kubernetes Infrastructure automation Enterprise GPU servers

Summary

What you’ll impact

Our organization is seeking a Staff GPU Infrastructure / Bare-Metal Platform Engineer. This role will design, deploy, automate, and maintain large‑scale GPU server infrastructure for AI/ML workloads. The focus is on Linux‑based bare‑metal provisioning, Kubernetes cluster management, and building automation tools using Golang or Python. The engineer will also ensure reliability, scalability, and performance of the GPU infrastructure in a hyperscale AI cloud environment.

Responsibilities

What you'll do

  • Design, deploy, automate and maintain large-scale bare-metal GPU server infrastructure.
  • Develop infrastructure automation tools to accelerate GPU server provisioning, deployment and operational readiness.
  • Design and scale infrastructure to support simultaneous deployment of 80 100 servers or more.
  • Configure and maintain Linux-based GPU servers and operating systems.
  • Implement automated bare-metal provisioning using PXE boot, MAAS or equivalent technologies.
  • Deploy, configure, administer and troubleshoot Kubernetes clusters supporting AI/ML workloads.
  • Manage GPU server hardware, BMC/IPMI, BIOS, firmware, operating systems and driver configurations.
  • Diagnose GPU hardware failures, server-level issues, provisioning failures and infrastructure performance problems.
  • Coordinate hardware troubleshooting, replacement and maintenance with enterprise server vendors.
  • Develop infrastructure automation and operational software using Golang or Python.
  • Implement infrastructure as code and configuration management using Ansible and Terraform.
  • Develop monitoring, health-checking, alerting and automated recovery capabilities.
  • Support infrastructure capacity planning, reliability, scalability and performance improvements.
  • Collaborate with GPU platform, networking, storage, security and data center operations teams.
  • Support the transition toward Kubernetes-based and containerized infrastructure.
  • Maintain technical documentation, troubleshooting procedures and operational runbooks.

Requirements

What you’ll bring

  • 8+ years of experience in infrastructure engineering, systems engineering, platform engineering, DevOps, SRE, or related roles.
  • Advanced hands-on Linux systems administration and troubleshooting.
  • Strong experience with bare-metal server provisioning and automated OS deployment.
  • Experience with PXE booting and large-scale server deployment.
  • Hands-on experience deploying, administering and troubleshooting Kubernetes clusters.
  • Strong programming experience with Python or Golang, preferably Golang.
  • Experience with infrastructure automation using Ansible and Terraform.
  • Strong understanding of BMC/IPMI, BIOS, firmware and enterprise server management.
  • Experience troubleshooting physical server hardware, preferably GPU-based systems.
  • Experience supporting large-scale infrastructure environments.
  • Strong understanding of server lifecycle management, infrastructure reliability and operational automation.
  • Experience with NVIDIA DGX, HGX, or equivalent GPU server infrastructure.
  • Experience with Canonical MAAS.
  • Familiarity with NVIDIA GPU drivers, CUDA environments and NVIDIA DCGM.
  • Experience with Docker and containerization technologies.
  • Experience with Prometheus, Grafana, or equivalent infrastructure monitoring tools.
  • Familiarity with Slurm, NVIDIA Run AI, or other GPU workload orchestration platforms.
  • Experience supporting distributed AI training infrastructure.
  • Experience with high-performance networking, InfiniBand, or RoCE.
  • Experience with AI cloud platforms and hyperscale infrastructure.
  • Bachelor's degree or equivalent combination of education and experience.

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.