Summary
What you’ll impact
The Vice President, AI Infrastructure Support Engineer will provide operational support for AI and GPU platforms, deploying and managing AI/ML workloads using Kubernetes and Docker. The role involves automation, incident management, cross‑functional collaboration, and on‑call duties to ensure high availability of AI infrastructure.
Responsibilities
What you'll do
- Deliver operational support and drive issue resolution for AI infrastructure, GPU platforms, and associated production services.
- Support the deployment, orchestration, and monitoring of AI/ML workloads across distributed environments using Kubernetes and Docker.
- Apply hands-on expertise with NVIDIA GPU infrastructure, including A100, H100, and B200 nodes, to support provisioning, utilization, performance tuning, and ongoing platform operations.
- Build, maintain, and optimize automation scripts and tooling using Ansible, Python and Shell scripting to improve efficiency and reliability.
- Manage incidents, service requests, and change activity through enterprise workflow and ticketing platforms such as JIRA and ServiceNow.
- Partner with cross-functional teams across engineering, platform, and operations to implement DevOps best practices, including CI/CD using GitLab, containerized deployment models, and secure access management through Azure AD.
- Create and maintain clear technical documentation, including runbooks, standard operating procedures, support guides, and infrastructure architecture artifacts.
- Diagnose and resolve networking, storage, and system performance issues affecting AI infrastructure and high-performance compute workloads.
- Participate in a rotating on-call support model, including occasional weekend coverage, to help maintain high availability and operational excellence.
Requirements
What you’ll bring
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
- Strong experience supporting AI infrastructure, GPU platforms, or high-performance compute environments in enterprise or production settings.
- Hands‑on experience working with NVIDIA GPU nodes, including A100, H100, and B200 systems.
- Solid understanding of Kubernetes and Docker for deploying, managing, and troubleshooting distributed workloads.
- Proficiency in Ansible, Python and Shell scripting for automation, tooling, and operational support tasks.
- Experience with DevOps practices and tools, including GitLab CI/CD, infrastructure automation, and container‑based delivery models.
- Familiarity with enterprise ticketing and service management platforms such as JIRA and ServiceNow.
- Good knowledge of networking, storage, Linux systems administration, and infrastructure troubleshooting in complex technical environments.
- Understanding of identity and access management, including secure access controls such as Azure AD.
- Strong problem‑solving skills, attention to detail, and the ability to manage multiple priorities in a fast‑paced environment.
- Effective written and verbal communication skills, with the ability to document processes and collaborate across technical and non‑technical teams.
- Willingness to participate in on‑call rotations and provide occasional weekend support as needed.