Summary
What you’ll impact
Our company is seeking a Member of Technical Staff to design and build next‑generation, power‑aware AI infrastructure systems that operate at scale across distributed cloud platforms. The role blends advanced research with hands‑on software engineering to create production‑quality software, prototypes, and experimental platforms for AI data centers, while collaborating with product, customers, and academic partners.
Responsibilities
What you'll do
- Design and build novel systems for power-aware AI infrastructure, distributed computing, and large-scale cloud platforms.
- Develop production-quality software, research prototypes, and experimental infrastructure that can be deployed in real-world AI data centers.
- Apply machine learning, optimization, systems, or control techniques to challenging problems in AI infrastructure and cloud operations.
- Design, implement, and evaluate algorithms using large-scale experimental platforms and production deployments.
- Partner with product and customer facing teams to transition research innovations into customer-facing products.
- Collaborate with partners across industry and academia on cutting-edge research initiatives.
- Publish high-impact research in leading systems and AI conferences when appropriate.
- Help shape Emerald AI's long-term technical roadmap and identify new research directions with commercial impact.
Requirements
What you’ll bring
- Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field.
- Machine learning systems
- AI infrastructure
- Distributed systems
- Cloud computing
- Systems for AI or HPC
- Performance optimization
- Excellent software engineering skills with experience developing large software systems in languages such as C++, Python, Go, or Rust.
- Experience building research prototypes or with large-scale production code.
- Strong publication record or demonstrated history of delivering impactful technical innovations.
- Ability to independently drive research from idea through implementation and evaluation.
- Experience deploying systems in production cloud or distributed environments.
- Experience working with large codebases, production software, or open-source infrastructure.
- Experience with Kubernetes, Slurm, distributed training/inference frameworks, or large-scale AI infrastructure.
- Experience with GPU systems, accelerators, or performance analysis tools.
- Experience in optimization, control systems, resource scheduling, or systems performance.
- Experience with power management, energy-efficient computing, sustainability, or data center infrastructure.
- Experience taking research innovations from prototype to production.