Featured Job

Senior LLM Inference Engineer

Menlo Park, California Full-time On-site 10/01/2026 Job ID: 000308
Apply Now
LLM inference optimization Distributed serving architectures for large language models Quantization techniques for transformer models Speculative decoding techniques with draft models Eagle speculative decoding approaches

Summary

What you’ll impact

The LLM Inference Engineer at the organization will own and optimize the serving infrastructure that delivers healthcare AI responses with sub-100ms latency. This role focuses on designing distributed inference architectures, applying quantization and speculative decoding techniques, and continuously benchmarking performance to ensure reliability and cost-effectiveness at scale.

Responsibilities

What you'll do

  • Design and implement multi-node serving architectures for distributed LLM inference
  • Optimize multi-LoRA serving systems
  • Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality
  • Implement speculative decoding and other latency optimization strategies
  • Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases
  • Continuously benchmark and improve system performance across various deployment scenarios and GPU types

Requirements

What you’ll bring

  • Experience optimizing LLM inference systems at scale
  • Proven expertise with distributed serving architectures for large language models
  • Hands-on experience implementing quantization techniques for transformer models
  • Strong understanding of modern inference optimization methods, including:
  • Speculative decoding techniques with draft models
  • Eagle speculative decoding approaches
  • Proficiency in Python and C++
  • Experience with CUDA programming and GPU optimization

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.