Summary
What you’ll impact
The LLM Inference Engineer at the organization will own and optimize the serving infrastructure that delivers healthcare AI responses with sub-100ms latency. This role focuses on designing distributed inference architectures, applying quantization and speculative decoding techniques, and continuously benchmarking performance to ensure reliability and cost-effectiveness at scale.
Responsibilities
What you'll do
- Design and implement multi-node serving architectures for distributed LLM inference
- Optimize multi-LoRA serving systems
- Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality
- Implement speculative decoding and other latency optimization strategies
- Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases
- Continuously benchmark and improve system performance across various deployment scenarios and GPU types
Requirements
What you’ll bring
- Experience optimizing LLM inference systems at scale
- Proven expertise with distributed serving architectures for large language models
- Hands-on experience implementing quantization techniques for transformer models
- Strong understanding of modern inference optimization methods, including:
- Speculative decoding techniques with draft models
- Eagle speculative decoding approaches
- Proficiency in Python and C++
- Experience with CUDA programming and GPU optimization