Featured Job

Senior Research Engineer

New York, NY Full-time Hybrid $165k — $310k per year 09/07/2026 Job ID: 000130
Apply Now
PyTorch transformer-based language models large language models fine-tuning supervised fine-tuning

Summary

What you’ll impact

Our company seeks a Senior Research Engineer to design, build, and optimize training pipelines and infrastructure for large language models. The role focuses on improving model quality, distributed training efficiency, and collaborating with customers and internal teams to deliver production-ready AI systems.

Responsibilities

What you'll do

  • Design, build, and optimize training and post-training pipelines for large language models.
  • Improve model quality through supervised fine-tuning, continued pretraining, preference optimization, reinforcement learning, evaluation, and experimentation.
  • Build and improve PyTorch-based training infrastructure, tooling, and developer workflows.
  • Optimize distributed training across multi-GPU environments by improving throughput, memory efficiency, scalability, and GPU utilization.
  • Investigate model training issues, including convergence, instability, communication overhead, and performance bottlenecks.
  • Design evaluation methodologies, benchmark models, analyze failure modes, and acheive model improvements through experimentation.
  • Collaborate directly with customers to understand real-world workloads and translate those learnings into improvements across Lightning AI's research platform.
  • Partner closely with research, infrastructure, and platform engineering teams to build production-ready AI systems.
  • Contribute to open-source projects through new features, tooling improvements, documentation, and community engagement

Requirements

What you’ll bring

  • Significant experience training, fine-tuning, evaluating, and/or optimizing transformer-based language models using PyTorch.
  • Experience with modern LLM training and post-training techniques such as continued pretraining, SFT, RLHF, preference optimization (DPO, PPO, GRPO), reward modeling, or similar approaches.
  • Strong understanding of distributed training and multi-node systems, with experience improving training performance, scalability, and/or efficiency.
  • Strong software engineering fundamentals, including building production-quality Python software and research tooling.
  • Experience designing experiments, evaluating model performance, and debugging complex training and/or optimization issues.
  • Excellent communication and collaboration skills, including the ability to work effectively across research, product, infrastructure, and customer-facing engagements.
  • Comfortable working in fast-moving, ambiguous environments where priorities evolve over time.
  • Master's degree, PhD, or equivalent industry experience in Machine Learning, AI, Computer Science, or a related field

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.