Summary
What you’ll impact
The role involves evaluating and improving large language models for security-focused coding tasks, including vulnerability assessment, exploit verification, and security patching. The contractor will create coding prompts, assess model outputs, document failures, and configure evaluation environments.
Responsibilities
What you'll do
- Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches.
- Create high-quality coding prompts and reference answers for benchmark-style problems.
- Evaluate model outputs for code generation, refactoring, debugging and implementation.
- Identify and document model failures, edge cases and reasoning gaps.
- Compare private language models with leading external models.
- Build or configure coding environments for evaluation and reinforcement learning.
- Follow detailed annotation and evaluation guidelines consistently.
Requirements
What you’ll bring
- At least five years of professional software-development experience and strong Python skills.
- Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches.
- The ability to apply structured evaluation criteria and write clear technical feedback.
- Fluency in written and spoken English.