Summary
What you’ll impact
Our organization is seeking a Senior Data Scientist to develop scalable data solutions and advanced AI capabilities for a classified, mission-critical program. The role involves end-to-end data science workflows, LLM integration, and collaboration with engineering and stakeholders to deliver actionable intelligence.
Responsibilities
What you'll do
- This position is fulltime on-site in Reston, VA.
- Automate and optimize data extraction, cleaning, processing, and analysis tasks using scripting languages and ETL pipelines.
- Process and analyze large-scale, structured and unstructured datasets using big data frameworks (e.g., Spark, Hadoop, or cloud-native equivalents).
- Independently conduct end-to-end data science/engineering workflows with minimal supervision, from hypothesis generation to delivery of actionable insights.
- Explore opportunities to combine graph technologies, machine learning, generative AI, and large language models to create new analytical capabilities.
- Clearly document methodologies and findings in formal written reports and present outcomes to both technical and non-technical audiences.
- Design, fine-tune, and implement Large Language Models (LLMs) into workflows to support knowledge management systems, including:
- Development of LLM-driven tools to automate summarization, tagging, and retrieval of internal data, reports, and documentation.
- Integration of LLMs into workflows to streamline knowledge capture, institutional memory, and information sharing.
- Deployment of LLM-based internal assistants to support decision-making and reduce redundant effort.
- Translate ambiguous business or customer problems into well-defined analytical approaches, experiments, and measurable outcomes.
- Present findings and recommendations through clear visualizations, reports, and presentations.
- Mentor junior data scientists and contribute to best practices in data science and analytics.
- Partner with engineering teams to operationalize machine learning models in production.
Requirements
What you’ll bring
- Current/active TS/SCI security clearance and be willing and able to obtain CI polygraph.
- 10 years of experience in data science, analytics, or a related technical field.
- Master’s degree in data science, computer science, engineering, statistics, GIS, or related discipline. Degree can be substituted with an additional 2 years of experience.
- Proficient in Python, SQL, and tools/libraries for data analysis, machine learning, and automation.
- Demonstrated ability to independently conduct full-cycle data projects, from processing to reporting.
- Strong verbal and written communication skills; able to explain complex findings clearly.
- Familiarity with foundational LLM/NLP concepts and experience applying pre-trained models (e.g., GPT) for internal automation or data summarization.
- Ability to work with minimal supervision and produce logically structured, actionable insights.
- Ability to collaborate with business stakeholders to define analytical requirements and deliver data-driven solutions.
- Demonstrated experience deploying graph databases or graph-based applications into production environments.
- Expertise in advanced statistical and machine learning techniques, including NLP and unsupervised learning.
- Experience developing or fine-tuning LLMs and integrating them into knowledge systems or workflows.
- Strong background with cloud-based platforms (e.g., AWS, Azure, GCP) and scalable data processing tools.
- Track record of designing or enhancing knowledge management solutions using AI or LLMs.
- Experience building internal tools or bots for search, summarization, or decision support using LLMs.
- Awareness of responsible AI practices and ethical considerations in LLM deployment.
- Knowledge of MLOps practices and tools (e.g., MLflow, Kubeflow, SageMaker, Vertex AI, or Azure ML).
- Experience with big data technologies such as Spark, Databricks, or Hadoop.
- Familiarity with generative AI, large language models (LLMs), retrieval-augmented generation (RAG), or AI agent frameworks.
- Experience with data visualization tools such as Power BI, Tableau, or Looker.
- Knowledge of containerization and orchestration technologies such as Docker and Kubernetes.
- Strong understanding of graph data modeling, graph algorithms, and graph query languages such as Cypher, Gremlin, or SPARQL.