Summary
What you’ll impact
The Principal Data Architect will lead data architecture for an AI platform, guiding senior engineers and shaping Azure‑based solutions. This remote, customer‑facing role involves architecture design, mentorship, and ensuring compliance with federal security standards.
Responsibilities
What you'll do
- Work with customers daily: lead architecture discussions, brief VA stakeholders, and turn their needs into clear designs and decisions
- Lead and mentor senior data engineers through design reviews, code reviews, and hands‑on problem solving
- Set architecture and standards for ingestion, storage, modeling, and serving with Azure Data Factory, Synapse Analytics, Databricks, and Delta Lake
- Architect the platform's move to Azure Commercial, with Databricks as the engineering platform and Unity Catalog governing data, access, lineage, and models
- Design metadata-driven frameworks that move thousands of tables through batch, incremental, change data capture (CDC), and streaming patterns
- Serve reporting and machine learning workloads through the Synapse Dedicated Pool, Databricks SQL, Power BI, and model training pipelines
- Protect PHI and PII through access policy, masking, and Authority to Operate (ATO) support under HIPAA and federal security requirements
- Write architecture and technical approach sections for proposals
Requirements
What you’ll bring
- 12+ years in data engineering and data architecture, starting as a hands‑on data engineer, with 5+ years as a lead or principal architect
- Proven leadership of engineering teams, including mentoring senior engineers and setting technical direction
- A demonstrated history of customer‑facing work: leading architecture discussions, briefing executives, and writing for technical and non‑technical readers
- Expert, hands‑on cloud data engineering and data warehousing on Azure: Azure Data Factory, Synapse Analytics (including Synapse Dedicated Pool), Databricks, Delta Lake, and ADLS Gen2
- Hands‑on Unity Catalog design: catalogs and grants, attribute‑based access control, row filters and column masks, and lineage
- Design of lakehouse (medallion) architectures, data warehouse models, and metadata‑driven frameworks across thousands of tables
- Strong proficiency in SQL and Python (PySpark)
- Experience with CDC, streaming, and incremental load patterns (e.g., Oracle GoldenGate, Azure Event Hubs, Apache Kafka)
- Experience migrating large on‑premises data warehouses to the cloud, including parallel runs and reconciliation
- Experience with Git, CI/CD, and infrastructure as code for data platforms (e.g., Terraform, Databricks Asset Bundles)
- Working knowledge of HIPAA, NIST 800‑53, and FISMA requirements for PHI and PII
- Bachelor's degree in computer science, information systems, or a related field
- US Citizen: Must be a citizen of the United States
- Security Clearance: Must be able to obtain a public trust clearance. Must be eligible to work in the United States