Featured Job

Lead Databricks Engineer

None Contract Remote $80 — $90 hourly 09/17/2026 Job ID: 000221
Apply Now
Databricks Delta Lake Apache Spark Databricks Workflows Unity Catalog

Summary

What you’ll impact

Our organization is hiring a hands‑on Databricks Lead to own and build the lakehouse and MLOps foundation for a large US enterprise client. The role involves technical leadership, data engineering, CI/CD, orchestration, and model operations, with a six‑month contract that may convert to full‑time based on performance.

Responsibilities

What you'll do

  • Lead the client’s Databricks and data platform team.
  • Set technical direction, engineering standards, and implementation patterns.
  • Plan and sequence delivery across data engineering, MLOps, and analytics initiatives.
  • Own platform reliability, scalability, security, cost efficiency, and delivery outcomes.
  • Work directly with client data science, IT, engineering, and business leadership.
  • Facilitate architecture discussions and working sessions.
  • Prepare clear technical decision documents and defend architectural recommendations.
  • Identify risks, make decisions with incomplete information, and keep delivery moving.
  • Mentor engineers, conduct code reviews, and remove technical blockers.
  • Create runbooks, operational documentation, and support procedures.
  • Design and build medallion architecture pipelines using Bronze, Silver, and Gold layers.
  • Develop scalable data pipelines using Databricks, Apache Spark, Delta Lake, Python, and SQL.
  • Integrate data from operational systems, CRM platforms, HRIS and payroll systems, web applications, finance systems, and other enterprise sources.
  • Implement incremental ingestion, change data capture, schema evolution, backfills, replayability, and historical data retention.
  • Build dimensional and analytical data models for reporting, experimentation, and machine learning.
  • Design reusable feature, training, and prediction tables with versioning and multi-year history.
  • Ensure datasets are reproducible, traceable, and suitable for model retraining.
  • Optimize Spark jobs, Delta tables, clusters, storage, partitioning, and data-processing costs.
  • Deliver trusted outputs to Power BI, downstream applications, APIs, and other consumer systems.
  • Implement CI/CD practices for data engineering and machine learning workloads.
  • Establish Git-based workflows, pull-request reviews, branching standards, and release processes.
  • Build automated unit, integration, data-quality, and end-to-end tests.
  • Parameterize code and configuration across environments.
  • Use pinned package versions and reproducible runtime environments.
  • Implement secure secrets handling and eliminate manual deployment steps.
  • Establish a controlled approval gate for production releases.
  • Automate deployment of notebooks, jobs, workflows, libraries, configurations, and infrastructure where appropriate.
  • Use infrastructure-as-code tools and practices to improve repeatability and operational control.
  • Design and operate time-based, trigger-based, event-driven, and manually initiated workflows.
  • Configure Databricks Workflows, Airflow, or comparable orchestration tools.
  • Provide DAG-level visibility into pipeline execution and dependencies.
  • Implement retries, failure handling, recovery procedures, and replay capabilities.
  • Define and monitor pipeline-level service-level agreements and operational targets.
  • Build alerts for failures, delays, SLA breaches, and abnormal processing behavior.
  • Maintain production runbooks and incident-response procedures.
  • Build the MLOps foundation using MLflow and Databricks machine learning capabilities.
  • Implement experiment tracking, model registration, model versioning, and lifecycle management.
  • Create batch inference pipelines and production prediction workflows.
  • Establish repeatable paths from experimentation to development, staging, and production.
  • Support model rollback and controlled promotion between environments.
  • Integrate feature tables, training datasets, model artifacts, and prediction outputs.
  • Monitor model performance, data drift, concept drift, pipeline health, and prediction quality.
  • Configure alerts for model and data anomalies.
  • Partner with data scientists to productionize models without unnecessary rewrites.
  • Prepare the platform for real-time and low-latency model serving as business requirements evolve.
  • Implement data-quality gates directly within ingestion, transformation, feature, and model pipelines.
  • Block downstream model runs when critical validation checks fail.
  • Detects schema changes, unexpected null increases, duplicate records, invalid values, and out-of-range metrics.
  • Establish validation rules for completeness, accuracy, consistency, uniqueness, timeliness, and referential integrity.
  • Alert stakeholders rather than allowing critical failures to occur silently.
  • Maintain quality metrics, issue history, and operational dashboards.

Requirements

What you’ll bring

  • At least 6 years of experience in data engineering, data platforms, analytics engineering, or related roles.
  • Demonstrated experience leading a technical team or owning a data platform end to end.
  • Deep hands-on experience with Databricks, Delta Lake, Apache Spark, Databricks Workflows, Unity Catalog, and MLflow.
  • Candidates with equivalent depth in Snowflake, Microsoft Fabric, BigQuery, or another modern cloud data platform may be considered if they can become productive on Databricks quickly.
  • Strong Python and SQL skills.
  • Production experience building and operating ETL or ELT pipelines at scale.
  • Experience with incremental data loads, schema evolution, historical data, backfills, data modeling, and performance tuning.
  • Hands-on experience implementing CI/CD for data and machine learning workloads.
  • Experience with Git workflows, automated testing, package and environment management, secrets handling, and controlled production releases.
  • Experience with orchestration and data-quality tooling, such as Databricks Workflows, Airflow, dbt tests, Great Expectations, or comparable technologies.
  • Working knowledge of at least one major cloud platform; Azure is preferred, but AWS or GCP experience is acceptable.
  • Ability to work directly with senior client stakeholders and technical leadership.
  • Ability to lead architecture sessions, write clear technical decisions, and maintain a position under pressure.
  • Strong communication, documentation, mentoring, and code-review skills.
  • Ability to work independently in a remote, cross-functional environment with US-based teams.
  • Comfort using AI-assisted development tools such as Claude Code, Cursor, GitHub Copilot, or similar tools to write, review, test, and accelerate engineering work.
  • Experience as a founder, founding engineer, principal engineer, or early technical leader.
  • Experience delivering data and MLOps platforms in a consulting, professional-services, or client-services environment.
  • Experience supporting US-based enterprise clients.
  • Experience with feature engineering, model training, model deployment, model monitoring, and drift detection.
  • Experience with Structured Streaming, Kafka, Azure Event Hubs, or other streaming technologies.
  • Experience with Power BI or another enterprise semantic and business-intelligence layer.
  • Experience with Azure Data Factory, Azure DevOps, Terraform, Bicep, or related Azure technologies.

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.