[Remote] Databricks Solutions Expert (Remote)
Note: The job is a remote job and is open to candidates in USA. GovCIO is hiring a Databricks Solutions Expert to support the Department of Veterans Affairs Data Modernization initiative. The role will architect, design, and implement Azure Databricks as a secure, governed, multi-tenant Lakehouse platform, while establishing engineering, governance, security, automation, and enablement standards for enterprise analytics teams.
Responsibilities
- The Databricks Solutions Expert will define platform patterns, build reference implementations, set standards (cluster policies, Unity Catalog governance, CI/CD), and drive enablement so product, analytics, and data science teams can deliver reliable, compliant, and performant insights at scale
- Blueprint the Azure Databricks landing zone: workspace topology (prod/non-prod), network architecture (VNet injection, Private Link, NAT), secure connectivity to ADLS Gen2, Azure SQL/MI, Event Hubs, and other data sources
- Governance with Unity Catalog: multi-region metastores, catalog/schema/table design, data classification tiers, row/column-level security, data lineage, and cross-domain data sharing patterns (Delta Sharing)
- Lakehouse foundations: Delta Lake storage design (bronze/silver/gold), medallion data flow standards, partitioning, Z‑ordering, OPTIMIZE/VACUUM policies, and performance best practices (Photon, AQE)
- Scalability for “thousands of teams”: multi-workspace strategy, tenancy model (shared vs. dedicated), workspace baselines, cluster policy tiers, and guardrails to prevent noisy-neighbor and cost runaways
- Reliability: HA/DR strategy, regional deployments, backup/restore, versioning, and repeatable environment provisioning via Terraform
- Identity & access: Entra ID (Azure AD) SSO, SCIM user/group provisioning, service principals/managed identities, attribute-based/role-based access controls mapped to Unity Catalog
- Secrets & credentials: Key Vault–backed secret scopes, credential passthrough patterns, token hygiene (PAT governance)
- Data protection: encryption at rest/in transit, private endpoints, data exfil and egress controls, policy-as-code, and audit logging to Log Analytics or secure storage
- Compliance: enforce enterprise policies (PII/PHI handling), data retention, legal hold, and regulatory reporting with auditable lineage
- Pipelines: build DLT (Delta Live Tables) pipelines and jobs for batch and streaming (Structured Streaming) with CDC (e.g., via Auto Loader) from enterprise sources
- Performance engineering: optimize notebooks/SQL/ETL (Photon, caching, skew mitigation), tune cluster sizing, and set standards for reliable, fast jobs
- MLOps: integrate MLflow for experiment tracking, model registry, feature store (Unity Catalog), and serving patterns where appropriate
- Observability: end-to-end monitoring (jobs, clusters, UC audits), dashboarding, alerting, and SLOs; integrate with Azure Monitor/Log Analytics
- Enablement: build reusable reference accelerators (templates, example notebooks, data products), run playbooks, and conduct office hours to uplevel 1000s of teams
- Infrastructure-as-Code: provision workspaces, catalogs, cluster policies, and permissions via Terraform (Databricks provider), with pipelines in Azure DevOps/GitHub
- CI/CD: notebook/package deployment, testing harnesses (dbx/pytest), environment promotion, and artifact versioning
- FinOps: cost modeling, budgets/alerts, instance pools, auto-termination, serverless SQL, tagging/chargeback, and usage analytics for executive reporting
- Partner with Security, Networking, Compliance, and FinOps to codify enterprise standards
- Establish a Lakehouse Platform Council to ratify patterns and review exceptions
- Create adoption metrics, business case narratives, TCO models, and executive updates
Skills
- Bachelor's degree in Information Technology or a related field. (or commensurate experience)
- 12+ years of experience in data engineering/analytics; 5+ years building on cloud data platforms (Azure preferred)
- 3+ years hands-on Azure Databricks (platform + pipelines) and Delta Lake
- Proven experience setting up Unity Catalog with granular governance (RLS/CLS)
- Deep knowledge of Azure networking (VNets, Private Link, NSGs), identity (Entra ID), Key Vault, ADLS Gen2, Event Hubs, Azure SQL/MI, and Data Factory
- Strong Spark expertise (PySpark/SQL), Structured Streaming, performance tuning, partitioning and storage optimization
- Practical Terraform experience for Databricks/Azure resources; CI/CD with Azure DevOps or GitHub Actions
- Security-first mindset; track record implementing audit logging, policy-as-code, and compliance controls
- Excellent communication skills; ability to standardize, teach, and influence at enterprise scale
- Ability to obtain and maintain a suitability/Public Trust
- Experience operating platforms for >500 concurrent users and 1000s of analytics teams
- Knowledge of Photon, DLT, Workflows, Lakehouse ML (MLflow, feature store), Delta Sharing
- Exposure to FinOps and chargeback models for data platforms
- Experience with Synapse/Fabric interoperability, and data virtualization patterns
- Background in data modeling (medallion, dimensional, domain‑driven design) and data quality (expectations, SLAs)
- + Databricks: Databricks Certified Data Engineer Professional, Lakehouse Fundamentals, Machine Learning Associate/Professional
- + Microsoft: Azure Solutions Architect Expert (AZ‑305), Azure Data Engineer (DP‑203), Azure Security Engineer (AZ‑500)
- + HashiCorp: Terraform Associate
- Scala (optional)
Benefits
- This position is fully remote within the United States.
Company Overview
Company H1B Sponsorship