Skip to main content
Posted 31 July, 2026

Senior Databricks Engineer

ExlService Holdings, Inc.
Chennai, Tamil Nadu, India Full Time
Reference: 218_689623_17929

We are looking for a skilled and passionate Senior Databricks Engineer to design, build, and optimize enterprise-scale data lakehouse solutions on the Databricks platform. The successful candidate will be responsible for creating Databricks pipeline delivering Financial Crime platforms covering Anti-Money Laundering (AML), Know Your Customer (KYC), Customer Risk Assessment (CRA), Sanctions Screening, Transaction Monitoring, Fraud Detection, and Regulatory Reporting

Education

  • Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or related field.

Experience

  • 6-8 years of total experience in data engineering or software engineering.
  • 4+ years of dedicated hands-on experience with the Databricks platform in production environments.
  • Strong background in big data engineering, cloud data platforms, and distributed computing.

Databricks Platform Engineering

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
  • Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage.
  • Optimize cluster configurations - instance types, auto-scaling policies, spot/preemptible nodes - for cost and performance.
  • Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
  • Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management.

Delta Lake & Lakehouse Architecture

  • Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM).
  • Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
  • Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations.
  • Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
  • Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).

Data Pipeline Development (PySpark / SQL)

  • Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
  • Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
  • Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning.
  • Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
  • Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines.

Sign up for Job Alerts