Skip to main content
Posted 11 July, 2026

Senior Data Engineer

Abacus Insights
Pune, India Full Time
Reference: 102_698163_8424242002

About the role

We are seeking an accomplished Data Engineer to join our dynamic and rapidly expanding Tech Ops division. With significant projected growth, this is an opportunity to drive meaningful technical impact. In this role, you will work directly with customers, data vendors, and internal engineering teams to design, implement, and optimize complex data integration solutions within a modern, largescale cloud environment.

You will leverage advanced skills in distributed computing, data architecture, and cloud-native engineering to enable scalable, resilient, and highperformance data ingestion and transformation pipelines. As a trusted technical advisor, you will guide customers in adopting Abacus's core data management platform and ensure high-quality, compliant data operations across the lifecycle.

Your day to day

  • Architect, design, and implement high-volume batch and real-time data pipelines using PySpark, SparkSQL, Databricks Workflows, and distributed processing frameworks.
  • Build endtoend ingestion frameworks integrating with Databricks, Snowflake, AWS services (S3, SQS, Lambda), and vendor data APIs, ensuring data quality, lineage, and schema evolution.
  • Develop data modeling frameworks, including star/snowflake schemas and optimization techniques for analytical workloads on cloud data warehouses.
  • Lead technical solution design for health plan clients, creating highly available, fault-tolerant architectures across multi-account AWS environments.
  • Translate complex business requirements into detailed technical specifications, engineering artifacts, and reusable components.
  • Implement security automation, including RBAC, encryption at rest/in transit, PHI handling, tokenization, auditing, and compliance with HIPAA and SOC 2 frameworks.
  • Establish and enforce data engineering best practices, such as CI/CD for data pipelines, code versioning, automated testing, orchestration, logging, and observability patterns.
  • Conduct performance profiling and optimize compute costs, cluster configurations, partitions, indexing, and caching strategies across Databricks and Snowflake environments.
  • Produce high-quality technical documentation including runbooks, architecture diagrams, and operational standards.
  • Mentor junior engineers through technical reviews, coaching, and training sessions for both internal teams and clients.

What you bring to the team

  • Bachelor's degree in Computer Science, Computer Engineering, or a closely related technical field.
  • 7+ years of handson experience as a Data Engineer working with largescale, distributed data processing systems in modern cloud environments.
  • Working knowledge of U.S. healthcare data domains-including claims, eligibility, and provider datasets-and experience applying this knowledge to complex ingestion and transformation workflows.
  • Strong ability to communicate complex technical concepts clearly across both technical and nontechnical stakeholders.
  • Expertlevel proficiency in Python, SQL, and PySpark, including developing distributed data transformations and performanceoptimized queries.
  • Demonstrated experience designing, building, and operating productiongrade ETL/ELT pipelines using Databricks, Airflow, or similar orchestration and workflow automation tools.
  • Proven experience architecting or operating largescale data platforms using dbt, Kafka, Delta Lake, and eventdriven/streaming architectures, within a cloudnative data services or platform engineering environment-requiring specialized knowledge of distributed systems, scalable data pipelines, and cloudscale data processing.
  • Experience working with structured and semistructured data formats such as Parquet, ORC, JSON, and Avro, including schema evolution and optimization techniques.
  • Strong working knowledge of AWS data ecosystem components-including S3, SQS, Lambda, Glue, IAM-or equivalent cloud technologies supporting highvolume data engineering workloads.
  • Proficiency with Terraform, infrastructureascode methodologies, and modern CI/CD pipelines (e.g., GitLab) supporting automated deployment and versioning of data systems.
  • Deep expertise in SQL and compute optimization strategies, including ZOrdering, clustering, partitioning, pruning, and caching for largescale analytical and operational workloads.
  • Handson experience with major cloud data warehouse platforms such as Snowflake (preferred), BigQuery, or Redshift, including performance tuning and data modeling for analytical environments.

What we would like to see but not required:

  • Experience in large-scale healthcare or payer data environments.

What you'll get in return :

  • Competitive Leave & Benefits
  • Comprehensive health coverage
  • Equity for every employee - share in our success
  • Growth-focused environment - your development matters here

Work arrangements

  • Standard hours: 9 hours/day, 5 days/week
  • Location: Pune, Hybrid (3 days a week in office)
  • Shift: Your standard working hours will be nine (9) hours per day within the Company's standard working hours. Specific working hours may vary based on business needs.

Sign up for Job Alerts