Skip to main content
Posted 19 July, 2026

Data engineer

ClifyX
Bangalore,Karnataka,India Full Time
Reference: 365_594563_25-03274

Job Title: Data engineer

Data Engineer responsible for building and optimizing our data pipelines and data architecture to support the collection, transformation, and storage of large-scale data. This role requires a strong ability to work autonomously while collaborating with data scientists, data analysts, and other stakeholders to ensure that data is available and structured in a way that meets the business's needs.

Responsibilities:

Independently design, develop, and maintain scalable data pipelines and ETL processes to collect, clean, and transform large-scale datasets from various sources.

Build and optimize data infrastructure, including data lakes, warehouses, and real-time data streaming solutions with minimal supervision.

Ensure data quality, integrity, and availability across all data systems, working autonomously to troubleshoot and resolve issues.

Collaborate with data scientists, analysts, and software engineers to ensure seamless integration between data pipelines and data models.

Continuously monitor and optimize the performance of data pipelines, databases, and infrastructure, ensuring high performance and scalability.

Design and implement solutions for data storage, including database schemas, indexing, and partitioning strategies.

Manage cloud-based data environments (AWS, GCP, Azure) and infrastructure as code using tools like Terraform or CloudFormation.

Qualifications:

Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field. A Master's degree is a plus.

5+ years of experience as a Data Engineer or in a similar role, building and maintaining data pipelines and infrastructure.

Primary Skill

Strong programming skills in Python, Pyspark, Spark/Scala for data processing tasks.

Proficiency in writing complex SQL queries for data manipulation and extraction.

Experience with ETL tools for building and orchestrating data workflows.

Hands-on experience with cloud platforms (AWS, GCP, Azure) and their data-related services such as S3, Redshift, BigQuery, Dataflow, Databricks, or Snowflake.

Familiarity with relational databases (MySQL, PostgreSQL, SQL Server) and NoSQL databases (MongoDB, Cassandra, HBase).

Experience with version control systems (e.g., Git) and familiarity with CI/CD pipelines for data pipelines and infrastructure.

Excellent communication and collaboration skills, with the ability to work effectively in a cross-functional team environment.

Implement data pipelines and workshops experience with Palantir foundry.

Javascript / Typescript

Preferred Skills:

Azure cloud knowledge, especially with Delta Lake, ADLS, Event Hubs, and Cosmos DB,

Experience with real-time data processing frameworks like Kafka Streams, Apache Flink, or Spark Streaming.

Experience with machine learning or data science pipelines

Experience with containerization and orchestration tools like Docker and Kubernetes.

Knowledge of APIs and microservices architectures for data integration.

Understanding of data lake architectures and experience with Delta Lake or similar technologies.

Sign up for Job Alerts