Key Responsibilities:<\/span>
<\/p>\n \n - Design, develop, and maintain ETL (Extract, Transform, Load) processes to ensure the seamless integration of raw data from various sources into our data lakes or warehouses.<\/span>
<\/li>\n - Utilize Python, PySpark, SQL and AirFlow etc., to process, analyze, and store large\-scale datasets efficiently.<\/span>
<\/li>\n - Write and maintain SQL queries for data retrieval, transformation, and storage in relational databases like Redshift or PostgreSQL.<\/span>
<\/li>\n - Support cloud\-based data platforms such as AWS, Azure, or GCP, with a focus on orchestrating AI retraining cycles, versioning, and automated pipeline monitoring.<\/span>
<\/li>\n - Familiarity in converting unstructured data into vectors using frameworks like LangChain or LlamaIndex and storing them.<\/span>
<\/li>\n - Collaborate with cross\-functional teams, including data scientists, ML engineers, and domain experts to design and implement scalable solutions.<\/span>
<\/li>\n - Troubleshoot and resolve performance issues, data quality problems, and errors in data pipelines.<\/span>
<\/li>\n - Document processes, code, and best practices for future reference and team training.<\/span>
<\/li>\n <\/ul> <\/span>
<\/p><\/span>\n \n
\n <\/div><\/span>
Requirements<\/h3>Additional Information:<\/span>
<\/p>\n \n - Experience level 3+ years.<\/span>
<\/li>\n - Strong understanding of data governance, security, and compliance principles is preferred.<\/span>
<\/li>\n - Ability to work independently and as part of a team in a fast\-paced environment.<\/span>
<\/li>\n - Excellent problem\-solving skills with the ability to identify inefficiencies and propose solutions.<\/span>
<\/li>\n - Experience with version control systems (e.g., Git) and scripting languages for automation tasks.<\/span>
<\/li>\n <\/ul><\/span>\n \n
\n <\/div><\/span>
\n <\/body>\n<\/html>