Data QA Automation Engineer
Streamhub is a data analytics and activation SaaS platform encompassing audience measurement, segmentation, and targeting. Our purpose is to understand how video shapes everyday life, and our mission is to make people's everyday lives better by being the most actionable data platform for the video business. Working at Streamhub offers unparalleled access to cutting-edge challenges in the complex adtech/media industry, within an international and rewarding entrepreneurial culture. We are in an exciting growth phase, making several key hires who will form the building blocks of our teams in India and Japan!
The Role
This is a challenging opportunity to contribute to an innovative data product as a Data Quality Automation Engineer on our core engineering team - joining a highly productive team and a culture built on individual growth, well-being, strong ethics, and an entrepreneurial approach.
We're looking for a strong, hands-on engineer - not a traditional manual QA hire - who believes automation is the default answer to quality at scale. You'll join the Quality Automation function as part of a dedicated squad, owning the accuracy of customer-facing analytics platforms, API testing, and AI eval. Day-to-day, that means identifying, scoping, and resolving data quality problems, and building the automated systems that catch issues before they reach customers.
This role is explicitly designed to grow: we expect it to evolve into a full-stack, AI-forward engineering role, contributing beyond data quality into other modules across the product as your range proves out.
Responsibilities & Required Skills
- Strong, hands-on coding skills in JavaScript or Python, with the ability to contribute to full-stack modules - including React, Python, API development, and data modeling - as the role expands beyond data quality.
- Develop and maintain API and web automation frameworks (Cypress, Playwright, Selenium, or similar).
- Design and maintain automated validation frameworks for ETL pipelines and large-scale data systems.
- Validate complex data transformations, aggregations, and business rules across high-volume datasets.
- Automate all workflows related to data quality, end-to-end testing, and agent eval.
- Leverage Generative AI (GenAI) to:
- Accelerate test case and script generation.
- Improve automation coverage and edge-case detection.
- Detect anomalies, regressions, and data drift.
- Enhance root cause analysis and the measurement of automation impact.
- Integrate automated validation and quality gates into CI/CD pipelines to enforce regression, schema, and contract checks before deployment.
- Perform advanced SQL and database validation testing, including reconciliation and data integrity checks.
Nice to Have
- Hands-on experience with AI-assisted development tools such as Cursor or similar.
- Experience with performance testing tools such as JMeter or Gatling.
- Knowledge of additional programming languages such as Python, Java, or Scala.
- Familiarity with distributed data systems, event-driven architectures, or large-scale data platforms.
The Package
- Competitive salary.
- Immense learning, exposure to niche data technologies, and handling TBs of data.
- Work with an open, diverse, and autonomous team.
- Entrepreneurial team culture.
- Flexible timing.
If you've read all the way to the bottom of this description, thank you for your interest in Streamhub!
We are committed to equal employment opportunities regardless of race, colour, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender, gender identity or expression, or veteran status. We are proud to be an equal opportunity workplace.
Employment Type: FULL_TIME