Skip to main content
Posted 15 August, 2026

Senior Research Engineer

Weekday AI
Bengaluru,Karnataka,India Full Time
Reference: 8_688697_0619619D6C

This role is for one of Weekday's clients
Salary range: Rs 5000000 - Rs 9000000 (ie INR 50 - 90 LPA)

Min Experience: 1+ years
Location: Bengaluru
JobType: full-time

We are looking for a highly motivated Senior Research Engineer with 1-8 years of experience to join our research and engineering team. The ideal candidate will work at the intersection of LLM research, model evaluation, and large-scale experimentation, helping develop rigorous methods to understand, measure, and improve the capabilities of modern language models.

You will design and implement evaluation frameworks, build high-quality datasets and test suites, analyze model behavior, and translate research findings into practical improvements. This role is ideal for someone who enjoys solving open-ended research problems while also being comfortable building production-quality systems.

Requirements

Key Responsibilities

  • Design, develop, and maintain comprehensive LLM evaluation (evals) frameworks to measure model capabilities, reliability, reasoning, instruction following, safety, and task performance.
  • Conduct research on large language models, including model behavior, capabilities, limitations, prompting, fine-tuning, and evaluation methodologies.
  • Develop novel evaluation methodologies and experiments for emerging LLM capabilities and use cases.
  • Create high-quality evaluation datasets, test cases, rubrics, and automated evaluation pipelines.
  • Analyze model outputs using quantitative and qualitative methods to identify performance gaps and behavioral patterns.
  • Design controlled experiments to compare models, prompts, training approaches, and inference strategies.
  • Build scalable tooling for running evaluations across large numbers of prompts, models, and datasets.
  • Collaborate with researchers, ML engineers, and product teams to convert research insights into measurable model improvements.
  • Investigate failures and edge cases and develop targeted evaluations to capture previously undetected model weaknesses.
  • Contribute to technical documentation, research reports, internal benchmarks, and presentations of findings.
  • Stay current with developments in LLM research, evaluation techniques, reasoning systems, and AI benchmarks.

Required Skills & Qualifications

  • 1-8 years of experience in machine learning, AI research, software engineering, data science, or a related technical field.
  • Strong hands-on experience designing and implementing LLM evals or model evaluation systems.
  • Solid understanding of LLM research, including model capabilities, prompting, fine-tuning, inference, and evaluation methodologies.
  • Strong Python programming and experience working with ML/AI frameworks and data-processing pipelines.
  • Ability to formulate research questions, design experiments, interpret results, and communicate technical findings clearly.
  • Strong analytical and problem-solving skills with attention to experimental rigor and reproducibility.
  • Experience working with large datasets, automated testing, and evaluation pipelines.

Good-to-Have Skills

  • Experience developing or working with LLM benchmarks and standardized evaluation suites.
  • Familiarity with benchmark design, dataset curation, scoring methodologies, and statistical analysis.
  • Experience with open-source LLMs, model APIs, Hugging Face, PyTorch, or similar frameworks.
  • Exposure to reinforcement learning, RLHF/RLAIF, fine-tuning, synthetic data generation, or agentic systems.
  • Research publications, technical blogs, open-source contributions, or demonstrated independent AI research work.

Sign up for Job Alerts