Senior Research Engineer
This role is for one of Weekday's clients
Salary range: Rs 5000000 - Rs 9000000 (ie INR 50 - 90 LPA)
Min Experience: 1+ years
Location: Bengaluru
JobType: full-time
We are looking for a highly motivated Senior Research Engineer with 1-8 years of experience to join our research and engineering team. The ideal candidate will work at the intersection of LLM research, model evaluation, and large-scale experimentation, helping develop rigorous methods to understand, measure, and improve the capabilities of modern language models.
You will design and implement evaluation frameworks, build high-quality datasets and test suites, analyze model behavior, and translate research findings into practical improvements. This role is ideal for someone who enjoys solving open-ended research problems while also being comfortable building production-quality systems.
Requirements
Key Responsibilities
- Design, develop, and maintain comprehensive LLM evaluation (evals) frameworks to measure model capabilities, reliability, reasoning, instruction following, safety, and task performance.
- Conduct research on large language models, including model behavior, capabilities, limitations, prompting, fine-tuning, and evaluation methodologies.
- Develop novel evaluation methodologies and experiments for emerging LLM capabilities and use cases.
- Create high-quality evaluation datasets, test cases, rubrics, and automated evaluation pipelines.
- Analyze model outputs using quantitative and qualitative methods to identify performance gaps and behavioral patterns.
- Design controlled experiments to compare models, prompts, training approaches, and inference strategies.
- Build scalable tooling for running evaluations across large numbers of prompts, models, and datasets.
- Collaborate with researchers, ML engineers, and product teams to convert research insights into measurable model improvements.
- Investigate failures and edge cases and develop targeted evaluations to capture previously undetected model weaknesses.
- Contribute to technical documentation, research reports, internal benchmarks, and presentations of findings.
- Stay current with developments in LLM research, evaluation techniques, reasoning systems, and AI benchmarks.
Required Skills & Qualifications
- 1-8 years of experience in machine learning, AI research, software engineering, data science, or a related technical field.
- Strong hands-on experience designing and implementing LLM evals or model evaluation systems.
- Solid understanding of LLM research, including model capabilities, prompting, fine-tuning, inference, and evaluation methodologies.
- Strong Python programming and experience working with ML/AI frameworks and data-processing pipelines.
- Ability to formulate research questions, design experiments, interpret results, and communicate technical findings clearly.
- Strong analytical and problem-solving skills with attention to experimental rigor and reproducibility.
- Experience working with large datasets, automated testing, and evaluation pipelines.
Good-to-Have Skills
- Experience developing or working with LLM benchmarks and standardized evaluation suites.
- Familiarity with benchmark design, dataset curation, scoring methodologies, and statistical analysis.
- Experience with open-source LLMs, model APIs, Hugging Face, PyTorch, or similar frameworks.
- Exposure to reinforcement learning, RLHF/RLAIF, fine-tuning, synthetic data generation, or agentic systems.
- Research publications, technical blogs, open-source contributions, or demonstrated independent AI research work.