AGENTIC AI
8-11 years in IT (Strong SDET) experience, including a minimum of 3 years dedicated to evaluating AI, LLM, and Agentic systems, preferably in a product-led organization or AI research lab. Define evaluation models for LLM and Agentic workflow safety and reliability, combining traditional QA foundations with next-gen AI testing paradigms to lead niche skills groups. Build automated evaluation frameworks for LLMs, RAG pipelines, and Agentic workflows. Implement "LLM-as-a-judge" benchmarks, HITL, and tools like LangSmith, TruLens, DeepEval, Phoenix, and Ragas into CI/CD. Deploy agentic frameworks (Perception/Reasoning/Action) in AWS, GCP, and Azure using containerization, service governance, and cloud orchestration. Validate functional output quality (hallucinations, relevance, tool-calling precision) and non-functional vectors (red-teaming, prompt injection, jailbreaking, toxicity, latency, cost). Mentor QA/SDETs into niche AI testing specialists, support presales, and drive market authority through articles and open-source benchmarks. Requires expert-level Java, Python, and Node.JS scripting alongside advanced hands-on automation with Selenium or Playwright, RestAssured, mobile testing, and strong manual testing foundations. AI/ML expertise must cover Transformer architectures, Vector DBs (Pinecone, Milvus), fine-tuning, and frameworks like LangChain, LlamaIndex, or AutoGen. Proven experience with BLEU, ROUGE, G-Eval, BERTScore, Context Precision, Faithfulness, MLflow, Weights & Biases, and cloud infrastructure-as-code. Requires a product mindset for testing non-deterministic systems, consultative exec-level communication, and management of onshore/offshore agile teams.