Skip to main content
Posted 24 August, 2026

MLOps / Cloud Deployment Engineer

Xenon7
Hyderabad,Telangana,India Full Time
Reference: 8_758585_697456F4B1

Our Client's Digital Finance IT is scaling AI and agentic systems in production. We need an MLOps / Cloud Deployment Engineer to own the deployment, reliability, observability, and operational scale of these systems in a regulated enterprise environment.

This is a cloud and platform engineering role with deep MLOps/LLMOps focus, not a model-building role. You will operate the runway that ML and GenAI systems run on, not build the models themselves.

What You'll Do

  • Own CI/CD pipelines for ML models, RAG applications, and agentic AI systems - from experiment to production
  • Deploy and operate AI workloads on cloud-native ML/AI platforms - AWS Bedrock/SageMaker, Azure AI Foundry / Azure Machine Learning, or equivalent
  • Build and maintain observability, tracing, and monitoring for LLM and agentic systems - latency, cost, hallucination rates, tool-call success, drift detection
  • Implement model governance and guardrails - approval gates, kill-switches, escalation paths, audit trails
  • Manage infrastructure-as-code (Terraform, Bicep, or equivalent) for reproducible AI/ML environments
  • Design cost and performance optimization strategies - token usage tracking, caching, model routing, autoscaling, warehouse/cluster right-sizing
  • Own security posture - RBAC, secret management (Key Vault / Secrets Manager), prompt-injection risk mitigation, auditability for regulated pharma
  • Partner with data engineers, AI engineers, and Finance business stakeholders to move systems from prototype to reliable production
  • Implement evaluation frameworks for AI systems in production - regression testing, adversarial testing, accuracy tracking, hallucination monitoring

Requirements

Must-Have Experience

  • 5+ years in cloud/DevOps/MLOps engineering on AWS, Azure, or GCP
  • Production deployment of ML or GenAI systems - CI/CD, containerization (Docker/Kubernetes), infrastructure-as-code (Terraform)
  • MLOps tooling - MLflow, SageMaker Pipelines, Azure ML Pipelines, or equivalent
  • LLM/GenAI operational experience - observability tools (LangSmith, Weights & Biases, or equivalent), cost monitoring, latency optimization, prompt/model versioning
  • Cloud-native AI platforms - hands-on with at least one of: AWS Bedrock, SageMaker, Azure AI Foundry, Azure OpenAI, Vertex AI
  • Python, Bash, and infrastructure scripting - strong
  • Security and governance in regulated environments - RBAC, secrets, audit, compliance

Nice to Have

  • Pharma, life sciences, or regulated financial services domain
  • Experience operating agentic AI systems in production - multi-agent orchestration, tool-calling, human-in-the-loop workflows
  • LangChain, LangGraph, CrewAI, AutoGen, or Semantic Kernel operational experience
  • Kubernetes-native ML platforms (Kubeflow, Ray)
  • Snowflake or Databricks operational experience (compute governance, cost management)
  • Certifications: AWS/Azure ML Engineer, Kubernetes CKA/CKAD, Terraform Associate

What We're NOT Looking For

  • Data Scientists or research engineers - this is a production platform role
  • Application developers with light DevOps exposure - need real MLOps/cloud engineering depth
  • Pure infra engineers with no AI/ML operational experience - need to understand what makes LLM systems different (evals, hallucinations, prompt versioning, RAG grounding)

Sign up for Job Alerts