Skip to main content
Posted 22 July, 2026

Lead ML Engineer

Societe Generale
India-Bangalore Full Time
Reference: 396_132173_260001GN

Responsibilities

  • Lead and own the endtoend production lifecycle of ML and LLM models (must have), ensuring models are deployable, scalable, observable, and maintainable.
  • Define and enforce ML engineering and MLOps standards across teams (must have).
  • Design and maintain CI/CD pipelines for ML workloads (must have).
  • Act as technical lead and mentor for ML engineers and contributors (must have).
  • Partner with Data Scientists to industrialize research into production systems (must have).
  • Collaborate with Platform, Cloud, and Data Engineering teams on infrastructure and runtime alignment (must have).
  • Own model monitoring, drift detection, testing, rollback, and incident analysis (must have).
  • Evaluate and introduce new ML, GenAI, and MLOps tools with a pragmatic, enterprise mindset (good to have).
  • Contribute to ML governance, reproducibility, and responsible AI practices (good to have).

Key Skills & Expertise

  • Cloud & DevOps: Azure (must have), CI/CD using Jenkins, GitHub Actions, ArgoCD (must have)
  • Container & Orchestration: Docker, Kubernetes (must have)
  • Workflow Orchestration: Airflow (must have)
  • Programming: Productiongrade Python (must have)
  • ML Engineering & GenAI: LLM integration, prompt engineering, model packaging and lifecycle management (must have)
  • Testing & Quality: Pytest, integration and system testing for ML systems (must have)
  • Data: SQL, relational databases, basic reporting and dashboards (must have)
  • ML Platforms: MLflow, Databricks (good to have)
  • LLM Frameworks: LangChain, LangGraph, agentbased patterns (good to have)
  • Data Science Awareness: ML algorithms, feature engineering, evaluation metrics, bias/leakage awareness (awareness required)
  • Specialized Use Cases: OCR and document processing pipelines (good to have)
  • Frontend / Visualization: Streamlit, widgets, lightweight UI layers (good to have)
  • Mindset: Awareness of emerging technologies and new tooling (good to have)

a { text-decoration: none; color: #464feb;}tr th, tr td { border: 1px solid #e6e6e6;}tr th { background-color: #f5f5f5;}

Responsibilities: ACE can help you write this job description (go/ACE)

Key Responsibilities

  • AI & ML Strategy: Advise, design, and execute ML- and LLM-driven transformations of business processes with clear, measurable outcomes.
  • End-to-End AI Solutions: Build and operationalize ML/LLM solutions including data pipelines, feature engineering, modelling, APIs, deployment, and monitoring.
  • AI Platform Engineering: Architect and scale distributed AI execution platforms capable of running thousands of concurrent models or agents
  • Agent & Model Lifecycle Management: Design services and APIs for agent/model registration, versioning, execution, and monitoring.
  • Scalability & Performance: Optimize distributed systems and microservices for low latency, high throughput, and cost efficiency.
  • Governance & Reliability: Establish standards for validation, monitoring, drift detection, bias mitigation, safety guardrails, and compliance.
  • LLMOps / MLOps: Implement CI/CD, automated evaluation, observability, and rollback strategies for ML and LLM systems.
  • Cross-Functional Leadership: Partner with Product, Engineering, and Business teams to identify opportunities and operationalize AI at scale.
  • Mentorship & Best Practices: Mentor engineers and data scientists; define best practices for AI platform, ML, and LLM development.

Required Skills & Experience

  • 12 years of experience in AI/ML systems, or/and AI platform engineering.
  • Expert proficiency in Python or similar programming language
  • Deep experience building and scaling distributed systems (e.g., Kafka, stream processing, HPC or large-scale compute frameworks).
  • Advanced knowledge of machine learning: regression, classification, clustering, tree-based models, ensembles, Bayesian/Markov methods.
  • Hands-on experience with Large Language Models (GPT, BERT, or similar), including prompt engineering, RAG pipelines, and evaluation.
  • Strong NLP expertise: text classification, summarization, question answering, and entity recognition.
  • Experience designing end-to-end AI pipelines: data ingestion, training, deployment, monitoring, and feedback loops.
  • Strong knowledge of cloud platforms (Azure or AWS), Kubernetes, autoscaling
  • Experience building high-throughput APIs (REST/gRPC) and platform service interfaces.
  • Strong understanding of observability (logging, metrics, tracing), security, and performance optimization

Familiarity with agentic frameworks, large language models (LLMs), agent protocols (MCP, A2A) and their unique deployment challenges will be plus

Profile Required: ACE can help you write this job description (go/ACE)

Key Responsibilities

  • AI & ML Strategy: Advise, design, and execute ML- and LLM-driven transformations of business processes with clear, measurable outcomes.
  • End-to-End AI Solutions: Build and operationalize ML/LLM solutions including data pipelines, feature engineering, modelling, APIs, deployment, and monitoring.
  • AI Platform Engineering: Architect and scale distributed AI execution platforms capable of running thousands of concurrent models or agents
  • Agent & Model Lifecycle Management: Design services and APIs for agent/model registration, versioning, execution, and monitoring.
  • Scalability & Performance: Optimize distributed systems and microservices for low latency, high throughput, and cost efficiency.
  • Governance & Reliability: Establish standards for validation, monitoring, drift detection, bias mitigation, safety guardrails, and compliance.
  • LLMOps / MLOps: Implement CI/CD, automated evaluation, observability, and rollback strategies for ML and LLM systems.
  • Cross-Functional Leadership: Partner with Product, Engineering, and Business teams to identify opportunities and operationalize AI at scale.
  • Mentorship & Best Practices: Mentor engineers and data scientists; define best practices for AI platform, ML, and LLM development.

  • Required Skills & Experience

  • 12 years of experience in AI/ML systems, or/and AI platform engineering.
  • Expert proficiency in Python or similar programming language
  • Deep experience building and scaling distributed systems (e.g., Kafka, stream processing, HPC or large-scale compute frameworks).
  • Advanced knowledge of machine learning: regression, classification, clustering, tree-based models, ensembles, Bayesian/Markov methods.
  • Hands-on experience with Large Language Models (GPT, BERT, or similar), including prompt engineering, RAG pipelines, and evaluation.
  • Strong NLP expertise: text classification, summarization, question answering, and entity recognition.
  • Experience designing end-to-end AI pipelines: data ingestion, training, deployment, monitoring, and feedback loops.
  • Strong knowledge of cloud platforms (Azure or AWS), Kubernetes, autoscaling
  • Experience building high-throughput APIs (REST/gRPC) and platform service interfaces.
  • Strong understanding of observability (logging, metrics, tracing), security, and performance optimization

Familiarity with agentic frameworks, large language models (LLMs), agent protocols (MCP, A2A) and their unique deployment challenges will be plus

Sign up for Job Alerts