Posted 11 August, 2026
Lead Software Engineer - MLOps Engineer
Societe Generale
India-Bangalore
Full Time
Reference: 396_132173_26000IAE
Job Description: ML OPS Engineer (LEAD)
You will be responsible for building, operating, and evolving the Data Science Delivery Platform (DSDP), enabling Data Scientists and Engineers to develop, deploy, and scale AI/ML solutions securely and efficiently. You will work closely with Data Science teams, infrastructure teams, and business stakeholders to maintain and evolve a reliable, self-service AI platform.
Key Responsibilities
- Build and maintain scalable platform services on Kubernetes.
- Build and maintain tools to facilitate the end-to-end data science lifecycle, covering experimentation, deployment, monitoring, and governance.
- Support and enhance enterprise AI tools such as Dataiku, MLFlow, Snowflake, and Spark.
- Implement automation, CI/CD pipelines, and platform engineering best practices.
- Ensure platform reliability, observability, security, and compliance.
- Provide technical support and enablement for Data Scientists and Data Engineers.
- Continuously improve platform capabilities and user experience.
Technical Environment
- Infrastructure: Kubernetes, Docker, Linux
- AI/ML: Dataiku, MLFlow, Kedro, Jupyter Notebooks
- Data: Snowflake, PostgreSQL, Spark, Hadoop
- Development: Python, FastAPI, SQLAlchemy, PyTest
- DevOps: GitHub Actions, Jenkins, Ansible, Terraform, Harbor, JFrog
- Observability: Grafana, Kibana, Elasticsearch, Zabbix
Required Skills
- Strong hands-on experience with Kubernetes and container technologies; debugging platform issues and operational anomalies in Kubernetes environments is a core day-to-day responsibility
- Strong server administration expertise across Linux environments, system operations, networking, performance troubleshooting, access management, and infrastructure reliability
- Practical Python development skills for platform automation, operational tooling, integration, and troubleshooting activities
- Hands-on experience with Terraform for infrastructure provisioning, configuration management, and infrastructure-as-code automation
- Experience building and maintaining CI/CD pipelines using GitHub Actions, Ansible, Jenkins, and related platform automation frameworks
- Good understanding of MLOps practices, AI/ML lifecycle management, model deployment, and Data Science workflows
- Knowledge of monitoring, observability, troubleshooting, and platform operations to ensure service reliability and operational continuity
- Experience with Dataiku, MLflow, or similar AI/ML platforms is required
Soft Skills
- Customer-focused and collaborative mindset - You will have to interact with data scientists and support them with platform issues on a day-to-day basis.
- Strong ownership and problem-solving abilities.
- Ability to communicate effectively with technical and non-technical stakeholders.
- A drive for innovation, simplification, and continuous improvement.
Preferred Experience
- 8 years in Platform Engineering and MLOps Engineering.
- Experience supporting enterprise AI/ML platforms at scale.
- Exposure to GenAI/LLMOps technologies is a plus.