Skip to main content
Posted 26 July, 2026

Specialist- Data Engineering

NR Consulting
Pune,Maharashtra Full Time
Reference: 365_463738_26-21913

Title: Specialist- Data Engineering
Location: Pune
Exp: 4-7 Years

Job Description:
Key Responsibilities

Architect and implement the three-tier data storage model: Azure SQL (live transactional SNPD DB) ADLS Gen2 Delta Lake (historical analytics / ML features, nightly ADF ETL) Azure AI Search (vector / RAG index).
Build and maintain Azure Data Factory (ADF) pipelines: nightly CDC-based ETL from the SNPD SQL database, SAP (cost/PO/BOM/vendor - masked at API layer), and Teamcenter PLM (BOM snapshots, part lifecycle, ECN).
Design and implement the unified SNPD data model: enforce project_id + part_number as universal pivot keys across all tables; ensure referential integrity across SNPD core domain tables and migrated portal tables (NVPC, RFQ, PPAP, CDMM, etc.).
Build the AI feature store on Delta Lake: dl_gate_cycle_times, dl_supplier_risk, dl_nvpc_benchmarks, dl_cost_variance, dl_deliverable_actuals - with incremental refresh, partitioning, and Z-ordering for query performance.
Conduct data audits on SAP S/4 HANA, Teamcenter PLM, and all 8 legacy homegrown portal databases; assess data quality, identify gaps, and remediate for ML readiness.
Implement SAP cost data masking at the API / pipeline layer - sensitive pricing data must be obfuscated before reaching any MCP server or AI agent.
Set up Azure AI Search vector index: embedding ingestion pipeline from the document store (SharePoint / Azure Blob), chunking strategy, metadata schema, and incremental re-indexing on document updates.
Establish data lineage, quality checks, and observability: row counts, null rates, schema drift alerts, and SLA monitoring for all ETL pipelines.
Support historical data migration: 5-7 years of legacy SNPD and portal data into the unified SNPD database; validate referential integrity and completeness post-migration.
Collaborate with the ML Engineer to serve training datasets from Delta Lake; optimize feature computation using Synapse Serverless or Databricks as compute.
Implement RBAC and data access controls at the data layer: ensure user-level and role-level scoping is enforced from Azure SQL through to Delta Lake reads and vector search results.
Maintain data catalogue and schema documentation; ensure all entities conform to the IATF 16949 audit traceability requirements.

TECHNICAL SKILLS REQUIRED
Azure Data Factory (ADF) - pipeline authoring Azure AI Search - index management, embedding pipelines
ADLS Gen2 / Delta Lake - storage & compute Data modelling - relational + lakehouse schemas
Azure SQL / SQL Server - T-SQL, stored procedures Azure Key Vault - secrets, connection string management
Azure Synapse Analytics or Databricks Data lineage & observability tools
Python - PySpark, pandas, data quality scripts Git / Azure DevOps for pipeline version control
SAP OData / RFC / BAPI integration patterns CDC (Change Data Capture) patterns in SQL Server

GOOD TO HAVE
Hands-on SAP S/4 HANA RISE data extraction experience (ACDOCA, Material Master, BOM, MM60).
Teamcenter PLM API familiarity (REST/SOA Gateway, BOM export, ECN feeds).
dbt (data build tool) for transformation layer on Delta Lake.
Apache Kafka / Azure Event Hubs for real-time streaming from SAP change events.
Experience with IATF 16949 or automotive quality data requirements.
Familiarity with Mahindra data platform (MDP) or Azure Purview for data governance.

Sign up for Job Alerts