Posted 26 July, 2026
Specialist- Data Engineering
NR Consulting
Pune,Maharashtra
Full Time
Reference: 365_463738_26-21913
Title: Specialist- Data Engineering
Location: Pune
Exp: 4-7 Years
Job Description:
Key Responsibilities
Architect and implement the three-tier data storage model: Azure SQL (live transactional SNPD DB) ADLS Gen2 Delta Lake (historical analytics / ML features, nightly ADF ETL) Azure AI Search (vector / RAG index).
Build and maintain Azure Data Factory (ADF) pipelines: nightly CDC-based ETL from the SNPD SQL database, SAP (cost/PO/BOM/vendor - masked at API layer), and Teamcenter PLM (BOM snapshots, part lifecycle, ECN).
Design and implement the unified SNPD data model: enforce project_id + part_number as universal pivot keys across all tables; ensure referential integrity across SNPD core domain tables and migrated portal tables (NVPC, RFQ, PPAP, CDMM, etc.).
Build the AI feature store on Delta Lake: dl_gate_cycle_times, dl_supplier_risk, dl_nvpc_benchmarks, dl_cost_variance, dl_deliverable_actuals - with incremental refresh, partitioning, and Z-ordering for query performance.
Conduct data audits on SAP S/4 HANA, Teamcenter PLM, and all 8 legacy homegrown portal databases; assess data quality, identify gaps, and remediate for ML readiness.
Implement SAP cost data masking at the API / pipeline layer - sensitive pricing data must be obfuscated before reaching any MCP server or AI agent.
Set up Azure AI Search vector index: embedding ingestion pipeline from the document store (SharePoint / Azure Blob), chunking strategy, metadata schema, and incremental re-indexing on document updates.
Establish data lineage, quality checks, and observability: row counts, null rates, schema drift alerts, and SLA monitoring for all ETL pipelines.
Support historical data migration: 5-7 years of legacy SNPD and portal data into the unified SNPD database; validate referential integrity and completeness post-migration.
Collaborate with the ML Engineer to serve training datasets from Delta Lake; optimize feature computation using Synapse Serverless or Databricks as compute.
Implement RBAC and data access controls at the data layer: ensure user-level and role-level scoping is enforced from Azure SQL through to Delta Lake reads and vector search results.
Maintain data catalogue and schema documentation; ensure all entities conform to the IATF 16949 audit traceability requirements.
TECHNICAL SKILLS REQUIRED
GOOD TO HAVE
Hands-on SAP S/4 HANA RISE data extraction experience (ACDOCA, Material Master, BOM, MM60).
Teamcenter PLM API familiarity (REST/SOA Gateway, BOM export, ECN feeds).
dbt (data build tool) for transformation layer on Delta Lake.
Apache Kafka / Azure Event Hubs for real-time streaming from SAP change events.
Experience with IATF 16949 or automotive quality data requirements.
Familiarity with Mahindra data platform (MDP) or Azure Purview for data governance.
Location: Pune
Exp: 4-7 Years
Job Description:
Key Responsibilities
Architect and implement the three-tier data storage model: Azure SQL (live transactional SNPD DB) ADLS Gen2 Delta Lake (historical analytics / ML features, nightly ADF ETL) Azure AI Search (vector / RAG index).
Build and maintain Azure Data Factory (ADF) pipelines: nightly CDC-based ETL from the SNPD SQL database, SAP (cost/PO/BOM/vendor - masked at API layer), and Teamcenter PLM (BOM snapshots, part lifecycle, ECN).
Design and implement the unified SNPD data model: enforce project_id + part_number as universal pivot keys across all tables; ensure referential integrity across SNPD core domain tables and migrated portal tables (NVPC, RFQ, PPAP, CDMM, etc.).
Build the AI feature store on Delta Lake: dl_gate_cycle_times, dl_supplier_risk, dl_nvpc_benchmarks, dl_cost_variance, dl_deliverable_actuals - with incremental refresh, partitioning, and Z-ordering for query performance.
Conduct data audits on SAP S/4 HANA, Teamcenter PLM, and all 8 legacy homegrown portal databases; assess data quality, identify gaps, and remediate for ML readiness.
Implement SAP cost data masking at the API / pipeline layer - sensitive pricing data must be obfuscated before reaching any MCP server or AI agent.
Set up Azure AI Search vector index: embedding ingestion pipeline from the document store (SharePoint / Azure Blob), chunking strategy, metadata schema, and incremental re-indexing on document updates.
Establish data lineage, quality checks, and observability: row counts, null rates, schema drift alerts, and SLA monitoring for all ETL pipelines.
Support historical data migration: 5-7 years of legacy SNPD and portal data into the unified SNPD database; validate referential integrity and completeness post-migration.
Collaborate with the ML Engineer to serve training datasets from Delta Lake; optimize feature computation using Synapse Serverless or Databricks as compute.
Implement RBAC and data access controls at the data layer: ensure user-level and role-level scoping is enforced from Azure SQL through to Delta Lake reads and vector search results.
Maintain data catalogue and schema documentation; ensure all entities conform to the IATF 16949 audit traceability requirements.
TECHNICAL SKILLS REQUIRED
| Azure Data Factory (ADF) - pipeline authoring | Azure AI Search - index management, embedding pipelines |
| ADLS Gen2 / Delta Lake - storage & compute | Data modelling - relational + lakehouse schemas |
| Azure SQL / SQL Server - T-SQL, stored procedures | Azure Key Vault - secrets, connection string management |
| Azure Synapse Analytics or Databricks | Data lineage & observability tools |
| Python - PySpark, pandas, data quality scripts | Git / Azure DevOps for pipeline version control |
| SAP OData / RFC / BAPI integration patterns | CDC (Change Data Capture) patterns in SQL Server |
GOOD TO HAVE
Hands-on SAP S/4 HANA RISE data extraction experience (ACDOCA, Material Master, BOM, MM60).
Teamcenter PLM API familiarity (REST/SOA Gateway, BOM export, ECN feeds).
dbt (data build tool) for transformation layer on Delta Lake.
Apache Kafka / Azure Event Hubs for real-time streaming from SAP change events.
Experience with IATF 16949 or automotive quality data requirements.
Familiarity with Mahindra data platform (MDP) or Azure Purview for data governance.