Posted 23 July, 2026
Specialist Data Engineering
NR Consulting - India
Pune, Maharashtra, IN
Full Time
Reference: 26-21913-2220-1
Title: Specialist— Data Engineering
Location: Pune
Exp: 4-7 Years
Job Description:
Key Responsibilities
• Architect and implement the three-tier data storage model: Azure SQL (live transactional SNPD DB) → ADLS Gen2 Delta Lake (historical analytics / ML features, nightly ADF ETL) → Azure AI Search (vector / RAG index).
• Build and maintain Azure Data Factory (ADF) pipelines: nightly CDC-based ETL from the SNPD SQL database, SAP (cost/PO/BOM/vendor — masked at API layer), and Teamcenter PLM (BOM snapshots, part lifecycle, ECN).
• Design and implement the unified SNPD data model: enforce project_id + part_number as universal pivot keys across all tables; ensure referential integrity across SNPD core domain tables and migrated portal tables (NVPC, RFQ, PPAP, CDMM, etc.).
• Build the AI feature store on Delta Lake: dl_gate_cycle_times, dl_supplier_risk, dl_nvpc_benchmarks, dl_cost_variance, dl_deliverable_actuals — with incremental refresh, partitioning, and Z-ordering for query performance.
• Conduct data audits on SAP S/4 HANA, Teamcenter PLM, and all 8 legacy homegrown portal databases; assess data quality, identify gaps, and remediate for ML readiness.
• Implement SAP cost data masking at the API / pipeline layer — sensitive pricing data must be obfuscated before reaching any MCP server or AI agent.
• Set up Azure AI Search vector index: embedding ingestion pipeline from the document store (SharePoint / Azure Blob), chunking strategy, metadata schema, and incremental re-indexing on document updates.
• Establish data lineage, quality checks, and observability: row counts, null rates, schema drift alerts, and SLA monitoring for all ETL pipelines.
• Support historical data migration: 5–7 years of legacy SNPD and portal data into the unified SNPD database; validate referential integrity and completeness post-migration.
• Collaborate with the ML Engineer to serve training datasets from Delta Lake; optimize feature computation using Synapse Serverless or Databricks as compute.
• Implement RBAC and data access controls at the data layer: ensure user-level and role-level scoping is enforced from Azure SQL through to Delta Lake reads and vector search results.
• Maintain data catalogue and schema documentation; ensure all entities conform to the IATF 16949 audit traceability requirements.
TECHNICAL SKILLS REQUIRED
GOOD TO HAVE
• Hands-on SAP S/4 HANA RISE data extraction experience (ACDOCA, Material Master, BOM, MM60).
• Teamcenter PLM API familiarity (REST/SOA Gateway, BOM export, ECN feeds).
• dbt (data build tool) for transformation layer on Delta Lake.
• Apache Kafka / Azure Event Hubs for real-time streaming from SAP change events.
• Experience with IATF 16949 or automotive quality data requirements.
• Familiarity with Mahindra data platform (MDP) or Azure Purview for data governance.
Location: Pune
Exp: 4-7 Years
Job Description:
Key Responsibilities
• Architect and implement the three-tier data storage model: Azure SQL (live transactional SNPD DB) → ADLS Gen2 Delta Lake (historical analytics / ML features, nightly ADF ETL) → Azure AI Search (vector / RAG index).
• Build and maintain Azure Data Factory (ADF) pipelines: nightly CDC-based ETL from the SNPD SQL database, SAP (cost/PO/BOM/vendor — masked at API layer), and Teamcenter PLM (BOM snapshots, part lifecycle, ECN).
• Design and implement the unified SNPD data model: enforce project_id + part_number as universal pivot keys across all tables; ensure referential integrity across SNPD core domain tables and migrated portal tables (NVPC, RFQ, PPAP, CDMM, etc.).
• Build the AI feature store on Delta Lake: dl_gate_cycle_times, dl_supplier_risk, dl_nvpc_benchmarks, dl_cost_variance, dl_deliverable_actuals — with incremental refresh, partitioning, and Z-ordering for query performance.
• Conduct data audits on SAP S/4 HANA, Teamcenter PLM, and all 8 legacy homegrown portal databases; assess data quality, identify gaps, and remediate for ML readiness.
• Implement SAP cost data masking at the API / pipeline layer — sensitive pricing data must be obfuscated before reaching any MCP server or AI agent.
• Set up Azure AI Search vector index: embedding ingestion pipeline from the document store (SharePoint / Azure Blob), chunking strategy, metadata schema, and incremental re-indexing on document updates.
• Establish data lineage, quality checks, and observability: row counts, null rates, schema drift alerts, and SLA monitoring for all ETL pipelines.
• Support historical data migration: 5–7 years of legacy SNPD and portal data into the unified SNPD database; validate referential integrity and completeness post-migration.
• Collaborate with the ML Engineer to serve training datasets from Delta Lake; optimize feature computation using Synapse Serverless or Databricks as compute.
• Implement RBAC and data access controls at the data layer: ensure user-level and role-level scoping is enforced from Azure SQL through to Delta Lake reads and vector search results.
• Maintain data catalogue and schema documentation; ensure all entities conform to the IATF 16949 audit traceability requirements.
TECHNICAL SKILLS REQUIRED
| • Azure Data Factory (ADF) — pipeline authoring | • Azure AI Search — index management, embedding pipelines |
| • ADLS Gen2 / Delta Lake — storage & compute | • Data modelling — relational + lakehouse schemas |
| • Azure SQL / SQL Server — T-SQL, stored procedures | • Azure Key Vault — secrets, connection string management |
| • Azure Synapse Analytics or Databricks | • Data lineage & observability tools |
| • Python — PySpark, pandas, data quality scripts | • Git / Azure DevOps for pipeline version control |
| • SAP OData / RFC / BAPI integration patterns | • CDC (Change Data Capture) patterns in SQL Server |
GOOD TO HAVE
• Hands-on SAP S/4 HANA RISE data extraction experience (ACDOCA, Material Master, BOM, MM60).
• Teamcenter PLM API familiarity (REST/SOA Gateway, BOM export, ECN feeds).
• dbt (data build tool) for transformation layer on Delta Lake.
• Apache Kafka / Azure Event Hubs for real-time streaming from SAP change events.
• Experience with IATF 16949 or automotive quality data requirements.
• Familiarity with Mahindra data platform (MDP) or Azure Purview for data governance.