• Full Time
  • Chennai

Website Zorba Ai

Data Engineer (PySpark, Databricks, AWS)

Opportunity Overview

Attribute Details
Position Title Data Engineer – PySpark, Databricks, AWS
Experience Requirement 3 – 8 Years in Data Engineering
Primary Location Chennai, Tamil Nadu, India (Open to Remote / Flexible)
Notice Period Immediate to 30 Days Preferred
Employment Type Full-Time, Permanent
Core Technical Stack PySpark, Databricks, Python, SQL, AWS (S3, Glue, EMR, Lambda), Delta Lake
Preferred Methodologies Git, CI/CD, Agile, Data Warehousing, Workflow Orchestration

Cloud Data Pipeline & Processing Architecture

The Data Engineer designs, optimises, and orchestrates scalable batch and streaming data pipelines, transforming complex multi-source datasets within an AWS and Databricks cloud environment:

┌───────────────────────────────────────────┐
│ 1. Multi-Source Ingestion & AWS Storage   │ ➔ Ingest raw datasets into AWS S3 using AWS Glue, Lambda, or automated feeds
└─────────────────────┬─────────────────────┘
                      ▼
┌───────────────────────────────────────────┐
│ 2. PySpark & Databricks Transformation    │ ➔ Execute distributed PySpark scripts and reusable Python frameworks on Databricks
└─────────────────────┬─────────────────────┘
                      ▼
┌───────────────────────────────────────────┐
│ 3. Delta Lake Storage & Data Quality      │ ➔ Store processed assets in Delta Lake tables with ACID compliance and quality validation
└─────────────────────┬─────────────────────┘
                      ▼
┌───────────────────────────────────────────┐
│ 4. Orchestration, CI/CD & Delivery        │ ➔ Automate pipelines via orchestration tools and manage code releases using Git & CI/CD
└───────────────────────────────────────────┘

Key Responsibilities

  • PySpark & Databricks ETL Development: Build, maintain, and optimize scalable data pipelines using PySpark, Databricks, and Python.

  • AWS Cloud Integration: Leverage AWS native data services—including S3, Glue, EMR, and Lambda—for serverless and cluster-based processing.

  • Delta Lake Management & Optimization: Enforce Delta Lake table formats, implement Z-Ordering/partitioning, and tune Spark performance for efficient querying.

  • Reusable Framework Design: Develop modular, object-oriented Python scripts and custom SQL query frameworks to standardize pipeline deployments.

  • DataOps & Quality Standards: Integrate automated validation procedures, continuous integration/continuous deployment (CI/CD) workflows, and version control using Git.

Qualification Matrix & Technical Skill Stack

Core Requirements

Category Specifications
Professional Experience 3 to 8 years of hands-on experience in enterprise data engineering.
Primary Spark Engine Expertise in PySpark, Databricks Platform, and Spark SQL.
AWS Cloud Stack Strong proficiency in AWS S3, AWS Glue, AWS EMR, and AWS Lambda.
Programming & Data Advanced Python scripting, complex SQL, and data modeling concepts.
Storage & Delta Lake Hands-on experience with Delta Lake storage layer, ACID transactions, and Time Travel.
Development Practices Familiarity with Git, CI/CD automation pipelines, Airflow/Databricks Workflows, and Agile.

Key Focus Areas for Interview Preparation

  1. Spark Engine Tuning & Performance: Practice explaining memory management, broadcast join conditions, data skew handling, and cluster configuration on AWS/Databricks.

  2. AWS Data Architecture: Be ready to design serverless vs cluster-based ingestion flows comparing AWS Glue, Databricks, Lambda, and S3 event triggers.

  3. Delta Lake & Data Warehousing: Review Delta Lake log mechanics, schema evolution, dynamic partitioning strategies, and dimensional modeling patterns.

To apply for this job please visit remotejobhiring.com.