• Full Time
  • Kochi

Website IBM

Data Engineer – Data Platforms (GCP) – IBM Consulting

Opportunity Overview

Attribute Details
Organisation IBM Consulting
Position Title Data Engineer – Data Platforms (Google Cloud Platform)
Primary Location Kochi, Kerala, India
Department / Practice Hybrid Cloud & AI Practice / IBM Consulting
Employment Type Full-Time, Permanent
Target Sector IT Services & Cloud Consulting
Academic Target Bachelor’s Degree required (Master’s Degree Preferred)
Core Technical Stack GCP Ecosystem (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Cloud Spanner), Apache Beam, Cloud Composer (Airflow), PySpark / Scala, dbt

GCP Data Engineering & Lakehouse Workflow

The GCP Data Engineer at IBM designs enterprise-grade batch and real-time streaming pipelines, migrating and optimizing client data assets across Google Cloud’s data ecosystem:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 1. Data Ingestion & Streaming             β”‚ βž” Ingest streaming events via Pub/Sub & batch files via Google Cloud Storage (GCS)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 2. Pipeline Processing & ETL              β”‚ βž” Transform data using Dataflow (Apache Beam) & Dataproc (Spark/Python/Scala)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 3. Data Layer Optimization & Serving      β”‚ βž” Store & query structured/NoSQL data across BigQuery, Bigtable, Spanner & AlloyDB
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 4. Workflow Orchestration & Data Ops      β”‚ βž” Schedule pipelines via Cloud Composer (Airflow), run dbt transformations & manage ops
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Key Responsibilities

  • GCP Pipeline Architecture: Design and deploy end-to-end batch and streaming pipelines for Data Warehouses and Data Lakes using Google Cloud native tools (Dataflow, Dataproc, Pub/Sub).

  • Data Processing & Scripting: Write scalable data transformation code using Apache Beam (Python/Java) on Dataflow, PySpark/Scala on Dataproc, and SQL transformations using dbt.

  • Orchestration & Workflow Scheduling: Automate, schedule, and monitor data platform workflows using Cloud Composer (managed Apache Airflow) and Google Cloud Scheduler.

  • Cloud Storage & Database Layer Design: Optimize multi-modal data storage systems utilizing BigQuery for analytics, Bigtable for NoSQL, Cloud Spanner for global transactional consistency, and AlloyDB/CloudSQL.

  • Enterprise Migration: Execute legacy-to-cloud data migration strategies, ensuring data integrity, schema conversion, and seamless cutover between systems.

Qualification Matrix & Technical Skill Stack

Core Requirements

Category Specifications
Educational Background Bachelor’s degree in Computer Science, IT, or quantitative fields (Master’s preferred).
Primary Cloud Engine Deep expertise in Google Cloud Platform (GCP) data services (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable).
Data Processing & Code Proficiency in Python, Scala, SQL, and frameworks like Apache Beam and Spark.
Workflow Orchestration Practical experience managing ETL pipelines with Cloud Composer (Apache Airflow) and dbt.
Database Paradigms Hands-on experience across Data Warehouse (BigQuery), NoSQL (Bigtable), and Relational (Cloud Spanner, AlloyDB, CloudSQL) layers.

Key Focus Areas for Interview Preparation

  1. Google Dataflow vs Dataproc: Be prepared to compare when to choose Apache Beam (Dataflow) vs Spark (Dataproc) based on batch vs streaming, autoscaling requirements, and unified pipeline architectures.

  2. BigQuery Performance & Storage Optimization: Practice scenario questions covering BigQuery partitioning, clustering, materialization strategies, slot allocation, and cost management.

  3. Cloud Composer DAG Design: Review building resilient Apache Airflow DAGs, custom operators, error handling, and XCom mechanics within GCP environments.

To apply for this job please visit remotejobhiring.com.