Website IBM
Data Engineer β Data Platforms (GCP) β IBM Consulting
Opportunity Overview
GCP Data Engineering & Lakehouse Workflow
The GCP Data Engineer at IBM designs enterprise-grade batch and real-time streaming pipelines, migrating and optimizing client data assets across Google Cloud’s data ecosystem:
βββββββββββββββββββββββββββββββββββββββββββββ
β 1. Data Ingestion & Streaming β β Ingest streaming events via Pub/Sub & batch files via Google Cloud Storage (GCS)
βββββββββββββββββββββββ¬ββββββββββββββββββββββ
βΌ
βββββββββββββββββββββββββββββββββββββββββββββ
β 2. Pipeline Processing & ETL β β Transform data using Dataflow (Apache Beam) & Dataproc (Spark/Python/Scala)
βββββββββββββββββββββββ¬ββββββββββββββββββββββ
βΌ
βββββββββββββββββββββββββββββββββββββββββββββ
β 3. Data Layer Optimization & Serving β β Store & query structured/NoSQL data across BigQuery, Bigtable, Spanner & AlloyDB
βββββββββββββββββββββββ¬ββββββββββββββββββββββ
βΌ
βββββββββββββββββββββββββββββββββββββββββββββ
β 4. Workflow Orchestration & Data Ops β β Schedule pipelines via Cloud Composer (Airflow), run dbt transformations & manage ops
βββββββββββββββββββββββββββββββββββββββββββββ
Key Responsibilities
-
GCP Pipeline Architecture: Design and deploy end-to-end batch and streaming pipelines for Data Warehouses and Data Lakes using Google Cloud native tools (Dataflow, Dataproc, Pub/Sub).
-
Data Processing & Scripting: Write scalable data transformation code using Apache Beam (Python/Java) on Dataflow, PySpark/Scala on Dataproc, and SQL transformations using dbt.
-
Orchestration & Workflow Scheduling: Automate, schedule, and monitor data platform workflows using Cloud Composer (managed Apache Airflow) and Google Cloud Scheduler.
-
Cloud Storage & Database Layer Design: Optimize multi-modal data storage systems utilizing BigQuery for analytics, Bigtable for NoSQL, Cloud Spanner for global transactional consistency, and AlloyDB/CloudSQL.
-
Enterprise Migration: Execute legacy-to-cloud data migration strategies, ensuring data integrity, schema conversion, and seamless cutover between systems.
Qualification Matrix & Technical Skill Stack
Core Requirements
Key Focus Areas for Interview Preparation
-
Google Dataflow vs Dataproc: Be prepared to compare when to choose Apache Beam (Dataflow) vs Spark (Dataproc) based on batch vs streaming, autoscaling requirements, and unified pipeline architectures.
-
BigQuery Performance & Storage Optimization: Practice scenario questions covering BigQuery partitioning, clustering, materialization strategies, slot allocation, and cost management.
-
Cloud Composer DAG Design: Review building resilient Apache Airflow DAGs, custom operators, error handling, and XCom mechanics within GCP environments.
To apply for this job please visit remotejobhiring.com.
