Our team is seeking an experienced Google Cloud Platform (GCP) Data Engineer with a strong background in building and optimising data pipelines, architectures, and data sets on GCP. This role involves close collaboration with data analysts, data scientists, and other stakeholders to ensure high-quality data availability for advanced analytics, machine learning, and real-time data processing.
Key Responsibilities:
Design, build, and maintain scalable ETL/ELT data pipelines on GCP to support analytics, reporting, and machine learning applications.
Implement and manage data warehouses and data lakes using GCP services such as BigQuery, Cloud Storage, Cloud Dataflow, and Cloud Composer (Airflow).
Develop and optimise data models to align with business needs, ensuring scalability and performance.
Collaborate with cross-functional teams to understand business requirements and deliver tailored data access solutions.
Ensure data quality, security, and governance through performance tuning, capacity management, and compliance best practices.
Implement CI/CD automation using GitHub Actions and Docker to streamline deployment of data pipelines.
Develop and maintain documentation for data architecture, models, and processes.
Support real-time data processing with tools such as Kafka Streaming and MLOps workflows for machine learning pipelines.
Qualifications:
Bachelor’s Degree in Computer Science, Information Systems, Engineering, or a related field.
4+ years of experience in data engineering, with at least 2 years focused on GCP.
Strong experience with SQL and ETL/ELT processes for transforming and processing large-scale structured and unstructured data.
Solid understanding of data warehousing concepts and dimensional modelling techniques.
Hands-on experience with BigQuery or Snowflake, as well as orchestration tools like Airflow or Argo.
Experience working with real-time data processing using Kafka Streaming.
Knowledge of machine learning pipelines and MLOps practices is a plus.
Technical Skills:
Google Cloud Platform (GCP) – BigQuery, Cloud Dataflow, Cloud Storage, Cloud Composer.
SQL & DBT – Data modelling and transformation within Medallion Architecture.
Python – Development of data workflows, automation, and integration.
Kafka Streaming – Real-time data processing.
GitHub & CI/CD (GitHub Actions) – Version control and automated deployment.
Docker & Kubernetes (GKE) – Containerisation and orchestration.
Airflow (Cloud Composer) – Data pipeline orchestration.
Preferred Competencies:
Google Cloud Certification (e.g., Professional Data Engineer) is a plus.
Experience with data governance and data quality frameworks.
Strong problem-solving skills and ability to work in a fast-paced, collaborative environment.
Familiarity with machine learning workflows and MLOps integration.
* In order to comply with the POPI Act, for future career opportunities, we require your permission to maintain your personal details on our database. By completing and returning this form you give PBT your consent
* If you have not received any feedback after 2 weeks, please consider you application as unsuccessful.