Location: Hybrid / Remote
Employment Type: Full-Time
Industry: Data Engineering | AI | Software Development | Data Platforms
WatersEdge Solutions is partnering with a number of leading South African and international organisations to recruit an experienced Python/Spark/AI Developer to join a technically driven team working on the modernisation of large-scale data platforms.
This role will focus on replatforming legacy T-SQL workloads into modern Spark and Delta Lake pipelines, while building the APIs and AI capabilities that sit around them. You’ll work primarily in Python, with a strong emphasis on type-safe, tested and spec-driven development.
This is an excellent opportunity for an intermediate-to-senior developer who enjoys solving complex data engineering problems and wants to work across Python, distributed data processing, lakehouse architecture, APIs and emerging AI technologies.
As Python/Spark/AI Developer, you’ll play a key role in migrating legacy data workloads onto a modern Spark-based architecture.
You’ll translate existing T-SQL logic into Spark SQL and PySpark pipelines, validate migrated workloads against legacy outputs, and ensure the new platform delivers reliable and demonstrable parity.
The engineering environment is heavily spec-driven. You’ll work autonomously from written requirements, design documentation and Architecture Decision Records (ADRs), with changes expected to include appropriate testing and documentation rather than code alone.
Alongside the core data engineering work, you’ll build REST APIs around pipeline jobs and contribute to AI capabilities, including provider-neutral LLM integrations and structured AI outputs.
Build production-grade Spark and PySpark pipelines on Delta Lake.
Replatform legacy T-SQL workloads and stored-procedure-era logic into Spark SQL.
Translate existing data logic accurately while maintaining expected business and data semantics.
Write modern, type-hinted Python with strict static analysis.
Develop and maintain automated pytest suites covering unit, integration and end-to-end testing.
Validate migrated pipelines against legacy systems and provide clear parity and regression evidence.
Build REST APIs using FastAPI or equivalent frameworks.
Implement API request validation, authentication, idempotency and job-status semantics.
Work within Docker-based development environments using Compose stacks.
Follow GitHub flow and PR-driven development practices.
Maintain linting, type-checking and automated test gates within CI pipelines.
Work from design documents, technical specifications and ADRs.
Document technical changes and decisions as part of the development process.
Contribute to AI and LLM-enabled functionality around data workflows.
Work independently within a hybrid or remote engineering environment.
4+ years of professional Python development experience.
2+ years of production experience building Spark or comparable distributed data pipelines.
Strong modern Python development skills, ideally using Python 3.12.
Experience writing type-hinted Python that passes strict static analysis using tools such as pyright or mypy.
Experience with Pydantic models and structured data validation.
Understanding of ABC-based provider patterns.
Experience with modern Python packaging tools such as uv or Poetry.
Familiarity with Click or similar CLI frameworks.
Strong production experience with Apache Spark and PySpark.
Experience working with open-source Spark environments rather than solely managed vendor platforms.
Strong knowledge of the DataFrame API and Spark SQL.
Understanding of Spark partitioning, performance tuning and driver/executor architecture.
Familiarity with Spark Connect.
Experience with Delta Lake or an equivalent lakehouse table format such as Apache Iceberg or Hudi.
Understanding of MERGE INTO, schema evolution, time travel and idempotent write patterns.
Strong, dialect-portable SQL skills.
Ability to understand legacy T-SQL and faithfully reproduce its behaviour in Spark SQL.
Strong understanding of hashing, surrogate keys, deduplication and set-based data logic.
Strong automated testing experience using pytest, including fixtures, markers and tiered test suites.
Experience developing REST APIs using FastAPI or an equivalent framework.
Understanding of authentication mechanisms, signed requests, idempotency and job-status patterns.
Daily experience working within Docker-based development environments.
Strong Git and CI discipline, including PR-driven development, linting, type-checking and automated test gates.
Ability to work autonomously against technical specifications, requirements documents and ADRs.
Experience engineering LLM integrations and AI-enabled data workflows.
Experience building provider-neutral AI abstractions.
Knowledge of prompt construction and structured output validation.
Exposure to local or self-hosted inference using Ollama, vLLM, llama.cpp or LM Studio.
Experience integrating hosted LLM APIs.
Understanding of LLM evaluation and guardrails.
Data governance and privacy engineering experience.
Knowledge of PII tokenisation, hashing, k-anonymity and re-identification risk.
Familiarity with POPIA and GDPR-related data handling.
Multi-tenant data isolation awareness.
Experience with lakehouse technologies including Hive Metastore, Trino and Apache Ranger.
Understanding of catalogs, views and row/column-level data policies.
Familiarity with the Azure data ecosystem, including ADLS Gen2, Synapse, ADF, Azure Key Vault and Service Bus.
Observability experience using structured logging, OpenTelemetry or OpenLineage.
Previous experience migrating legacy pipelines and performing detailed output parity comparisons.
Databricks Certified Developer for Apache Spark or a comparable Spark certification.
Bachelor’s degree in Computer Science, Engineering or a related discipline, or equivalent practical experience.
Proven track record delivering production Python and Spark pipelines.
Spark certification or equivalent credential is advantageous but not required.
Flexible hybrid and remote working options.
Opportunity to work on a substantial legacy-to-modern data platform transformation.
Hands-on exposure to Python, Spark, Delta Lake, APIs and AI engineering.
Opportunity to work with modern lakehouse and distributed data technologies.
Wellness initiatives and home-office support.
Continuous learning and professional development opportunities.
Supportive and inclusive working environment.
Regular team activities and social events.
A culture centred around transparency, accountability and sustainable work-life balance.
Competitive remuneration and performance recognition.
Employee share ownership opportunities.
You’ll be joining a technically ambitious team that values quality engineering, autonomy and continuous improvement. The environment is collaborative but expects developers to take ownership of their work, understand the specifications behind what they are building and deliver code together with the tests and documentation needed to support it.
The culture combines technical rigour with flexibility, encouraging people to keep learning, explore better ways of solving problems and contribute meaningfully to the evolution of the platform.
Please Note: If you have not been contacted within 10 working days, consider your application unsuccessful.
You have successfully created your alert.
You will receive an email when a new job matching your criteria is posted.
Please check your email. It looks like you haven't verified your account yet. Here's what you're missing out on:
Didn't receive the link? Resend Verification Link