Data Engineer
- Blue Pearl
- Johannesburg, South Africa
- ZAR 600,000 – ZAR 900,000
Key Responsibility:
You will be assigned a portfolio of client engagements where you will be expected to:
•
Design and build scalable data platforms using modern cloud-native and Lakehouse architectures
•
Develop and optimise data pipelines using Python, SQL, and tools such as Azure Data Factory, AWS Glue, Google Cloud Dataflow, Databricks, and dbt
•
Modernise legacy data environments, migrating from on-premises solutions to cloud-native platforms such as Microsoft Fabric, Azure Synapse Analytics, AWS Redshift, Google BigQuery, or Databricks
•
Engage with clients to conceptualize data solutions aligned to their business strategy
•
Support our sales team with pre-sales activities, proof-of-concept deliveries, and technical proposals
•
Provide technical guidance and mentorship to junior and intermediate consultants
•
Lead technical reviews and contribute to consultants' growth plans
•
Identify opportunities to automate manual processes, optimise data delivery, and improve infrastructure scalability
•
Work with stakeholders, including executive, product, and analytics teams, to address data infrastructure needs
•
Drive knowledge sharing through technical blogs, internal forums, and workshops
•
Balance billable project work with team support responsibilities
Requirements
Data Engineer – Candidate Requirements
Intermediate Level
3–5 years' experience
3–5 years of hands-on experience in data engineering.
Strong proficiency in Python and/or SQL , including query optimisation.
Experience working with both relational and non-relational databases.
Experience designing and building data pipelines and data models.
Understanding and practical experience with lakehouse architectures , including the medallion pattern.
Practical experience with at least one major cloud platform, including:
- Microsoft Azure
- AWS
- Google Cloud Platform (GCP)
Familiarity with:
- Databricks
- Snowflake
- Delta Lake
- PySpark
Understanding of data transformation frameworks such as dbt .
Experience with version control using Git .
Understanding of CI/CD practices for data workflows.
Strong analytical and problem-solving skills.
Ability to perform root-cause analysis on complex data issues.
Good communication and stakeholder engagement skills.
Senior Level
6–8+ years' experience
6–8+ years of hands-on experience in data engineering.
All intermediate-level technical requirements, together with demonstrable experience in:
- Leading end-to-end data platform delivery.
- Architecting enterprise-grade lakehouse environments.
- Implementing data mesh patterns.
- Infrastructure-as-code using tools such as Terraform, Bicep, AWS CDK or Pulumi.
- DevOps and CI/CD pipelines.
- Working effectively with cross-functional teams in a dynamic consulting environment.
- Mentoring junior engineers.
- Contributing to technical strategy and solution direction.
Qualifications
Bachelor's degree in:
- Computer Science
- Information Systems
- Information Technology
- or a related field.
Master's degree in a relevant field is advantageous.
Certifications
One or more of the following certifications would be advantageous:
- Microsoft Fabric Data Engineer Associate
- Microsoft Azure Data Engineer Associate
- Databricks Certified Data Engineer Associate
- Google Professional Data Engineer
- AWS Certified Data Engineer – Associate
- Databricks Certified Data Engineer Professional
Technology Experience
Languages & Frameworks
- Python
- PySpark
- SQL
- dbt
Microsoft Fabric & Azure
- Microsoft Fabric Lakehouses
- Fabric Pipelines
- Fabric Semantic Models
- Direct Lake
- Azure Data Factory
- Azure Data Lake Storage Gen2
- Azure Synapse Analytics
- Azure Databricks
- Azure Event Hubs
Google Cloud Platform
- BigQuery
- Cloud Storage
- Dataflow
- Dataproc
- Pub/Sub
Amazon Web Services
- Amazon S3
- AWS Glue
- Amazon Redshift
- Amazon EMR
- Amazon Kinesis
Databricks & Data Platforms
- Databricks
- Delta Lake
- Unity Catalog
- MLflow
- Databricks Workflows
Databases
- Azure SQL
- Azure Cosmos DB
- PostgreSQL
- Snowflake
- BigQuery
- Amazon Redshift
DevOps & Infrastructure as Code
- Git
- Azure DevOps
- GitHub Actions
- Terraform
- Bicep
- AWS CDK
- CI/CD pipelines
Streaming & Messaging
- Azure Event Hubs
- Azure Stream Analytics
- Apache Kafka
- Amazon Kinesis
- Google Pub/Sub
Visualisation & Analytics
- Microsoft Power BI
- Microsoft Fabric Real-Time Dashboards
- Looker / Looker Studio
- Amazon QuickSight
Skills
- Python
- SQL
- Azure Data Factory
- Databricks
- dbt
- Cloud Data Warehousing
- Data Pipeline Optimization









