Start Your Search Here

Job Search

BASIL TECHNOLOGIES PTE. LTD.

Singapore / Global

Data Engineer (Databricks)

Job Description

Job Summary

The Data Engineer will design, develop, and maintain scalable, reliable data pipelines on Databricks and cloud platforms, integrating diverse data sources to support analytics, reporting, and machine learning. Collaborate with cross-functional teams to enhance data platform governance, monitoring, and reliability.

Responsibilities

Design, develop, and maintain ETL pipelines for centralized data storage systems such as Delta Lake to ensure efficient data ingestion and transformation

Integrate data from databases, APIs, log files, streaming platforms, and external providers to support diverse analytics needs

Develop data transformation routines to clean, normalize, and aggregate complex or inconsistent datasets for accurate analysis

Apply data processing techniques to handle large-scale and varied data sources effectively

Contribute to the development and enforcement of frameworks and best practices for code development and deployment to ensure quality and consistency

Implement data governance policies aligned with company standards to maintain data integrity and compliance

Collaborate with analytics and product teams to design and operationalize data pipelines that meet business requirements

Work with infrastructure teams to advance cloud-based data platforms leveraging Azure and Databricks technologies

Explore and evaluate new tools and techniques to enhance data platform capabilities on Azure, Databricks, or related platforms

Monitor data pipelines continuously to detect, troubleshoot, and resolve issues promptly, ensuring high availability

Develop monitoring tools, alerts, and automated error-handling mechanisms to maintain pipeline reliability

Analyze business requirements and translate them into data extraction and pipeline development tasks

Participate in requirement grooming and refinement sessions with users to clarify data needs

Optimize pipeline performance and batch scheduling to maximize efficiency and minimize latency

Develop dashboards, reports, scorecards, and data visualizations to support data-driven decision-making

Perform system integration testing (SIT), data validation, and profiling to confirm data accuracy and completeness

Validate completeness and consistency of ETL loads to ensure reliable data delivery

Support user acceptance testing (UAT) and production implementation activities to ensure smooth deployment

Required competencies and certifications

Databricks Certified Data Engineer Associate (strongly preferred)

Databricks Certified Data Engineer Professional (strongly preferred)

Preferred competencies and qualifications

3 or more years of experience in data engineering with scalable pipelines

Strong experience designing data solutions including data modelling and distributed computing architectures

Hands-on experience with data processing jobs using PySpark, Spark SQL, and Databricks notebooks/jobs

Experience orchestrating data pipelines with Azure Data Factory (ADF), Airflow, or similar tools

Experience with both real-time and batch data processing

Experience building pipelines on Azure; AWS experience is beneficial

Proficiency in SQL including window functions and performance optimization

Understanding of DevOps tools, Git workflows, and CI/CD pipelines

Familiarity with Scrum methodology and experience working in Scrum teams

Ability to apply Scrum practices in a practical project context

Strong problem-solving and collaborative mindset

Experience with streaming technologies such as Apache Kafka, Apache Flink, or AWS Kinesis

Ability to design and implement real-time data processing pipelines

#J-18808-Ljbffr

Apply Now

Similar Opportunities

View all jobs