Senior Technical Lead
Technology, Data & Digital · Data, AI & Analytics · Data Engineering
In short
We are seeking an experienced ETL Data Engineer to design, develop, and optimize scalable data pipelines supporting analytics and reporting. The role involves working with Python, PySpark, Apache Spark, and Apache Airflow to build robust, high-performance ETL solutions for large-scale datasets.
Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Python and PySpark.
- Build and optimize distributed data processing applications using Apache Spark.
- Develop data ingestion frameworks for structured and semi-structured data sources.
- Transform, cleanse, validate, and enrich data to meet business requirements.
- Implement data quality checks and monitoring mechanisms across data pipelines.
- Design and manage workflow orchestration using Apache Airflow.
- Create, schedule, monitor, and troubleshoot DAGs for reliable data processing.
- Ensure pipeline resiliency through automated recovery and alerting mechanisms.
- Work with columnar storage formats such as Parquet and Avro.
- Optimize data partitioning, compression, and storage strategies.
- Improve query performance and processing efficiency for large datasets.
- Tune Spark jobs for performance and resource utilization.
- Analyze bottlenecks and optimize distributed processing workloads.
- Ensure scalability, availability, and reliability of data platforms.
- Collaborate with Data Architects and Business Analysts to translate requirements into technical solutions.
- Follow data governance, security, and compliance standards.
- Document technical designs, workflows, and operational procedures.
Requirements
- Strong programming experience in Python.
- Hands-on experience with PySpark and Apache Spark.
- Expertise in ETL/ELT design and implementation.
- Experience with Apache Airflow for workflow orchestration.
- Knowledge of Avro and Parquet file formats.
- Strong understanding of distributed data processing concepts.
- Experience working with SQL and relational databases.
- Familiarity with data modeling concepts and best practices.
- Experience with Git and CI/CD practices.
Desired Qualifications
- Bachelor’s or master’s degree in computer science, Information Technology, Engineering, or a related field.
- 4+ years of experience in Data Engineering or ETL development.
- Proven experience developing enterprise-scale data pipelines and data platforms.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Knowledge of Delta Lake, Iceberg, or Hudi.
- Experience with Kafka or other streaming technologies.
- Familiarity with containerization technologies such as Docker and Kubernetes.
- Understanding of DataOps and MLOps practices.
- Databricks Certified Data Engineer
- Azure Data Engineer Associate (DP-203)
- AWS Certified Data Analytics
- Google Professional Data Engineer
Benefits
- Global technology company with over 223,000 people across 60 countries.
- Industry-leading capabilities centered around digital, engineering, cloud, and AI.
- Work with clients across all major verticals including Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services.
#ETL#Data Engineering#Data Pipelines#Python#PySpark#Apache Spark#Apache Airflow#Avro#Parquet#SQL#Cloud#Big Data