About this role
The team is seeking a skilled Data Engineer to join their dynamic team. The successful candidate will be responsible for designing, developing, and maintaining scalable ETL/ELT data pipelines using Python and PySpark. This role involves working with large volumes of structured and semi-structured data to ensure efficient data processing and analysis.
Key Responsibilities:
- Design and implement scalable ETL/ELT data pipelines.
- Develop PySpark applications for data processing and transformation.
- Write efficient Python scripts and reusable functions for data validation and automation.
- Create complex SQL queries for data extraction, transformation, and analysis.
- Collaborate with cross-functional teams to understand data requirements and deliver solutions.
- Work with Apache Spark, Spark SQL, DataFrames, Hive, and Hadoop for distributed data processing.
Required Skills & Qualifications:
- Strong experience with Python and PySpark.
- Proficiency in SQL and experience with complex queries.
- Familiarity with data processing frameworks like Apache Spark and Hadoop.
- Ability to work with structured and semi-structured data.
- Excellent problem-solving skills and attention to detail.
Experience:
- 5-8 years of experience in data engineering or a related field.
What we offer:
- Opportunity to work in a collaborative and innovative environment.
- Chance to work with cutting-edge technologies and tools.
- Professional development and growth opportunities.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.