Responsibilities
- Design, construct, and sustain resilient ETL/ELT processes for data ingestion, transformation, and delivery across a contemporary Data Lake setup
- Create and enhance distributed data processing workflows utilizing Python and PySpark to manage extensive datasets effectively
- Apply and improve partitioning approaches for data lake storage systems like Delta Lake or Apache Iceberg to optimize query speed and storage expenses
- Compose, refine, and convert intricate SQL queries incorporating CTEs, window functions, conditional expressions, and aggregations
- Transfer and update data pipelines from older RDBMS systems to cloud-based analytics platforms
- Utilize object-oriented programming concepts to add to internal libraries for code reuse and standardization
- Operate proficiently with AWS services such as Glue (Jobs, Catalog, Triggers, Workflows), Athena, Redshift, S3, Lambda, EventBridge, and associated data tools
- Cooperate with infrastructure and DevOps groups to provision and oversee data resources via Infrastructure as Code (IaC) utilities like CloudFormation, CDK, or Terraform
- Oversee data pipeline status and efficiency using CloudWatch and other monitoring instruments, preemptively resolving problems and boosting dependability
- Guarantee data accuracy, uniformity, and adherence across pipelines and storage tiers
- Set up metric tracking and observability systems to offer clarity into data workflows and service level agreements
- Aid in data catalog administration and metadata governance procedures
- Collaborate with data analysts, scientists, and business partners to grasp needs and convert them into expandable technical answers
- Add to technical guides, code evaluations, and information exchange among the team
- Keep abreast of new data engineering methods, instruments, and cloud-native advancements
Requirements
- Substantial background in ETL processes and data pipeline creation with AWS
- High competence in Python as the main programming language, with proven skill in writing and refining PySpark code for distributed data handling
- Comprehensive knowledge of SQL, including complex queries (CTEs, window functions, aggregations, conditional expressions) and background in converting workloads from legacy RDBMS systems
- Practical involvement with AWS Glue (Jobs, Catalog, Triggers, Workflows), Athena, and Redshift
- Firm grasp of Data Lake structures and partitioning tactics to enhance performance and cost
- Adequate understanding of object-oriented programming (OOP) concepts and experience with reusable code libraries
- Proficient use of Git, Shell scripts, and Linux settings
- Acquaintance with observability, monitoring, and metric tracking methods
- Advanced or fluent English proficiency
Nice to Have
- Background with open table formats like Delta Lake or Apache Iceberg
- Understanding of TypeScript
- Hands-on experience with Infrastructure as Code (IaC) tools such as CloudFormation (preferred), CDK, or Terraform
- Familiarity with extra AWS services including SageMaker AI, ECS, RDS, DynamoDB, IAM, or EventBridge
- Experience employing Pandas for data manipulation and analysis
- Exposure to machine learning workflows or AI-focused data projects
Company
CI&T