Job Description
Job Description
We are looking for an experienced Data Engineer with strong Python and Google Cloud Platform (GCP) expertise to design, develop, and maintain scalable data pipelines and cloud-based data solutions. The candidate will work closely with data analysts, data scientists, and business stakeholders to build reliable data platforms and enable data-driven decision-making.
\nKey Responsibilities
\nDesign, develop, and maintain scalable ETL/ELT data pipelines using Python and GCP services.
\nDevelop high-quality, reusable, and efficient Python-based data processing solutions.
\nBuild and manage data pipelines using Google Cloud Dataflow, Cloud Composer, BigQuery, and Cloud Storage.
\nDesign data models and optimize queries and workloads in Google BigQuery.
\nDevelop data ingestion processes from APIs, databases, files, and other data sources.
\nImplement data validation, quality checks, monitoring, and error-handling mechanisms.
\nOptimize data pipelines for performance, scalability, reliability, and cost.
\nWork with stakeholders to understand data requirements and translate them into technical solutions.
\nImplement CI/CD processes and follow DevOps best practices for data engineering solutions.
\nTroubleshoot production data pipeline issues and provide root-cause analysis.
\nEnsure data security, governance, access control, and compliance requirements are followed.
\nPrepare technical documentation for data pipelines, architecture, and operational processes.
\nRequired Skills
\nStrong hands-on experience with Python for data engineering and automation.
\nStrong experience with Google Cloud Platform (GCP).
\nHands-on experience with BigQuery and SQL.
\nExperience with Google Cloud Storage (GCS).
\nExperience with Cloud Composer / Apache Airflow.
\nExperience with Dataflow / Apache Beam.
\nStrong understanding of ETL/ELT concepts and data pipeline architecture.
\nGood knowledge of relational databases and SQL optimization.
\nExperience working with REST APIs and different data formats such as JSON, CSV, and Parquet.
\nExperience with Git and CI/CD tools.
\nUnderstanding of cloud security, IAM, and data governance.
\nPreferred Skills
\nExperience with Pub/Sub and event-driven data pipelines.
\nKnowledge of Cloud Functions / Cloud Run.
\nExperience with Dataproc / Spark.
\nKnowledge of data warehousing and dimensional modeling.
\nExperience with Terraform or other Infrastructure-as-Code tools.
\nExposure to data visualization tools such as Looker or Power BI.
\nExperience working in Agile/Scrum environments