HirePortal

Data Engineer - GCP/Databricks

  • AuxoAI
  • India
  • INR 3,000,000 – INR 5,000,000

AuxoAI is seeking a Senior Data Engineer to lead the design, development, and optimisation of modern data pipelines and cloud-native platforms. This role is ideal for someone with deep experience building scalable batch and streaming data workflows across cloud and lakehouse environments, strong hands-on engineering skills, and a drive to mentor junior engineers.

You will work closely with AI engineers, solution architects, and cross-functional teams to build production-grade pipelines spanning ingestion, transformation, and curated data delivery — enabling high-quality data for AI and analytics use cases at scale.

Location: Bangalore / Mumbai / Hyderabad / Gurgaon (Hybrid — 3 days in office)

Responsibilities

  • Design and build scalable batch and streaming data pipelines across bronze, silver, and gold medallion layers.
  • Build historical and incremental ingestion using Auto Loader/Spark Structured Streaming/Kafka feeds, with GCS/Azure/AWS storage and Databricks Jobs orchestration.
  • Develop and maintain Databricks-based pipelines using Spark and Delta Lake for lakehouse architecture, including migration of legacy or on-premises data sources.
  • Design and maintain analytical data layers in BigQuery or Databricks SQL, applying best practices in partitioning, clustering, and performance tuning.
  • Implement SQL/PySpark transformations for wide and semi-structured data, including wide-to-long processing and typed or hybrid models suited to consumer requirements.
  • Collaborate with AI engineers and data scientists to build pipelines that feed ML models, AI agents, and analytical systems.
  • Implement data governance, quality controls, and security best practices including schema enforcement, lineage tracking, and access controls.
  • Drive engineering best practices across CI/CD, testing, monitoring, and pipeline observability.
  • Partner with solution architects to translate data requirements into technical designs.
  • Mentor junior data engineers and contribute to documentation, code reviews, and agile ceremonies.

Requirements

  • 5+ years of hands-on experience in data engineering, building and operating production-grade pipelines.
  • Hands-on experience with Databricks on GCP, including BigQuery, GCS, Databricks, Spark, Delta Lake, and structured streaming.
  • Hands-on experience with Databricks and Apache Spark, including Delta Lake and end-to-end lakehouse implementations.
  • Strong programming skills in Python and/or Scala, with solid SQL for modelling and transformation.
  • Experience with data modelling, ETL/ELT, pipeline orchestration, and data warehousing concepts, including experience working with large, evolving JSON/map/array payloads, wide-to-long transformations, event-time context joins and schema-change handling.
  • Familiarity with Git, CI/CD pipelines, and data quality monitoring frameworks.
  • Solid understanding of data architecture, schema design, and performance tuning.
  • Experience with Unity Catalog, source reconciliation, schema evolution, correction handling and replay/recovery testing.
  • Strong problem-solving and collaboration skills.

Bonus Skills

  • GCP Professional Data Engineer certification.
  • Experience with Vertex AI, Cloud Functions, Dataproc, or real-time streaming architectures.
  • Experience with factory or industrial data sources — MES systems, IoT sensor streams, or operational telemetry.
  • Familiarity with data governance and cataloguing tools such as Dataplex, Unity Catalog, Atlan, or Collibra.
  • Exposure to Docker, Kubernetes, API integration, and infrastructure-as-code (Terraform).

Skills

  • Databricks
  • Apache Spark
  • PySpark
  • SQL
  • Google BigQuery
  • Delta Lake
  • Data Pipeline Orchestration

Related jobs

AuxoAIApply for this job