Job Description & Scope
Job description Job Description Job Summary We are seeking a Senior IAC Engineer to architect, develop, and automate DA GCP PAAS services and Databricks platform provisioning using Terraform, Spacelift, and GitHub. This role combines the depth of platform engineering with the principles of reliability engineering, enabling resilient, secure, and scalable cloud environments. The ideal candidate has 6+ years of hands-on experience with IaC , CI/CD, infrastructure automation, and driving cloud infrastructure reliability. Key Responsibilities Infrastructure Automation Design, implement, and manage modular, reusable Terraform modules to provision GCP resources (BigQuery, GCS, VPC, IAM, Pub/Sub, Composer, etc.). Automate provisioning of Databricks workspaces, clusters, jobs, service principals, and permissions using Terraform. Build and maintain CI/CD pipelines for infrastructure deployment and compliance using GitHub Actions and Spacelift. Standardize and enforce GitOps workflows for infrastructure changes, including code reviews and testing. Integrate infrastructure cost control, policy-as-code, and secrets management into automation pipelines. Architecture Reliability Lead the design of scalable and highly reliable infrastructure patterns across GCP and Databricks. Implement resiliency and fault-tolerant designs, backup/recovery mechanisms, and automated alerting around infrastructure components. Partner with SRE and DevOps teams to enable observability, performance monitoring, and automated incident response tooling. Develop proactive monitoring and drift detection for Terraform-managed resources. Contribute to reliability reviews, runbooks, and disaster recovery strategies for cloud resources. Collaboration Governance Work closely with security, networking, FinOps, and platform teams to ensure compliance, cost-efficiency, and best practices. Define Terraform standards, module registries, and access patterns for scalable infrastructure usage. Provide mentorship, peer code reviews, and knowledge sharing across engineering teams. Required Skills Experience 6+ years of experience with Terraform and Infrastructure as Code (IaC), with deep expertise in GCP provisioning. Experience in automating Databricks (clusters, jobs, users, ACLs) using Terraform. Strong hands-on experience with Spacelift (or similar tools like Terraform Cloud or Atlantis) and GitHub CI/CD workflows. Deep understanding of infrastructure reliability principles: HA, fault tolerance, rollback strategies, and zero-downtime deployments. Familiar with monitoring/logging frameworks (Cloud Monitoring, Stackdriver, Datadog, etc.). Strong scripting and debugging skills to troubleshoot infrastructure or CI/CD failures. Proficient with GCP networking, IAM policies, folder/project structure, and Org Policy configuration. Nice to Have HashiCorp Certified: Terraform Associate or Architect. Familiarity with SRE principles (SLOs, error budgets, alerting). Exposure to FinOps strategies: cost controls, tagging policies, budget alerts. Experience with container orchestration (GKE/Kubernetes), Cloud Composer is a plus No Relocation support available Business Unit Summary Job Type Regular Analytics Modelling Analytics Data Science Role: Site Reliability Engineer Industry Type: FMCG Department: Engineering - Software & QA Employment Type: Full Time, Permanent Role Category: DevOps Education UG: Any Graduate PG: Any Postgraduate Key Skills AutomationUsageorchestrationNetworkingGCPDebuggingPackagingInfrastructureMonitoringAnalytics