About this role
The Role
As an Azure Infrastructure Engineer, you will design, build, and operate resilient cloud infrastructure that supports Databricks workloads and engineering teams. You will work across Azure, Kubernetes, infrastructure as code, observability, networking, security, and automation to create scalable platforms that are reliable, efficient, and easy to operate. You will partner with software engineers, SREs, security teams, and product stakeholders to improve platform reliability, developer productivity, deployment velocity, and operational excellence. The role combines hands-on engineering with systems thinking and a strong ownership mindset.
The Impact You Will Have
- Design and operate secure, highly available infrastructure on Microsoft Azure.
- Build and manage Azure resources, networking, identity, storage, compute, and monitoring services.
- Develop reusable infrastructure-as-code modules using Terraform or similar tools.
- Manage and scale containerized workloads using Kubernetes, including Azure Kubernetes Service (AKS).
- Build reliable CI/CD workflows using GitHub Actions, Azure DevOps, or comparable platforms.
- Automate provisioning, configuration, deployments, upgrades, and operational runbooks.
- Improve observability through structured logging, metrics, tracing, alerting, and actionable dashboards.
- Establish practical standards for security, access control, secrets management, backup, disaster recovery, and compliance.
- Diagnose complex issues across infrastructure, applications, networking, data pipelines, and cloud services.
- Optimize plaƞorm performance, availability, scalability, and cost.
- Collaborate with engineering teams to improve developer experience and reduce operational friction.
- Contribute to architecture documentaƟon, technical standards, incident reviews, and reliability practices.
What We Look For
- 4+ years of experience in cloud infrastructure, site reliability engineering, platform engineering, DevOps, or a related role.
- Strong hands-on experience with Microsoft Azure and core services such as Azure Virtual Network, private endpoints, identity, storage, compute, Azure Monitor, and Log Analytics.
- Production experience with Kubernetes, Docker, and container orchestration; experience with AKS is strongly preferred.
- Advanced experience with Terraform, including reusable modules, state management, testing, and secure deployment practices.
- Strong scripting and software engineering skills in Python, Bash, or PowerShell.
- Experience designing and operating CI/CD pipelines with GitHub Actions, Azure DevOps, or Jenkins.
- Practical understanding of networking, DNS, load balancing, TLS, firewalls, IAM, secrets, and service-to-service authentication.
- Experience implementing observability using logs, metrics, traces, and tools such as Azure Monitor, Prometheus, Grafana, Datadog, or ELK.
- Experience with high-availability design, incident response, capacity planning, performance tuning, and disaster recovery.
- Ability to troubleshoot distributed systems and communicate technical issues clearly to varied audiences.
- Bachelor’s degree in computer science, engineering, or a related discipline—or equivalent practical experience.
Nice to Have
- Experience supporting Azure Databricks workspaces and associated cloud infrastructure.
- Familiarity with Databricks Terraform Provider, Declarative Automation Bundles, Unity Catalog, Lakeflow, or Databricks compute.
- Experience supporting data-intensive, machine learning, or AI workloads.
- Knowledge of Azure Policy, Microsoft Entra ID, Azure Key Vault, Private Link, VNets, and managed identities.
- Experience with GitOps, Helm, Argo CD, service meshes, or policy-as-code.
- Relevant Azure, Kubernetes, Terraform, or Databricks certifications.
- Experience participating in on-call rotations and leading incident or post-incident improvements.
