Lead Site Reliability Engineer
About the Company
Our client's purpose is helping people live longer, healthier, happier lives and making a better world. As a global healthcare company, we combine innovation, technology, and customer-centric thinking to deliver exceptional experiences for our customers and colleagues.
Roles and Responsibilites
As a Lead Site Reliability Engineer (Hyperscale Cloud), you will be responsible for ensuring the stability, reliability, scalability, and performance of mission-critical services deployed across Microsoft Azure and Google Cloud Platform (GCP).
You will lead a team of SRE professionals while partnering closely with Engineering, DevOps, Platform, and Product teams to drive cloud-native best practices, automate operations, and improve service reliability.
Key Responsibilities
Lead and mentor the Site Reliability Engineering team, fostering a culture of reliability, automation, and continuous improvement.
Own the operational health, uptime, and performance of production and non-production cloud environments.
Drive adoption of SRE principles including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
Act as a Developer Advocate, promoting cloud-native architectures and standardized OneSDLC practices across engineering teams.
Design and implement proactive monitoring, observability, alerting, and dashboards for cloud services and platforms.
Develop and maintain smoke tests for APIs and critical services to identify and resolve issues before they impact customers.
Partner with development and DevOps teams to ensure environment readiness aligns with CI/CD pipelines and release schedules.
Lead incident response, troubleshooting, root cause analysis, and post-incident remediation activities.
Implement automation for environment validation, self-healing, recovery, and operational efficiencies.
Optimize cloud infrastructure for reliability, scalability, performance, and cost effectiveness.
Maintain architecture documentation, operational runbooks, monitoring standards, and support procedures.
Support risk management initiatives through secure-by-design cloud engineering practices.
Participate in on-call support rotations to ensure 24/7 service availability and resilience.
Skills and qualifications
Required Experience
Extensive hands-on experience with Microsoft Azure and Google Cloud Platform (GCP).
Strong expertise in Site Reliability Engineering, cloud operations, and large-scale production environments.
Proven experience leading technical teams and influencing engineering best practices.
Experience designing and operating highly available, resilient cloud-native architectures.
Strong background in incident management, troubleshooting, and root cause analysis.
Technical Skills
Infrastructure as Code (IaC) using Terraform Enterprise.
Scripting and automation using Python, PowerShell, or Bash.
Deep understanding of cloud networking including DNS, VPNs, load balancing, firewalls, and connectivity patterns.
Experience with containerization and orchestration technologies such as
Kubernetes
Azure Kubernetes Service (AKS)
Google Kubernetes Engine (GKE)
Azure Container Apps (ACA)
Expertise with observability and monitoring tools such as
Azure Monitor
Grafana
Logging and alerting platforms
Strong understanding of CI/CD tooling including:
GitHub Actions
Azure DevOps
Preferred Qualifications
Experience working within enterprise, healthcare, financial services, or other highly regulated industries.
Knowledge of security, risk management, governance, and compliance practices in cloud environments.
Excellent stakeholder management, communication, and collaboration skills.
What's On Offer
Opportunity to lead reliability engineering for large-scale cloud platforms supporting critical customer services.
Exposure to cutting-edge technologies across Azure, GCP, Kubernetes, observability, and cloud automation.
Collaborative, innovative, and inclusive work environment.
Career growth opportunities within a global organization.
Opportunity to influence cloud strategy, engineering standards, and platform modernization initiatives.
Flexible and supportive culture focused on wellbeing and professional development.
Competitive compensation and comprehensive benefits package.
Work with high-performing teams committed to delivering exceptional customer outcomes.
Lagrange Point For Job Seekers

Haven’t found your desired role?
Join us to discover opportunities that match your skills, aspirations, and career goals.
For Talent.
For Opportunity.
For Growth.
Why someone should join us?










