For Talent/Open roles
Team Lead/EM – SRE
Technology
Hyderabad
Technology
About the Company
- Our client is a fast-growing, product-led SaaS company building a next-generation finance operations platform for high-growth businesses. The platform helps finance teams automate processes, analyse business data, and make faster, more informed decisions.
- The company operates in a high-ownership, technology-driven environment, solving complex data and infrastructure challenges at scale. With an expanding enterprise customer base and multiple deployment environments, the organisation is focused on building highly reliable, secure, and scalable technology.
- This is an opportunity to join a growing engineering organisation where you will work closely with senior stakeholders, enterprise customers, and cross-functional teams while helping shape the systems and processes that support the company's next phase of growth.
Roles and Responsibilities
- Lead, mentor, and manage a team of SRE and DevOps engineers, including planning, prioritization, and performance management.
- Define operational standards, runbooks, on-call practices, and escalation processes.
- Own enterprise customer onboarding across customer-hosted and managed environments.
- Gather infrastructure, security, and deployment requirements and coordinate with customer IT, security, and infrastructure teams.
- Drive provisioning, deployment, configuration, access, and approvals end-to-end.
- Own enterprise application, product, and platform support, including triage, resolution, escalation, and SLA management.
- Design and implement monitoring, alerting, and observability across customer environments.
- Build dashboards and alerts covering infrastructure, application, and data-pipeline health.
- Own the complete incident management lifecycle, from detection and response through resolution and communication.
- Lead Root Cause Analyses (RCAs) and post-incident reviews, ensuring corrective actions are closed.
- Establish and manage on-call rotations, severity frameworks, and escalation paths.
- Track and improve reliability metrics including uptime, MTTR, and incident trends.
Skills and qualifications
- 7–10 years of experience in SRE, DevOps, or infrastructure engineering, with at least 2 years of team leadership experience.
- Strong hands-on experience with Kubernetes and containerized production environments.
- Hands-on experience with at least one major cloud platform: AWS, GCP, or Azure.
- Strong experience with monitoring and observability tools such as Prometheus, Grafana, ELK, Datadog, or similar.
- Proven experience in incident management, RCA, and post-incident improvements.
- Experience deploying and supporting software in customer-controlled enterprise environments.
- Experience with Infrastructure as Code, particularly Terraform, Helm, or similar tools.
- Strong understanding of CI/CD pipelines and release management.
- Scripting proficiency in Python or Bash.
- Working knowledge of networking, security, IAM, VPNs, firewalls, and certificates.
- Strong stakeholder management and communication skills, with comfort working directly with enterprise customers.
What's On Offer
- Opportunity to lead a critical SRE & DevOps function within a fast-growing, product-led technology company.
- Own platform reliability, enterprise customer onboarding, monitoring, and incident management across multiple deployment environments.
- Work closely with enterprise customers and cross-functional engineering teams to solve complex production challenges.
- Build and scale operational processes, automation, and reliability practices while developing a strong SRE/DevOps team.
