A
IT(Software/Hardware)

Site Reliability Engineer (Cloud & Platform Reliability)

Alter Domus

Role Summary

Alter Domus is hiring a Site Reliability Engineer (Cloud & Platform Reliability) for its technology operations in Hyderabad. Operating on a hybrid model aligned with US West Coast hours (PST/PDT), you will drive reliability, observability, and infrastructure automation across EKS and AKS Kubernetes clusters on AWS and Azure. You will build OpenTelemetry pipelines, enforce SLI/SLO error budgets, manage Terraform infrastructure, automate toil reduction scripts, and participate in a follow-the-sun on-call rotation using PagerDuty.

Key Responsibilities

  • Observability & Reliability Infrastructure: Construct and maintain OpenTelemetry collectors, Prometheus metrics, and Grafana dashboards; define SLIs, SLOs, and error budgets for platform services.
  • Cloud & Kubernetes Operations: Manage, auto-scale (Karpenter, KEDA), and troubleshoot AWS EKS and Azure AKS clusters, cluster networking (CNI, ingress), and Terraform infrastructure as code.
  • Incident Response & On-Call: Participate in a follow-the-sun on-call rotation (including periodic weekend coverage), act as incident commander during outages, and drive blameless postmortems.
  • GitOps & Automation: Write Python, Go, or Bash scripts to eliminate operational toil and maintain GitOps-based deployment workflows using Helm, Flux, or ArgoCD.

Key Qualifications

  • Education: Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • Experience: 4+ years of professional experience in Site Reliability Engineering (SRE), DevOps, or Cloud Infrastructure Engineering with hands-on production ownership.
  • Mandatory Technical Stack: Deep hands-on experience with Kubernetes (EKS/AKS), AWS/Azure cloud, Terraform, OpenTelemetry, Prometheus, Grafana, PagerDuty, and Linux shell automation (Python/Bash/Go).
  • Preferred Technical Exposure: Experience with GitOps tools (ArgoCD/Flux/Helm), policy-as-code (Kyverno/OPA), CI/CD pipelines (Azure DevOps/GitHub Actions), and regulated financial services compliance (SOC 2).
  • Core Soft Skills: Calm under pressure during critical outages, clear technical documentation skills, and readiness for PST-aligned working hours.

About Alter Domus

Alter Domus is a global leader in integrated tech-enabled solutions for the alternative investment industry, supporting 90% of the top 30 private market asset managers worldwide.

Important Candidate Notice

This job posting and its descriptions are owned by Alter Domus . We provide this overview to help developers discover opportunities. For applications, responses, or follow-ups, please connect with Alter Domus directly. Job Central Hub does not process applications or review candidate queries.

Profile Match Score

Sign in to see how well your skills and experience align with this job.

Sign In
Overview
Expected Salary ₹14 - ₹28 Lakhs/Annum
Experience Required 4 - 8 Years
Last Date to Apply Continuous Recruitment
Posted On 31-08-2026
Mandatory Skills
Site Reliability Engineering (SRE) Kubernetes (EKS / AKS) AWS Azure Terraform OpenTelemetry Prometheus Grafana PagerDuty Python / Bash / Go
Preferred Skills
GitOps Karpenter KEDA Kyverno / OPA Azure DevOps GitHub Actions