We are seeking a Lead Site Reliability Engineer to embed with the team and drive reliable operations across a broad infrastructure landscape. You will own monitoring, observability, and logging with Dynatrace and Splunk, and help evolve dashboards, alerting, and analysis practices. You will also join the on-call rotation and partner on incident response. Apply now.
Responsibilities
-
Monitor and sustain the health, performance, and reliability of the client's applications and services
-
Operate and continuously enhance Dynatrace and Splunk, including dashboards, alerting, anomaly detection, and log analysis
-
Analyze alerts, determine likely root causes, and deliver actionable recommendations to engineering and incident management teams
-
Support incident response activities and participate in the on-call rotation
-
Perform Root Cause Analysis (RCA) and drive measurable post-incident improvements
-
Define and monitor SLOs, SLAs, and error budgets
-
Detect and remediate observability and monitoring gaps across services
-
Maintain operational runbooks and supporting documentation
Requirements
-
Proven background with 5+ years in Site Reliability Engineering, Production Operations, DevOps, or a closely related role
-
Deep expertise in Dynatrace and Splunk, including APM, alerting, dashboards, RUM, synthetic monitoring, service flow analysis, SPL queries, and log analysis
-
Hands-on experience with production incident management and on-call support, covering alert triage, incident response, RCA, and post-incident reviews
-
Strong troubleshooting skills and RCA capability across distributed applications and services
-
Solid understanding of application architecture, service dependencies, integrations, performance analysis, dependency mapping, and bottleneck identification
-
Practical experience with AWS services, including CloudWatch, ECS, EC2, ALB, Route53, RDS, and VPC
-
Working knowledge of CI/CD pipelines and release validation processes
-
English proficiency at B2 (Upper-Intermediate) level or higher
Nice to have
-
Familiarity with AI-assisted observability capabilities
-
Experience optimizing monitoring and alerting strategies
-
Exposure to Infrastructure as Code (Terraform or equivalent)
-
Travel/Airline industry experience
-
Background supporting modernization and cloud transformation initiatives
We offer
-
International projects with top brands
-
Work with global teams of highly skilled, diverse peers
-
Healthcare benefits
-
Employee financial programs
-
Paid time off and sick leave
-
Upskilling, reskilling and certification courses
-
Unlimited access to the LinkedIn Learning library and 22,000+ courses
-
Global career opportunities
-
Volunteer and community involvement opportunities
-
EPAM Employee Groups
-
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn