SRE Engineer
Atlanta - Georgia - USAOn-siteFull-timeInformation Technology
Description
#Careers JC 1483593 Qualifications · Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS. · Hands-on experience with incident management and 24/7 production support models. · Proficiency with monitoring and observability tools such as CloudWatch, Dynatrace, and Quantum Metric. ·� � � � � � � � Experience building and maintaining monitoring dashboards.
·� � � � � � � � Strong troubleshooting skills across infrastructure, networking, and application layers.
· Working knowledge of CI/CD pipelines and AWS deployment processes. · Experience working with databases and Unix/Linux environments. Key Responsibilities Incident Management and Production Support · Provide Level 1 and Level 2 support for production incidents across AWS-hosted applications and infrastructure. ·� � � � � � � � Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.
·� � � � � � � � Escalate code-level defects to development teams with clear diagnostics, supporting logs, and impact assessments.
· Participate in on-call rotations, major incident bridges, and post-incident reviews. · Investigate application defects, configuration issues, and infrastructure anomalies reported through monitoring tools or user incidents. Monitoring and Operational Health · Perform regular health checks across applications, infrastructure, and AWS services. · Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes. · Respond proactively to s related to resource utilization, latency, errors, and availability. · Maintain and improve monitoring and observability dashboards.
Requirements
Mandatory Skills : CloudWatch, Dynatrace, Git, Observability, Reliability Patterns
Good to Have Skills : Chaos Testing, Shell Scripting