SRE (Site Reliability Engineer) / Platform Engineer
Job Description
What You’ll Be Doing:
What You’ll Be Doing:
- Managing and optimizing large-scale AWS cloud infrastructure
- Driving Kubernetes (EKS) adoption and workload migration
- Automating infrastructure using Terraform and GitOps practices
- Enhancing platform reliability, observability, and operational excellence
- Defining and implementing SLI/SLO frameworks
- Improving monitoring, alerting, and incident response processes
- Collaborating with development teams to build highly available and resilient systems
- AWS (EKS, EC2, RDS Aurora, ElastiCache, Control Tower)
- Kubernetes (Multi-Cluster & Multi-Environment)
- Terraform (Infrastructure as Code)
- ArgoCD, Atlantis, GitOps
- Karpenter & KEDA
- Datadog, Prometheus, Monitoring & Alerting
- CI/CD Pipelines
- SLI/SLO Implementation and Site Reliability Engineering Practices
Recommended Jobs
Principal Product Manager, AI Build Governance
Posted just now
Director, Cash-In Engineering Partner
Posted just now
Sr. QA Engineer 1 - MyWP Core SIT
Posted just now
Assoc. Mgr., Engineering - MyWP Core SIT
Posted just now
Social Media Marketing Intern
Posted just now

