Senior Site Reliability Engineer
Cloud & Infrastructure Deploy and manage cloud-native infrastructure across AWS, Azure, or GCP. Automate infrastructure provisioning using Infrastructure as Code (IaC). Implement scalable and secure infrastructure solutions. Support Kubernetes-based platforms and containerized workloads. Observability & Monitoring Build monitoring, logging, tracing, and alerting solutions. Implement observability frameworks using industry-standard tools. Monitor application health, performance metrics, and infrastructure utilization. Drive continuous improvements in platform visibility and diagnostics.
Reliability Engineering Design, build, and maintain highly available and fault-tolerant production systems. Define and monitor SLIs, SLOs, and SLAs for critical services. Drive reliability improvements through automation and proactive engineering. Conduct capacity planning and performance optimization activities. Production Support & Operations Manage production environments and ensure service uptime. Lead incident response, troubleshooting, and root cause analysis (RCA). Develop runbooks, operational playbooks, and disaster recovery procedures.
Automation & DevOps Automate deployments, infrastructure management, and operational workflows. Improve CI/CD pipelines and release processes. Implement self-healing, auto-scaling, and operational automation solutions. Promote DevOps and SRE best practices across engineering teams. Security & Compliance Ensure production environments meet security and compliance requirements. Manage secrets, access controls, and vulnerability remediation. Partner with security teams to implement security best practices.
Reliability Engineering Design, build, and maintain highly available and fault-tolerant production systems. Define and monitor SLIs, SLOs, and SLAs for critical services. Drive reliability improvements through automation and proactive engineering. Conduct capacity planning and performance optimization activities. Production Support & Operations Manage production environments and ensure service uptime. Lead incident response, troubleshooting, and root cause analysis (RCA). Develop runbooks, operational playbooks, and disaster recovery procedures.
Automation & DevOps Automate deployments, infrastructure management, and operational workflows. Improve CI/CD pipelines and release processes. Implement self-healing, auto-scaling, and operational automation solutions. Promote DevOps and SRE best practices across engineering teams. Security & Compliance Ensure production environments meet security and compliance requirements. Manage secrets, access controls, and vulnerability remediation. Partner with security teams to implement security best practices.
Recommended Jobs
Principal Product Manager, AI Build Governance
Posted just now
Director, Cash-In Engineering Partner
Posted just now
Sr. QA Engineer 1 - MyWP Core SIT
Posted just now
Assoc. Mgr., Engineering - MyWP Core SIT
Posted just now
Social Media Marketing Intern
Posted just now

