AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you
ABOUT THE ROLE
We are looking for a DevOps Engineer to manage application migrations between environments and maintain production and staging system reliability. This person handles incident response and on-call support, while improving observability, automation, and Kubernetes-based infrastructure. Weekend availability for migrations and comfort with CI/CD pipelines and SLAs are essential.
WHAT YOU WILL DO
- - Migrate applications between environments; these migrations typically take place on weekends, so weekend availability is required.
- - Monitor and support production and staging environments in real time, ensuring high availability, performance, and stability.
- - Respond to incidents, perform triage and root cause analysis, and contribute to post-incident reviews and remediation efforts.
- - Participate in an on-call rotation with defined SLAs.
- - Handle ad-hoc and unplanned operational requests from Product, Support, and internal teams.
- - Maintain and enhance monitoring, alerting, dashboards, logs, and metrics; improve signal-to-noise ratio and standardize observability practices.
- - Support CI/CD pipelines, production releases, and GitOps workflows.
- - Contribute to automation efforts to reduce operational toil.
- - Maintain and improve Kubernetes-based infrastructure and containerized workloads.
- - Support Infrastructure as Code practices and ongoing environment improvements.
MUST HAVES
- - At least 2 years of experience in Site Reliability Engineering, DevOps, or Production Operations
- - Demonstrable AWS experience supporting production