Unleash your potential: What you will be doing and owning:
• Monitor and support the reliability, performance, and scalability of PAR’s cloud platforms and SaaS services
• Help build and maintain monitoring, alerting, dashboards, and observability solutions
• Support tracking of Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics
• Participate in incident response, troubleshooting, escalation, and service restoration activities
• Contribute to root cause analysis and blameless postmortems that drive continuous improvement
• Develop and maintain automation scripts to reduce manual effort and eliminate operational toil, using AI-assisted tools where useful
• Help identify monitoring and observability gaps through data analysis
• Assist in creating and maintaining runbooks, operational procedures, and troubleshooting workflows
• Support Kubernetes platform operations and containerized application deployments
• Assist with Infrastructure as Code initiatives using Terraform and related tooling
• Collaborate with Engineering teams to improve application reliability and operational readiness
• Participate in change management reviews and production readiness assessments
• Help analyse system trends, performance metrics, and operational data to identify improvements
• Support capacity planning, performance optimization, and disaster recovery activities
• Partner with Security teams on vulnerability remediation and operational compliance activities
• Participate in rotational shift coverage and on-call support for critical production systems