1. 24/7 Eyes-on Monitoring, Log Troubleshooting & Incident Triage
• Maintain continuous monitoring of production telemetry across bare-metal RHEL servers, EKS clusters, and AWS cloud environments.
• Act as first responder to high-priority alerts generated by Prometheus and Alertmanager, performing immediate triage and root cause investigation.
• Perform deep-dive log troubleshooting using Elasticsearch and Kibana (building queries, analyzing container/application stdout logs, systemd/journald logs, and ingress traffic logs) to isolate error patterns and service failures.
• Coordinate incident resolution bridges in Slack, engaging secondary on-call engineers, security teams, or infrastructure engineers as required.
• Maintain rigorous adherence to Incident Response SLAs and document timeline logs for post-incident reviews (RCAs).
2. Bare-Metal On-Premises RHEL Server Support & Hardware Operations
• Provide operational support for bare-metal on-premises Linux servers (RHEL), including OS configuration, system maintenance, and day-to-day lifecycle administration.
• Diagnose physical network interface bonding, link aggregation, VLAN tagging, and local server connectivity issues on bare-metal systems.
• Coordinate vendor hardware dispatch requests for failing server components and oversee physical parts replacements.
3. FedRAMP Security & Vulnerability Remediation
• Execute day-to-day security operational duties aligned with FedRAMP High/Moderate (NIST SP 800-53) standards.
• Perform routine vulnerability patching (CVE remediation) across bare-metal RHEL servers, cloud AMIs, Kubernetes worker nodes, and container images within strict regulatory timelines.
• Apply operational system hardening based on DISA STIG guidelines and maintain FIPS 140 compliance configurations across both cloud and on-premise operating systems.
• Ensure strict Role-Based Access Control (RBAC), SSH key management, and security boundary enforcement across operational environments.
4. Kubernetes (EKS), AWS Infrastructure & Networking Operations
• Support day-to-day operations and node maintenance for Amazon EKS clusters across AWS Commercial and AWS GovCloud environments.
• Perform Layer 4 (L4) and Layer 7 (L7) secure load balancing troubleshooting, including AWS Application Load Balancers (ALB), Network Load Balancers (NLB), ingress controllers, SSL/TLS certificate termination, and traffic routing issues.
• Perform Linux administration tasks including kernel parameter tuning (sysctl), storage expansion, log rotation, and system troubleshooting.
• Execute infrastructure changes and updates using Infrastructure as Code (Terraform) in alignment with change control procedures.
• Monitor key AWS cloud infrastructure components including VPCs, Security Groups, EC2, IAM, S3, and KMS.
5. CI/CD & GitOps Deployment Operations (GitLab, ArgoCD, Argo Workflows)
• Execute, monitor, and troubleshoot automated deployment pipelines using GitLab CI/CD.
• Manage application state, synchronization, and rollouts across Kubernetes clusters using ArgoCD (GitOps paradigm).
• Monitor, execute, and troubleshoot operational batch processes, system maintenance tasks, and automated pipelines using Argo Workflows.
• Facilitate application releases and configuration rollouts using Helm charts, Kustomize, and GitOps workflows.
• Validate build pipeline compliance, container security scanning results, and image signature verifications prior to production deployment.
6. Documentation & Operational Runbooks
• Maintain accurate, step-by-step incident runbooks, standard operating procedures (SOPs), and triage workflows for both cloud and bare-metal environments.
• Write detailed post-incident reports and post-mortems for production impact events.
• Identify manual operational overhead (toil) and implement shell scripts (Bash/Python) to streamline routine monitoring, hardware checks, and maintenance tasks.