1. L3 On-Call Escalation & Production Support
• Act as the final technical escalation tier (L3 support) for critical platform and production incidents, troubleshooting complex issues escalated by L1/L2 operations or customer support teams.
• Lead high-priority incident response bridges, coordinating across development, security, and networking teams to drive rapid issue containment and resolution under strict service-level agreements (SLAs).
• Troubleshoot deep, transient infrastructure errors (e.g., routing, network package drops, Kubernetes control-plane failures) that go beyond standard SOPs.
• Establish clear, documented escalation pathways, translating complex L3 resolutions into actionable runbooks and empower L1/L2 teams
2. FedRAMP Compliance, Audit Readiness & Log Auditing
• Implement and enforce security controls aligned with the NIST SP 800-53 framework to achieve and maintain our FedRAMP Authorization to Operate (ATO).
• Act as the technical lead for SRE during annual FedRAMP 3PAO audits, gathering evidence, demonstrating compliance, and proving operational control implementation.
• Build and maintain secure, tamper-proof audit logging pipelines—forwarding application logs, system logs, API call records, and Kubernetes audit trails securely into Elasticsearch (and central SIEMs) with strict, compliant retention and index-lifecycle management policies.
3. Deep Log Troubleshooting & Analysis (ELK Stack)
• Utilize Elasticsearch and Kibana as primary investigative tools to perform deep-dive troubleshooting of complex, distributed system anomalies and application errors across AWS and edge environments.
• Build, customize, and curate high-signal Kibana dashboards, search queries (KQL/Lucene), and visualizations to provide real-time operational visibility and dramatically reduce Mean Time to Resolution (MTTR).
• Troubleshoot log-ingestion pipelines (Vector, Fluentd) to resolve bottlenecks, parsing errors, or missing metadata in high-volume production environments.
4. Secure Architectural Design & Hybrid Infrastructure
• Design and document highly resilient, secure-by-default architectural topologies for hybrid networks spanning AWS GovCloud and on-premises edge servers.
• Architect and scale multi-tenant Kubernetes (EKS) clusters enforcing strict physical or logical boundaries, zero-trust network policies, and identity federation.
• Configure and optimize high-availability L4/L7 load balancer networking, ingress controllers, and FIPS 140-3 validated traffic routing for secure, low-latency edge-to-cloud communication.
• Manage physical bare-metal edge nodes, defining operating system hardening baselines (STIG compliance), secure boot processes, and automated physical host provisioning.
5. Infrastructure as Code & Deployments
• Author clean, modular, and secure Terraform code to provision cloud infrastructure, networking topologies, and security boundaries.
• Build declarative deployment pipelines using GitLab CI/CD and ArgoCD to achieve true GitOps-driven delivery, ensuring all code modifications are traceable, signed, and fully audited.
• Develop custom automation tools, controllers, and CLI utilities in Python or Go (Golang) to eliminate repetitive toil and automate continuous compliance reporting.
6. Observability & Operational Excellence
• Design end-to-end monitoring and dashboarding systems utilizing Prometheus, Grafana, and Alertmanager.
• Build actionable alerting pipelines integrated with Slack, ensuring notifications are high-signal and low-noise.
• Facilitate blameless post-mortem reviews following high-severity incidents, utilizing evidence harvested from Kibana logs to document timelines, determine root causes, and programmatically prevent recurrences.