Sr Technical Consultant - Azure Cloud, DevOps, Sql & Site Reliability Engineering(SRE)
Location2 Locations
Work modeon-site
Typefull-time
Company size5,001–10,000 people
First seen1w ago
Last seen3d ago
Overview:
•
The BYX Platform provides shared cloud provisioning, compute, storage, observability and runtime capabilities for enterprise applications and implementation teams.
•
We are seeking an astute Cloud Platform with a strong technical foundation and hands-on experience in Microsoft Azure, Kubernetes, containers, observability, networking and production automation.
Our current technical environment:
•
Cloud and Compute: Microsoft Azure, Kubernetes, node pools, container workloads, Azure Container Registry, KEDA, Dask, autoscaling, CPU, memory and GPU capacity.
•
Observability: Logs, metrics and traces; collectors and exporters; Elastic/Elasticsearch, Logstash, Kibana, dashboards, alerts and correlation identifiers.
•
Data and Storage: MongoDB/Atlas, SQL, Redis, NFS, managed storage, replicas, regional capacity and service quotas.
•
Networking and Security: DNS, CIDR, private endpoints, firewall rules, allowlists, TLS/certificates, SPNs, API keys, tokens, image-pull credentials and Git-hosted secrets.
•
Operations and Automation: Linux, Git, CI/CD deployment workflows, infrastructure automation and scripting using Bash, Python or PowerShell.
What you’ll do:
•
Review and act on incidents, service requests, infrastructure requests and provisioning failures logged by implementation teams and platform users.
•
Own L2/L3 cloud-platform issues from initial triage through recovery, validation, communication, root-cause analysis and closure.
•
Troubleshoot Kubernetes pods, deployments, replicas, services, events, health checks, node pools, scheduling and resource constraints.
•
Diagnose container image-pull, startup, shutdown, registry-authentication, rollout and workload-reconciliation failures.
Experience with a container registry; Azure Container Registry experience is highly relevant.
•
Practical observability experience across logs, metrics, traces, dashboards and alerts.
•
Experience with Elastic Stack components—Elasticsearch, Logstash and Kibana—or a closely comparable platform.
•
Working knowledge of DNS, TLS/certificates, CIDR, firewalls, proxies, load balancing and private networking.
•
Strong Linux administration, application log analysis and production incident-troubleshooting skills.
•
Scripting experience with Bash, Python or PowerShell for diagnostics and operational automation.
•
Experience troubleshooting CI/CD deployments, configuration changes and failed rollouts.
•
Working knowledge of Git, service identities and secrets-management fundamentals.
•
Working knowledge of MongoDB/Atlas, Redis, SQL or NFS operations is preferred.
•
Experience with infrastructure as code such as Terraform, Bicep or ARM templates and Kubernetes packaging such as Helm is preferred.
•
Experience with incident response, Root Cause Analysis, post-incident reviews and controlled production changes.
•
Strong collaboration and communication skills with the ability to work across application, security, networking and vendor teams.
•
Willingness to participate in a scheduled production on-call rotation and planned out-of-hours changes when required.
Our Values
If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values
Legal
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.