When something goes wrong in production overnight, the Production Support & Incident Response Lead is the person leading the response.
EviSmart is looking for a Production Support & Incident Response Lead to own incident response and platform continuity during our night operations.
This is not a role where you simply monitor dashboards, create a ticket, and wait for Engineering.
You will be the lead incident response person on shift. You are expected to investigate first, understand what is happening, determine the safest way to restore operations, bring in the right technical people when necessary, and remain accountable for the incident until the platform is stable.
Our platform supports 2,000+ dental labs, so an issue in production can quickly become a real business problem for our customers. The goal is simple: keep cases moving and minimize disruption.
What you’ll own • You will be responsible for the health and continuity of the platform during your coverage.
•
Lead the response to production incidents during the night shift
•
Continuously monitor platform health and act on warning signs before they become customer-impacting problems
•
Personally perform first-line investigation using logs, dashboards, monitoring tools, diagnostic commands and available system access
•
Determine the impact and likely source of an issue before escalating
•
Look for safe workarounds or restoration options when a permanent fix is not immediately available
•
Decide when an issue can be handled at your level and when Engineering, DevOps or another specialist needs to be brought in
•
Command the incident even after technical teams become involved: keep people aligned, decisions moving and communication clear
•
Keep Application Support and other stakeholders informed during active incidents
•
Document incidents, root causes, workarounds and follow-up actions
•
Make sure recurring issues don’t simply become accepted problems
•
Provide a complete handoff to the daytime team with nothing dropped overnight
•
Develop one Application Support teammate into a reliable backup who can eventually handle routine night triage independently
The role’s ownership of monitoring, incident command, proactive customer communication, post-mortems and backup development is explicit in the operating playbook.
What this role is NOT • This is not a traditional Service Delivery Manager or ITIL governance position.
It is also not a pure DevOps or Software Engineering role. You don’t need to be the person who writes the permanent code fix for every problem. But you do need enough technical depth to investigate intelligently before asking someone else to solve it.
If your normal incident process is: Alert → Create ticket → Escalate → Wait … then this probably isn’t the right role.
We’re looking for someone whose instinct is closer to:
Detect → Investigate → Isolate → Restore or Work Around → Escalate Intelligently → Command Through Resolution → Prevent Recurrence