Engineering-Led Issue Resolution
• Act as the final escalation point for complex, cross-system production issues impacting integrations or customer-facing features.
• Lead incident bridges for high-severity events, acting as technical incident commander and driving timely root cause analysis (RCA), postmortems, and recovery.
• Investigate campaign and loyalty defects end to end — from a brand’s report through segment, scheduling, delivery, and reporting layers — to a defensible root cause.
Code-Level Debugging & Platform Stewardship
• Perform deep root cause analysis across APIs, integrations, and microservices using logs, telemetry, SQL queries, and source-level debugging.
• Contribute directly to the codebase with fixes, refactors, and stability improvements, collaborating with development teams for safe rollouts.
• Maintain a comprehensive understanding of Punchh’s platform architecture, data flows, and infrastructure to support ongoing reliability efforts.
Data Remediation & Script Execution
• Design, review, and execute remediation scripts to correct guest-level data — missed gifting, incorrect point balances, duplicate transactions, and reward corrections.
• Validate scope and counts before execution and verify results afterwards; work within platform constraints such as job runtime limits using checkpointing and resumable approaches.
• Exercise appropriate caution and secure sign-off before any operation that modifies production guest data at scale.
Technical Leadership for Integrations
• Serve as the engineering lead for third-party integrations (POS, online ordering, payment processors, marketing platforms), ensuring compliance, reliability, and quality.
• Validate and certify partner integrations through code review, API validation (leveraging Postman where appropriate), and architecture alignment.
• Define and enforce integration best practices and standards across the platform ecosystem.
Operational Excellence, ITIL & Automation
• Leverage Databricks for scheduled job analysis, performance tuning, and large-scale data troubleshooting.
• Use New Relic and other observability tools to monitor application health, detect anomalies, and proactively identify system issues.
• Maintain and evolve technical runbooks, diagnostic playbooks, and integration guidelines to reduce future escalations.
• Apply ITIL best practices for incident, problem, and change management to improve process consistency.
• Identify opportunities to automate triage, alerting, and recovery workflows, reducing manual effort and improving mean-time-to-recovery (MTTR).
Brand-Facing Technical Communication
A distinctive part of this role: technical findings frequently need to reach the restaurant brands we serve. You will be expected to:
• Translate technical root causes into clear, accurate explanations that Customer Success can take to a brand — including drafting the customer-facing wording.
• Be precise about what is known, what is not known, and what cannot be determined, rather than over-claiming under pressure.
• Distinguish clearly between platform defects, configuration issues on the brand’s side, and expected behavior — and communicate that distinction diplomatically.
• Contribute to formal RCA documents delivered to brands following significant incidents.
Cross-Functional Collaboration & Knowledge Sharing
• Collaborate with Engineering, Product, DevOps, and Customer Success teams to identify systemic trends and implement platform-level improvements.
• Represent Sustaining Engineering in release planning, readiness reviews, and go-live support.
• Mentor Tier 2 teams, share knowledge, and help build organizational capability to reduce escalations over time.