Job Summary
First point of contact for web tier code releases. They monitor the continuous push pipeline, triage breakages, perform routine remediation, manage release scheduling, and escalate complex issues
Key Responsibilities
Release Monitoring - Continuously monitor pipeline health via FlightDeck and Conveyor, track build progress and test signals, acknowledge alerts within 15 minutes, and keep pushctl updated.Triage & Classification - Investigate failures by reading logs and exceptions, classify as infrastructure/product/capacity/flaky, and compare against historical runs to identify patterns.Remediation - Execute standard remediation steps per runbook, including retrying failed steps and skipping documented flaky tests.Release Schedule Management - Stay aware of upstream dependencies (drain testing, HHVM builds, config pushes), coordinate around code freezes and maintenance windows, and flag release drift early.Runbook Maintenance - Keep runbooks current as failure patterns evolve, document new remediation steps, and flag gaps to L3.Rollback Readiness - Maintain awareness of rollback procedures and ensure rollback paths are clear before high-risk releases proceed.Reporting & Analysis - Produce structured shift handoff reports, prepare daily/weekly on-call summaries, and surface recurring trends to L3.Escalation - Escalate infrastructure-level issues to Tree Hugger (L2) with full context: timeline, logs reviewed, classification, and actions taken.