Program Manager, Site Reliability Engineering

Hydrolix
Hydrolix

Software Engineering, Operations

United States · Remote

Posted on Jul 24, 2026
Careers

Program Manager, Site Reliability Engineering

LOCATION
|
Hydrolix | United States, Remote
Job Description

Hydrolix is looking for a Program Manager to embed in our Site Reliability Engineering organization. Reporting directly to the Head of SRE, you will own coordination, process, and delivery across SRE’s key programs: incident management, feature readiness, cross-functional intake, and operational health.

The SRE team is globally distributed and moves fast. We need someone who can bring structure without adding bureaucracy, is comfortable working alongside engineers without needing things translated, and wants to have a real impact on how a growing SRE organization operates.

Key Responsibilities
  • Track and drive delivery across multiple SRE programs and workstreams, keeping milestones, owners, and risks clear and visible.
  • Lead cross-functional coordination between SRE and adjacent teams, managing intake, surfacing dependencies, and ensuring commitments are tracked.
  • Own SRE process discipline: incident management cadence, post-incident review quality and completion, on-call health metrics, and program-level reporting.
  • Drive feature readiness and release coordination, working with SRE, Engineering, and Product to define go-live criteria and structured handoff processes.
  • Identify process gaps within the SRE organization and design solutions that work for a globally distributed team.
  • Keep stakeholders and leadership informed through clear status reporting, risk tracking, and dependency management.
  • Maintain comprehensive program documentation: project plans, runbooks, retrospectives, and post-implementation reviews.
Qualifications and Skills
  • 5+ years of Technical Program Management experience in a SaaS or technology company; 7+ years if coming from a traditional PM background. Startup experience is a strong plus.
  • Experience working in or alongside an SRE, Platform Engineering, or DevOps organization, with familiarity with concepts like incident lifecycle, SLOs, on-call operations, and release pipelines.
  • Strong communication and facilitation skills, with a track record of driving alignment across engineering, product, and customer-facing teams across multiple time zones.
  • Able to earn trust with technical stakeholders and ask the right questions to reach clarity, without needing to write code.
  • Proven ability to take an undefined problem, structure it, and turn it into a repeatable process.
  • Proactive on risk: spots issues early and drives them to resolution.
  • Confident in influencing without direct authority, holding ICs and leadership accountable to our processes and promises.
  • PMP or equivalent certification is a plus; track record of delivery matters more than credentials.
Bonus Qualifications
  • Experience with incident management platforms, observability tooling, or ITSM processes in a production SaaS environment.
  • Background supporting an infrastructure or reliability team through a scaling phase.
  • Experience building processes from scratch in a startup or scale-up.

To apply, please fill out the form below and attach your resume. Requirement of cover letter / relevant references are subject to role, please see job description.

We look forward to seeing how you can make an impact at Hydrolix.