Regístrese para acceder a todas las funciones de nuestro servicio
  • Búsqueda de ofertas de trabajo
  • Favoritos
  • Crear CV
    Nuevo
  • Alertas de empleo

Staff Site Reliability Engineer

Jornada completa

Garner Health

Responsibilities

  • Own the end-to-end reliability, performance, and resilience strategy for Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads.
  • Architect the SLO framework and set technical direction for production quality and reliability at organizational scale.
  • Lead the incident response program, complex escalations, on-call improvements, root-cause analysis, and post-incident corrective actions.
  • Own the monitoring, alerting, and observability platform across engineering teams.
  • Convert ambiguous scaling and reliability requirements into automated, composable Terraform infrastructure-as-code deliverables.
  • Drive cloud cost-efficiency, performance optimization, technical debt reduction, and operational automation.
  • Establish deployment and observability standards, mentor engineers, and raise operational rigor across the organization.
  • Ensure infrastructure and operations meet security and HIPAA compliance obligations.

Requirements

  • 7+ years of hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering roles.
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment, preferably AWS.
  • Experience designing reliability practices including SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews.
  • Strong Python or Go skills applied to infrastructure automation; Kubernetes API experience is a plus.
  • Demonstrated experience optimizing cloud costs and performance across compute, storage, and networking.
  • Experience mentoring engineers and setting technical direction as a senior reliability leader.
  • Strong communication skills with both technical and non-technical stakeholders.
  • Fluency with AI tools such as Claude in engineering and operations workflows, or strong motivation to develop that capability.
  • Experience supporting AI/ML or data-intensive production workloads is preferred.
  • Experience in security-conscious or regulated environments, including HIPAA or SOC 2, is preferred.

Benefits

  • Remote work is available for candidates comfortable with occasional travel to Garner’s New York City headquarters.
  • Flexible paid time off.
  • Medical, dental, and vision plan options.
  • 401(k) plan.
  • Teladoc Health and other competitive benefits.
  • Eligibility for equity incentive participation.
Vacante publicada el 12 días atrás
Empleos similares que podrían interesarleBasado en la vacante Staff Site Reliability Engineer en Teletrabajo
  •  ...and define Windows-agent product parity with macOS. Raise engineering standards through code review, design feedback, and direct mentorship...  ...failure handling, cleanup, least privilege, and production reliability. Ability to evaluate kernel-level tradeoffs and explain... 

    CloudZero

    Teletrabajo
    1 día atrás
  •  ...deploy, and maintain reusable platform capabilities, reference architectures, deployment templates, integration components, and engineering standards for enterprise AI. Build secure integrations with the Anthropic API, Claude models, enterprise APIs, data sources, workflow... 

    NewRocket

    Teletrabajo
    4 días atrás
  •  ...drive architectural decisions with lasting platform impact. Turn ambiguous technical challenges into scalable solutions that engineering teams can execute. Build shared systems, patterns, and platform capabilities that improve engineering delivery. Partner with... 

    Onebrief

    Teletrabajo
    6 días atrás
  •  ...implement solutions to resolve or mitigate vulnerabilities and security risks. Drive the product security roadmap with product, engineering, and the broader security team. Research relevant threats and attack methods and convert findings into secure-coding education... 

    Grow Therapy

    Teletrabajo
    10 días atrás
  •  ...creativity. About the Role We are seeking an experienced Staff Data Engineer to architect, scale, and operationalize the data systems...  ...media delivery, and business outcomes. That requires scalable, reliable data platforms for analytics, machine learning, customer-... 

    Vidmob

    Teletrabajo
    9 días atrás
  •  ...ll Be Doing : Lead a small team of engineers, providing technical mentorship, code review...  ...deployment, including performance, reliability, and monitoring considerations....  ...products are intuitive, impactful, and reliable. Who We Hope You Are: ~10+ years of... 

    Federato

    Teletrabajo
    14 días atrás
  • Responsibilities Design, automate, validate, and deliver the complete lifecycle of customer compute environments from provisioning through upgrades, recovery, and decommissioning. Use AI to automate and accelerate infrastructure delivery and operations. Provision...

    fal

    Teletrabajo
    16 días atrás
  • Responsibilities Design and implement scalable pipelines for telemetry, logs, sensor data, and cloud-software ingestion and processing. Develop and maintain robot-side data ingestion agents for high-volume collection in low-connectivity environments. Develop data...

    Agility Robotics

    Teletrabajo
    15 días atrás
  •  ...GitHub Actions and Kubernetes to improve reliability, scalability, and developer velocity....  ...high-confidence releases. Partner with engineering teams on SDLC and developer-experience improvements...  ...experience operating at a senior or staff level. ~ Strong experience with AWS... 

    Phantom

    Teletrabajo
    16 días atrás
  •  ...nice thing about growing 30%+ at ~$2B ARR is that we need great engineers everywhere, and we have the room to shape the role around the...  ...metrics on a dashboard. It shows up on the loading dock, the job site, and the dispatch floor. You want to build, not maintain.... 

    Samsara

    Teletrabajo
    21 días atrás

¿Desea recibir más vacantes?

Suscríbase y reciba vacantes similares a Staff Site Reliability Engineer. ¡Sea el primero en aplicar!