Staff Site Reliability Engineer
Jornada completa
Garner Health
Responsibilities
- Own the end-to-end reliability, performance, and resilience strategy for Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads.
- Architect the SLO framework and set technical direction for production quality and reliability at organizational scale.
- Lead the incident response program, complex escalations, on-call improvements, root-cause analysis, and post-incident corrective actions.
- Own the monitoring, alerting, and observability platform across engineering teams.
- Convert ambiguous scaling and reliability requirements into automated, composable Terraform infrastructure-as-code deliverables.
- Drive cloud cost-efficiency, performance optimization, technical debt reduction, and operational automation.
- Establish deployment and observability standards, mentor engineers, and raise operational rigor across the organization.
- Ensure infrastructure and operations meet security and HIPAA compliance obligations.
Requirements
- 7+ years of hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering roles.
- Deep expertise with Kubernetes and Terraform in a cloud-first environment, preferably AWS.
- Experience designing reliability practices including SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews.
- Strong Python or Go skills applied to infrastructure automation; Kubernetes API experience is a plus.
- Demonstrated experience optimizing cloud costs and performance across compute, storage, and networking.
- Experience mentoring engineers and setting technical direction as a senior reliability leader.
- Strong communication skills with both technical and non-technical stakeholders.
- Fluency with AI tools such as Claude in engineering and operations workflows, or strong motivation to develop that capability.
- Experience supporting AI/ML or data-intensive production workloads is preferred.
- Experience in security-conscious or regulated environments, including HIPAA or SOC 2, is preferred.
Benefits
- Remote work is available for candidates comfortable with occasional travel to Garner’s New York City headquarters.
- Flexible paid time off.
- Medical, dental, and vision plan options.
- 401(k) plan.
- Teladoc Health and other competitive benefits.
- Eligibility for equity incentive participation.
Vacante publicada el 12 días atrás
Empleos similares que podrían interesarleBasado en la vacante Staff Site Reliability Engineer en Teletrabajo
- ...and define Windows-agent product parity with macOS. Raise engineering standards through code review, design feedback, and direct mentorship... ...failure handling, cleanup, least privilege, and production reliability. Ability to evaluate kernel-level tradeoffs and explain...
- ...deploy, and maintain reusable platform capabilities, reference architectures, deployment templates, integration components, and engineering standards for enterprise AI. Build secure integrations with the Anthropic API, Claude models, enterprise APIs, data sources, workflow...
- ...drive architectural decisions with lasting platform impact. Turn ambiguous technical challenges into scalable solutions that engineering teams can execute. Build shared systems, patterns, and platform capabilities that improve engineering delivery. Partner with...
- ...implement solutions to resolve or mitigate vulnerabilities and security risks. Drive the product security roadmap with product, engineering, and the broader security team. Research relevant threats and attack methods and convert findings into secure-coding education...
- ...creativity. About the Role We are seeking an experienced Staff Data Engineer to architect, scale, and operationalize the data systems... ...media delivery, and business outcomes. That requires scalable, reliable data platforms for analytics, machine learning, customer-...
- ...ll Be Doing : Lead a small team of engineers, providing technical mentorship, code review... ...deployment, including performance, reliability, and monitoring considerations.... ...products are intuitive, impactful, and reliable. Who We Hope You Are: ~10+ years of...
- Responsibilities Design, automate, validate, and deliver the complete lifecycle of customer compute environments from provisioning through upgrades, recovery, and decommissioning. Use AI to automate and accelerate infrastructure delivery and operations. Provision...
- Responsibilities Design and implement scalable pipelines for telemetry, logs, sensor data, and cloud-software ingestion and processing. Develop and maintain robot-side data ingestion agents for high-volume collection in low-connectivity environments. Develop data...
- ...GitHub Actions and Kubernetes to improve reliability, scalability, and developer velocity.... ...high-confidence releases. Partner with engineering teams on SDLC and developer-experience improvements... ...experience operating at a senior or staff level. ~ Strong experience with AWS...
- ...nice thing about growing 30%+ at ~$2B ARR is that we need great engineers everywhere, and we have the room to shape the role around the... ...metrics on a dashboard. It shows up on the loading dock, the job site, and the dispatch floor. You want to build, not maintain....
¿Desea recibir más vacantes?
Suscríbase y reciba vacantes similares a Staff Site Reliability Engineer. ¡Sea el primero en aplicar!
Búsquedas relacionadas
- personal-para-hotel Teletrabajo
- ripley personal Teletrabajo
- personal jumbo Teletrabajo
- personal supermercado Teletrabajo
- personal para trabajar Teletrabajo
- personal para comidas rapidas Teletrabajo
- seguridad personal Teletrabajo
- personal aseo Teletrabajo
- personal minero Teletrabajo
- personal para packing Teletrabajo

