Systems Reliability Engineer

Local Public, a Public Benefit Organization · Remote (US)

Software Engineering
Public Infrastructure
Public Service & Civic Engagement
$100,000 - $110,000 Per Year
Posted 1 hour ago
Report an Issue
Featured Job

Role Overview

We are seeking a skilled proactive Systems Reliability Engineer (SRE) to maintain and enhance the stability, scalability, and performance of our next-generation streaming infrastructure. In this role, you will bridge the gap between development and operations, working closely with the Technology and Engineering teams to maintain a robust multi-cloud environment. You will be responsible for the reliability of our streaming infrastructure, apps and releases. Ensuring that our interfaces and serverless backends operate at peak efficiency for audiences distributed across the United States and Canada to maintain constant uptime for our broad audiences.

Remote Work & Hours

This is a remote position open to candidates based in the United States. We offer flexible working hours, with some scheduling shaped by streaming operations and viewer patterns. While schedules may vary, work will generally take place during the afternoon and evening hours to align with audience demand and service needs.

Key Responsibilities

Infrastructure Management: Manage and optimize a multi-cloud footprint across AWS (ECS, CloudFront, Aurora Serverless) and GCP (Cloud Functions, Firebase), ensuring high availability and low latency for streaming services.

Deployment: Maintain backend containerized workloads across AWS and GCP to support flexible, scalable deployment patterns and framework delivery using a variety of tools (Firebase, FlightControl, Terraform) either directly or via IaC systems.

CI/CD & Automation: Own and refine the deployment pipeline using AWS, GitHub, and GCP Cloud Build to enable frequent, low-risk releases.

System Health & Monitoring: Implement comprehensive monitoring and alerting strategies to proactively identify and resolve bottlenecks within our streaming layers and database clusters.

Scalability & Performance: Engineer and monitor automated scaling solutions to handle fluctuating traffic patterns inherent in streaming media.

Security & Compliance: Ensure rigorous security standards across all cloud services and deployment stages.

App release to various stores: Work with the product team on app releases.

Required Technical Skills

Cloud Platforms: Experience with AWS (Aurora Serverless) and GCP (Functions, Cloud Build, Firebase).

Deployment: Experience with Firebase and container orchestration platforms on both AWS and GCP. IaC concepts and monitoring.

Development Frameworks: Experience with ORMs and SQL, NodeJS GraphQL.

CI/CD & DevOps: Hands-on experience with GitHub Actions, automated pipeline orchestration and infrastructure as code (IaC)

Databases: Strong SQL knowledge and experience managing PostgreSQL and MongoDB in a serverless context.

Content Delivery Network (CDN): Networking fundamentals, caching & proxies, observability & monitoring

App Release: General understanding of app release processes to various app stores (Apple, Google, Roku, LG, Samsung) and coordinating release activities with product teams.

Qualifications

  • Experience in SRE or DevOps roles, preferably within a streaming or high-traffic digital media environment.
  • Hands-on experience with containerization strategies on AWS and GCP.
  • Deep understanding of distributed systems and microservices architecture.
  • Strong problem-solving skills and the ability to work independently in a fast-paced, remote-first culture.
  • Excellent communication skills, with the ability to provide technical updates.
View more remote jobs
Be the first to see new Systems Reliability Engineer jobs

Save this search to get an email when new jobs match this search.

Create Email Alert