Senior Site Reliability Engineer

Senior Site Reliability Engineer

Company, Product: Exygy, Bloom 

Date Posted: August 18, 2026

Application Deadline: September 3, 2026

Function: Engineering

Level: P3

Job Purpose 


Remote, US | Engineering | full-time, termed

Exygy has been building technology for the public good for more than 20 years. Founded in 2003 as an impact-focused technology agency, Exygy has delivered life-changing solutions for nonprofits, mission-driven organizations, and governments across the globe. Today, the company is focused on scaling Bloom Housing and CiviForm, two digital tools that help government teams better serve their communities. Our mission has remained the same from the start: to design and build user-centered technology that improves lives.


Bloom Housing is an open-source platform, built by Exygy, that helps governments and housing partners modernize access to affordable housing. By centralizing the housing search and application process into a single portal, Bloom simplifies how residents find and apply for affordable housing, while giving property managers and housing staff a unified way to manage listings, applications, and housing stock all in one place. Bloom currently supports affordable housing access across the 9 Bay Area counties, Detroit, and Los Angeles.


Over the last decade, Bloom has developed from an open-source framework into a housing platform used by cities, counties, and regional housing partners. As Bloom expands to additional jurisdictions, we need reliable, secure, and repeatable infrastructure that can support a growing number of government partners without requiring every implementation to become a custom deployment. 


The Senior Site Reliability Engineer will lead the technical work required to make Bloom’s environments, deployment systems, and operating practices more dependable, scalable, secure, and efficient. You will own and evolve Bloom’s infrastructure and deployment capabilities, manage production and staging environments, establish measurable reliability practices, improve monitoring and incident response, and partner with engineers to strengthen how Bloom is tested, released, and operated.


Operating as a senior individual contributor within a small-scale team, you will merge high-level architecture and future planning with practical execution, troubleshooting, documentation, and operational maintenance. This position demands sound technical discernment, adaptability within a shifting landscape, and a commitment to constructing the frameworks and methodologies you suggest. Furthermore, there may be instances where standalone contributions and feature development are necessary.


Employment Terms


This is a full-time, termed, remote, U.S.-based position. This is a funding-contingent position with an anticipated duration through August 2027, based on the current term of the applicable client contract or grant and the portion of the scope of that contract or grant that requires specifically the work of this role. If the contract or funding is not awarded, is delayed, reduced, modified, terminated, expires, or is not renewed, the position may be modified or eliminated, and employment may end earlier than anticipated, subject to applicable law and any legally required processes. The role may occasionally require travel for team meetings or other business needs. Because this position supports production systems, participation in an on-call rotation is required.


Who Does This Role Report To

Principal Engineer, Bloom


Supervisory Responsibilities


None.


Responsibilities 


Site Reliability & Production Operations


  • Own the reliability, availability, performance, and operational health of Bloom’s production and staging environments.

  • Participate in a sustainable on-call rotation and respond to production incidents, service interruptions, and urgent operational issues.

  • Identify recurring operational problems and implement durable solutions rather than relying on repeated manual intervention.

  • Establish clear escalation paths and operating procedures for outages, degraded service, deployment failures, and security events.

  • Evaluate system capacity, dependencies, failure modes, and operational risks as Bloom grows.

  • Identify and mitigate security risks in Bloom’s infrastructure, deployment systems, and production environments.


Infrastructure, Deployment, and Continuous Integration/Delivery


  • Own and evolve Bloom’s deployment systems from development and prototyping through production delivery.

  • Develop a more standardized, repeatable, and scalable approach to deploying Bloom across jurisdictions and hosting environments.

  • Build and maintain containerized and Kubernetes-based infrastructure where appropriate using infrastructure as code, including Terraform.

  • Reduce deployment complexity, configuration drift, and reliance on manual processes.

  • Evaluate existing and proposed infrastructure against reliability, security, scalability, cost, portability, and maintainability needs.

  • Improve Bloom’s continuous-integration and continuous-delivery systems and practices.

  • Partner with engineers to strengthen automated testing, release procedures, deployment validation, and rollback capabilities.

  • Increase the safety, consistency, and frequency of releases while reducing avoidable production risk.

  • Automate repetitive infrastructure, deployment, and operational tasks.

  • Improve visibility into deployment status, release health, and production impact.

  • Promote shared ownership of reliability and operability throughout the software-development lifecycle.


Observability and Reliability Engineering


  • Define, document, and implement service-level indicators and service-level objectives for Bloom’s critical services.

  • Develop and maintain monitoring, logging, alerting, dashboards, and other observability capabilities.

  • Gather and analyze deployment and production metrics to identify reliability, performance, and configuration improvements.

  • Develop disaster-recovery, backup, restoration, and service-continuity practices.

  • Create clear operational measures that help Engineering and Product make informed decisions about reliability risks and priorities.


Engineering Collaboration & Technical Leadership


  • Collaborate with software engineers, product leaders, delivery team members, and organizational leadership to ensure new features are designed and implemented with reliability and operability in mind.

  • Contribute code to Bloom when needed to improve reliability, performance, security, operability, or platform capabilities.

  • Participate in code reviews and technical refinement of work affecting infrastructure and production systems.

  • Translate reliability, security, and infrastructure risks into clear options and recommendations for technical and nontechnical audiences.

  • Support government partners and internal teams in resolving issues related to Bloom environments and deployments.

  • Create and maintain deployment guides, incident playbooks, architecture documentation, troubleshooting materials, and operational runbooks.

  • Mentor teammates and help build broader organizational knowledge of site reliability and production operations.

  • Work effectively with external vendors, contractors, or government technology teams when responsibilities cross organizational boundaries.


Required Skills, Knowledge, and Abilities


  • Deep experience building and operating infrastructure in AWS

  • Strong experience with Kubernetes and container technologies such as Docker.

  • Strong experience with infrastructure-as-code tools, preferably Terraform.

  • Strong understanding of networking, DNS, load balancing, firewalls, certificates, and service-to-service communication.

  • Experience designing CI/CD systems and automated deployment workflows.

  • Experience in monitoring, logging, alerting, SLIs/SLOs, incident response, root cause analysis, disaster recovery, high availability, and resilient deployment strategies.

  • Understanding of public-cloud and container security, including identity and access management, role-based access control, network security, process isolation, secrets management, and certificate management.

  • Experience using Git and GitHub.

  • Experience working with Javascript and/or Typescript

  • Ability to diagnose complex production issues across application, infrastructure, networking, and data layers.

  • Ability to make pragmatic technical decisions that balance reliability, security, cost, speed, and organizational capacity.

  • Strong written and verbal communication skills, including the ability to explain technical risks and tradeoffs clearly.

  • Ability to work independently and collaboratively in an entrepreneurial, ambiguous, and resource-constrained environment.

  • Commitment to Exygy’s mission and to advancing technology that improves access, equity, and trust in public services


Preferred Qualifications


  • Experience operating multi-tenant SaaS platforms or platforms deployed across multiple customer environments.

  • Experience supporting open-source products or products built through collaboration with government agencies.

  • Experience using node.js and NestJS, or similar frameworks.

  • Experience using frontend frameworks such as React and Next.js

  • Experience supporting systems used by local, regional, state, or federal government organizations.

  • Experience improving infrastructure that has accumulated technical debt or grown through multiple implementation models.

  • Experience operating systems that handle sensitive personal information

  • Experience working effectively outside of a large enterprise organization and understanding what it takes to succeed in a small, hands-on team.


Education and Experience


  • Bachelor’s degree in computer science, engineering, information systems, or a related field, or equivalent professional experience.

  • Demonstrated experience performing senior-level site reliability, infrastructure, platform engineering, DevOps, or production-operations work.

  • Experience independently leading complex infrastructure or reliability initiatives from problem definition through implementation and ongoing operation.

  • Experience working cross-functionally with software engineers, product teams, organizational leaders, and external stakeholders.


Metrics for Success


  • Bloom’s critical services have clearly defined and actively monitored SLIs and SLOs.

  • Production and staging environments are more observable, dependable, secure, and consistently configured.

  • Deployment processes require less manual effort, produce fewer avoidable errors, and enable engineers to release, validate, and roll back with greater confidence.

  • Production incidents are detected and resolved more quickly, with fewer recurring incidents.

  • Bloom has a documented infrastructure direction that can support additional jurisdictions without unnecessary custom deployment systems.

  • Infrastructure costs, performance, operational complexity, and reliability risks are visible enough to support informed prioritization.

  • Operational documentation and incident playbooks are current, usable, and regularly referenced.

  • Reliability becomes a shared engineering practice rather than the responsibility of one individual


Payment & Timeline


The annual salary range for this full-time position is $112,000 - $140,000. Exygy’s compensation is benchmarked to national industry averages, not geographic location, ensuring equitable pay across all roles. Our hiring targets align with the median of the salary range for this position, with offers based on role requirements and candidate experience. 

Employees are eligible to enroll in our 401K benefits. We match 100% of the first 3% you defer plus 50% of the next 2% of each paycheck’s eligible compensation (maximum 4% match per paycheck). Healthcare benefits package with options up to 100% coverage toward select medical, dental, and vision plans. Exygy provides up to 10 days of sick leave and flexible PTO. 

Exygy employees may work remotely across the US. Exygy employees' main residence must be within the US. Exygy proudly embraces work/life balance. Full-time employees are expected to work 32-40 hours in a typical week. We aim to hold all internal meetings between 10AM and 3PM PT, Monday-Thursday. We strive to keep Fridays meeting-free, allowing employees flexibility to catch up on their workload or take needed time to rest.


EEO & Commitment to Equity, Diversity, and Inclusion

We are actively seeking to create a diverse and equitable work environment because we believe that it creates a stronger team.


Exygy values a diverse workplace and strongly encourages women, people of color, LGBTQIA individuals, people with disabilities, members of ethnic minorities, foreign-born residents, older members of society, and others from minority groups and diverse backgrounds to apply. Exygy is an equal opportunity employer. We will not discriminate against applicants because of race, color, sex (including pregnancy), sexual orientation, gender identity or expression, age, religion, national origin, disability, ancestry, marital status, veteran status, medical condition, or any protected category prohibited by local, state, or federal laws. All employees and contractors of Exygy are responsible for maintaining a work atmosphere free from discrimination and harassment by treating others with dignity and respect.

Job Purpose 


Remote, US | Engineering | full-time, termed

Exygy has been building technology for the public good for more than 20 years. Founded in 2003 as an impact-focused technology agency, Exygy has delivered life-changing solutions for nonprofits, mission-driven organizations, and governments across the globe. Today, the company is focused on scaling Bloom Housing and CiviForm, two digital tools that help government teams better serve their communities. Our mission has remained the same from the start: to design and build user-centered technology that improves lives.


Bloom Housing is an open-source platform, built by Exygy, that helps governments and housing partners modernize access to affordable housing. By centralizing the housing search and application process into a single portal, Bloom simplifies how residents find and apply for affordable housing, while giving property managers and housing staff a unified way to manage listings, applications, and housing stock all in one place. Bloom currently supports affordable housing access across the 9 Bay Area counties, Detroit, and Los Angeles.


Over the last decade, Bloom has developed from an open-source framework into a housing platform used by cities, counties, and regional housing partners. As Bloom expands to additional jurisdictions, we need reliable, secure, and repeatable infrastructure that can support a growing number of government partners without requiring every implementation to become a custom deployment. 


The Senior Site Reliability Engineer will lead the technical work required to make Bloom’s environments, deployment systems, and operating practices more dependable, scalable, secure, and efficient. You will own and evolve Bloom’s infrastructure and deployment capabilities, manage production and staging environments, establish measurable reliability practices, improve monitoring and incident response, and partner with engineers to strengthen how Bloom is tested, released, and operated.


Operating as a senior individual contributor within a small-scale team, you will merge high-level architecture and future planning with practical execution, troubleshooting, documentation, and operational maintenance. This position demands sound technical discernment, adaptability within a shifting landscape, and a commitment to constructing the frameworks and methodologies you suggest. Furthermore, there may be instances where standalone contributions and feature development are necessary.


Employment Terms


This is a full-time, termed, remote, U.S.-based position. This is a funding-contingent position with an anticipated duration through August 2027, based on the current term of the applicable client contract or grant and the portion of the scope of that contract or grant that requires specifically the work of this role. If the contract or funding is not awarded, is delayed, reduced, modified, terminated, expires, or is not renewed, the position may be modified or eliminated, and employment may end earlier than anticipated, subject to applicable law and any legally required processes. The role may occasionally require travel for team meetings or other business needs. Because this position supports production systems, participation in an on-call rotation is required.


Who Does This Role Report To

Principal Engineer, Bloom


Supervisory Responsibilities


None.


Responsibilities 


Site Reliability & Production Operations


  • Own the reliability, availability, performance, and operational health of Bloom’s production and staging environments.

  • Participate in a sustainable on-call rotation and respond to production incidents, service interruptions, and urgent operational issues.

  • Identify recurring operational problems and implement durable solutions rather than relying on repeated manual intervention.

  • Establish clear escalation paths and operating procedures for outages, degraded service, deployment failures, and security events.

  • Evaluate system capacity, dependencies, failure modes, and operational risks as Bloom grows.

  • Identify and mitigate security risks in Bloom’s infrastructure, deployment systems, and production environments.


Infrastructure, Deployment, and Continuous Integration/Delivery


  • Own and evolve Bloom’s deployment systems from development and prototyping through production delivery.

  • Develop a more standardized, repeatable, and scalable approach to deploying Bloom across jurisdictions and hosting environments.

  • Build and maintain containerized and Kubernetes-based infrastructure where appropriate using infrastructure as code, including Terraform.

  • Reduce deployment complexity, configuration drift, and reliance on manual processes.

  • Evaluate existing and proposed infrastructure against reliability, security, scalability, cost, portability, and maintainability needs.

  • Improve Bloom’s continuous-integration and continuous-delivery systems and practices.

  • Partner with engineers to strengthen automated testing, release procedures, deployment validation, and rollback capabilities.

  • Increase the safety, consistency, and frequency of releases while reducing avoidable production risk.

  • Automate repetitive infrastructure, deployment, and operational tasks.

  • Improve visibility into deployment status, release health, and production impact.

  • Promote shared ownership of reliability and operability throughout the software-development lifecycle.


Observability and Reliability Engineering


  • Define, document, and implement service-level indicators and service-level objectives for Bloom’s critical services.

  • Develop and maintain monitoring, logging, alerting, dashboards, and other observability capabilities.

  • Gather and analyze deployment and production metrics to identify reliability, performance, and configuration improvements.

  • Develop disaster-recovery, backup, restoration, and service-continuity practices.

  • Create clear operational measures that help Engineering and Product make informed decisions about reliability risks and priorities.


Engineering Collaboration & Technical Leadership


  • Collaborate with software engineers, product leaders, delivery team members, and organizational leadership to ensure new features are designed and implemented with reliability and operability in mind.

  • Contribute code to Bloom when needed to improve reliability, performance, security, operability, or platform capabilities.

  • Participate in code reviews and technical refinement of work affecting infrastructure and production systems.

  • Translate reliability, security, and infrastructure risks into clear options and recommendations for technical and nontechnical audiences.

  • Support government partners and internal teams in resolving issues related to Bloom environments and deployments.

  • Create and maintain deployment guides, incident playbooks, architecture documentation, troubleshooting materials, and operational runbooks.

  • Mentor teammates and help build broader organizational knowledge of site reliability and production operations.

  • Work effectively with external vendors, contractors, or government technology teams when responsibilities cross organizational boundaries.


Required Skills, Knowledge, and Abilities


  • Deep experience building and operating infrastructure in AWS

  • Strong experience with Kubernetes and container technologies such as Docker.

  • Strong experience with infrastructure-as-code tools, preferably Terraform.

  • Strong understanding of networking, DNS, load balancing, firewalls, certificates, and service-to-service communication.

  • Experience designing CI/CD systems and automated deployment workflows.

  • Experience in monitoring, logging, alerting, SLIs/SLOs, incident response, root cause analysis, disaster recovery, high availability, and resilient deployment strategies.

  • Understanding of public-cloud and container security, including identity and access management, role-based access control, network security, process isolation, secrets management, and certificate management.

  • Experience using Git and GitHub.

  • Experience working with Javascript and/or Typescript

  • Ability to diagnose complex production issues across application, infrastructure, networking, and data layers.

  • Ability to make pragmatic technical decisions that balance reliability, security, cost, speed, and organizational capacity.

  • Strong written and verbal communication skills, including the ability to explain technical risks and tradeoffs clearly.

  • Ability to work independently and collaboratively in an entrepreneurial, ambiguous, and resource-constrained environment.

  • Commitment to Exygy’s mission and to advancing technology that improves access, equity, and trust in public services


Preferred Qualifications


  • Experience operating multi-tenant SaaS platforms or platforms deployed across multiple customer environments.

  • Experience supporting open-source products or products built through collaboration with government agencies.

  • Experience using node.js and NestJS, or similar frameworks.

  • Experience using frontend frameworks such as React and Next.js

  • Experience supporting systems used by local, regional, state, or federal government organizations.

  • Experience improving infrastructure that has accumulated technical debt or grown through multiple implementation models.

  • Experience operating systems that handle sensitive personal information

  • Experience working effectively outside of a large enterprise organization and understanding what it takes to succeed in a small, hands-on team.


Education and Experience


  • Bachelor’s degree in computer science, engineering, information systems, or a related field, or equivalent professional experience.

  • Demonstrated experience performing senior-level site reliability, infrastructure, platform engineering, DevOps, or production-operations work.

  • Experience independently leading complex infrastructure or reliability initiatives from problem definition through implementation and ongoing operation.

  • Experience working cross-functionally with software engineers, product teams, organizational leaders, and external stakeholders.


Metrics for Success


  • Bloom’s critical services have clearly defined and actively monitored SLIs and SLOs.

  • Production and staging environments are more observable, dependable, secure, and consistently configured.

  • Deployment processes require less manual effort, produce fewer avoidable errors, and enable engineers to release, validate, and roll back with greater confidence.

  • Production incidents are detected and resolved more quickly, with fewer recurring incidents.

  • Bloom has a documented infrastructure direction that can support additional jurisdictions without unnecessary custom deployment systems.

  • Infrastructure costs, performance, operational complexity, and reliability risks are visible enough to support informed prioritization.

  • Operational documentation and incident playbooks are current, usable, and regularly referenced.

  • Reliability becomes a shared engineering practice rather than the responsibility of one individual


Payment & Timeline


The annual salary range for this full-time position is $112,000 - $140,000. Exygy’s compensation is benchmarked to national industry averages, not geographic location, ensuring equitable pay across all roles. Our hiring targets align with the median of the salary range for this position, with offers based on role requirements and candidate experience. 

Employees are eligible to enroll in our 401K benefits. We match 100% of the first 3% you defer plus 50% of the next 2% of each paycheck’s eligible compensation (maximum 4% match per paycheck). Healthcare benefits package with options up to 100% coverage toward select medical, dental, and vision plans. Exygy provides up to 10 days of sick leave and flexible PTO. 

Exygy employees may work remotely across the US. Exygy employees' main residence must be within the US. Exygy proudly embraces work/life balance. Full-time employees are expected to work 32-40 hours in a typical week. We aim to hold all internal meetings between 10AM and 3PM PT, Monday-Thursday. We strive to keep Fridays meeting-free, allowing employees flexibility to catch up on their workload or take needed time to rest.


EEO & Commitment to Equity, Diversity, and Inclusion

We are actively seeking to create a diverse and equitable work environment because we believe that it creates a stronger team.


Exygy values a diverse workplace and strongly encourages women, people of color, LGBTQIA individuals, people with disabilities, members of ethnic minorities, foreign-born residents, older members of society, and others from minority groups and diverse backgrounds to apply. Exygy is an equal opportunity employer. We will not discriminate against applicants because of race, color, sex (including pregnancy), sexual orientation, gender identity or expression, age, religion, national origin, disability, ancestry, marital status, veteran status, medical condition, or any protected category prohibited by local, state, or federal laws. All employees and contractors of Exygy are responsible for maintaining a work atmosphere free from discrimination and harassment by treating others with dignity and respect.

APPLICATIONS DUE WEDNESDAY, SEPT 3 AT 12 noon PT