Senior Platform Engineer (Backend & Infra)
Our backend works. It was built by a small team moving fast, and it now needs to scale from 100 to 1,000 facilities and beyond. It also has to become safe enough to act on its own. Dosing changes, pump sequencing, and setpoints are increasingly executed by software on live plants, which means every action must be observable and every deploy reversible. This role owns that safety rail.
You'll be our first dedicated platform hire, writing Node.js services and the Terraform they run on, and setting the standard for how infrastructure, delivery, and reliability work here.
You'll inherit real systems in production (alert engine, event log, device management, edge-to-cloud ingestion) and a clear mandate: make them boring to operate, then make them scale 50x.
Infrastructure and delivery (core):
Our AWS footprint as code — Terraform or CDK, no more click-ops — with environments that are reproducible and provably identical
CI/CD from commit to production: build, test gates, staged rollout, and rollback
Containers, orchestration, and the deployment story for services and edge components alike
Secrets, IAM, network boundaries, and a cloud bill that is measured and managed line by line
Backend services (with the team):
Node.js/TypeScript services and APIs alongside other backend engineers; you review their PRs and they review yours
The edge-to-cloud ingestion path: MQTT at fleet scale, buffering, replay, backpressure, and what happens when a plant's link drops for six hours
The alert and event systems that decide what a plant operator sees, and what wakes someone up
The action path: how a recommendation from our intelligence stack becomes a verified, reversible command on a plant, with guardrails at every hop
Contracts between edge, cloud, and our intelligence stack: agreed before build, versioned, enforced
Cloud and edge deployment workflows, including deployment of sensor configuration files across a fleet of edge devices currently managing ~35,000 sensors
Reliability and operations:
Observability as a product surface, not an afterthought: metrics, traces, and logs that answer "is the fleet healthy?" immediately
SLOs on the paths that matter, alerts that fire on symptoms rather than noise, and an on-call rotation you'll help design so it stays humane as the fleet grows
Incident response and blameless postmortems that actually change the code
Load and failure testing against 50x device volume before the contract lands, not after
Elevating how we ship:
Test coverage and CI gates on all of our services
Runbooks and architecture notes that let anyone on the team operate what you build
Paved paths: the easy way to add a service should also be the correct way
KPI Expectations
This role is not for you if:-
You're looking for a greenfield stack. This is real code with real history, and the job is making it excellent.
You want to stay purely on infrastructure and never touch product code, or purely on product and never be accountable for it in production
You'd rather move fast now and add tests later
You're not curious about how a wastewater plant actually works. You'll be visiting one in your first month.
Recommended Jobs
Posted just now
Posted 3 hours ago
Posted 3 hours ago
Posted 21 hours ago

