FluidstackHome
All open roles

OperationsSan Francisco, CAFull Time

Production Engineer, Facilities

San Francisco, CA · New York, NY · Austin, TX · Seattle, WAOnsite

About Fluidstack

Technology is the most important tool humans have found for improving the human condition, but it is not a default good. It is a lever for ideology.

Today the most powerful technology in history is being built: machines that solve problems better than humans can. The most important mission of our generation is to imbue AI with democratic values: error-correcting institutions, freedom of speech, individual liberties. AI trains and runs on massive compute clusters. Whichever ideology builds the infrastructure the fastest is the ideology that will endure. Democracy is delicate and won't survive this technology by default.

We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate

  • Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.

  • Insane urgency. We drive everything forward as fast as possible.

  • Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

  • Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.

The Production Engineering Team

Examples of key problems the team is working on

  • Scaling the systems which make the physical plant observable and operable as one fleet. Building the telemetry, alarm, topology, and health systems that allow operators and automations to see the true state of power, cooling, and environmental infrastructure across every site.

  • Turn facility incidents into a closed repair loop. Automate the path from detection and diagnosis through maintenance, remediation, validation, and return to service so failures do not disappear into handoffs between software, engineering, vendors, and site operations.

  • Bring new sites and equipment into production safely at construction speed. Build repeatable readiness gates, commissioning signals, staged deployments, canaries, and rollback mechanisms for a fleet growing by multiple sites at once.

  • Keep operators ahead of power and cooling risk. Build capacity views, safeguards, anomaly detection, service-health reviews, and operational tooling that identify problems before they affect customers.

Role Scope

  • Carry the facilities production on-call pager and lead incidents involving facility software, telemetry, controls integrations, and automation. Diagnose the failure, coordinate the responsible teams, restore service, and drive the systemic fix.

  • Own the production reliability of the facilities telemetry and alarm platform end to end. Build and operate ingestion, storage, APIs, data-quality checks, actionable alerts, retention, backups, failover, and recovery across industrial protocols and site integrations.

  • Turn diagnosis and repair into pipelines rather than procedures. Build Python or Go tooling for fleet-wide debugging, maintenance workflows, automated validation, incident response, and safe return to service.

  • Own production deployment and runtime management for facilities services, including BMS and EPMS integrations, SCADA platforms such as Ignition, virtual PLCs, demand management.

  • Define and enforce production-readiness standards for new sites, equipment, APIs, telemetry integrations, and controls deployments. Build the tests, canaries, release gates, staged promotion, and rollback mechanisms that define what healthy looks like before launch.

  • Own the operational maturity of every in-scope service. Establish SLOs, capacity plans, health dashboards, runbooks, escalation paths, incident drills, and regular service reviews with product owners and partner teams.

  • Partner with Facilities Software Automation, Controls and Design Engineering, Field Engineering, and Facilities Operations. You make the systems these other teams build in and consume reliable, observable, scalable, and supportable in production.

What We're Looking For

  • You have carried a pager for production infrastructure and can run an incident from first alert through restoration, postmortem, and systemic fix.

  • You have written production automation in Python, Go, or a similar language that replaced a manual operational workflow other teams depended on.

  • You understand observability as an operating system, not a collection of dashboards. You have defined meaningful service health, alerts, SLOs, and review rhythms.

  • You debug across system boundaries. You can follow a failure from a physical sensor or controller through an industrial protocol, data pipeline, API, dashboard, and operator workflow.

  • You treat toil as a bug. If a repair or deployment requires repeated manual steps, you build the safe, repeatable path.

  • You are comfortable with infrastructure as code, Kubernetes, GitOps, deployment pipelines, and production data systems.

  • You move toward ambiguous, high-impact failures and build enough domain knowledge to make good decisions quickly.

  • You work effectively with software engineers, controls and design engineers, field teams, vendors, and site operators without blurring ownership.

  • You use modern AI-assisted engineering tools to investigate systems, write and review code, and reduce time from diagnosis to resolution.

  • For senior or lead-level scope, you have set technical direction, built a reliability roadmap, grown engineers, and balanced interrupt-driven operations with sustained engineering delivery.

  • Bonus: Experience with BMS, EPMS, SCADA, Ignition, virtual PLCs, BACnet, Modbus, OPC UA, time-series databases, data center power or cooling, alarm rationalization, repair automation, or industrial control security.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email careers@fluidstack.io with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Compensation

$173K – $279K • Offers Equity

Ready to help build civilization-scale infrastructure for AI?

Apply for this role

Applying? See how we handle applicant data in our privacy policy, or make a privacy request.