All roles

Infrastructure · Contract-to-hire

Senior DevOps / SRE Engineer

Remote, USA preferred, CET overlap$150k–$200k/yr + potential for equity

About LuxAlgo

LuxAlgo is a full trading & charting platform: state-of-the-art charts, Quant (our coding agent, the intelligence inside LuxAlgo), a growing family of open-source projects (the Vela charting engine, and PineTS, the open-source runtime Pine Script® runs on), and the market-data infrastructure underneath all of it. For over five years our tools have been used by millions of traders.

We're bootstrapped and profitable, small by design, and fully remote. We ship fast with modern tooling, and AI coding agents are everyday collaborators here. We're scaling toward a platform that serves millions of users in real time.

The role

You'll be our first dedicated infrastructure engineer. You own the reliability, performance, security, and cost of a distributed SaaS and market-data platform, and you design the platform layer our next-generation runtime infrastructure runs on.

It's a hands-on role spanning cloud infrastructure, database operations, CI/CD, observability, incident response, and backend reliability. The stack: AWS, Vercel, PostgreSQL and TimescaleDB (on Tiger Cloud), Node.js/TypeScript services, Redis, and Cloudflare.

Our philosophy: managed services first, automation over manual operations, and infrastructure that removes pagers rather than adds them. Expect roughly 60% building and 40% operating, shifting toward building as automation matures.

What you'll do

  • Run production. Operate and improve our distributed production infrastructure across multiple cloud providers.
  • Build what comes next. Design and build our next-generation runtime and data infrastructure.
  • Own the cloud. Manage cloud resources, identities, networking, databases, secrets, and security.
  • Own the database. Operate and optimize PostgreSQL and TimescaleDB: replication, WAL, IOPS, continuous aggregates, compression, connection pooling, and query performance.
  • Make recovery boring. Maintain tested backup, recovery, failover, and disaster-recovery procedures.
  • Ship safely. Own infrastructure as code, GitHub Actions pipelines, deployment safety, health checks, canary releases, and rollbacks.
  • See everything. Build centralized monitoring, logging, dashboards, and actionable alerts.
  • Lead incidents. Drive incident investigation, root-cause analysis, and the long-term corrective work that follows.
  • Go deep in the services. Diagnose and improve Node.js and TypeScript services.
  • Improve continuously. Proactively raise availability, capacity, performance, security, and cost efficiency.
  • Share the pager. Join a three-engineer production on-call rotation: one week in three, with an on-call stipend.

What we're looking for

  • Strong production experience with AWS, Linux, networking, and infrastructure automation.
  • Strong PostgreSQL administration and performance-tuning experience. You've been the person responsible for a production database.
  • Hands-on TimescaleDB knowledge: hypertables, continuous aggregates, compression, background jobs, and replication.
  • Terraform (or similar infrastructure as code) and GitHub Actions in production.
  • Monitoring, alerting, centralized logging, and real incident-response experience.
  • Ability to read and debug Node.js, TypeScript, SQL, HTTP, and WebSocket services.
  • A managed-first mindset: you reach for a managed service before a server, and you can explain when not to.
  • Comfort using AI coding agents daily, and designing automation and runbooks that agents can execute safely.
  • A track record of eliminating operational toil, not just handling it.
  • Clear written communication: your incident updates and postmortems will be read by non-engineers.

Nice to have

  • Vercel, Tiger Cloud, or other managed cloud platforms; bare-metal infrastructure.
  • High-volume time-series systems.
  • Nginx, PM2, Redis, and PostgreSQL connection poolers.
  • OpenTelemetry, Prometheus, Grafana, Datadog, or CloudWatch.
  • Python or Go.

Your first months

  • First 30 days. Join the on-call rotation and ship the observability foundation: dashboards, health checks, alerting, runbooks.
  • By 60 days. Tested backup and disaster-recovery procedures, hardened replication, existing infrastructure captured in Terraform, and cost visibility.
  • By 90 days. Services running on managed containers with canary deploys and rollback, and the platform ready for our next-generation runtime.

What we offer

  • Compensation. $150,000–$200,000 USD per year, with the possibility of equity.
  • Contract-to-hire. We start with a contract engagement and convert to a long-term role together. We say this up front so there are no surprises.
  • Remote, with overlap. Fully remote. USA preferred; wherever you are, you need to comfortably overlap with Central European Time for part of your day, since much of the team works from Europe.
  • Budgets. Equipment and home-office budget; learning and conference budget.
  • A real mandate. A senior team, full ownership of the infrastructure surface, and direct impact on a product traders depend on every day.

How to apply

Apply below. We read every application as it comes in. Include a short note about a production system you made more reliable, and a pager you retired. Questions before applying: joseph@luxalgo.com.

Apply

Apply for this role.

A resume and a few links, that's it. Applications go straight to the team.

Your application goes straight to the team, with no ATS in between.