Improvado is an AI-powered marketing intelligence platform trusted by enterprise brands like ASUS, Activision, Docker, and H&R Block. We raised $34M Series A and are scaling fast — which means our infrastructure needs to be rock-solid. We're looking for a DevOps/SysOps engineer who takes ownership and keeps things running.
What you'll do
-
Manage and evolve our cloud infrastructure on AWS/Azure (Kubernetes) - ensure cluster reliability: capacity planning, autoscaling, incident response, and post-mortems
-
Design and architect scalable, reliable infrastructure to support AI-driven analytics and data processing at scale
-
Build and maintain Helm charts, Terraform, and multi-environment setups
-
Own monitoring and alerting across the stack (Prometheus/Mimir, Grafana, CloudWatch)
-
Administer and optimize storages: PostgreSQL, ClickHouse, Redis
-
Support message brokers: RabbitMQ, AWS SQS, Temporal
-
Drive infrastructure security: secrets management, IAM policies, vulnerability scanning, and encryption
-
Own network design, configuration — VPCs, subnets, firewalls, load balancers, VPNs
-
Monitor and optimize cloud costs — identify waste, right-size resources, and report on spend efficiency
-
Participate in on-call rotation
What we're looking for
-
5+ years in a DevOps/SRE role
-
Solid hands-on experience with most of our stack (80%): AWS/Azure/GCP, Kubernetes (cluster design, capacity planning, and reliability at scale), Helm charts, CI/CD pipelines (GitHub CI), Terraform, PostgreSQL, ClickHouse, Redis, RabbitMQ, AWS SQS, Temporal, and monitoring/alerting stacks (Prometheus/Mimir, Grafana, CloudWatch)
-
AI-assisted development in practice — actively uses Claude Code / agentic coding to ship, knows where AI-generated code needs validation
-
Strong Linux fundamentals
-
Bash and/or Python scripting
-
Comfortable with ambiguity and high-velocity, fast-shifting priorities — we move like a startup, not a committee
-
Detail-oriented and pedantic, with sharp attention to detail in high-volume, high-stakes systems — accountable and genuinely invested in the quality of your work
-
EST timezone availability
-
Fluency in English
Nice to have
-
Golang or Python development background
-
HashiStack: Vault, Packer
-
ClickHouse administration
-
Experience supporting data-intensive pipelines
What we offer
-
Remote-first environment
-
20 days PTO + US holidays
-
Optional relocation support to Latin America
-
Modern AI-native tech stack
-
Stock options
-
Professional development reimbursement
-
A genuinely fun, transparent startup culture
dOPxGCqsbY