-
Location: Anywhere in LATAM
-
Work Model: 100% Remote
-
Contract Type: Contractor
-
Project: AI Platform Infrastructure & Observability
-
Industry: Healthtech
-
Time Zone: Availability to collaborate with US and LATAM teams
-
English Level: B2 / C1
-
Seniority: Senior
At Darwoft, we build digital products that create real impact. We are a Latin American technology company working with international organizations to develop scalable, innovative, and human-centered solutions.
Our remote-first culture is based on trust, collaboration, continuous learning, and technical excellence.
We are looking for a Senior Site Reliability Engineer with strong experience in observability, cloud infrastructure, Kubernetes, and production reliability.
You will be responsible for the operational health and visibility of AI-powered systems, including LLM applications, AI API gateways, token-consumption pipelines, and model-serving infrastructure.
The role combines traditional SRE practices with emerging AI platform operations. You do not need experience with every AI tool available, but you should understand the current AI ecosystem and the challenges of operating LLM-enabled services in production.
-
Design and operate gateway infrastructure for LLM and AI API traffic.
-
Manage routing, authentication, throttling, rate limiting, retries, and traffic shaping.
-
Build observability for token consumption, model latency, API usage, costs, errors, retries, quotas, and provider availability.
-
Create dashboards and alerts using Prometheus, Grafana, Loki, CloudWatch, or similar tools.
-
Instrument LLM applications to capture prompt, completion, and request telemetry securely.
-
Define and maintain SLIs, SLOs, alerting strategies, and error budgets.
-
Build resilient architectures using fallback models, provider redundancy, caching, circuit breakers, and graceful degradation.
-
Lead incident response for provider outages, gateway saturation, latency issues, quota exhaustion, and unexpected cost increases.
-
Automate infrastructure and operational workflows using Terraform, CI/CD, and scripting.
-
Maintain AWS and Kubernetes environments supporting AI-powered applications.
-
Collaborate with AI, engineering, platform, security, and FinOps teams.
-
Mentor engineers and promote observability and reliability best practices.
-
7+ years of experience in SRE, DevOps, Platform Engineering, Cloud Infrastructure, or similar roles.
-
Strong experience operating distributed and highly available production systems.
-
Advanced hands-on experience with AWS, ideally including Bedrock, SageMaker, Lambda, API Gateway, EKS, CloudWatch, IAM, and networking services.
-
Strong production experience with Kubernetes and containerized workloads.
-
Experience building observability stacks with Prometheus, Grafana, Loki, CloudWatch, or equivalent technologies.
-
Experience defining application-level and API-level metrics.
-
Solid understanding of API gateway patterns, routing, authentication, throttling, retries, and traffic monitoring.
-
Experience defining SLIs, SLOs, alerts, and error budgets.
-
Understanding of LLM applications, AI APIs, token-based consumption, latency, quotas, and cost monitoring.
-
Proficiency in Python, Go, Bash, or a similar language.
-
Experience with Terraform or another Infrastructure as Code tool.
-
Experience building and maintaining CI/CD pipelines.
-
Strong troubleshooting skills across AWS, Kubernetes, networking, APIs, and distributed systems.
-
Experience participating in or leading production incident response.
-
English proficiency at B2 level or higher.
-
Experience with Kong AI Gateway, Portkey, LiteLLM, AWS Bedrock, or similar platforms.
-
Familiarity with OpenTelemetry and distributed tracing.
-
Experience monitoring prompts, completions, token usage, model latency, and provider errors.
-
Knowledge of AI FinOps, cost allocation, token budgets, usage forecasting, or anomaly detection.
-
Experience with multi-provider routing, fallback models, semantic caching, or service mesh technologies.
-
Experience in healthtech or another regulated industry.
-
Contractor agreement with payment in USD
-
100% remote work
-
Argentina's public holidays
-
English classes
-
Referral program
-
Access to learning platforms
Explore this and other opportunities at:
www.darwoft.com/careers