We're looking for a Senior Agentic AI Engineer to design, build, evaluate, and continuously improve sophisticated AI systems for a fast-moving enterprise AI platform. You will work across multi-agent architectures, context graphs, memory systems, model routing, and automated evaluation frameworks to build production systems that support complex, real-world business decisions.
This is not an AI API integration or prompt-engineering role. We are looking for someone who thinks deeply about how models, agents, tools, context, memory, and evaluation systems work together and can turn that understanding into reliable production systems. The ideal engineer combines strong software engineering fundamentals with deep curiosity, experimentation, and hands-on experience building modern agentic AI systems. If that describes you, this role is a strong fit.
Why You'll Want to Join
-
You will be paid in USD (bi-monthly: every 15th and 30th)
-
Paid Time Off in accordance with company policy
-
Observance of Holidays per company guidelines
-
100% remote setup so you can work wherever you're most productive
-
This role requires availability during US business hours
-
High ownership over AI architecture and technical direction at an early, high-leverage stage
-
Work directly with enterprise customers and see the real impact of what you build
What You'll Work On
Agentic AI Architecture
-
Design and build production-grade agentic and multi-agent systems
-
Architect specialized agents with clearly defined roles, tools, context, permissions, and decision boundaries
-
Design orchestration and routing strategies across agents, tools, models, and workflows
-
Build systems that intelligently use multiple LLMs and model providers depending on the task
-
Determine which models should power specific agents based on quality, reasoning capability, reliability, latency, and cost
-
Design failure handling, fallbacks, guardrails, and human-in-the-loop mechanisms
Agent and Context Graph Interactions
-
Design and manage how agents interact with context graphs, knowledge graphs, memory systems, and other context sources
-
Determine what information and which portions of a context graph individual agents should be able to access
-
Design context retrieval, filtering, ranking, and permission strategies based on each agent's role and task
-
Build and evaluate interactions between agents and nested or interconnected context graphs
-
Determine how much context an agent needs without unnecessarily increasing tokens, latency, or noise
-
Design systems that dynamically provide agents with the most relevant context for a specific task
-
Evaluate how changes to context availability affect agent accuracy, reasoning, reliability, and performance
AI Evaluation and Reliability
-
Design sophisticated evaluation frameworks for agentic AI systems
-
Build evals across model, agent, and context graph interactions
-
Systematically evaluate which models perform best for specific agents, tasks, and workflows
-
Evaluate which context, memory, or graph information should be available to different agents
-
Implement adversarial testing, model-to-model critique, LLM-as-judge, regression testing, and systematic evaluation
-
Identify failure modes and use evaluation results to continuously improve system architecture
-
Optimize AI systems across quality, accuracy, reliability, latency, token usage, and cost
-
Prefer measurable experimentation and evaluation over assumptions when making architectural decisions
Context, Memory, and Knowledge Systems
-
Design context and memory architectures for autonomous and semi-autonomous agents
-
Build with RAG, embeddings, retrieval systems, vector databases, memory systems, context graphs, and knowledge graphs
-
Design how information flows between context systems and individual agents
-
Determine how context should be retrieved, structured, filtered, and updated
-
Build context systems that support different agents, tasks, and levels of access
-
Balance context quality and completeness against token usage, latency, reliability, and cost
AI Harness Engineering
-
Design, customize, and extend AI harnesses and agent development environments
-
Work deeply with tools such as Claude Code, Claude Agent SDK, Codex, Cursor, MCP, agent frameworks, and similar technologies
-
Build custom instructions, skills, tools, context systems, feedback loops, and evals around AI models
-
Understand when existing AI tooling is sufficient and when customized harnesses or workflows are needed
-
Use AI extensively throughout the engineering lifecycle while maintaining strong technical understanding, ownership, and code quality
AI Research and Experimentation
-
Maintain an active, hands-on approach to staying current with rapidly evolving AI capabilities
-
Regularly explore new models, research, agent architectures, frameworks, and development tools
-
Turn promising research and emerging capabilities into experiments and prototypes
-
Evaluate whether new approaches can meaningfully improve existing production systems
-
Continuously evolve engineering practices as the AI ecosystem changes
-
Demonstrate curiosity and openness to challenging existing approaches rather than relying only on established patterns
Production Engineering
-
Turn ambiguous business problems into reliable production systems
-
Build and maintain backend services, APIs, databases, data pipelines, and asynchronous workflows
-
Contribute across the product stack when required including integrations and product-facing applications
-
Deploy and operate AI systems with appropriate testing, observability, logging, and monitoring
-
Design for enterprise requirements including security, permissions, data isolation, privacy, and reliability
What You Bring
-
5 or more years of professional software engineering experience including ownership of production systems
-
Deep, hands-on experience building production agentic AI and LLM systems beyond prototypes or simple API wrappers
-
Strong experience designing agents, multi-agent workflows, orchestration, and tool use
-
Production experience working across multiple LLMs and model providers with a strong understanding of model selection, routing, quality, reliability, latency, and cost tradeoffs
-
Advanced understanding of AI evaluation systems, experimentation, and failure analysis
-
Experience with adversarial evaluation, model-to-model critique, automated evals, or comparable approaches
-
Strong understanding of agent, model, and context graph interactions
-
Experience designing or managing how agents interact with context graphs, memory systems, knowledge graphs, or complex retrieval architectures
-
Deep understanding of context engineering, RAG, memory, retrieval, and context management
-
Experience with AI harness engineering or deeply customized AI development workflows
-
Hands-on experience with tools such as Claude Code, Codex, Cursor, Agent SDKs, MCP, or similar
-
Strong software engineering fundamentals with the ability to build reliable production systems independently
-
Strong systems thinking and ability to reason about second-order effects across interconnected AI systems
-
High ownership and comfort operating in ambiguous, fast-moving environments
-
Strong curiosity and willingness to continuously experiment with new AI approaches
-
Clear written and verbal English communication
Nice to Have
-
Hands-on experience with knowledge graphs, context graphs, Neo4j, or other graph databases
-
Experience building dynamic model-routing or agent-routing systems
-
Experience building automated AI evaluation infrastructure at scale
-
Experience with PostgreSQL, Redis, vector databases, and advanced search and retrieval
-
Experience with Python and or TypeScript
-
Frontend experience with React, Next.js, or similar frameworks
-
Experience with Docker, Kubernetes, cloud infrastructure, and infrastructure as code
-
Experience deploying AI systems inside enterprise VPCs or private environments
-
Experience with enterprise security, permissions, and sensitive data
-
Founding engineer or highly autonomous early-stage engineering experience
How to Apply
Please include:
-
Your updated resume
-
A GitHub link, portfolio, or examples of sophisticated agentic AI systems you have built in production
-
A short Loom video (1 to 2 minutes) introducing yourself and walking through a sophisticated agentic AI system you personally designed and owned — covering the agent architecture, model selection and routing, context and memory design, evaluation methodology, failure modes encountered, and what you personally built versus what was inherited
Only candidates who submit both a portfolio and Loom video will be moved to the next step of the hiring process.
If you think deeply about how models, agents, context, memory, and evaluation systems work together, challenge assumptions through experimentation rather than intuition, and want to own the architecture of enterprise AI systems that influence real decisions at scale, this role gives you the ownership and the direct impact to do your best work.