About Echelon
Echelon is building the AI platform for Business Operations. Our goal is to automate the knowledge work of BizOps so one exceptional operator can deliver the leverage of an entire team, as Ramp and Rippling have done for Finance and HR.
Our platform combines passive process mining with a living ontology that maps how work happens across people, systems, documents, decisions, and outcomes. It connects structured and unstructured data trapped in fragmented, duplicative, and legacy enterprise systems, then turns that context into automated workflows and AI agents.
We are an early-stage company tackling a difficult technical problem at high speed. Engineers work directly with founders and customers, make decisions with incomplete information, ship production systems, and own the results. The pace, rate of change, and standards are high.
The mandate
Build the agent architecture that turns enterprise context into reliable action. You will design agent loops, tool-use patterns, evaluation systems, and multi-agent coordination that work on ambiguous, high-stakes operational problems rather than controlled demos.
What you'll own
Design and ship production agents that plan, retrieve context, call tools, generate artifacts, and complete multi-step work.
Build orchestration patterns for specialized agents and agent swarms, including task decomposition, delegation, shared state, recovery, and human approval.
Create evaluation suites for task success, factuality, tool selection, safety, latency, and cost.
Improve context engineering across prompts, retrieval, memory, structured state, documents, and the Echelon ontology.
Build guardrails for permissions, tenant isolation, sensitive data, unsafe actions, and unreliable model output.
Compare models and prompting strategies with production traces and measured outcomes rather than intuition.
Partner with Product Engineering and domain experts to turn real BizOps workflows into dependable agent capabilities.
Own failures in production and improve the architecture until the same class of failure cannot recur.
What you bring
4+ years building production software, including substantial hands-on work with LLM or ML-powered products.
Evidence that you have shipped an agentic system beyond a prototype and improved it using production feedback or evaluations.
Strong Python and/or TypeScript engineering, including APIs, tests, data modeling, asynchronous work, and observability.
Experience with tool calling, structured generation, retrieval, model evaluation, and multi-step workflows.
Strong judgment about when to use a model, deterministic code, a human checkpoint, or a simpler product interaction.
Comfort debugging nondeterministic behavior and reducing it to measurable, reproducible failure modes.
High agency, direct communication, and comfort owning outcomes under changing requirements.
Useful experience
Multi-agent systems, long-running agents, computer use, code execution, or planning and reasoning systems.
Vercel AI SDK, OpenAI, Anthropic, Mistral, AI Gateway, Braintrust, or comparable platforms.
Knowledge graphs, ontologies, process mining, entity resolution, or temporal data.
Enterprise security, least-privilege tool execution, approvals, audit trails, or regulated workflows.
Sandboxed execution, document intelligence, browser automation, or multimodal models.
Benefits
Fully covered health, dental, and vision insurance.
401(k) plan.
Team lunches and dinners in the office.
Unlimited PTO.