Type: Full-time permanent contract or an Internship with a potential follow-up offer
Location: San Francisco or remote with future relocation to San Francisco (sponsored)
Start: ASAP
About us
We're building state-of-the-art context compression. Our mission is to become the "Cloudflare for LLMs" — a compression layer embedded into most LLM pipelines by default.
We're a team of ex-EPFL MSc/PhDs from dlab. We started by publishing papers, then got into YC and started making money helping companies cut their LLM costs.
We run the business like a research lab: form hypotheses, kill the ones that don't work, double down on the ones that do.
About you:
A cracked full-stack engineer who enjoys a high-paced startup environment, takes pride in what they build and owns it end to end.
Python/basic ML Ops skills. Experience in scaling AI infra products is a plus.
Proactive, strong communicator with fast response time, team player
Tech Requirements
Strong backend engineering fundamentals
Experience with concurrency and distributed systems
Ability to work across systems (Python + light frontend)
Excellent Claude Code (or similar) user
Nice to have
Open-source contributions
Startup experience
OAuth / API auth flows
Stack
Backend: Python, FastAPI, PostgreSQL (Supabase), Redis, AWS
Frontend: Next.js, React, TypeScript
Tools: GitHub, Docker, Sentry, GitHub Actions
Interview process
1. Intro call (20 min)
2. Practical technical interview (60 min)
3. Culture interview (30 min)