Every dispute is one move in a repeated game against an adaptive opponent. Build the system that gets better at playing it.
Why This Exists
A federal arbitration system called Independent Dispute Resolution, or IDR, now determines billions of dollars in healthcare payments each year. Providers win the vast majority of disputes, yet most eligible claims are never filed. The process is manual, fragmented, and resource-intensive, and most providers don't have the infrastructure to pursue what they're owed.
The No Surprises Act created the framework, and the market already exists. Today it runs on spreadsheets, consultants, and static playbooks. We're building the first intelligent system designed to operate inside it.
Automating the paperwork is table stakes. The interesting part is the second half: IDR is baseball-style arbitration, where each side submits one number and an arbitrator picks one. No splitting the difference. That means every submission is a bet, and every outcome is a signal about how to bet better next time.
This role owns that.
Why This Is Hard (and Interesting)
You're building a decision system in an adversarial, regulated, sparse-data environment.
The same payers appear repeatedly and they change behavior when they notice patterns. Arbitrators rotate and have their own tendencies. Regulations shift under you. You have to make a call on every claim before you have statistically comfortable data, and being wrong costs a provider real money.
The system also has to be defensible. When a customer asks why we submitted a particular number, "the model said so" is not an acceptable answer. You need reasoning you can explain to a hospital CFO.
Concretely: which claims to contest, what to offer, what evidence to include, how to write the argument, and how all of that should change based on the payer, the arbitrator, the procedure codes, and everything we've learned so far. Deterministic rules and LLM reasoning both have a role. Figuring out which does what is your call.
Who We Are
Recourse is being built in partnership with 25M Health, a healthtech venture firm. We have institutional backing, a shared platform team spanning engineering, strategy, design, and back-office, and early access to large provider systems.
We are actively filing disputes for real customers, including a large multi-facility health system and a litigation-finance partner with hundreds of millions in claim value. This is a funded, validated opportunity with real customers and real data.
We are a small, nimble team. We move quickly and we value clarity over theater. We want this to be the best work of your career. The stretch you look back on as the one where you shipped real things, with people who raised your game, on something that mattered.
We care about clear thinking, high ownership, intellectual honesty, and direct communication. We believe operations, product, and engineering should operate as one pod, not three functions. We want the machines to do machine work, and the humans to do their best work.
The Role
You'll own the intelligence layer of the product. You will:
- Build the decision engine that determines which claims to contest and at what offer amount
- Design the LLM systems that generate arguments: medical necessity, patient acuity, market comparables, the full submission
- Build evaluation infrastructure. Golden datasets, offline evals, and the ability to know whether a change actually improved outcomes
- Architect the split between deterministic rules and model reasoning, and defend where you drew the line
- Close the feedback loop from arbitration outcomes back into the system, so every decision makes the next one better
- Build for explainability. Every recommendation needs reasoning a customer would accept
- Partner with operations and payer strategy, who see patterns in the claims before the data does
We build with Claude Code daily and run a Codex review pass on every slice. The stack is TypeScript, Next.js, Prisma, and PostgreSQL on Google Cloud, with LLM reasoning throughout.
Who You Are
You've shipped LLM systems that people depend on. Not demos. Production systems with evals, guardrails, and a real answer for what happens when the model is wrong.
You think about systems, not prompts. Prompt engineering is a component. The interesting work is architecture: what's deterministic, what's learned, how they interact, and how the whole thing improves over time.
You're rigorous about evaluation. You know that "it seems better" isn't evidence. You build the measurement before you build the feature.
You are AI-pilled and current. The landscape moves monthly. You track it because you want to see the next shift before anyone else, and you have opinions about what's real versus hype.
You have a bias to action. You don't default to no. Speed of iteration over polish of iteration. You start, you learn, you fix things in motion. Most decisions are reversible and do not need extensive study.
You are intellectually honest. You seek out evidence that disconfirms your approach. You say so when you're wrong. You use plain language. You respectfully challenge decisions you disagree with, and once a decision is made, you commit.
You put the team first. You are reliable and fully invested. You take your vacations. You check on your teammates. You help build a culture where people do their best work because they are supported, not squeezed.
What You Bring
- 5+ years engineering, with meaningful recent time building LLM-powered or ML-driven products in production
- Real experience with evaluation: building golden datasets, running offline evals, measuring whether changes helped
- Comfort across the stack. You can ship the thing end to end, not just the model layer
- Strong instincts for where probabilistic reasoning helps and where it introduces unacceptable risk
- Comfort with TypeScript or Python. The specific stack matters less than the ability to pick things up
- Experience building in regulated or security-sensitive environments (HIPAA, SOC 2, PCI, financial controls) is a plus
- A preference for small teams and early-stage chaos over mature org charts
Strong plus, not required: healthcare data, claims, or any adversarial decision domain (fraud, risk, pricing, trading). If you have it, you'll move faster. If you don't, we'll teach you.
Sound judgment, technical depth, and ownership mindset are required. Grit matters more than pedigree.
Why This Role
Most applied AI jobs are wrapping a model around an existing workflow. This one is different: the intelligence is the product, the feedback loop is real and measurable, and the outcome is dollars recovered rather than an engagement metric.
You'd also be first here. Nobody in this market has built a genuine decision-intelligence layer on arbitration outcomes. The data exists and almost nobody is doing the work. You'd have unusual latitude to decide what this becomes.
If you want this to be the most memorable stretch of your career, where you shipped something real, with a team you respected, in a domain that actually matters, this is the seat. Please apply even if you don't fit 100% of these requirements. We would like to talk.
Equal Opportunity
Recourse is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other characteristic protected by law. We believe the best teams are built from people with different backgrounds and perspectives, and we're committed to creating an environment where everyone can do their best work.