Agents that do the work. A platform that learns from every run.
Recursion is the managed agent platform for enterprise work. Set a goal; a coordinator plans it, specialists work it in parallel across your systems, and a grader scores the outcome. Every graded run becomes memory, skills and training signal, so the fleet gets better every week.
- Tracker0
- Cluster0
- Evals0
- GitHub0
- Notion0
- Slack0
- Docs0
One goal in. A fleet of specialists out.
Recursion Managed Agents are persistent, long-horizon, asynchronous agents that run as one fleet. Hand them a goal, not a script.
Multi-agent by design
A coordinator plans the work and spawns specialists in waves. Hard problems get more agents, not longer prompts.
Built for long-horizon work
Sessions run for hours or days, pause on dependencies, and pick up where they left off. Nothing is lost between turns.
Every outcome graded
A grader scores each session against your rubric. Scores decide what the fleet remembers and what gets trained.
Any model, per agent
Claude, GPT, Gemini and open-weight models side by side. Swap in your fine-tuned specialist when it wins.
A coordinator, specialists, and one graded outcome
Every agent runs on its own track. A coordinator plans the goal and posts tasks; specialists claim them in waves, work across your systems, and report back. A grader agent scores the session. Sessions persist for hours or days and pick up where they left off.
- Tracker0
- Cluster0
- Evals0
- GitHub0
- Notion0
- Slack0
- Docs0
Every session in the fleet, live
The fleet view shows every session this week as one board: completed, running, waiting, and how each was scored. Persistent, long-horizon, asynchronous agents running as one fleet, each on the model it is best at.
Real work is too big for one agent, one prompt, one try
Enterprise work spans dozens of systems, runs for hours, and rarely fits in a single context window. A general agent in a chat box can look strong in a demo and still fail in production: it loses the thread across long workflows, misses edge cases, and treats every run as the first.
The core issue isn't model capability. It's that nothing around the model is built for the work: no way to split it across agents, no memory between runs, no grader saying whether the outcome was actually good.
Without a tight loop between execution, measurement, and learning, every run is just another task completed, not another signal the system gets to learn from.
A unified RL platform for specialist models
Recursion connects environments, evaluation, and training into a closed-loop reinforcement learning system. Rather than treating deployment as the endpoint, production becomes the training surface.
RL environments that reflect real work
Recursion turns workflows, tools, policies, and edge cases into executable environments for RL training and evaluation. Powered by WorldSim, these environments recreate the full enterprise software stack — with configurable world effects that generate diverse, realistic scenarios at scale.
Explore RL environments0554-crude-pipeline-lbo-waterfall
Task runs
Show all 10 runs
# Pipeline LBO Returns Model with Multi-Tier Waterfall
## Context
You are an associate at a private equity fund preparing an investment committee update for a crude oil pipeline project acquired through an LBO. The task is to build an analyst-ready workbook that connects operating performance, the pre-computed debt schedule, exit valuation, and sponsor-management waterfall.
Save the completed workbook to /workspace/output/model.xlsx.
## Workbook to build
- 1. Financing Assumptions
- 2. Operating Assumptions
- 3. Debt Schedule
- 4. Model
- 5. Returns
Task files
OptionalFiles accessible to the model at the container mount path.
inputs.xlsx
/workspace/files/problem/inputs.xlsx
source.xlsx
/workspace/files/problem/source.xlsx
Evaluation systems that measure real execution
The difference between agents that stagnate and agents that improve is measurement. Recursion builds evaluation systems that score intelligence and skill at every level — final outcomes, intermediate decisions, and execution quality — so every run generates the signal your models need to get better.
specialist-model-evals
rankedAccuracy, efficiency, and latency across model candidates
Leader
GLM 5.1 FTv2
Accuracy
Pass@4 score on held-out financial analysis tasks
GLM 5.1 FTv2
GPT 5.5
Claude 4.8
GLM 5.1 FTv1
GLM 5.1 base
Token efficiency
Normalized useful work per token
GLM 5.1 FTv2
GLM 5.1 FTv1
GLM 5.1 base
Claude 4.8
GPT 5.5
Latency
Normalized responsiveness, raw latency shown
GLM 5.1 FTv2
GLM 5.1 FTv1
GLM 5.1 base
GPT 5.5
Claude 4.8
A training loop that compounds from real work
Every rollout produces graded trajectories that can feed fine-tuning and reinforcement learning. The result is a specialist model that improves from enterprise execution signals instead of synthetic benchmarks alone.
financial-agent-glm-5.1-v7-fits6h
completedGRPO run - GLM 5.1
Run result
86% Pass@4 score (+30%)
Training run summary
Base
8h 20m
Training
11h 45m
Tuned
8h 55m
| Metric | Value |
|---|---|
| Source | 4,300 financial analysis tasks |
| Task family | DCF, LBO, acquisition, projection |
| Base model | GLM 5.1 |
| Training method | GRPO |
| Compute | GKE, H100 cluster |
| Endpoints | Baseline and tuned |
Evaluations
4 attempts per task
Problem
Baseline
Tuned
Delta
0195-acme-industrials-dcf-irr
0.55 -> 0.85
0248-pumptech-acquisition-rollup
0.60 -> 0.89
0515-hotel-renovation-noi-yield
0.51 -> 0.83
0843-hartwell-plant-expansion-dcf
0.54 -> 0.83
0120-smasco-revenue-forecast
0.52 -> 0.85
0631-proactivate-lbo-exit-returns
0.57 -> 0.87
From goal to specialist model — continuously
Launch
Set a goal. A coordinator plans it, spawns specialists, and works across your systems.
Grade
A grader scores every session against your rubric: outcome, process, and execution quality.
Remember
Graded sessions become memories, notes, and skills the fleet recollects before acting. Improvement without retraining.
Generate
Turn the hardest, lowest-scoring sessions into RL environments with long-tail scenarios, powered by Horizon.
Train
Run RL and fine-tuning on graded rollouts to produce a specialist model for your tasks.
Serve
Swap the specialist into the fleet, per agent. Every run from here feeds the next cycle.
Every cycle captures additional expertise. Every execution generates a new signal. Every improvement compounds your proprietary edge.
The specialist always beats the generalist
Fine-tuned on the fleet's graded runs, specialist models out-resolve frontier generalists on your tasks at a fraction of the cost. Labelbox Horizon's frontier-class RL environments are trusted by over 90% of leading US AI labs; fine-tuned models deliver a 5–10× lower cost per task.
A customer service agent resolved more tickets with fewer hallucinations, faster responses, and lower inference cost
Resolution rate
Finetuned OS model
84%GPT-5.5
76%Claude Opus 4.8
73%Reduction in hallucinations
Finetuned OS model
72%GPT-5.5
46%Claude Opus 4.8
41%Cost per 1,000 support tickets
Finetuned OS model
$32GPT-5.5
$158Claude Opus 4.8
$176Time to first token
Finetuned OS model
0.42sGPT-5.5
1.10sClaude Opus 4.8
1.28sPrivacy and security for enterprise agents

Your advantage is the learning loop, not the model
As foundation models improve and become accessible to everyone, sustainable advantage shifts away from the model itself. The organizations that win won't simply deploy intelligence; they'll own the learning loops that transform their expertise, workflows, and decisions into durable competitive advantage.
Every organization has unique workflows, judgments, and domain expertise embedded in its people. The question is whether that expertise remains trapped inside individuals or becomes a compounding organizational asset.
Recursion is designed to close that gap.
Try Recursion on your most ambitious tasks
Tell us the goal. We'll show you what a Recursion session looks like on it.