Agents that do the work. A platform that learns from every run.

Recursion is the managed agent platform for enterprise work. Set a goal; a coordinator plans it, specialists work it in parallel across your systems, and a grader scores the outcome. Every graded run becomes memory, skills and training signal, so the fleet gets better every week.

Recursion· Managed AgentsOne goal, a team of agents, a graded outcome
A team of eight agents reads out a finished 48-run training sweep.
Every agent on its own track, on the model it is best at, across the tracker, cluster, evals, GitHub, Notion, Slack and Docs. A grader agent scores the outcome.
sweep-2291 · curriculum-v3 · 48 runs · 16 configs × 3 seeds · 1.3B
Readout: does curriculum-v3 beat the baseline on reasoning evals at 1.3B?
Check run health, verify the evals, compare to baseline across seeds, and promote the winning configs to the 7B run.
running
8 agents · 4 models
outcome 0.91 · 2 configs promoted · 7B launched
  • Tracker0
  • Cluster0
  • Evals0
  • GitHub0
  • Notion0
  • Slack0
  • Docs0
Sweep coordinatorClaudecoordinatorrunningRun-health checkerGeminispecialistwaitingEval analystClaudespecialistwaitingPrior-work mineropen weightsspecialistwaitingConfig differGPTspecialistwaitingReviewerGPTreviewerwaitingReport authorClaudespecialistwaitingOutcome graderClaudegraderwaiting
coordinator · Tracker: 48 runs, 16 configs × 3 seeds · plan: 7 tasks across 4 specialists
Managed Agents

One goal in. A fleet of specialists out.

Recursion Managed Agents are persistent, long-horizon, asynchronous agents that run as one fleet. Hand them a goal, not a script.

Multi-agent by design

A coordinator plans the work and spawns specialists in waves. Hard problems get more agents, not longer prompts.

Built for long-horizon work

Sessions run for hours or days, pause on dependencies, and pick up where they left off. Nothing is lost between turns.

Every outcome graded

A grader scores each session against your rubric. Scores decide what the fleet remembers and what gets trained.

Any model, per agent

Claude, GPT, Gemini and open-weight models side by side. Swap in your fine-tuned specialist when it wins.

Works withClaudeGPTGeminiLlamaQwenyour fine-tuned specialist

A coordinator, specialists, and one graded outcome

Every agent runs on its own track. A coordinator plans the goal and posts tasks; specialists claim them in waves, work across your systems, and report back. A grader agent scores the session. Sessions persist for hours or days and pick up where they left off.

Recursion· Managed Agentssession · sweep-2291 readout
sweep-2291 · curriculum-v3 · 48 runs · 16 configs × 3 seeds · 1.3B
Readout: does curriculum-v3 beat the baseline on reasoning evals at 1.3B?
Check run health, verify the evals, compare to baseline across seeds, and promote the winning configs to the 7B run.
running
8 agents · 4 models
outcome 0.91 · 2 configs promoted · 7B launched
  • Tracker0
  • Cluster0
  • Evals0
  • GitHub0
  • Notion0
  • Slack0
  • Docs0
Sweep coordinatorClaudecoordinatorrunningRun-health checkerGeminispecialistwaitingEval analystClaudespecialistwaitingPrior-work mineropen weightsspecialistwaitingConfig differGPTspecialistwaitingReviewerGPTreviewerwaitingReport authorClaudespecialistwaitingOutcome graderClaudegraderwaiting
coordinator · Tracker: 48 runs, 16 configs × 3 seeds · plan: 7 tasks across 4 specialists

Every session in the fleet, live

The fleet view shows every session this week as one board: completed, running, waiting, and how each was scored. Persistent, long-horizon, asynchronous agents running as one fleet, each on the model it is best at.

Recursion· Sessions144 sessions this week · 127 completed · 10 running

Real work is too big for one agent, one prompt, one try

Enterprise work spans dozens of systems, runs for hours, and rarely fits in a single context window. A general agent in a chat box can look strong in a demo and still fail in production: it loses the thread across long workflows, misses edge cases, and treats every run as the first.

The core issue isn't model capability. It's that nothing around the model is built for the work: no way to split it across agents, no memory between runs, no grader saying whether the outcome was actually good.

Without a tight loop between execution, measurement, and learning, every run is just another task completed, not another signal the system gets to learn from.

A unified RL platform for specialist models

Recursion connects environments, evaluation, and training into a closed-loop reinforcement learning system. Rather than treating deployment as the endpoint, production becomes the training surface.

RL environments that reflect real work

Recursion turns workflows, tools, policies, and edge cases into executable environments for RL training and evaluation. Powered by WorldSim, these environments recreate the full enterprise software stack — with configurable world effects that generate diverse, realistic scenarios at scale.

Explore RL environments

0554-crude-pipeline-lbo-waterfall

Task runs

Show all 10 runs

Prompt6798 chars

# Pipeline LBO Returns Model with Multi-Tier Waterfall

## Context

You are an associate at a private equity fund preparing an investment committee update for a crude oil pipeline project acquired through an LBO. The task is to build an analyst-ready workbook that connects operating performance, the pre-computed debt schedule, exit valuation, and sponsor-management waterfall.

Save the completed workbook to /workspace/output/model.xlsx.

## Workbook to build

  1. 1. Financing Assumptions
  2. 2. Operating Assumptions
  3. 3. Debt Schedule
  4. 4. Model
  5. 5. Returns
AttachmentsGradingIssuesQAFormsSynthesizersSolver

Task files

Optional

Files accessible to the model at the container mount path.

inputs.xlsx

/workspace/files/problem/inputs.xlsx

source.xlsx

/workspace/files/problem/source.xlsx

Evaluation systems that measure real execution

The difference between agents that stagnate and agents that improve is measurement. Recursion builds evaluation systems that score intelligence and skill at every level — final outcomes, intermediate decisions, and execution quality — so every run generates the signal your models need to get better.

specialist-model-evals

ranked

Accuracy, efficiency, and latency across model candidates

Leader

GLM 5.1 FTv2

GLM 5.1 base
GLM 5.1 FTv1
GLM 5.1 FTv2
Claude 4.8
GPT 5.5

Accuracy

Pass@4 score on held-out financial analysis tasks

GLM 5.1 FTv2

86%

GPT 5.5

82%

Claude 4.8

79%

GLM 5.1 FTv1

78%

GLM 5.1 base

56%

Token efficiency

Normalized useful work per token

GLM 5.1 FTv2

84%

GLM 5.1 FTv1

76%

GLM 5.1 base

54%

Claude 4.8

42%

GPT 5.5

38%

Latency

Normalized responsiveness, raw latency shown

GLM 5.1 FTv2

1.1s

GLM 5.1 FTv1

1.7s

GLM 5.1 base

2.4s

GPT 5.5

2.6s

Claude 4.8

2.9s

A training loop that compounds from real work

Every rollout produces graded trajectories that can feed fine-tuning and reinforcement learning. The result is a specialist model that improves from enterprise execution signals instead of synthetic benchmarks alone.

financial-agent-glm-5.1-v7-fits6h

completed

GRPO run - GLM 5.1

Run result

86% Pass@4 score (+30%)

Training run summary

Base

8h 20m

Training

11h 45m

Tuned

8h 55m

MetricValue
Source4,300 financial analysis tasks
Task familyDCF, LBO, acquisition, projection
Base modelGLM 5.1
Training methodGRPO
ComputeGKE, H100 cluster
EndpointsBaseline and tuned

Evaluations

4 attempts per task

Problem

Baseline

Tuned

Delta

0195-acme-industrials-dcf-irr

+0.30

0.55 -> 0.85

0248-pumptech-acquisition-rollup

+0.29

0.60 -> 0.89

0515-hotel-renovation-noi-yield

+0.32

0.51 -> 0.83

0843-hartwell-plant-expansion-dcf

+0.29

0.54 -> 0.83

0120-smasco-revenue-forecast

+0.33

0.52 -> 0.85

0631-proactivate-lbo-exit-returns

+0.30

0.57 -> 0.87

From goal to specialist model — continuously

01

Launch

Set a goal. A coordinator plans it, spawns specialists, and works across your systems.

02

Grade

A grader scores every session against your rubric: outcome, process, and execution quality.

03

Remember

Graded sessions become memories, notes, and skills the fleet recollects before acting. Improvement without retraining.

04

Generate

Turn the hardest, lowest-scoring sessions into RL environments with long-tail scenarios, powered by Horizon.

05

Train

Run RL and fine-tuning on graded rollouts to produce a specialist model for your tasks.

06

Serve

Swap the specialist into the fleet, per agent. Every run from here feeds the next cycle.

Every cycle captures additional expertise. Every execution generates a new signal. Every improvement compounds your proprietary edge.

The specialist always beats the generalist

Fine-tuned on the fleet's graded runs, specialist models out-resolve frontier generalists on your tasks at a fraction of the cost. Labelbox Horizon's frontier-class RL environments are trusted by over 90% of leading US AI labs; fine-tuned models deliver a 5–10× lower cost per task.

A customer service agent resolved more tickets with fewer hallucinations, faster responses, and lower inference cost

Resolution rate

Finetuned OS model

84%

GPT-5.5

76%

Claude Opus 4.8

73%

Reduction in hallucinations

Finetuned OS model

72%

GPT-5.5

46%

Claude Opus 4.8

41%

Cost per 1,000 support tickets

Finetuned OS model

$32

GPT-5.5

$158

Claude Opus 4.8

$176

Time to first token

Finetuned OS model

0.42s

GPT-5.5

1.10s

Claude Opus 4.8

1.28s

Privacy and security for enterprise agents

Privacy and security illustration for the security page and Agent Studio privacy section.

Your advantage is the learning loop, not the model

As foundation models improve and become accessible to everyone, sustainable advantage shifts away from the model itself. The organizations that win won't simply deploy intelligence; they'll own the learning loops that transform their expertise, workflows, and decisions into durable competitive advantage.

Every organization has unique workflows, judgments, and domain expertise embedded in its people. The question is whether that expertise remains trapped inside individuals or becomes a compounding organizational asset.

Recursion is designed to close that gap.

Try Recursion on your most ambitious tasks

Tell us the goal. We'll show you what a Recursion session looks like on it.