Deep RL Crew Pairing

A transformer DQN building a month of crew pairings, one leg at a time Source and full write-up

A pairing is a legal multi-day tour of duty: a sequence of flights that starts and ends at the same crew base, obeying duty-hour limits, rest rules, legs per duty, time away from base, per-base credit caps and crew availability. The agent below builds them one leg at a time, and every rule is enforced in its action mask, so it can only ever produce a feasible schedule.

Start here: pick a recording from the drop-down below — a training run, to see how the policy learned over 500 episodes, or a search run, to see it solve a month it has never seen. Each one opens as a finished month, with every connection already drawn. Press Play at the bottom of the page to replay it from the first decision, one connection at a time; you can pause, drag the slider to any step, or skip straight back to the finished solution.

Everything here is a recording. The searches were run offline and baked into this page, so nothing on this site talks to a training machine or runs a model in your browser.

Episode network

crew base airport sit short (crew follows aircraft) layover critical NCC / stranded deadhead

Agent decision (top actions by Q-value at this step)

Coverage — train vs held-out

train held-out eval

Episode reward

train eval

Reward components

coverage cost robustness preference penalty

Loss / epsilon

loss epsilon (0–1)

Solution so far

coverage
pairings
short conns
critical
NCCs
stranded
deadheads

Credit per base

Flight details (Esc to clear)

Click any connection in the network to see the flight, the pairing that operates it, how it joined, and the reward it earned.
0 / 0