What the published runs say
These runs are the ones published on this site, all computed from one snapshot of the review table so they're directly comparable. They cover headline estimates at three aggregation levels, one-knob sensitivity checks against the occupation-level headline, and leave-one-out validation of rank recovery. Full tables are on Imputation runs.
Findings
Headline estimates
Default parameters: β = 2.5, SOC majors 37/45/47/49/51/53 pruned, activity threshold 10, and the AIOE baseline for speed only. Lists show the five largest estimated effects. Speed is a log-ratio (higher = faster with AI); quality is Hedges' g. OBS marks an observed (not imputed) node.
Does the ranking survive a changed knob?
Each row changes one setting and compares the resulting estimates with the matching occupation-level headline run. The comparison uses Spearman ρ over the nodes both runs share, with coarser levels rolled up before joining. ρ near 1 means that choice doesn't change the ordering.
Can the graph recover held-out effects?
Each observed node is held out and re-imputed. Kendall τ-b compares the order of held-out predictions with the order of the actual values. The bars are bootstrap 95% CIs over folds. An interval entirely right of 0 means the order is recovered; entirely left of 0 means it's reversed.
How these runs were chosen
- Headline: both metrics at occupation, SOC-minor and SOC-major level, with default parameters.
- Sensitivity: each run changes one knob from the occupation-level headline: β (1, 10), SOC pruning (keep all 22 majors; also drop Management, SOC 11), activity threshold (5), and, for speed only, removing the AIOE baseline. The baseline is speed-only, so a quality run without it would be identical to the headline.
- Validation: leave-one-out at each level for both metrics, plus speed without the baseline to test whether the AIOE prior helps recover held-out effects.
- The set is defined in scripts/generate_site_runs.py, which regenerates it from the live review table.