Reference data
Everything the model runs on is data, not code. ReferenceData (packages/core/src/reference/types.ts) holds every table and constant; administrators edit it and publish numbered versions; each dataset is pinned to the version it was scored with. The user-facing side is in the user guide's Reference data.
The document
ReferenceData is plain JSON with schemaVersion: 1 and fifteen sections (REFERENCE_SECTIONS):
| Section | Contents |
|---|---|
occupations | usMajorGroups (two-digit code → title); us (US SOC-2018 code → { title, major }, the codes the classifier and overrides may choose); aioe (US code → { score, percentile }); uk (UK SOC-2020 code → { title, median, p25, p75 } from ONS ASHE); ukProxyMedians; ukDefaultCode; fallbackMedian. |
taxonomy | rules: ordered title rules { pattern, category, subcategory, division? } (case-insensitive regex, first match wins, optional exact-division restriction); roles: { category, subcategory, ukSoc2020, anchor: { usSoc2018, socTitle, aioePercentile } | null, rubric }; defaultRubric for roles a rule or override names but the roles table lacks (unmatched titles are Other / Uncategorised, scored by that row). |
exposure | augFloor, augCeil (percent), aeiAutoShare (0–1), horizonNowFrom, horizonNearFrom (percentiles). |
classification | reviewThreshold (0–1). |
costs | employerCostMultiplier (the default for new datasets). |
locations | regions: ordered { key, label, defaultFactor, countryAliases, keywords, nativeSalary }; baselineRegion (factor 1). |
scenarios | Exactly conservative, moderate, aggressive with automationRealised and augmentationProductivity. |
programme | scurve, reinvestment (components and totals), investmentSplit, roiBands per scenario, four investmentPhases, paybackMonths. |
gcc | offshoreBelow, onshoreFrom, defaultOffshoreFactor, migratableSocMajors, migratableSubcategories. |
org | narrowSpan, wideSpan. |
lifecycles | generic (phases, categoryPhase, defaultPhase); defined lifecycles (id, label, phases, per-role phase weights); autoLifecycleId; meaningfulShare. |
benchmark | aeiGroups (AEI group, SOC major, theoretical and observed %, taxonomy categories); sources (lines shown in reports). |
copilot | defaults (CopilotOptions); tiers (labels and notes for the four fixed tiers); fitBySocMajor; defaultTier; highSubcategories, lowSubcategories, unclassifiedSubcategories. |
packs | Industry packs { id, label, description, overlay, lifecycleId?, copilotFit? }; general is required. |
sources | Source id → SourceRef. |
The built-in version lives in reference/builtin/ (split by topic: soc-reference.ts, aioe-by-soc.ts, ashe-salaries.ts, salary-map.ts, taxonomy-rules.ts, exposure-rubric.ts, exposure-anchor-data.ts, locations.ts, programme.ts, lifecycles.ts, benchmark.ts, copilot.ts, packs.ts, sources.ts), assembled by builtinReferenceData(). It's about 300 KB, so it's exported for the server and tests but the web app never imports it: the browser always fetches a published version.
How the larger built-in tables were produced, so they can be regenerated:
aioe-by-soc.ts: Felten, Raj & Seamans' Language-Modeling AIOE (the "LM AIOE" sheet, 774 occupations keyed by SOC 2010), re-keyed to SOC 2018 with the BLSsoc_2010_to_2018_crosswalk.xlsx. Each 2018 code takes the unweighted mean score of the 2010 codes that map to it; a 2010 code that splits passes its score to every 2018 code it becomes. The percentile is rank ÷ (n − 1) × 100 over the crosswalked set, to one decimal place (the rule that reproduces the original 774 percentiles exactly). 800 of the 867 SOC 2018 codes have a row; the rest are military, "All Other" residuals, codes derived only from residuals (such as 15-2051 Data Scientists) and Legislators.exposure-anchor-data.ts: each anchor's title and percentile copysoc-reference.tsandaioe-by-soc.ts;test/reference-sources.test.tsfails if any anchor or pack code isn't a SOC 2018 occupation with an AIOE row.ashe-salaries.ts: the "Full-Time" sheet of ONS ASHE Table 14.7a (2025 provisional), median and 25th/75th percentiles of gross annual pay. Suppressed cells (x,..) becomenull; a code with a suppressed median is left out and priced by a proxy insalary-map.tsor the fallback. ASHE 2026 provisional is due on 22 October 2026: load it the same way and publish it as a new reference version.benchmark.ts: theoretical and observed coverage, both employment-weighted, from Figure 2 of Massenkoff & McCrory (2026): figures stated in the text exactly, the rest read from the chart (±1 point). Occupation-level detail is inlabor_market_impacts/job_exposure.csv.
Validation
validateReferenceData(raw) (reference/schema.ts) returns { ok, issues, data }:
-
Shape with zod (
referenceDataSchema): types, ranges, string lengths, SOC code format, valid case-insensitive regex patterns (≤300 characters), region and pack id format, array bounds, exactly three scenarios and four investment phases. A shape failure returnsdata: nulland error issues. -
Cross-checks (
crossCheck) between tables, each anerroror awarningwith a dot path such astaxonomy.rules.12:- errors: unknown major groups; UK codes without pay (default code, roles, pack rules); duplicate roles; rules pointing at missing roles; no Other / Uncategorised role; band and horizon ordering; duplicate regions or country aliases; baseline missing or ≠ 1; scenario order; missing or inverted ROI bands; inverted reinvestment ranges; GCC and span thresholds; lifecycle phase keys, weights referencing undefined phases, unknown auto lifecycle; duplicate packs, unknown pack lifecycles, pack rules with unknown US codes, no
generalpack; source ids not matching their keys; - warnings: ASHE percentiles on the wrong side of the median; anchor US codes missing from the AIOE table or percentiles drifting ≥0.05 from it; rubric automation above augmentation; investment split or phase shares not summing to 100%; the S-curve going down; lifecycle weights not summing to 1.
- errors: unknown major groups; UK codes without pay (default code, roles, pack rules); duplicate roles; rules pointing at missing roles; no Other / Uncategorised role; band and horizon ordering; duplicate regions or country aliases; baseline missing or ≠ 1; scenario order; missing or inverted ROI bands; inverted reinvestment ranges; GCC and span thresholds; lifecycle phase keys, weights referencing undefined phases, unknown auto lifecycle; duplicate packs, unknown pack lifecycles, pack rules with unknown US codes, no
ok is false when any error exists. Saving a draft needs only a valid shape; publishing needs ok. The built-in data has no errors or warnings.
Compilation
compileReference(data, version) (reference/compile.ts) turns valid data into a Reference ready to score with: compiled RegExps for rules and pack overlays, roles keyed category||subcategory, a sorted list of taxonomy roles for pickers, region and normalised country-alias maps, packs joined to their lifecycles, sets for Copilot and GCC lists. Results are memoised per data object and version.
Shared lookups live beside it: isValidSoc2018, socCategory, isKnownUkSoc, getIndustryPack (falls back to general), isIndustryPackId, defaultModelSettings, copilotTierLabel.
Versioning
Stored in reference_versions (see Data model):
- Seeding.
seedBuiltinReferenceinserts version 1 frombuiltinReferenceData()withON CONFLICT DO NOTHING. It runs after migrations innpm run db:migrate, the Node dev server and the test harness. The Worker doesn't seed: a database that has never been migrated fails with Reference data hasn't been set up. Run the database migrations (npm run db:migrate). - Current. The newest published version (
currentReferenceVersion). - Immutability. RLS allows updates and deletes only while
status = 'draft'. Publishing is an update from draft to published, after which the row can't change. - One draft. A unique partial index on
status = 'draft'. - Pinning.
POST /engagements/:id/datasetssetsdatasets.reference_versionto the current version;finalizeDatasetwrites the version it scored with;referenceForDatasetloads it (null reads as 1, for datasets older than migration 0002). - Caching.
loadReference(tx, version)re-validates a published version (so a damaged row fails loudly), compiles it, and keeps up to 16 compiled versions per isolate. - Rescoring.
POST /datasets/:id/reference→rescoreWithReference(see Upload and classification), audited asdataset_reference_changedwithfromandto.
Overrides are validated against the current version (PUT /engagements/:id/overrides), then applied through each dataset's pinned version.
API
| Endpoint | Purpose |
|---|---|
GET /api/reference | The live version number, publish date and notes. |
GET /api/reference/versions/:v | A version's full data (published for everyone; drafts for admins, under RLS). |
GET /api/admin/reference | Admin: current, draft and version history metadata. |
POST /api/admin/reference/draft | Admin: start the draft, optionally { basedOn }. |
GET /api/admin/reference/draft | Admin: the draft with live validation issues. |
PUT /api/admin/reference/draft | Admin: save { data } (the whole document), { section, value } (one section) or { notes }; up to 4 MB. |
POST /api/admin/reference/draft/publish | Admin: { notes } (at least three characters). |
DELETE /api/admin/reference/draft | Admin: discard. |
Draft creation, publishing and discarding are audited (reference_draft_created, reference_published with notes and the warning count, reference_draft_discarded). Saves aren't audited individually.
In the browser
lib/reference.tsx:CurrentReferenceProvider(in the app shell) loads the live version for everything outside a dataset;ReferenceProvider version={dataset.referenceVersion}wraps dataset pages.useReference()returns the compiledReference. Published versions never change, so each is fetched once per session (staleTime: Infinity).pages/admin/ReferenceTab.tsx: the editor shell, working copy, live validation (validateReferenceDataon a deferred copy), change markers (per view, comparing against thebasedOnversion), publish dialog (changedSections), version history and JSON import/export. The draft is saved as a whole document.pages/admin/reference/sections.tsx:REFERENCE_VIEWS, one per navigation entry, each with agroup,label, the dotpathsit edits (for issue links and change markers) and anEditor. Tables useDataGrid; small groups of constants are forms; nested structures (programme,lifecycles, the default rubric) use a validated JSON editor.pages/admin/reference/DataGrid.tsxandlib/referenceEdit.ts: typed cells (text,optionalText,number,nullableNumber,list,boolean,select), dot-path get/set, record ⇄ rows conversion, CSV export with a BOM and formula guard, and CSV import keyed by column path with per-line errors.referenceEdit.tsis pure and unit-tested (apps/web/test/referenceEdit.test.ts).
To add a field, see Add a reference data field.