Provenance
Every metric the Workbench shows is either rooted in a citable source of record or explicitly labelled an assumption. The chain is shown beside the figure in the UI (<ProvenanceBadge>), in the methodology panel, and in every export's methodology and sources section. This is the product's audit backbone: a figure without provenance doesn't ship.
The pieces
| Piece | Where | What |
|---|---|---|
SourceRef | packages/core/src/provenance.ts (type); ReferenceData.sources (data) | One source of record: id, name, publisher, citation, optional url, licence, version (edition), and kind (dataset, classification, framework, statutory, vendor, research, model). |
Provenance | provenance.ts | { basis, sourceIds, note? } attached to a figure or group of figures. |
ProvenanceBasis | provenance.ts | measured, classified, framework, calibration, derived or assumption, labelled by BASIS_LABEL. |
MetricKey, MetricProvenance | metric-provenance.ts | Every report metric (34 keys, from classification to copilotValueModel, including savingMultiplier, fteSaved, spanThresholds, orgDisconnected and managementCost) with a label, what, an honest qualifier and its provenance. |
metricProvenance(ref), metricProvenanceList(ref) | metric-provenance.ts | Build the chains with numbers from the given reference data, memoised per Reference; the list is in appendix order. |
resolveSources, provenanceSummary | provenance.ts | Resolve source ids against ref.data.sources (unknown ids are dropped); one-line attributions such as Measured · ONS ASHE (2025). |
The sources register is reference data, so administrators record new editions (for example a new ASHE release) when they load new data, and any dataset's methodology shows the editions of the version it was scored with.
Bases
| Basis | Use for | Example |
|---|---|---|
measured | Taken from a published dataset, statute or vendor price. | AIOE percentile; ASHE median. |
framework | An established external methodology. | BCG 10-20-70. |
classified | Produced by a deterministic rule or a human review. | The role classification; Copilot fit tier. |
calibration | A disclosed, tuned constant. | The 30–85% augmentation band; horizon cut-offs. |
derived | Computed from other provenanced figures. | Capacity gain; FTE-equivalent; saving multiplier; FTE saved; ROI; spans and layers; disconnected managers; cost of managers. |
assumption | A modelling judgement with no external measurement; sourceIds are supporting references, and only sources that actually contain the figure or its context belong there. | Location factors; realism discount; employer-cost multiplier; Copilot licence price; narrow and wide span thresholds (spanThresholds, with no sources). |
isGrounded(p) is true for everything except assumption; the UI marks assumptions in salmon.
Per-figure provenance in the model
Some provenance is computed per person, not just per metric. For example scoreExposureAnchored returns measured(["aioe", "aei", "soc2018", "wbExposure"], note) (the exposure metric chain also cites blsSocCrosswalk) with a note quoting the percentile, occupation, band and AEI share actually used, or assumption(["wbExposure"], …) when a role falls back to the rubric. Notes quote live reference data values, so an edited constant can never contradict its own explanation.
Rendering
- UI:
apps/web/src/components/Provenance.tsx.<ProvenanceBadge metric="exposure" />renders a chip with the basis, opening a popover withwhat, the qualifier, the basis headline (Measured — from published data), the note and each source with its edition, publisher, licence and link.<MethodologyPanel />renders the full list. Both read theReferencefrom context, so on a dataset's pages they describe the dataset's pinned version. - Exports:
apps/web/src/exports/provenance.tsrenders the same list into the HTML reports and decks (methodologySection), plusbasisLine(ref, key), a one-line basis chip, description and sources that the spans and layers report puts at the end of each section; andpackages/core/src/workbook.tsbuilds the finance workbook's Methodology & Sources sheet frommetricProvenanceList.
AI citations
AI output has its own, stricter form of provenance (citations.ts): every narrative must carry citations, validated against an allow-list of sources the model was actually shown. validCitations(raw, allowed) drops invalid, duplicate and unlisted citations (keeping at most 50). For narratives the only allowed source is the dataset itself (dataset:<id>); the classifier cites the two SOC codes it chose (us-soc-2018:…, uk-soc-2020:…). A response without a valid citation is rejected in favour of the rule-based text. See AI integration.
Adding a metric
- Add a key to the
MetricKeyunion and an entry inbuild(ref)inmetric-provenance.ts: alabel,what, an honestqualifierand aProvenancerooted in source ids that exist in the reference data'ssources(if the source is new, add it to the built-in data inreference/builtin/sources.tsfor new databases, and publish a reference data version that includes it for existing deployments). If it's a judgement, useassumption(...). - Quote any constant from
ref.data, never a literal. - Add it to
ORDERso it appears in the methodology appendix. - Show it in the UI with
<ProvenanceBadge metric="…" />beside the figure. - Make sure the exports' methodology sections include it (they read
metricProvenanceList). - Extend
packages/core/test/provenance.test.ts.
Metric chains refer to sources by id, listed in CITED_SOURCE_IDS (provenance.ts): aioe, blsSocCrosswalk, aei, aeiLabourMarket, ashe, soc2020, soc2018, hmrcNi, tprPension, bcg102070, mckinseyGenai, eloundou, msCopilotPrice, wbExposure, wbTaxonomy, wbClassifierAI, wbOverride, wbIndustryPack, clientData, wbScenarios. validateReferenceData refuses to publish a version that lacks any of them, and a test checks that every id a chain cites is in the list. resolveSources still drops unknown ids silently, so a new cited id must be added to the list, to the built-in sources and to a published reference data version for existing deployments.