Upload and classification
A staff list goes from a workbook on the consultant's machine to a scored dataset through a sequence of small, idempotent steps, each resumable.
browser: parse workbook (Web Worker) → map columns → filter rows → resolve manager emails
→ derive tenants from mail domains → build rows (+ optional identities)
│
├─ POST /engagements/:id/datasets create (status uploading, pinned reference version)
├─ POST /datasets/:id/rows × n 2,000-row chunks, 3 in parallel, idempotent per chunk
├─ POST /datasets/:id/complete distinct titles extracted (status classifying)
└─ POST /datasets/:id/classify × until done up to 600 pending titles per call
override → industry pack → shared cache → AI (batches of 150, 4 at a time) → rules
when none are pending: score every row, build and store the DatasetSummary (ready)
In the browser
apps/web/src/lib:
| Module | Does |
|---|---|
spreadsheet.ts, parse.worker.ts, spreadsheet-core.ts | Parse .xlsx (read-excel-file) or .csv off the main thread, with limits: 20 MB file, 64 MB uncompressed XLSX, 1,024 archive entries, 100,001 rows, 256 columns, 5 million cells, 32,768 characters a cell. Returns headers and rows. |
ingest.ts | FIELDS (the mappable fields), guessMapping (exact header matches, then substrings, never reusing a column, and never taking a "manager …" header for a non-manager field), buildRows (filter, skip rows without a title, generate row-N IDs, resolve manager emails, derive tenants, parse salaries like £42,500 or 42.5k, read FTE with readFte, carry the chosen attribute columns), attributeProblem (why a column can't be an extra grouping), domainCounts, tenantLabelFor. |
upload.ts | uploadDataset and completeAndClassify: drive the server pipeline, with retries. |
buildRows resolves manager emails to employee IDs using the mapped email column. An email held by more than one person maps to null and any reporting line through it is left unresolved (reported as ambiguous). An email that matches nobody in the list becomes an opaque stand-in, unmatched-manager-N (UNMATCHED_MANAGER_PREFIX in core; each distinct email takes the next number from a running counter that skips any employee ID of the same form), shared by everyone who named that manager, so the spans report can count them as disconnected. It's still counted as unresolved. The emails themselves are dropped unless identities are stored.
FTE. readFte accepts 0.6, 0,6, 60% or 1, rounded to two decimal places; a blank is null (counts as 1). Anything else, such as hours, is rejected, counted in stats.implausibleFte and stored as null. The API also stores a value that rounds to 0 as null (unknown, counts as 1), and a value the database refuses answers 400, not 500. FTE counts FTE saved only; it never scales cost.
Extra groupings. The page offers every unmapped column as a client attribute, up to MAX_ATTRIBUTES (5). attributeColumnProblem (core, attributes.ts) refuses a column:
- whose heading, normalised to plain lower-case words first (NFKC for look-alike characters; camelCase, snake_case, hyphens and dots split; plurals matched), looks like personal data or a person: names, emails, phones, addresses, postcodes, dates of birth, NI numbers, passports, usernames and logins, payroll, employee/staff/person/worker/colleague IDs or numbers, managers, supervisors, advisers, mentors, line managers, reports to, team leads and leaders, owners and colleagues. A heading that is only
<group> name(s)(Firm name, Team name, Cost centre name…) passes; Client name doesn't; - with any value (after NFKC, so a full-width @ counts) containing
@or longer than 120 characters; - with more than 100 distinct values that are also more than a quarter of the rows (
attributeVarietyProblem).
The page also disables a second column with the same heading (compared trimmed, ignoring case), and the limit of five counts only the columns that will be sent. Chosen columns travel on each row as attributes: { label: value }. The API's employeeRowSchema runs attributesProblem on every row (headings, values, at most five, no two headings equal after trimming), and completeUpload runs attributeVarietyProblem per grouping over the stored rows, so a modified client can't bypass either check.
The previous ready dataset's column_mapping (stored by header name) is reused when at least two fields, including the job title, match the new file's headers.
On the server
Create
POST /engagements/:id/datasets (routes/engagements.ts) checks write access (and PII access when withIdentities), and inserts the dataset with reference_version = the current published version, settings (the model assumptions, defaulted in the browser from that version) and expected_rows.
Rows
POST /datasets/:id/rows → insertChunk (services/datasets.ts):
- registers the chunk number in
dataset_chunkswithON CONFLICT DO NOTHING; a retried chunk is recognised and returns{ duplicate: true }without re-inserting; - increments
uploaded_rows, refusing to exceedexpected_rows; - inserts rows at
row_no = chunk × 2000 + i, includingfteandattributes(migration 0006); - with identities, encrypts names and emails under the engagement's data key and inserts them into
employee_identities(auditedidentities_uploaded).
employeeRowSchema refuses any employee or manager ID containing an @ (look-alikes folded with NFKC): the browser never sends one, because the upload page blocks a mapping whose ID column holds emails or names (idMappingProblems in ingest.ts, using idColumnProblem from core).
The browser retries a failed chunk up to three times (not for 4xx errors other than 408 and 429).
Complete
POST /datasets/:id/complete checks every expected row has arrived and that no extra grouping has so many distinct values that it identifies people (400 attribute_identifies_people; the dataset stays uploading, so it must be deleted and uploaded again without that column), groups rows by exact title into dataset_titles (with headcount and normalised form, source = 'pending'), counts duplicate employee IDs and moves the dataset to classifying.
Classify
POST /datasets/:id/classify { useAi? } → classifyStep (services/classify.ts) advances by one bounded step. AI is used only if the dataset was created with useAi (stored as datasets.use_ai) and the call doesn't pass useAi: false; omitting it, as Continue does, uses the stored choice.
- In a short
asUsertransaction: load up to 600 pending titles (most people first) and the dataset's pinnedReference, and resolve deterministically: overrides, then the industry pack overlay, then the shared cache (for engagements that share, keyed by normalised title and pack id). - Outside any transaction, if titles remain and AI is allowed: check the budget (
aiUnavailableReason), load the published classifier prompt (activePrompt), and callclassifyTitleswith the industry context (sector and non-general pack label) and region. It batches 150 titles, runs four batches at a time, and never throws: a failed batch leaves its titles for the rules. - Titles the AI didn't place fall back to
classificationByRules. - If the model was called, record AI usage (
feature: classify) in a transaction of its own, so the tokens are metered even if the next step fails. - In another transaction: write the classifications only to titles still
pending(an override set while the model was working is kept), and if nothing is pending, finalise: score every row, build the summary, setready, auditdataset_finalised. - On the service connection: write new AI picks to the cache (never overwriting a verified entry) and bump cache hit counts.
The response is progress: { total, pending, classifiedThisCall, aiUsed, aiUnavailableReason }. The browser loops until pending === 0; Continue on the Datasets tab calls the same loop. No transaction is held open while waiting for the model.
Each step is safe to repeat. A dropped connection during rows leaves the chunks already stored, but an uploading dataset can only be continued from the same browser tab, because the rows live there; otherwise delete it and upload again (Continue on the Datasets tab only completes an upload whose rows have all arrived, then classifies). A dropped connection during classification leaves titles pending; calling classify again carries on.
Finalise
finalizeDataset(tx, ref, dataset, engagement, pack):
- loads every title's classification and every row;
scoreDataset(ref, rows, titles, dataset.settings);buildDatasetSummary(ref, scored, titles, { pack, horizonStart }), with the horizon starting at the half-year the dataset was created;- updates
summary,headcount,total_cogs,status = 'ready',finalized_atandreference_version; - deletes cached AI outputs (
dataset_ai), whose figures would now be stale.
It's also called after an override is set or removed, after assumptions are saved, and by rescoreWithReference (including after a pack change).
Re-scoring
| Trigger | What happens |
|---|---|
| Override set | First checked against every reference version pinned by a dataset holding the title: if one doesn't know the code or role, 400 override_unknown_in_version names the version and asks for a rescore. Then every dataset holding the normalised title gets the override classification (through its own pinned reference) and is finalised. |
| Override removed | Those titles are re-resolved without AI (pack → cache → rules), then finalised. |
Assumptions saved (PATCH /datasets/:id) | Finalised with the new settings. |
Rescore with version N (POST /datasets/:id/reference) | rescoreWithReference: re-resolve every title deterministically under version N. Overrides and pack rules win; a previous AI pick whose codes N still knows is kept even when the shared cache now holds something else; other titles take the cache, a kept cache pick, or N's rules. Then finalise. No AI is called, so it's reproducible. |
Industry pack changed (PATCH /engagements/:id) | rescoreWithReference on every ready dataset in the engagement, each on its pinned version, with the new pack: its rules, value chain and Copilot fit lists. |
Reading back
GET /datasets/:id/rowspages rows in upload order (up to 10,000 a page); the web client fetches all pages four at a time (fetchAllRows) and re-scores them withuseWorkforce.GET /datasets/:id/titlesreturns each title's classification withreviewRequired.- The browser loads the dataset's pinned reference version (
useReferenceVersion) and scores with it, so its figures match the stored summary.