Security and privacy
The Workbench handles client staff lists: job titles, reporting lines, salaries and, optionally, names and emails. The design goal: each person sees only the engagements they're staffed on, names and emails are handled as little as possible, and every sensitive action leaves a trace. Sign-in and access control are covered in Authentication and authorisation; this page covers the rest. The canonical statement is docs/SECURITY.md.
Personal data: minimisation first
- The staff-list workbook is parsed in the browser, in a Web Worker (
apps/web/src/lib/parse.worker.ts). Only mapped columns, and any client groupings the consultant ticks, are sent. - Manager emails are resolved to employee IDs in the browser (
lib/ingest.ts), and Microsoft 365 tenants are derived from mail domains there. The emails themselves aren't uploaded. - Where two people share an email, reporting lines through it are left unlinked rather than guessed.
- IDs are never emails or names.
idColumnProblem(packages/core/src/identifiers.ts) refuses, in the browser, a column mapped asemployeeIdormanagerEmployeeIdthat holds any email address or is mostly people's names, and the upload page blocks the upload until the mapping changes. On the server,employeeRowSchemarefuses any employee or manager ID containing an @ after NFKC (looksLikeEmail), so a modified client can't store one (apps/api/test/security.test.ts). - Where a manager is named by email but isn't in the list, only an opaque numbered stand-in (
unmatched-manager-1…) is uploaded, shared by everyone who named them. The email stays in the browser. - Client groupings (up to five extra columns, such as an IFA firm) are categories carried on each row in
dataset_employees.attributes.attributeColumnProblemin the browser refuses a column whose heading (normalised: NFKC, camelCase and snake_case split, plurals) looks like personal data, an identifier or a person (manager, adviser, owner…), whose values look like emails (after NFKC), or which has so many distinct values it identifies individuals. On the server,employeeRowSchemarunsattributesProblemon every uploaded row andcompleteUploadre-checks each grouping's variety (400 attribute_identifies_people), so a modified client can't store one either. They never enter summaries, logs, AI prompts or benchmarks (apps/api/test/savings-notes.test.tschecks the summary). - An engagement's report wording (
engagements.report_notes) is the engagement's own text, not meant for personal data, but it is free text that may name people: it changes only while the engagement is active (409otherwise), its audit entry records only that it was updated or cleared, never the text, and a purge clears it. - A missing manager reference that might be an email or a name (it has an
@or whitespace, or is over 40 characters) is never printed in the spans and layers report: it reads Unrecognised manager reference (#n). - Names and emails are uploaded only when a member with PII access ticks Store names and emails, encrypted, and then only to
employee_identities, neverdataset_employees.
Encryption at rest
Identities use envelope encryption (apps/api/src/services/crypto.ts), all with WebCrypto so it runs identically in the Worker and Node:
IDENTITY_KEK (Worker secret, 32 bytes)
└─ wraps one random 256-bit data key (DEK) per engagement → engagements.dek_wrapped
├─ AES-256-GCM encrypts each name and email, with additional authenticated data
│ binding the ciphertext to its dataset, employee and field
└─ HKDF-SHA-256 ("workbench/email-index/v1") derives an HMAC-SHA-256 key
for the email blind index → employee_identities.email_hash
- The AAD means a ciphertext can't be swapped between rows or fields.
- The blind index lets a licence list be matched without decrypting anything: the browser sends emails, the server HMACs them under the engagement's key and compares hashes.
services/identities.tsrefuses clearly when no KEK is configured (503 pii_unconfigured) or the engagement's key has been destroyed (409 pii_unavailable).
If IDENTITY_KEK is lost or changed, every stored name and email becomes permanently unreadable (the analysis is unaffected). Key rotation (re-wrapping each engagement's DEK) isn't implemented. Generate it once per environment and keep it in the Davies password vault.
Access to identities
- Reveal (
POST /api/datasets/:id/identities/reveal) requires PII access and a statedpurpose, and is audited with the purpose and the number of people. - Match (
POST /api/datasets/:id/identities/match) is audited the same way. - Exports never contain names or emails, except the Copilot allocation workbook, which includes them only after an authorised reveal and is audited as a named export.
- Administrators are not exempt from the PII rule;
app.pii_engagement_ids()ignores the admin flag. - Nobody grants PII access to themselves, administrators included.
PUT /engagements/:id/membersrefuses it with403, and theengagement_membersinsert and update policies (migration0003) refuse it for the application role too. Keeping access you already hold, or dropping it, is allowed. The engagement creator's first row (lead with PII access) is the one exception, throughapp.can_bootstrap_members.
Retention and deletion
When an engagement is closed, its data is kept for retention_days (default 90) and then purged by the retention job (the Worker's nightly cron at 03:00 UTC; hourly under the Node dev server). A lead can purge immediately by typing the engagement name. purgeEngagement (services/retention.ts):
- deletes identities and destroys the engagement's wrapped data key (
dek_wrapped = null) in the live database (crypto-shredding). Backups and point-in-time history from before the purge still hold the wrapped key and the ciphertext until they age out; - deletes employee rows, title lists, upload chunks and rollout plans (which reference employee IDs), and clears the engagement's report wording (free text that may name people);
- marks datasets and the engagement
purged, keeping each dataset's anonymised summary; - writes an
engagement_purgedaudit entry (reason: manualorretention).
Neon's point-in-time restore window is the only place row data and identities outlive a purge; keep it short.
AI data flows
AI is used only when configured (AI_PROVIDER=gemini with a key). What's sent:
| Feature | Sent | Never sent |
|---|---|---|
| Title classification | Distinct job titles (trimmed to 300 characters), the client's sector, the pack label and region | Names, emails, employee IDs, salaries, reporting lines |
| Executive summary, value-chain map, Ask the data | The dataset name, sector and aggregate figures (headcounts, costs, percentages by category and scenario; role type names and headcounts for the value chain) | Any row-level data |
- Validation. Classifier codes must exist in the reference catalogues and each row must cite exactly its two codes. Narratives must cite the dataset (the only allowed source) or they're discarded in favour of rule-based text.
- Labelling. AI output is always labelled as AI-written, with a confidence; low-confidence narratives are marked Review before sharing; low-confidence classifications go to the review queue.
- Cost. A monthly budget guardrail pauses AI features when spent.
- Sharing. Engagements can opt out of the cross-engagement classification cache, which holds titles and codes only.
- Prompts. Administrators can edit prompts, but the response contracts are fixed in code, so an edited prompt can't widen what the platform accepts.
See AI integration.
Benchmarks
Cross-engagement benchmarks use anonymised aggregates from each participating engagement's most recently uploaded finished dataset (by created_at; finalized_at moves on every rescore). Two rules stop a figure being traced back to a client:
- Five contributors. Any figure contributed by fewer than
MIN_BENCHMARK_ENGAGEMENTS(5) is withheld: the whole page, each occupation group, and spans and layers separately. An engagement is excluded from the benchmark it's compared against, and the threshold applies after that exclusion. With three contributors, linearly interpolated quartiles invert exactly (the lowest value is 2 × p25 − median), and with four a member comparing their own dataset could recover the other three. - Rounding (
BENCHMARK_ROUNDING). Every published quartile is rounded: pounds to the nearest £500, percentages to whole numbers, spans and layers to 0.5. With five contributors the quartiles are contributors' own values, so rounding is what stops them being read back exactly.
Audit
Every change, upload, override, export, reveal, match, purge and administrative action (including reference data and prompt publishing and prompt tests) is appended to audit_log with audit(), in the same transaction as the change.
workbench_apphasINSERTandSELECTonaudit_logbut noUPDATEorDELETEgrant or policy, so the trail can't be edited through the application.- The insert policy requires
actor_id = app.current_user_id(): no one can write entries as someone else. - Leads see their engagement's trail; admins see everything. Rows survive an engagement purge; their details never contain names or emails.
Web and API hardening
- Same-origin SPA and API, so there's no CORS surface. Every API response is
no-store, with Hono's secure headers. - Request bodies are size-capped while streaming (
body()inroutes/helpers.ts: 256 KB by default, larger for row chunks and reference data drafts) and validated with zod. - Upload chunks are limited to 2,000 rows; datasets to 50,000.
- Route IDs must be UUIDs, or the route answers
404. - Database errors are mapped to generic messages, never internals. Unexpected errors are logged as their class, Postgres code and route only (
loggableErrorinerrors.ts): a failed query's message includes its parameters, which can be row data. - Spreadsheet parsing is bounded (file size, rows, columns, cells, cell length, and the XLSX archive's uncompressed size and entry count) against zip bombs and oversized input.
- CSV exports neutralise formula-like cells (CSV injection).
- Model output is rendered with
react-markdown, which never renders raw HTML. Deliverables are self-contained HTML with escaped content and no external scripts. - Model names are restricted to
[A-Za-z0-9._-]before being placed in the Gemini URL, and a prompt may only name a model inGEMINI_ALLOWED_MODELS. - CSRF. Middleware in
app.tsrefuses POST, PUT, PATCH and DELETE with403 cross_site_requestwhenSec-Fetch-Siteiscross-siteor theOriginhost differs fromHost, and with415 unsupported_media_typewhen a body isn'tapplication/json. - Static headers.
apps/web/public/_headersgives the SPAnosniff,Referrer-Policy: same-origin,X-Frame-Options: DENY, a permissions policy andframe-ancestors 'none'. Scripts aren't restricted by CSP because export previews run the export's inline script in a same-origin window. - Worker. Development sign-in is refused in the Worker whatever
ENVIRONMENTsays; Access tokens must be RS256 withexp;workers_devand preview URLs are off. - Local server. The Node server binds to
HOST(default127.0.0.1) and, under development sign-in, answers403 forbidden_hostunless theHostislocalhost,127.0.0.1or[::1](DNS rebinding). - Regular expressions.
titlePatternProblemrefuses administrator patterns with nested repetition such as(a+)+, backreferences, lookbehind, or more than three open-ended wildcards, so a title rule can't take exponential time. - Prompt injection. Titles are sent to the classifier as JSON string literals and the prompt treats them as data; a batch where one US code takes more than 60% of eight or more titles isn't written to the shared cache (
isAnomalousBatch). - Spend. Gemini thinking tokens are metered, each call in its own transaction so a failed write can't lose the record, and "Ask the data" is capped per person per UTC day (
AI_ASK_DAILY_LIMIT). - AI opt-out. A dataset created with
useAi: falsestores that choice (datasets.use_ai); no laterclassifycall, including Continue, can send its titles to the model.
Reporting a vulnerability
Raise it privately with the Workbench owners in Davies Technology. Don't open a public issue.