Skip to main content

Security and privacy

The Workbench handles client staff lists: job titles, reporting lines, salaries and, optionally, names and emails. The design goal: each person sees only the engagements they're staffed on, names and emails are handled as little as possible, and every sensitive action leaves a trace. Sign-in and access control are covered in Authentication and authorisation; this page covers the rest. The canonical statement is docs/SECURITY.md.

Personal data: minimisation first​

  • The staff-list workbook is parsed in the browser, in a Web Worker (apps/web/src/lib/parse.worker.ts). Only mapped columns, and any client groupings the consultant ticks, are sent.
  • Manager emails are resolved to employee IDs in the browser (lib/ingest.ts), and Microsoft 365 tenants are derived from mail domains there. The emails themselves aren't uploaded.
  • Where two people share an email, reporting lines through it are left unlinked rather than guessed.
  • IDs are never emails or names. idColumnProblem (packages/core/src/identifiers.ts) refuses, in the browser, a column mapped as employeeId or managerEmployeeId that holds any email address or is mostly people's names, and the upload page blocks the upload until the mapping changes. On the server, employeeRowSchema refuses any employee or manager ID containing an @ after NFKC (looksLikeEmail), so a modified client can't store one (apps/api/test/security.test.ts).
  • Where a manager is named by email but isn't in the list, only an opaque numbered stand-in (unmatched-manager-1…) is uploaded, shared by everyone who named them. The email stays in the browser.
  • Client groupings (up to five extra columns, such as an IFA firm) are categories carried on each row in dataset_employees.attributes. attributeColumnProblem in the browser refuses a column whose heading (normalised: NFKC, camelCase and snake_case split, plurals) looks like personal data, an identifier or a person (manager, adviser, owner…), whose values look like emails (after NFKC), or which has so many distinct values it identifies individuals. On the server, employeeRowSchema runs attributesProblem on every uploaded row and completeUpload re-checks each grouping's variety (400 attribute_identifies_people), so a modified client can't store one either. They never enter summaries, logs, AI prompts or benchmarks (apps/api/test/savings-notes.test.ts checks the summary).
  • An engagement's report wording (engagements.report_notes) is the engagement's own text, not meant for personal data, but it is free text that may name people: it changes only while the engagement is active (409 otherwise), its audit entry records only that it was updated or cleared, never the text, and a purge clears it.
  • A missing manager reference that might be an email or a name (it has an @ or whitespace, or is over 40 characters) is never printed in the spans and layers report: it reads Unrecognised manager reference (#n).
  • Names and emails are uploaded only when a member with PII access ticks Store names and emails, encrypted, and then only to employee_identities, never dataset_employees.

Encryption at rest​

Identities use envelope encryption (apps/api/src/services/crypto.ts), all with WebCrypto so it runs identically in the Worker and Node:

IDENTITY_KEK (Worker secret, 32 bytes)
└─ wraps one random 256-bit data key (DEK) per engagement → engagements.dek_wrapped
├─ AES-256-GCM encrypts each name and email, with additional authenticated data
│ binding the ciphertext to its dataset, employee and field
└─ HKDF-SHA-256 ("workbench/email-index/v1") derives an HMAC-SHA-256 key
for the email blind index → employee_identities.email_hash
  • The AAD means a ciphertext can't be swapped between rows or fields.
  • The blind index lets a licence list be matched without decrypting anything: the browser sends emails, the server HMACs them under the engagement's key and compares hashes.
  • services/identities.ts refuses clearly when no KEK is configured (503 pii_unconfigured) or the engagement's key has been destroyed (409 pii_unavailable).
The identity key

If IDENTITY_KEK is lost or changed, every stored name and email becomes permanently unreadable (the analysis is unaffected). Key rotation (re-wrapping each engagement's DEK) isn't implemented. Generate it once per environment and keep it in the Davies password vault.

Access to identities​

  • Reveal (POST /api/datasets/:id/identities/reveal) requires PII access and a stated purpose, and is audited with the purpose and the number of people.
  • Match (POST /api/datasets/:id/identities/match) is audited the same way.
  • Exports never contain names or emails, except the Copilot allocation workbook, which includes them only after an authorised reveal and is audited as a named export.
  • Administrators are not exempt from the PII rule; app.pii_engagement_ids() ignores the admin flag.
  • Nobody grants PII access to themselves, administrators included. PUT /engagements/:id/members refuses it with 403, and the engagement_members insert and update policies (migration 0003) refuse it for the application role too. Keeping access you already hold, or dropping it, is allowed. The engagement creator's first row (lead with PII access) is the one exception, through app.can_bootstrap_members.

Retention and deletion​

When an engagement is closed, its data is kept for retention_days (default 90) and then purged by the retention job (the Worker's nightly cron at 03:00 UTC; hourly under the Node dev server). A lead can purge immediately by typing the engagement name. purgeEngagement (services/retention.ts):

  1. deletes identities and destroys the engagement's wrapped data key (dek_wrapped = null) in the live database (crypto-shredding). Backups and point-in-time history from before the purge still hold the wrapped key and the ciphertext until they age out;
  2. deletes employee rows, title lists, upload chunks and rollout plans (which reference employee IDs), and clears the engagement's report wording (free text that may name people);
  3. marks datasets and the engagement purged, keeping each dataset's anonymised summary;
  4. writes an engagement_purged audit entry (reason: manual or retention).

Neon's point-in-time restore window is the only place row data and identities outlive a purge; keep it short.

AI data flows​

AI is used only when configured (AI_PROVIDER=gemini with a key). What's sent:

FeatureSentNever sent
Title classificationDistinct job titles (trimmed to 300 characters), the client's sector, the pack label and regionNames, emails, employee IDs, salaries, reporting lines
Executive summary, value-chain map, Ask the dataThe dataset name, sector and aggregate figures (headcounts, costs, percentages by category and scenario; role type names and headcounts for the value chain)Any row-level data
  • Validation. Classifier codes must exist in the reference catalogues and each row must cite exactly its two codes. Narratives must cite the dataset (the only allowed source) or they're discarded in favour of rule-based text.
  • Labelling. AI output is always labelled as AI-written, with a confidence; low-confidence narratives are marked Review before sharing; low-confidence classifications go to the review queue.
  • Cost. A monthly budget guardrail pauses AI features when spent.
  • Sharing. Engagements can opt out of the cross-engagement classification cache, which holds titles and codes only.
  • Prompts. Administrators can edit prompts, but the response contracts are fixed in code, so an edited prompt can't widen what the platform accepts.

See AI integration.

Benchmarks​

Cross-engagement benchmarks use anonymised aggregates from each participating engagement's most recently uploaded finished dataset (by created_at; finalized_at moves on every rescore). Two rules stop a figure being traced back to a client:

  • Five contributors. Any figure contributed by fewer than MIN_BENCHMARK_ENGAGEMENTS (5) is withheld: the whole page, each occupation group, and spans and layers separately. An engagement is excluded from the benchmark it's compared against, and the threshold applies after that exclusion. With three contributors, linearly interpolated quartiles invert exactly (the lowest value is 2 × p25 − median), and with four a member comparing their own dataset could recover the other three.
  • Rounding (BENCHMARK_ROUNDING). Every published quartile is rounded: pounds to the nearest £500, percentages to whole numbers, spans and layers to 0.5. With five contributors the quartiles are contributors' own values, so rounding is what stops them being read back exactly.

Audit​

Every change, upload, override, export, reveal, match, purge and administrative action (including reference data and prompt publishing and prompt tests) is appended to audit_log with audit(), in the same transaction as the change.

  • workbench_app has INSERT and SELECT on audit_log but no UPDATE or DELETE grant or policy, so the trail can't be edited through the application.
  • The insert policy requires actor_id = app.current_user_id(): no one can write entries as someone else.
  • Leads see their engagement's trail; admins see everything. Rows survive an engagement purge; their details never contain names or emails.

Web and API hardening​

  • Same-origin SPA and API, so there's no CORS surface. Every API response is no-store, with Hono's secure headers.
  • Request bodies are size-capped while streaming (body() in routes/helpers.ts: 256 KB by default, larger for row chunks and reference data drafts) and validated with zod.
  • Upload chunks are limited to 2,000 rows; datasets to 50,000.
  • Route IDs must be UUIDs, or the route answers 404.
  • Database errors are mapped to generic messages, never internals. Unexpected errors are logged as their class, Postgres code and route only (loggableError in errors.ts): a failed query's message includes its parameters, which can be row data.
  • Spreadsheet parsing is bounded (file size, rows, columns, cells, cell length, and the XLSX archive's uncompressed size and entry count) against zip bombs and oversized input.
  • CSV exports neutralise formula-like cells (CSV injection).
  • Model output is rendered with react-markdown, which never renders raw HTML. Deliverables are self-contained HTML with escaped content and no external scripts.
  • Model names are restricted to [A-Za-z0-9._-] before being placed in the Gemini URL, and a prompt may only name a model in GEMINI_ALLOWED_MODELS.
  • CSRF. Middleware in app.ts refuses POST, PUT, PATCH and DELETE with 403 cross_site_request when Sec-Fetch-Site is cross-site or the Origin host differs from Host, and with 415 unsupported_media_type when a body isn't application/json.
  • Static headers. apps/web/public/_headers gives the SPA nosniff, Referrer-Policy: same-origin, X-Frame-Options: DENY, a permissions policy and frame-ancestors 'none'. Scripts aren't restricted by CSP because export previews run the export's inline script in a same-origin window.
  • Worker. Development sign-in is refused in the Worker whatever ENVIRONMENT says; Access tokens must be RS256 with exp; workers_dev and preview URLs are off.
  • Local server. The Node server binds to HOST (default 127.0.0.1) and, under development sign-in, answers 403 forbidden_host unless the Host is localhost, 127.0.0.1 or [::1] (DNS rebinding).
  • Regular expressions. titlePatternProblem refuses administrator patterns with nested repetition such as (a+)+, backreferences, lookbehind, or more than three open-ended wildcards, so a title rule can't take exponential time.
  • Prompt injection. Titles are sent to the classifier as JSON string literals and the prompt treats them as data; a batch where one US code takes more than 60% of eight or more titles isn't written to the shared cache (isAnomalousBatch).
  • Spend. Gemini thinking tokens are metered, each call in its own transaction so a failed write can't lose the record, and "Ask the data" is capped per person per UTC day (AI_ASK_DAILY_LIMIT).
  • AI opt-out. A dataset created with useAi: false stores that choice (datasets.use_ai); no later classify call, including Continue, can send its titles to the model.

Reporting a vulnerability​

Raise it privately with the Workbench owners in Davies Technology. Don't open a public issue.