Testing
npm test # Vitest in every workspace
npm run typecheck # strict TypeScript in every workspace
npm run e2e # Playwright journey against a throwaway database
npm test and npm run typecheck must pass before every commit.
Unit and integration tests (Vitest)
Each workspace has its own vitest.config.ts and runs test/**/*.test.ts. Run one workspace with:
npm test --workspace @workbench/core
npm test --workspace @workbench/api
npm test --workspace @workbench/web
# or
npx vitest run --root packages/core
npx vitest run --root apps/api test/security.test.ts
Core (packages/core/test)
Pure model tests: classification and SOC validation, the pipeline, reference data (validation, cross-checks, compilation, the built-in data's warnings), prompts (rendering and validation), provenance (every metric covered and rooted), reports, summaries, org graphs (cycles, dangling managers), the spans and layers report model (org-report.test.ts: disconnected managers and the sensitivity), savings by group (savings.test.ts: multipliers, FTE saved, and roll-ups that add up to the scenario total), report wording (report-notes.test.ts), the Copilot case and rollout, ASHE salaries and the finance-workbook model (formulas reproduce the platform's scenarios exactly, and the group sheets add up). identifiers.test.ts covers the rule that IDs are never emails or names.
Every model change needs a test here, including edge cases: empty input, a missing SOC, cycles, zero salaries.
API (apps/api/test)
The API tests run the real app, migrations, seed data and row-level security on an in-memory PGlite database per file, so they exercise the actual policies rather than mocks. test/harness.ts provides:
| Helper | Does |
|---|---|
createHarness({ ai?, env?, now? }) | PGlite + migrations + seedPlatformData, a test config (ENVIRONMENT=test, dev sign-in, a fixed KEK), a scripted FakeAi, and call(method, path, { as, body }) which sends requests through app.request as any user (x-dev-user). |
ADMIN, LEAD, ANALYST, VIEWER, OUTSIDER | Standard test identities (ADMIN is the bootstrap admin). |
provision(h, ...emails) | Signs users in once so they exist, returning their IDs. |
engagementFixture(h, { lead, pack, members, share }) | A client and engagement with a team. |
ingest(h, engagementId, rows, …) | Creates, uploads, completes and classifies a dataset. |
FakeAi, classifierResponder(table) | A scripted model that records calls; a responder that answers classifier prompts from a title → codes table. |
| File | Covers |
|---|---|
auth.test.ts | Access JWT validation, config refusal rules, provisioning, deactivation, the last admin, /me. |
security.test.ts | Forbidden things against real RLS: cross-engagement reads, identities without PII access (admins included), granting yourself PII access (through the API and directly as the app role), viewer and outsider writes, the append-only audit log, the service-only cache, RLS user ID spoofing, closed engagements, retention and purge rules. Also refuses an email (look-alikes included) as an employee or manager ID. |
pipeline.test.ts | Upload, chunk idempotency, classification ladder, the cache and opt-out, AI off, overrides and rescoring. |
features.test.ts | Clients, AI narratives and budget, Copilot plans, benchmarks, admin endpoints, audit. |
reference.test.ts | Seeding, admin-only drafts, validation, immutable publishing, pinning and rescoring, pack checks, audit. |
prompts.test.ts | Seeded prompts, drafts and publishing, dry runs, test runs judged by the parsers, the development stand-in. |
services.test.ts | Crypto (wrapping, AAD binding, blind index), the classifier parser and batching, narratives and their fallbacks. |
ai.test.ts | The Gemini client, choosing a client, the development stand-in model, metering and the monthly budget. |
hardening.test.ts | Cross-site request defences, localhost-only development sign-in, cache poisoning by prompt injection, AI spend controls, purged datasets' cached AI outputs, stale draft saves and racing publishes, administrators. |
savings-notes.test.ts | FTE and client attributes stored and returned with rows but never in the cached summary; an FTE that rounds to zero stored as unknown; personal-data headings however written, look-alike @ signs and duplicate headings refused; a grouping with as many values as people refused at completion; report notes set, read back and cleared by a lead only, leaving other settings alone, malformed notes rejected, audited without the text, fixed once closed and cleared on purge. |
review-fixes.test.ts | The stored AI choice, overrides set mid-classification, AI picks kept on rescore, overrides against pinned versions, the value chain without AI, pack changes, stale AI text, and what error logs contain. |
Security-relevant changes need a test that tries the forbidden thing.
The harness's FakeAi is a different thing from the development stand-in (AI_PROVIDER=fake): the harness's is scripted per test; the stand-in is a deterministic model for running the app.
Web (apps/web/test)
Pure helpers only: the API client's session handling, ingest mapping and row building (including FTE and extra groupings, ingest-attributes.test.ts), spreadsheet parsing limits, reference editor helpers, and exports (including the spans and layers report, org-report.test.ts, and report wording in the report, Markdown and deck, report-notes.test.ts). Components aren't unit-tested; journeys are covered end to end.
End-to-end (Playwright)
apps/web/e2e/workbench.spec.ts walks one journey with a small fictional staff list: create an engagement for a new client, upload a CSV keeping only active staff, see the report, the organisation, the Copilot case and an export. Other specs cover particular areas; set-up that isn't under test goes through the API helpers in e2e/support.ts.
| Spec | Covers |
|---|---|
dataset.spec.ts | The report, overrides, assumptions, rollout plans, every export, the Explorer, compare and benchmarks. |
savings.spec.ts | The Savings by group tab adding up to the report by department and by an extra grouping, the Explorer CSV's new columns, the lead's report wording in the downloaded report, and an analyst seeing the wording read-only. |
spans-report.spec.ts | The Organisation tab's narrow-spans count and link, the spans report's managers, disconnected managers and sources, and the manager CSV being recorded. Also checks that an email column mapped as the employee ID is refused before upload. |
branding.spec.ts | Client brands and branded exports. |
permissions.spec.ts | Viewers, audited name reveals, and people outside the team. |
auth.spec.ts, admin.spec.ts, reference.spec.ts, prompt-lab.spec.ts | Sign-in, the admin screens, reference data and the Prompt Lab. |
mobile.spec.ts | The main pages, including Explorer, Savings by group and engagement Settings, fitting a phone screen. |
npx playwright install chromium # once, for the bundled browser
npm run e2e
# or use an installed browser instead of downloading one
PW_CHANNEL=msedge npm run e2e # bash
$env:PW_CHANNEL = "msedge"; npm run e2e # PowerShell
apps/web/playwright.config.ts starts its own servers, on ports that don't collide with npm run dev:
- the Node API on port 8797, with
DATA_DIRset to a fresh temporary folder (a throwaway PGlite database),AI_PROVIDER=fake(the development stand-in) and an emptyGEMINI_API_KEY; - Vite on port 5183, proxying
/apito 8797.
Tests run serially with one worker and a 90-second timeout; traces are kept on failure. On CI the reporter is github.
Where downloading Playwright's Chromium isn't possible, PW_CHANNEL=msedge (or chrome) uses the installed browser.
What a change needs
| Change | Tests |
|---|---|
| Model logic in core | Core unit tests, with edge cases. |
| A new or changed metric | Provenance test coverage. |
| A new API route | Harness tests for the happy path, each role, and the forbidden cases. |
| A schema change or new table | Harness tests that the RLS policies hold (read and write, each role). |
| Reference data shape | Core validation tests; referenceEdit tests if the editor changes. |
| A prompt or parser | Core prompt tests; API prompt tests with FakeAi. |
| An export | Pure-part tests in apps/web/test/exports.test.ts. |
| A user journey | Extend the relevant Playwright spec, or add one. |