Deployment and operations
The canonical runbook is docs/OPERATIONS.md; this page follows it.
Environments
| Environment | Runs on | Database | Sign-in |
|---|---|---|---|
| Local | npm run dev (Node + Vite) | Embedded PGlite in .data/ (or DATABASE_URL) | Development sign-in as dev@davies-group.com |
| Staging | Worker ai-pob-workforce-workbench-staging | Neon branch staging | Cloudflare Access |
| Production | Worker ai-pob-workforce-workbench | Neon branch main | Cloudflare Access |
The Worker serves the built SPA and the API (/api/*) from one hostname, and runs the retention job nightly at 03:00 UTC (triggers.crons in wrangler.jsonc).
First-time set-up
1. Database (Neon)
-
Create a Neon project in an EU region with a
mainbranch (production) and astagingbranch. -
Use the project's owner role in the connection string. The first migration creates the
workbench_approle the API switches to for every request and grants it to the owner, so the owner needsCREATEROLE(Neon's default owner has it). -
Apply migrations and seed reference data and prompts:
DATABASE_URL='postgres://…/neondb?sslmode=require' npm run db:migrate -
Set the branch's history retention (point-in-time restore window) as short as is comfortable (for example one day). Purged row data persists in history until the window passes; identities are unreadable there regardless.
2. Cloudflare Access
- In Zero Trust, add a self-hosted application covering the whole Workbench hostname (not just
/api). - Add a policy allowing the Davies identity provider group that should use the Workbench.
- Note the application's AUD tag and the team domain (
<team>.cloudflareaccess.com).
3. Worker configuration
Non-secret variables go in apps/api/wrangler.jsonc under vars (per environment), not in the dashboard: without keep_vars, every deploy replaces the dashboard's variables with the file's. Secrets set with wrangler secret put are kept.
| Variable | Example | Purpose |
|---|---|---|
ENVIRONMENT | production | production or staging; development sign-in is refused in both. |
ACCESS_TEAM_DOMAIN | davies.cloudflareaccess.com | JWT issuer and JWKS location. |
ACCESS_AUD | 3f1c… | The Access application's AUD tag. |
ALLOWED_EMAIL_DOMAINS | davies-group.com | Domains provisioned on first sign-in. |
BOOTSTRAP_ADMIN_EMAILS | someone@davies-group.com | Become administrators when first provisioned. |
GEMINI_MODEL | gemini-3.1-flash-lite | The default AI model. |
AI_MONTHLY_BUDGET_USD | 50 | AI pauses when month-to-date spend reaches this. |
AI_INPUT_PRICE_PER_MTOK, AI_OUTPUT_PRICE_PER_MTOK | 0.1, 0.4 | Prices for the spend estimate. |
Secrets:
cd apps/api
npx wrangler secret put DATABASE_URL # Neon owner connection string
npx wrangler secret put IDENTITY_KEK # base64 of 32 random bytes
npx wrangler secret put GEMINI_API_KEY # optional; omit to run without AI
# add --env staging for staging
Generate the identity key once per environment and store it in the Davies password vault:
node -e "console.log(require('crypto').randomBytes(32).toString('base64'))"
If IDENTITY_KEK is lost, every stored name and email becomes permanently unreadable (the analysis is unaffected). Changing it has the same effect, because key rotation isn't implemented.
AI_PROVIDER doesn't need setting in staging or production: it defaults to gemini when GEMINI_API_KEY is present and none otherwise, and fake is refused there.
4. Deploy
npm ci
npm test && npm run typecheck
npm run deploy --workspace @workbench/api # production: builds the SPA, then wrangler deploy
npm run build && npx wrangler deploy --env staging -c apps/api/wrangler.jsonc # staging
The custom domain is declared under routes in wrangler.jsonc (custom_domain: true), and the deploy attaches it. Create the Access application covering the hostname first.
Production deploys from GitHub through Workers Builds: build command npm run build, deploy command npx wrangler deploy -c apps/api/wrangler.jsonc, root directory /. The config file sits in apps/api, so the deploy command must point at it; run from the root without -c, Wrangler stops at the workspace root. Workers Builds deploys under the name of the Worker it's connected to, so keep name in wrangler.jsonc the same (ai-pob-workforce-workbench).
wrangler.jsonc serves apps/web/dist as static assets with single-page-application fallback, runs the Worker first only for /api/*, enables observability, and sets nodejs_compat.
Routine operations
- Schema changes. Add a new migration file, never edit an applied one. Run
npm run db:migrateagainst staging, deploy staging and test, then repeat for production. Each migration runs in its own transaction. Migrate before deploying code that needs the change. - Reference data and prompts. Changed through Admin, not deployments.
npm run db:migrateseeds version 1 only if it's missing; it never overwrites published versions. Changing the built-in data in code affects only new databases, so publish the same change as a new version through Admin → Reference data (the JSON import helps) for existing ones. - Retention. The nightly cron purges engagements closed for longer than their retention period. Results appear in the Worker logs (Retention purged N engagement(s)) and the audit trail (Purged personal data, System, reason: retention).
- AI spend. Admin → AI usage shows spend against the budget by month and feature. At 100% AI features pause until the 1st; rules and the cache carry on. Raise
AI_MONTHLY_BUDGET_USDto resume sooner. - Releases reach open tabs. Each build publishes
/version.jsonwith its build id (the commit, under Workers Builds) and carries the same id in the app. An open tab checks the file every minute while it's visible, and when it's brought back into view, and offers Refresh now once a different build is deployed. A tab still running the old build that moves to a page, or starts an export, whose file the new build replaced reloads once onto the new build; a second failure within 30 seconds shows the error page instead, so a broken file can't cause a reload loop. Nothing to configure. - Logs. Workers observability is enabled; use the dashboard or
npx wrangler tail. - Health.
GET /api/health(behind Access, but unauthenticated by the app) returns{ ok, version, environment }.
Runbooks
| Situation | Action |
|---|---|
| Someone needs access | They sign in; if their domain is allowed, they're provisioned as a member. A lead adds them to the engagement team, granting PII access only if needed. |
| Someone needs to be an administrator | An existing admin changes their role in Admin → Users. Admins change other people's roles, not their own, and the API refuses to let the last active administrator give up the role. |
| Someone leaves | Admin → Users → Deactivate. Access ends immediately everywhere; their audit history stays. |
| An upload or classification stopped part-way | Classification: the engagement's Datasets tab → Continue; it's resumable and idempotent, and keeps the upload's AI choice. An upload that stopped before every row arrived can't be continued: delete the dataset and upload the file again. |
| A classification is wrong everywhere | Admin → Classification cache: remove the entry so it's re-classified next time. For the current engagement, add an override. If a title rule is wrong, fix it in reference data. |
| A client asks for their data to be deleted | Engagement → Settings → Purge personal data now (lead or admin; type the engagement name). Summaries remain; delete the engagement as an admin to remove those too. |
| Reference data published with a mistake | Admin → Reference data → Version history → Draft from this on the last good version, then publish it as a new version. Datasets rescored onto the bad version need rescoring again. |
| A prompt change made AI output worse | Admin → Prompt Lab → Version history → Draft from this on the previous version, test, publish. |
| "Reference data hasn't been set up" | The database was never migrated and seeded: run npm run db:migrate against it. |
| Restore from backup | Before restoring, list the purges made since the restore point (Admin → Audit trail, Purged personal data entries), because the restore rolls the audit trail back too. Restore the Neon branch to the point in time; the Worker is stateless and needs no redeploy. Then purge those engagements again. |
Dependency notes
npm audit reports advisories in two export libraries; neither affects the shipped application:
- image-size (via pptxgenjs): pptxgenjs declares
"browser": { "image-size": false }, so it's never bundled; decks are built in the browser from data URIs. - uuid (via exceljs): the advisory concerns
v3,v5andv6with a caller-supplied buffer; exceljs's browser build calls onlyv4.
npm audit fix --force would downgrade both libraries by several major versions, so don't run it.
This documentation site
The site in docusaurus/ builds to static files (npm run build there, output in docusaurus/build). Host it behind the same Cloudflare Access policy as the Workbench (it describes internal systems), for example as a separate Pages project or a path on an internal docs host, and set url in docusaurus.config.ts to match.