Skip to main content

AI integration

AI is optional and bounded. Four features use it: title classification, the executive summary, the value-chain map and Ask the data. Each has a deterministic fallback, validates the model's output against a fixed contract, requires citations, is metered against a monthly budget, and is labelled in the UI. With AI off, every feature still works and says so.

The client​

apps/api/src/services/ai.ts is the only door to a language model:

interface AiRequest {
system: string;
user: string;
temperature?: number;
maxOutputTokens?: number;
model?: string | null; // a prompt version may name its own model
hints?: AiHints; // structured inputs, for the development stand-in
}

interface AiClient {
readonly model: string;
generateJson(req: AiRequest): Promise<{ text: string; inputTokens: number; outputTokens: number }>;
}

createAiClient(config) returns:

AI_PROVIDERClient
geminigeminiClient: Google Gemini generateContent with responseMimeType: application/json, the system instruction, temperature and token limit; a 30-second timeout (AI_TIMEOUT_MS); the API key in the x-goog-api-key header; model names restricted to [A-Za-z0-9._-]{1,100}.
fakefakeAiClient (ai-fake.ts): the deterministic development stand-in (workbench-dev-stand-in), refused outside development and test.
nonenull: AI off.

The key never reaches the browser.

The development stand-in​

fakeAiClient answers from req.hints rather than the prompt text, so it keeps working however prompts are edited in the Prompt Lab:

  • classifier: maps titles by keyword (solicitor → 23-1011/2412, software → 15-1252/2134, claims → 13-1031/4159, and so on) with plausible confidences, and a 0.45-confidence fallback so the review queue is exercised;
  • exec-summary and ask: a clearly labelled development stand-in markdown, citing the dataset;
  • lifecycle: a four-phase chain (Engage, Deliver, Support, Govern) with each role in one phase, at low confidence.

Token counts are estimated from text length, so budget and usage views have data. The e2e suite runs on it.

Budget and metering​

  • Every call records ai_usage (recordAiUsage: user, engagement, feature, model, input and output tokens), skipping zero-token calls.
  • monthToDateAiSpendUsd sums the calendar month's (UTC) tokens across the platform, priced at AI_INPUT_PRICE_PER_MTOK and AI_OUTPUT_PRICE_PER_MTOK.
  • aiUnavailableReason(ai, db, config, now) returns AI is not configured on this environment. or This month's AI budget (US$N) has been reached., or null.
  • assertAiAvailable throws 503 ai_unavailable or 429 ai_budget_exceeded (used by Ask).

Features: classify, exec-summary, lifecycle, ask, prompt-lab. The admin AI usage view groups by month and feature.

Pricing per prompt model

A prompt version can choose another model, but spend is always estimated at the configured prices. If a prompt uses a pricier model, raise the configured prices or the estimate will be low.

Prompts​

Every call uses the published version of its prompt, read with activePrompt(db, key) on the service connection (prompts are admin-only under RLS). Rendering is renderPrompt(text, vars), which fills {{variables}} verbatim. See Prompts.

Title classification​

services/classifier.ts:

  • classifierVariables(batch, { industry, region }) builds {{context}} (Industry/sector: … Region: …. plus a blank line, or empty) and {{titles}} (a 1-based numbered list, whitespace collapsed, each title cut to 300 characters).
  • classifyTitles(ai, ref, prompt, titles, ctx) de-duplicates, batches 150 titles, runs 4 batches at a time and never throws: a failed batch is counted and its titles fall back to the rules.
  • parseClassifierResponse(ref, text, batch) accepts rows { i, u, k, c, z } only when:
    • i indexes a title in this batch (titles are matched by index, so a model that rewords a title can't misattribute a code), and each title is used once;
    • u is in the US catalogue and k has UK pay (ASHE or proxy) in the dataset's reference version;
    • z cites exactly us-soc-2018:<u> and uk-soc-2020:<k>;
    • c is clamped to 0–1 (default 0.5).

Accepted picks become source: "ai" classifications; low confidence goes to the review queue. The step meters the call in its own transaction before writing, so tokens are recorded even if the write fails, and writes picks only to titles still pending, so an override set while the model was working wins. The classification step and cache are described in Upload and classification.

Narratives​

services/narratives.ts. The model only ever sees the dataset's anonymised aggregates and may not introduce new figures.

OutputVariablesAccepted whenFallback
Executive summary (executiveSummary)facts (JSON from facts(name, sector, summary): headcount, cost, averages, scenarios, net of reinvestment, moderate investment, top five roles, top ten categories, offshore, organisation), citationSourceIdValid JSON; markdown ≥ 40 characters; a valid confidence; at least one citation of dataset:<id>A deterministic summary with the position, where value concentrates, what it takes, and caveats. It has no heading (the card is titled already), opens with the dataset name in bold, and formats figures with the shared helpers (pct to one decimal, gbpRange, timesRange) so they match the stats beside it.
Value chain (industryLifecycle)sector, roles (the 120 largest role types: name (category, headcount)), citationSourceIdparseLifecycleResponse: 3–7 phases kept, weights normalised and keyed by exact role names, cited; population > 0. Roles the model didn't place (including any beyond the 120) are spread evenly and reported in lifecycle.mapping, never counted as mapped.The summary's own value chain (summary.report.lifecycle), the lens every export uses.
Ask (askDataset)facts, question, citationSourceIdAs for the summaryA polite refusal: it couldn't ground an answer.

Citations are validated with validCitations against an allow-list containing only dataset:<id>. The result carries generated, confidence and citations; the API adds reviewRequired when an AI result has low confidence.

Routes and caching​

routes/ai.ts, mounted at /api/datasets/:id/ai:

  • GET /:kind (exec-summary, lifecycle): returns the cached AI version from dataset_ai if there is one, otherwise computes the rule-based version on the fly, so opening a report never spends tokens.
  • POST /:kind (leads and analysts): generates with the published prompt (or explains why not, returning aiUnavailableReason with the fallback), meters the call in its own transaction, then caches the result in dataset_ai with the prompt version and audits ai_generated. If the dataset was re-scored while the text was being written (finalized_at changed since the summary was read), nothing is cached and the request fails with 409: the text describes figures that have since moved.
  • POST /ask (any member): asserts availability, answers, meters and audits ai_generated with the question's length (not its text). Answers aren't cached.

finalizeDataset deletes a dataset's dataset_ai rows, so any re-score (override, assumptions, reference version) clears AI text that would quote stale figures.

Datasets that are Summary only

The AI routes require status = 'ready'. For a purged dataset (status = 'purged'), even GET /ai/exec-summary answers 409 (This dataset hasn't finished classifying yet.), so the Overview shows no executive summary for purged datasets, although the summary and any cached AI text still exist.

In the web app​

components/ui/Ai.tsx frames every AI surface (AiSurface): an AI-generated (or AI-mapped value chain) pill or Rule-based — no AI-written text, a confidence pill, Review before sharing when required, and the citations. Markdown is rendered with react-markdown and remark-gfm, which never render raw HTML. Controls are hidden or disabled when /api/me reports aiEnabled: false.

Adding an AI feature​

Follow the pattern: a prompt key and definition with a fixed output contract (see Add a prompt), a parser that validates and requires citations, a deterministic fallback, metering with a new feature name, an audit action, and labelling in the UI.