services/api/product_app.py and its routers, or frontend/src/app/(console) and admin/(portal)) needs before their first change. Every term also appears in the Glossary in alphabetical order.
Tenant
A tenant is one customer organization. It is a row incontrol.tenants — slug, display name, industry, region, a lifecycle_state, and pointers to its own PostgreSQL schema (schema_name, always tenant_<slug>) and object-storage prefix. Every tenant-scoped table across the platform (sources, datasets, records, model_checkpoints, caps, reviews, …) also carries a tenant_id column and lives inside that tenant’s own schema, never in a shared table keyed by tenant.
The frontend calls a tenant a workspace in end-user-facing copy (see GET /workspaces, the pre-login directory of live tenants for the login page’s picker) — same entity, different word depending on audience.
Tenant lifecycle
A tenant’slifecycle_state is constrained by an explicit CHECK constraint in the control-plane migration (migrations/versions/20260722_c001_control_schema.py) to exactly these eight values:
draft, onboarding, review, live, offboarded. POST /admin/tenants/{tenant_id}/provision never sets provisioning itself — it just enqueues a provision_tenant job (ensure_provision_job) and the tenant stays draft until that job’s activate_onboarding step (services/api/tenancy/provisioner.py) moves it straight from draft/provisioning to onboarding. Nothing writes training, and nothing writes or reads out of suspended — both are reserved states allowed by the constraint but not wired to any code path yet:
POST /admin/tenants/{tenant_id}/go-live is a real checklist gate, not a status flip on request: it refuses with 409 and names exactly what’s missing — dataset ingested, checkpoint activated, user invited — until all three are true, then sets lifecycle_state to live. Key onboarding milestones are written to control.audit_log — tenant_created, dataset_ingested, checkpoint_activated, and tenant_go_live (services/api/onboarding_api.py), plus employee login and impersonation — but this is not exhaustive: the provisioner’s draft/provisioning -> onboarding transition and routes like PUT /training-config, PUT /data-policy, POST /train, POST /invite, and POST /sources/seed write no audit-log row today.
services/api/onboarding_api.py requires every route to have an authenticated employee; only the onboarding_admin control role can execute mutations, while support_readonly can view everything but change nothing.Control plane vs. tenant schema isolation
Causeloop runs two Alembic branches against the same PostgreSQL cluster:- Control branch (
controlschema): includingemployees,employee_sessions,tenants,tenant_invitations,audit_log,platform_jobs,password_reset_tokens,runtime_heartbeats. One copy, shared across all tenants, owned by the onboarding portal. tenant_templatebranch: applied once per tenant into its owntenant_<slug>schema — includingusers,user_sessions,user_roles,role_profiles,sources,datasets,records,ingest_jobs,model_checkpoints,llm_settings,insight_collections,reviews,caps,settings,pii_field_policies,password_reset_tokens,audit_log,source_uploads.
tenant_id and a FORCE ROW LEVEL SECURITY policy on top of schema separation — even an unscoped app_rw connection sees zero rows on any tenant table, so a bug that forgets to filter by tenant fails closed instead of leaking cross-tenant data. get_tenant_session/get_admin_session (services/api/tenancy/deps.py) are the only way a request handler gets a database connection: get_tenant_session yields a TenantSession, always resolved from a TenantRecord, never a bare string; get_admin_session yields a control-plane ControlSession (services/api/db/session.py) directly, with no tenant involved at all. This is deliberately two independent layers, not one: schema separation means a correctly-scoped query can never even see another tenant’s rows, and row-level security means that even if a call site forgot to scope a query, the database itself refuses the cross-tenant read. See Tenancy and data model for the enforcement details and the isolation test suite (tests/test_tenancy_isolation.py).
Platform jobs and the two workers
Long-running work — tenant provisioning, training, insight materialization, outbound email — runs as a row incontrol.platform_jobs, not as inline request-handler work or a daemon thread. A job has a job_type, a status (queued, running, succeeded, failed, cancelled), progress, and steps_completed. Job creation commits atomically with the application write that triggered it (for example, a source sync and its materialize_insights job commit together), so a crash never leaves a dangling side effect.
Two separately privileged worker processes both run scripts/run_pipeline_worker.py, distinguished by the CAUSELOOP_WORKER_JOB_TYPES job types each claims, a CAUSELOOP_WORKER_COMPONENT name, and the database credentials they’re started with (infra/docker-compose.product.yml):
Workers claim jobs with
FOR UPDATE SKIP LOCKED, retry with bounded backoff, and recover jobs whose lease expired without a heartbeat. GET /health/ready on the product API refuses to report ready unless both workers have heartbeated recently in control.runtime_heartbeats and advertise the job types the API depends on.
Model checkpoints
A checkpoint is a row in a tenant’smodel_checkpoints table: an immutable, versioned snapshot of a trained clustering/embedding model, produced by POST /admin/tenants/{tenant_id}/train against the tenant’s latest ingested dataset. It records the algorithm, embedding provider, metrics, and an is_active flag — POST /admin/tenants/{tenant_id}/activate picks which checkpoint version powers the tenant’s live insights. See Training and checkpoints.
Insight snapshot and materialization
Materialization is the act of computing a tenant’s insight collections and writing them toinsight_collections, keyed by (tenant_id, collection, model_version, dataset_version). It runs as a materialize_insights platform job whenever a connector sync produces a new dataset version, and reuses the same computation the legacy file-based client API used (compute_model_collection/compute_llm_collection in causegraph.onboarding.insights, refactored into storage-agnostic *_from_data functions). Five model-backed collections always attempt — summary, themes, issues, severity, lens — and five LLM-backed collections are skipped, not failed, with a clear reason when no LLM is configured or reachable: fishbone, theme_summaries, narratives, recommendations, cause_remediation. One of those, cause_remediation, is export-only (services/api/insights_api.py) — it materializes but never appears in the client-visible snapshot, leaving four LLM-backed collections a tenant can actually see.
An insight snapshot is what the tenant console actually reads: GET /insights/snapshot picks one coherent (dataset_version, model_version) pair — never a mix of collections from different pipeline runs — and returns every client-visible collection materialized for that pair, each with its own ready/generating/disabled status plus model_ready/all_ready flags. A row from an onboarding-era snapshot only surfaces if the employee turned share_onboarding_outputs on for that tenant; rows produced by a real post-go-live ingest (source="ingest") are always visible. See Insights console.
Invitations
An invitation is a row incontrol.tenant_invitations: an email, a target role-profile key, a hashed single-use token, and a status (pending, accepted, revoked, expired). POST /admin/tenants/{tenant_id}/invite creates one and queues a send_email platform job carrying an absolute application link; accepting it (POST /accept-invite) creates the tenant-schema users row, assigns the role profile, and logs the new user in. See Invitations and impersonation.
Impersonation
Impersonation is the onboarding portal’s “open portal” action (POST /admin/tenants/{tenant_id}/impersonate): it mints a real tenant session for that tenant’s earliest-created active user, so an onboarding admin sees exactly what the tenant’s own admin sees. It only works once a tenant is live, and every use is unconditionally written to control.audit_log — this is treated as a real access grant into customer data, not a read-only preview.
Data policy
Each tenant has a small data policy preference set — residency (US/EU/APAC), retention (5 years/7 years/10 years), and a PII-redaction toggle — set through PUT /admin/tenants/{tenant_id}/data-policy and stored in the tenant’s own settings row alongside training config. See Data policy.