> ## Documentation Index
> Fetch the complete documentation index at: https://docs.causeloop.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Client onboarding pipeline

> The end-to-end tenant lifecycle: create, provision, ingest, train, activate, invite, and go live — every API call, validation, worker job, and database effect.

Onboarding a new client is a single linear pipeline driven entirely through `services/api/onboarding_api.py`'s `onboarding_router` (mounted at `/admin/tenants`), executed either step-by-step from a tenant's [detail page](/features/staff-portal) or in one guided pass from the [new-tenant wizard](/features/staff-portal). Both UIs call the exact same endpoints — the wizard just sequences them with extra client-side state (a slug preview, a step indicator, source-kind checkboxes) that never reaches the backend as its own concept.

Every mutating step requires the `onboarding_admin` control role; `support_readonly` employees can load every `GET` in this router but every mutation 403s for them (`Depends(require_employee_role("onboarding_admin"))`, aliased in the router as `_MUTATE`). Every mutation except `POST /admin/tenants` also requires the CSRF double-submit header (`Depends(require_csrf)`) — see [Security model](/architecture/security). `create_tenant` is declared with `dependencies=[_MUTATE]` only, so a create-tenant request without the CSRF header currently succeeds; that's a gap relative to every other mutation in the router, not an intentional exemption.

## Lifecycle state machine

`control.tenants.lifecycle_state` is `CHECK`-constrained to eight values: `draft, provisioning, onboarding, training, review, live, suspended, offboarded` (see [Tenancy and data model](/architecture/tenancy-and-data-model)). Only four of them are ever actually written by code running today:

```mermaid theme={null}
stateDiagram-v2
    [*] --> draft: POST /admin/tenants (create_draft_tenant)
    draft --> onboarding: POST .../provision → provision_tenant() completes
    onboarding --> review: POST .../activate (checkpoint_activated)
    review --> live: POST .../go-live (all three prerequisites met)
    live --> offboarded: scripts/offboard_tenant.py (operator-run, no API route)
```

`provisioning`, `training`, and `suspended` are reserved in the schema's `CHECK` constraint but no route or worker in the current codebase ever sets them: `provision()` accepts a tenant whose state is `draft`, `provisioning`, *or* `onboarding` as a valid input (so a stuck or re-run provision is still acceptable), but `create_draft_tenant()` only ever inserts `draft`, and `provision_tenant()`'s last step writes `onboarding` directly — there is no intermediate persisted `provisioning` row in practice. Training likewise never flips the tenant record itself; only the underlying `platform_jobs`/`model_checkpoints` rows change. Treat those three as forward-reserved states, not states you'll observe in `GET /admin/tenants/{tenant_id}`.

## 1. Create — `draft`

**Call:** `POST /admin/tenants` with `CreateTenantRequest` (`slug`, `name`, `industry?`, `region?`, `data_residency_notes?`, `primary_contact_email?`).

**Validation:** `validate_tenant_slug()` (`services/api/tenancy/db.py`) enforces the pattern `^[a-z0-9][a-z0-9-]{1,61}[a-z0-9]$` — lowercase alphanumeric and hyphens, 3–63 characters, must start and end with an alphanumeric — before anything else runs; a mismatch is a `422` with the pattern echoed in the error detail. A slug that already exists in `control.tenants` is a `409`.

**Effect:** `create_draft_tenant()` (`services/api/tenancy/provisioner.py`) inserts one `control.tenants` row with `lifecycle_state='draft'` (the column default), `schema_name = f"tenant_{slug}"`, and a temporary `storage_prefix` immediately overwritten with `tenants/<tenant uuid>` once the row's real id is known. A concurrent duplicate-slug race is resolved by catching the resulting `IntegrityError` and returning the durable winner rather than a `500`. The route then writes a `control.audit_log` row (`action="tenant_created"`) in a second transaction. **No data-plane schema exists yet** — the tenant is a control-plane record only.

## 2. Provision — `draft`/`onboarding` → `onboarding`

**Call:** `POST /admin/tenants/{tenant_id}/provision`, `202 Accepted`.

**Validation:** the route rejects any tenant whose `lifecycle_state` isn't `draft`, `provisioning`, or `onboarding` with a `409` — you cannot re-provision a `live` tenant through this route.

**Job:** `ensure_provision_job()` synchronously inserts a `provision_tenant` row in `control.platform_jobs` (status `queued`, idempotency key `provision:{tenant_id}`) *before* returning `202`, specifically so a client that immediately polls `GET .../provision/status` never observes `not_started`. (The wizard's own poller — `tenants/new/page.tsx` — keeps refetching on both `not_started` and `running`, so it isn't caught out by the initial `queued` status either.) The dedicated `provisioner-worker` process (distinct credentials from the product worker — see [Workers and jobs](/architecture/workers-and-jobs)) claims it and runs `provision_tenant()`, six named, independently idempotent steps tracked in `steps_completed`:

| Step                    | What it does                                                                            |
| ----------------------- | --------------------------------------------------------------------------------------- |
| `create_schema`         | `CREATE SCHEMA IF NOT EXISTS tenant_<slug>`                                             |
| `run_tenant_migrations` | Runs the `tenant_template` Alembic branch to head inside that schema                    |
| `grant_app_rw`          | `GRANT`/`ALTER DEFAULT PRIVILEGES` so `app_rw` can read/write the new schema            |
| `seed_role_profiles`    | Inserts the four default role profiles (`org_admin`, `risk_admin`, `analyst`, `viewer`) |
| `verify_object_storage` | Fails provisioning if the configured object store isn't reachable                       |
| `activate_onboarding`   | `UPDATE control.tenants SET lifecycle_state='onboarding'`                               |

A crash mid-provision is safe to resume: `run-again` skips every step already recorded in `steps_completed`, and every step is independently safe to repeat (`CREATE SCHEMA IF NOT EXISTS`, `ON CONFLICT DO NOTHING`, Alembic's own per-schema version tracking). A failed job records `error_detail`; re-`POST`ing `/provision` picks the same job back up (`retry_terminal=True` on the enqueue) rather than creating a duplicate.

**Effect:** the tenant's isolated `tenant_<slug>` schema exists, is migrated to head, is grantable to `app_rw`, has its four seeded role profiles, and `control.tenants.storage_prefix`'s object-storage prefix has been verified reachable. `lifecycle_state` becomes `onboarding`.

<Note>
  Full mechanics — the provisioner/product-worker credential split, heartbeats, retry/backoff, stale-job recovery — are covered in [Workers and jobs](/architecture/workers-and-jobs); this page only covers the tenant-facing transition.
</Note>

## 3. Historical data ingest — upload, then commit

Ingest is two calls on purpose: the object lands in durable storage first, and only a separate, explicit commit turns it into a queryable dataset. A file that fails validation never produces a half-written `datasets`/`records` row.

**Upload — `POST /admin/tenants/{tenant_id}/ingest/upload`** (multipart, one `file` field):

1. `store_raw_upload()` streams the file to the tenant's object-storage prefix at key `tenants/<tenant_id>/uploads/<upload_id>/<safe filename>` (S3-compatible; MinIO locally, `causeloop-uploads` bucket) **before** any database write.
2. Only after the object is durably stored does it insert a `source_uploads` row (`status='staged'`) in the tenant schema, in the same transaction as nothing else — if that insert fails, the just-written object is best-effort deleted so an upload never becomes an orphaned, undiscoverable object.
3. `parse_upload()` (`causegraph.onboarding.ingestion`) parses the bytes (`.xlsx`/`.csv`/`.json`/`.jsonl`), `detect_mapping()` suggests a column→field mapping, and `_detect_pii_columns()` flags columns whose *name* matches a PII heuristic (`email`, `phone`, `ssn`, `name`, `address`, `dob`, `account_number`, `card_number`, …) with a suggested strategy (`redact`/`partial`).
4. The response returns `upload_id`, a 5-row preview, `suggested_mapping`, and `detected_pii_columns` for the operator to confirm or edit before committing.

A file too large (`UploadTooLargeError`) is `413`; an unavailable object store is `503`; a file that fails to parse (`IngestError`) is `422` and marks the just-created upload row `failed`.

**Commit — `POST /admin/tenants/{tenant_id}/ingest/commit`** with `{upload_id, mapping}`:

1. Re-loads the upload's raw bytes from object storage, re-parses, and calls `normalize_rows()` with the operator-confirmed mapping.
2. `lock_tenant_dataset_versions()` takes an advisory lock so two concurrent commits for the same tenant can't race on the next dataset version number; the new `Dataset` row gets `version = max(existing) + 1`.
3. `DatasetRecord` rows insert in batches of 500 with `ON CONFLICT (tenant_id, dataset_id, content_hash) DO NOTHING` — `content_hash` is `sha256(issue_id:text)`, so a byte-identical row re-committed into the same dataset version is a no-op, not a duplicate.
4. The same detected-PII columns from the upload step are written as `pii_field_policies` rows (`ON CONFLICT DO NOTHING`, so a manually-adjusted policy from a prior commit is never silently overwritten).
5. A `materialize_insights` job enqueues in the **same transaction** as the dataset/record writes (idempotency key `materialize:{tenant_id}:{dataset_id}:{version}`) — see [Training and checkpoints](/features/training-and-checkpoints) for what that job does once a checkpoint exists.
6. `source_uploads.status` flips to `committed`, and a `control.audit_log` row (`action="dataset_ingested"`) records the version and row count.

Re-committing an `upload_id` that's already `committed` is **idempotent by design**: the route returns the existing dataset's `{dataset_id, version, row_count}` immediately rather than creating dataset version N+1 again. Re-committing an upload whose status is `failed` is rejected with `409`.

## 4. Training configuration (not itself a lifecycle transition)

**Call:** `PUT /admin/tenants/{tenant_id}/training-config`. Covered in full, including per-tenant LLM key handling, in [Training and checkpoints](/features/training-and-checkpoints) — it only writes `settings.key_values.training_config` and an `llm_settings` row; it never touches `lifecycle_state`.

## 5. Train — produces a checkpoint

**Call:** `POST /admin/tenants/{tenant_id}/train`, `202`. Requires a committed dataset to exist (`409` otherwise: *"No ingested dataset to train on -- complete the ingest step first."*).

**Job:** enqueues `train_tenant_model` (idempotency key includes a hash of the current training config, so changing the config before re-training produces a genuinely new job rather than colliding with a stale one; `retry_terminal=True` lets a previously-failed training job be re-run from the same button). The `product-worker` runs `run_training_job()` (`services/pipeline/training_job.py`): loads every row of the tenant's latest dataset, embeds and clusters them per `TrainingConfig`, and writes a new `model_checkpoints` row (`version = max(existing) + 1`, `is_active=false`). See [Training and checkpoints](/features/training-and-checkpoints) for the embedding/clustering algorithm and checkpoint metrics.

**Effect:** a new, inactive checkpoint exists. Nothing is client-visible yet, and `lifecycle_state` is unchanged.

## 6. Activate — `onboarding` → `review`

**Call:** `POST /admin/tenants/{tenant_id}/activate` with `{checkpoint_version, share_onboarding_outputs}`. An unknown version is `404`.

**Effect (one tenant transaction):**

* Every checkpoint's `is_active` is cleared, then the named version is set `is_active=true` — exactly one active checkpoint per tenant at a time.
* A `materialize_insights` job enqueues for the tenant's latest dataset (idempotency key `materialize-activation:{tenant_id}:{dataset_id}:{checkpoint_version}`), and the matching `insight_collections` rows for that `(model_version, dataset_version)` pair have their `source`/`client_visible` updated to reflect the operator's `share_onboarding_outputs` choice.
* `control.tenants.lifecycle_state` becomes `review`, `share_onboarding_outputs` is persisted on the tenant record itself (read by the tenant console to decide whether onboarding-era data is visible before go-live), and a `control.audit_log` row (`action="checkpoint_activated"`) records which version and sharing choice.

## 7. Invite — at least one user, any time after provisioning

**Call:** `POST /admin/tenants/{tenant_id}/invite` with `{email, role_profile_key}` (default `org_admin`), `201`.

**Effect:** `create_invitation()` (`services/api/auth/tenant_auth.py`) writes a `control.tenant_invitations` row (`status='pending'`, 7-day expiry, `invite_token_hash` — never the raw token) and enqueues an encrypted `send_email` job in the same transaction. Full delivery mechanics, statuses, and the accept flow are covered in [Invitations and impersonation](/features/invitations-and-impersonation). Inviting is independent of `activate` — you can invite before, during, or after activation — but `go-live` will refuse without at least one invitation in `pending` or `accepted` status.

## 8. Go live — `review` → `live`

**Call:** `POST /admin/tenants/{tenant_id}/go-live`.

**Validation:** three independent prerequisite checks, each queried directly rather than inferred from `lifecycle_state`:

| Prerequisite                 | Query                                                                  |
| ---------------------------- | ---------------------------------------------------------------------- |
| Historical data ingested     | `count(datasets) > 0`                                                  |
| A model checkpoint activated | `count(model_checkpoints WHERE is_active) > 0`                         |
| At least one user invited    | `count(tenant_invitations WHERE status IN ('pending','accepted')) > 0` |

Any unmet prerequisite is a single `409` listing every missing item by name (e.g. *"Cannot go live yet: a model checkpoint activated, at least one user invited."*) — not just the first one found.

**Effect:** `control.tenants.lifecycle_state = 'live'` and a `control.audit_log` row (`action="tenant_go_live"`). This is also the gate `login_tenant_user()` checks — no tenant user can sign in, and no [impersonation](/features/invitations-and-impersonation) is possible, until this flips.

## The wizard's shortcut through steps 6–8

The [new-tenant wizard](/features/staff-portal)'s final step calls `activateCheckpoint()`, then `inviteUser()`, then `goLive()` sequentially from the browser — three independent HTTP calls, not one atomic operation. If `inviteUser()` fails after `activateCheckpoint()` already succeeded, the tenant is left in `review` with an active checkpoint and no invitation; the wizard surfaces the error, but the tenant's own [detail page](/features/staff-portal) can complete the remaining steps (invite, then go-live) individually from where the wizard left off, exactly as if the operator had run every step there from the start.

## Recap: every transition at a glance

| Step            | Endpoint                  | Job enqueued                            | Primary DB effect                                                          |
| --------------- | ------------------------- | --------------------------------------- | -------------------------------------------------------------------------- |
| Create          | `POST /admin/tenants`     | —                                       | `control.tenants` row, `lifecycle_state='draft'`                           |
| Provision       | `POST .../provision`      | `provision_tenant` (provisioner-worker) | tenant schema created + migrated + granted; `lifecycle_state='onboarding'` |
| Ingest upload   | `POST .../ingest/upload`  | —                                       | object store write, then `source_uploads` row (`staged`)                   |
| Ingest commit   | `POST .../ingest/commit`  | `materialize_insights`                  | `datasets` + `records` rows; `source_uploads` → `committed`                |
| Training config | `PUT .../training-config` | —                                       | `settings.key_values.training_config`, `llm_settings`                      |
| Train           | `POST .../train`          | `train_tenant_model` (product-worker)   | new `model_checkpoints` row (`is_active=false`)                            |
| Activate        | `POST .../activate`       | `materialize_insights`                  | one checkpoint flips `is_active=true`; `lifecycle_state='review'`          |
| Invite          | `POST .../invite`         | `send_email` (product-worker)           | `tenant_invitations` row (`pending`)                                       |
| Go live         | `POST .../go-live`        | —                                       | `lifecycle_state='live'`                                                   |

### Minimal end-to-end example

```bash theme={null}
# 1. Create (draft)
curl -sX POST http://localhost:18000/admin/tenants \
  -H 'content-type: application/json' -H "x-csrf-token: $CSRF" \
  --cookie "cl_admin_session=$SESSION; cl_csrf=$CSRF" \
  -d '{"slug":"aurora-savings","name":"Aurora Savings Bank"}'

# 2. Provision (draft -> onboarding)
curl -sX POST http://localhost:18000/admin/tenants/$TENANT_ID/provision \
  -H "x-csrf-token: $CSRF" --cookie "cl_admin_session=$SESSION; cl_csrf=$CSRF"

# ...ingest, train, activate, invite...

# 8. Go live (review -> live)
curl -sX POST http://localhost:18000/admin/tenants/$TENANT_ID/go-live \
  -H "x-csrf-token: $CSRF" --cookie "cl_admin_session=$SESSION; cl_csrf=$CSRF"
```

Port `18000` and the `cl_admin_session`/`cl_csrf` cookie pair are the local product stack's conventions — see [Local product stack](/onboarding/local-product-stack). Every mutating call except `POST /admin/tenants` (see above) needs the CSRF header to match the `cl_csrf` cookie or it's rejected with `403` before the route body ever runs.

## Operator-only routes not on the published surface

`onboarding_api.py` also defines `rename_tenant`, `materialize_now`, `run_tenant_migrations_now`, and `import_checkpoint` — each explicitly marked `TEMPORARY` in its own docstring as a workaround for a current infrastructure gap (no direct external Postgres access to a Render-hosted database from a local machine, no paid worker plan yet, no rename UI). None of these are mounted onto the deployed product surface: `services/api/product_app.py`'s `ONBOARDING_OPERATIONS` allowlist — the definitive list of the 20 onboarding operations actually served — does not include them, so `_selected_router()` silently drops any route with no allowlisted operations (`if not matched: continue`) rather than mounting it. The app only fails to build if a route mixes allowed and unapproved methods on the same path, or if an allowlisted operation never shows up in the router at all — not merely for serving an out-of-allowlist route, which is exactly why these four `TEMPORARY` routes are able to exist in the source at all. They exist for direct, ops-authorized backend access only, not as part of the pipeline any staff user reaches through the portal.

## Related

* [Staff portal](/features/staff-portal)
* [Training and checkpoints](/features/training-and-checkpoints)
* [Invitations and impersonation](/features/invitations-and-impersonation)
* [Workers and jobs](/architecture/workers-and-jobs)
