services/api/product_app.py is the deployed tenant/staff product — 44 allowlisted operations, described in API reference overview. services/api/main.py is everything that came before it: a much larger, process-global research application that predates tenancy, row-level security, and the allowlist model entirely. This section of the docs (/research/*) is the canonical reference for that second application — what it is, why it still lives in the repository, and the one rule that keeps it from becoming a liability: product code must not import it.
What the research workbench is
services/api/main.py and the causegraph package underneath it (packages/causegraph/src/causegraph/) are a single experimentation platform built to answer three questions before the tenant product existed: what does a causal-issue graph model look like, how do you cluster and interpret large volumes of unstructured issue text, and what can process-mining techniques (event logs, directly-follows graphs, conformance checking) tell you about the workflows that produce those issues. Three components make up the surface:
- A causal-graph schema and process-global store.
causegraph.graph_schema.full_graph_schema()defines the full provenance graph model — as of this writing, 31 node types and 34 edge types across nine populated layers (source_evidence,observation_normalization,process_workflow,issue_failure,remediation_outcome,knowledge_taxonomy,shared_pattern_anonymization,provenance_governance,data_science_rag_ml); a tenth layer,technical_lineage, is defined in theGraphLayerenum but not yet used by any node or edge type.GET /graph/schemaserves it directly. Every research route operates againstcausegraph.storage.memory.CauseLoopState, an in-process, dataset-id-keyed store — not the tenant-scoped, row-level-secured PostgreSQL schemas the product uses. See Datasets for what gets loaded into that store. - Clustering and theme interpretation.
/clusters/runand/clusters/algorithmsexpose a materially larger algorithm set than the product’s fixedhdbscandefault —kmeans,leiden,bm25,tfidf, and hybrid combinations — plus/themes/*for theme-level exploration, LLM-generated summaries, and fishbone diagrams once a run has produced clusters. The strict, governed variant of this same idea — an audited 250-issue severity, clustering, and root-cause pipeline — is documented separately in MVP1 pipeline. - Process discovery and conformance.
causegraph.tasksdefines a 16-task catalog (TASK_SPECS) spanning process mining, issue intelligence, and safety/cyber/reliability classification. Five of those tasks — process discovery, conformance/deviation, predictive process monitoring, process anomaly detection, and process-stub stitching — are the event-log/process-mining family and each has an implemented baseline algorithm (directly-follows graphs, PM4Py Inductive Miner, dominant-path deviation, a Markov next-activity baseline, overlap-graph stitching) plus named future providers (Split Miner, alignment-based checking, ProcessTransformer, graph autoencoders) that are cataloged but not wired. The remaining 11 tasks target other datasets (OSHA, MAUDE, NASA C-MAPSS, VCDB, and more); six of them have no implemented baseline yet, onlyadapter_readyprovider slots.GET /taskslists all 16 catalog entries;POST /tasks/{task_id}/runexecutes a task’s implemented baseline against imported data (BPIC13, Helpdesk, and other adapters, depending on the task). See API surface for the full grouped tour, including the file-based/clientsonboarding API that predates the real tenant pipeline.
Why it is preserved
This is not dead code kept out of inertia. Three concrete reasons keep it in the repository:- It is the historical basis for real product systems. The product’s remediation feature (
services/api/remediation_api.py) is a direct, tenant-scoped successor to the process-global/loops/*and legacy/mvp1/caps//mvp1/reviewssystems still hosted here; the product’s onboarding pipeline (services/api/onboarding_api.py) is the successor to the file-based/clientsAPI inservices/api/clients_api.py. Keeping the predecessor in place gives engineers working on the current feature a working reference implementation to compare against. - It hosts research directions the product hasn’t absorbed yet. Process discovery, conformance checking, predictive process monitoring, and the wider public-dataset acquisition program (CISA KEV, NVD, OSHA, MAUDE, NHTSA, and more — see Datasets) are active research areas with no tenant-product equivalent. Removing the workbench would delete the only place that work currently runs.
- It hosts a stable, audited analysis contract that the product intentionally has not absorbed. The MVP1 relational pipeline (strict 250-issue severity/clustering/fishbone analysis, reviewed nine-sheet Excel workbook) is deliberately kept as a separate, explicitly versioned sub-contract (
services/api/mvp1_contract.py) rather than folded into the product API, specifically so its transport shape and its analysis logic can each evolve without the other silently drifting. See MVP1 pipeline.
The hard boundary
The dependency direction is one-way and enforced structurally, not by convention alone:- The repository’s
Dockerfilestartsuvicorn services.api.product_app:app. Nothing in the deployed container ever importsservices.api.main. product_app.py’s own module docstring states the reasoning directly: importingmain’s app “creates stores and exposes operations the customer/staff product does not need.”product_app.pyimports exactly seven product routers —auth_api,onboarding_api,insights_api,members_api,metrics_api,remediation_api,sources_api— and neverservices.api.main,services.api.clients_api,services.api.mvp1_contract, orcausegraph.storage.memory.CauseLoopState.apps/webis a separate Vite project from the Next.js tenant console infrontend/. It talks toservices/api/main.pyon its own local port (5173by default) and is never built into a deployed artifact.
If you are working on a product feature and find yourself reaching for something that only exists in
main.py or clients_api.py, treat that as a signal, not a shortcut: either a product-side equivalent already exists (check onboarding_api.py, remediation_api.py, sources_api.py first) or it genuinely needs to be built on the product side — never imported from the research app. See Research API for the full statement of this rule.Where to go next
- To run the workbench locally and choose the right mode for the work at hand, see Research modes.
- For a domain-by-domain tour of the ~279-path surface with representative endpoints, see API surface.
- For the strict, OpenAI-governed 250-issue pipeline and its Postgres-backed reuse contract, see MVP1 pipeline.
- For what data is actually loaded versus what must be downloaded or generated, see Datasets.