# Architecture

## Purpose

A local-first Node.js service that:

1. **Ingests** `html/**/*.html` and `html/data/*.json` into searchable chunks.
2. **Embeds** chunks (OpenAI embeddings) into a vector store.
3. **Retrieves** the top-k chunks for a user query (vector + light metadata filters).
4. **Generates** an answer with **GPT-4o-mini**, using only retrieved context plus a fixed system prompt.

```
┌─────────────┐     ingest      ┌──────────────┐     embed      ┌─────────────────┐
│  html/      │ ───────────────►│  Chunker     │ ─────────────► │ Vector store    │
│  HTML+JSON  │                 │  + metadata  │                │ (Chroma / Qdrant│
└─────────────┘                 └──────────────┘                │  or LanceDB)    │
                                                                └────────┬────────┘
                                                                         │ retrieve
┌─────────────┐     query       ┌──────────────┐     context     ┌───────▼────────┐
│  Client     │ ───────────────►│  Express API │ ───────────────►│ GPT-4o-mini    │
│  (curl/UI)  │ ◄───────────────│  /v1/query   │ ◄───────────────│ + citations    │
└─────────────┘     answer      └──────────────┘                 └────────────────┘
```

## Components

| Component | Responsibility |
|-----------|----------------|
| **Corpus loader** | Read files from `html/` with stable `source_id` (relative path). |
| **HTML parser** | Strip CSS/JS; keep headings, tables, interview trees, callouts. |
| **JSON normalizer** | Flatten `disease_trees.json` nodes, questions, BMI rows, products into text. |
| **Chunker** | Split by semantic unit (section, tree node, BMI row group), not raw character dump. |
| **Embedder** | `text-embedding-3-small` (default) via OpenAI. |
| **Vector store** | Persistent local store; collection `uw_knowledge_v1`. |
| **Retriever** | Query embedding + optional `module` / `doc_type` filter; return k=8 default. |
| **Composer** | Build system + user prompt with numbered context blocks. |
| **Generator** | Chat Completions: `gpt-4o-mini`. |
| **API** | Express, JSON, CORS optional, health check. |

## Trust and safety

- The model **must not** invent carrier-specific rates or claim a bind.
- If retrieval score is below threshold, respond with “insufficient corpus coverage” rather than guessing.
- Always attach **source paths** and chunk ids used.
- Do not send PII in queries if this later sits behind a production CRM (v1 assumes internal use).

## Model choices (v1)

| Role | Model | Why |
|------|--------|-----|
| Chat | `gpt-4o-mini` | Cost, latency, instruction following for grounded Q&A |
| Embeddings | `text-embedding-3-small` | Strong retrieval, cheap at corpus size (~dozens of HTML + JSON) |

Corpus size is small. In-memory LanceDB or Chroma is enough; no need for a managed vector DB in v1.

## Learning vs hosted models

GPT-4o-mini and the embedding model are **not trained in this repo**. Local learning is the vector index and optional feedback files. Full paths and the feedback loop: **[learning-model.md](./learning-model.md)**.

## Runtime

- **Node.js** 20 LTS or 22 LTS.
- **ESM** (`"type": "module"`).
- TypeScript recommended (`strict`).
- Single process: ingest CLI + HTTP server. Ingest can run as `npm run ingest`; server as `npm start`.
