# RAG pipeline

## 1. Ingest (`npm run ingest`)

1. Walk `html/` (skip nothing except optional `*.css` if added later).
2. Compute file SHA-256. Compare `data/ingest-manifest.json`. Unchanged files can skip re-chunk (optional optimization).
3. Produce `KnowledgeChunk[]` per [data-sources.md](./data-sources.md).
4. Embed `chunk.text` with `text-embedding-3-small` in batches of 64.
5. Upsert into vector table `uw_knowledge_v1` with payload = full chunk JSON.
6. Write manifest: `{ ingested_at, file_hashes, chunk_count }` to `data/ingest-manifest.json`.
7. Persist vectors to **`data/lancedb/`** (this is the saved learning index). Optionally dump `data/chunks/chunks.jsonl`.
8. Also ingest `data/learning/corrections/*.md` if present.

Where every file is stored: [learning-model.md](./learning-model.md).

Idempotency: `chunk.id = sha256(source_path + "|" + section_key + "|" + text_hash_16)`.

## 2. Query

```
query text
  → embed query
  → vector search k=8 (optional filter module / doc_type / knowledge_layer)
  → drop hits below MIN_RELEVANCE
  → build prompt
  → gpt-4o-mini
  → { answer, citations[], usage }
```

### Retrieval

- Metric: cosine similarity (LanceDB default for OpenAI embeddings).
- Default `k = 8`. Allow `k` 1–20 on the API.
- Optional filters (metadata):
  - `module`: e.g. `cancer`, `bmi`, `mortgage_process`
  - `doc_type`
  - `knowledge_layer`: `synthesis_2026` | `legacy_2022_dump`

Hybrid search (keyword + vector) is **v2**. v1 is dense retrieval only. Queries like exact `ans_code` may miss; add a cheap JSON lookup in v1.1: if the query matches `/^[A-Z0-9]{4,6}$/`, also fetch tree nodes by `ans_code`.

### Prompt composition

**System** (fixed):

- You are an assistant over a **local underwriting knowledge corpus**.
- Answer **only** from the numbered CONTEXT blocks.
- If context is insufficient, say so. Do not invent carrier names, premiums, or wait periods not in context.
- Distinguish **2026 synthesis HTML** vs **legacy product dump JSON**.
- This is not medical, legal, or insurance advice; not a bind.
- Prefer structured answers: short verdict, then drivers, then caveats.
- End with a **Sources** list using the provided `source_path` values.

**User**:

```
QUESTION:
{user question}

CONTEXT:
[1] source={path} module={module} layer={knowledge_layer}
{text}

[2] ...
```

Temperature: `0.2`. Max tokens: `1200`.

### Answer contract

The API returns JSON, not only chat text:

```ts
{
  answer: string;
  citations: { id: string; source_path: string; title: string; score: number }[];
  model: "gpt-4o-mini";
  retrieved: number;
  insufficient_context: boolean;
}
```

`insufficient_context` is true when fewer than 2 chunks pass `MIN_RELEVANCE`, or the model is instructed to set a flag if it cannot answer (parse a trailing JSON line, or a second structured output call — v1 can use a simple heuristic: if all scores < threshold).

## 3. Chat vs tools

v1: **single-shot** RAG (no function calling).

v1.1 optional tools:

- `lookup_bmi(height_in, weight_lb)` — compute BMI and find nearest narrative class + legacy max.
- `lookup_tree_option(option_id)` — exact dump stats.

Keep tools thin; they still ground in local files, not the web.

## 4. Observability

Log (no PII assumed): `request_id`, `k`, top chunk ids, token usage, latency_ms.

Do not log full user questions in production if they may contain health data; v1 local logs may include them with a config flag `LOG_QUERIES=false` by default.
