# Data sources → chunks

All paths relative to repo root. Ingest walks `CORPUS_DIR` (`html/` by default).

## A. Narrative HTML (embed as sections)

Parse with Cheerio. Drop `<style>` and `<script>`. Split on `h1`/`h2`/`h3`. Keep tables as markdown-like text.

| File | `doc_type` | `module` | Chunking rule |
|------|------------|----------|----------------|
| `html/bmi-build-chart-review.html` | `bmi_table` + `html_narrative` | `bmi` | One chunk per H2; **build table** may be split by height bands (e.g. 4'8–5'4, 5'5–6'0, 6'1–6'10) so embeddings stay specific. |
| `html/diseases/disease-*.html` | `html_narrative` | slug from filename (`cancer`, `diabetes`, …) | One chunk per H2 (Actuarial view, interview logic, decision patterns, software note). Interview `<details>` trees: one chunk per question + options. |
| `html/vertical-processes/process-mortgage-protection.html` | `process` | `mortgage_process` | One chunk per `article.step` (steps 0–11) plus journey/overview cards. |
| `html/vertical-datasets/dataset-mortgage-protection.html` | `dataset` | `mortgage_dataset` | One chunk per H2; field checklist table as one or two chunks (demographics/BMI/tobacco vs mortgage-specific fields). |

Shared metadata on every HTML chunk:

- `vertical`: `mortgage_protection` when the page is mortgage-specific; `general_uw` otherwise.
- `disclaimer`: `true` for pages that state illustrative synthesis.

### Disease slug map

| File | module |
|------|--------|
| disease-cancer.html | cancer |
| disease-diabetes.html | diabetes |
| disease-digestive.html | digestive |
| disease-disabled.html | disabled |
| disease-immune-neurological.html | immune_neurological |
| disease-joint-muscle.html | joint_muscle |
| disease-kidney.html | kidney |
| disease-heart-circulatory.html | heart_circulatory |
| disease-liver.html | liver |
| disease-lung.html | lung |
| disease-mental-nervous.html | mental_nervous |
| disease-other.html | other |

## B. Structured JSON

### `html/data/disease_trees.json`

Nested trees keyed by top-level option id (`1` Cancer … `14` Other). Flatten **leaf and near-leaf nodes** into chunks:

```
Module: Cancer
Question: When was client medically diagnosed...
Answer: 2-3 years
Decline: 25 / PASS: 10 (of 35 products)
Path: Cancer > [timing question] > 2-3 years
```

- `doc_type`: `interview_tree`
- `module`: from tree label
- Include `ans_code`, `ss_code`, `option_id` in metadata (not necessarily in embed text).
- Parent nodes with only children and no useful stats: skip or merge into child text to avoid empty chunks.
- Use `summaries` for one **overview chunk per module** (pass/decline totals).

### `html/data/health_questions.json`

Chunk groups of related questions (by `health_cat_id`) or one chunk per question if text is long.

- `doc_type`: `interview_tree`
- Fields: `id`, `question`, `que_code`, `sscode`

### `html/data/health_options.json`

Do **not** ingest every option as a standalone tiny chunk (408 rows). Prefer trees. Optionally attach option labels as metadata on tree chunks.

Ingest **top-level disease names** from options whose `health_questions_id` is `1` as a single “module index” chunk.

### `html/data/product_answer_stats.json`

One summary chunk (totals: 14267 rows, Decline 5808, empty=PASS 8459). Do not embed the entire `by_option` map as one blob. Per-option stats already live on tree leaves.

### `html/data/products.json`

One chunk listing products as `company_id + name + product_code` (batch ~20 products per chunk).

- `doc_type`: `product`

### `html/data/meta.json`

One chunk: companies list, tobacco classes, question/option counts. Companies 16–19 duplicate names of earlier ids — keep as-is and note duplication in text.

- `doc_type`: `stats`

### `html/data/health_bmi.json` and `bmi_enriched.json`

Prefer **`bmi_enriched.json`** (has `display`, approx BMI at max). Chunk by height ranges (e.g. 10 rows per chunk).

```
Height 5'8" (68 in): min≈122, table2 max 266 lb (~BMI 40.5), ...
```

- `doc_type`: `bmi_table`
- `module`: `bmi`
- Note: legacy table2/table4 are **max weight** style, distinct from the multi-class narrative BMI page. Chunk text must say **legacy dump table** vs **2026 actuarial class chart**.

### `html/data/disease_index.json`

One index chunk mapping slug → file → score. Helps retrieval of “which disease pages exist”.

## C. Dual corpus (important)

The HTML dossiers are **modern actuarial framing**. The JSON trees are **legacy product knockouts** (Decline vs empty=PASS across ~35 products).

Chunk text must label origin:

- `knowledge_layer`: `synthesis_2026` | `legacy_2022_dump`

The generator prompt must tell GPT-4o-mini: if both layers retrieve, **prefer synthesis for “how to design software / typical SI path”** and **use dump stats for “how many products decline this answer”**. Never mix them as if they were one ruleset.

## D. Files linked but missing

Do not invent content for:

- `health-uw-index.html`
- `uw-knowledge-for-carriers.html`
- `insurance-verticals.html`
- dataset/process indexes
- `shared-general-dataset.html`

If a user asks about “carrier ruleset API page”, retrieval will be empty — answer should say that page is **not in the corpus**.

## E. Target chunk volume (estimate)

| Source | Approx chunks |
|--------|----------------|
| 12 disease HTML | 60–120 |
| BMI HTML | 8–15 |
| Mortgage process + dataset | 20–30 |
| Disease trees (leaves) | 250–400 |
| BMI JSON | 4–8 |
| Products + meta | 5–15 |
| **Total v1** | **~400–600** |

Well within a single local LanceDB table and a few dollars of embedding cost.
