# Learning model and where data is saved

**Current on-disk index:** three independent verticals under `data/verticals/{health,bmi,mortgage}/`. See [vertical-learning-data.md](./vertical-learning-data.md).

This document is the **learning model**. GPT-4o-mini is not trained in this project. “Learning” means: **turn `html/` into stored vectors, retrieve them, optionally record corrections, and re-ingest**.

```
                    ┌─────────────────────────────────────┐
                    │  OpenAI (remote, we do not train)   │
                    │  • gpt-4o-mini          = talker    │
                    │  • text-embedding-3-small = mapper  │
                    └─────────────────────────────────────┘
                                      ▲
                                      │ API calls only
                                      │
 html/  ──ingest──►  data/lancedb/  ──retrieve──►  answer
 (source            (LEARNED INDEX                  (not saved
  knowledge)         = what the app                  unless
                     “knows” how to find)             feedback)
                                      │
                                      ▼
                            data/learning/   (optional)
                            questions, ratings,
                            human corrections → next ingest
```

## 1. What the learning model is (and is not)

| Piece | Role | Trained here? | Saved on disk? |
|-------|------|----------------|----------------|
| **Source knowledge** | HTML + JSON in `html/` | No — authored content | Yes, repo: `html/` |
| **Embedding model** | Turns text into vectors | No — OpenAI hosted | No (weights stay at OpenAI) |
| **Chat model** | GPT-4o-mini writes the answer | No — OpenAI hosted | No |
| **Vector index** | This app’s memory of the corpus | Built locally at ingest | **Yes: `data/lancedb/`** |
| **Feedback / corrections** | Humans teach better answers | Applied on next ingest | **Yes: `data/learning/`** |

There is **no** custom `.pt` / `.gguf` / fine-tune checkpoint in v1. If someone asks “where is the trained model file?”, the honest answer is: **there isn’t one**. The local “brain” is the **vector table**.

### Learning loop (v1 + v1.1)

1. **Cold start:** ingest `html/` → embeddings written to `data/lancedb/`.
2. **Query time:** question → embed → nearest chunks → GPT-4o-mini. Chat weights do not change.
3. **Optional learning:** store Q&A and corrections under `data/learning/`. A later ingest **promotes** approved corrections into chunks so retrieval improves.

That is RAG learning: **index + feedback files**, not gradient descent on GPT-4o-mini.

## 2. Disk map — every learning artifact

All runtime learning data lives under **`f:\agent\data\`** (gitignored except placeholders). Source knowledge stays in **`f:\agent\html\`**.

```
f:/agent/
├── html/                            # SOURCE (not generated). Edit here to “teach” facts.
│   ├── diseases/*.html
│   ├── data/*.json
│   └── ...
│
└── data/                            # LEARNING / RUNTIME STORE
    ├── .gitkeep
    │
    ├── verticals/                   # INDEPENDENT learning stores (do not combine)
    │   ├── index.json
    │   ├── health/                  # disease + interview + products
    │   ├── bmi/                     # build charts only
    │   └── mortgage/                # process + mortgage dataset only
    │
    └── learning/                    # HUMAN / SESSION LEARNING (v1.1)
        ├── queries.jsonl            # each question + retrieved chunk ids (if LOG_QUERIES)
        ├── feedback.jsonl           # thumbs up/down on an answer
        └── corrections/             # approved new knowledge
            └── *.md                 # operator-written facts → next ingest
```

### Exact save locations

| What | Path | Written by | Purpose |
|------|------|------------|---------|
| Original UW knowledge | `html/**` | Authors | Ground truth corpus |
| Parsed chunks (debug/audit) | `data/chunks/chunks.jsonl` | `npm run ingest` | Inspect text before/without vectors |
| **Vector learning data** | `data/lancedb/` (`VECTOR_DIR`) | `npm run ingest` | Similarity search |
| Ingest bookkeeping | `data/ingest-manifest.json` | `npm run ingest` | Skip unchanged files; prove ingest ran |
| Fallback vectors | `data/embeddings.fallback.json` | ingest if LanceDB fails | Same role as LanceDB |
| Query log | `data/learning/queries.jsonl` | API if `LOG_QUERIES=true` | Offline eval; **health data — keep off git** |
| Ratings | `data/learning/feedback.jsonl` | `POST /v1/feedback` | Improve retrieval later |
| New facts | `data/learning/corrections/*.md` | Operators | Merged into chunks on next ingest |

**Git:** commit `html/` and `docs/`. Do **not** commit `data/lancedb/`, `data/learning/queries.jsonl`, or `.env`. Commit `data/.gitkeep` and `data/learning/corrections/.gitkeep` only.

## 3. What is stored inside the vector table

Each row is one learned unit:

```
id            stable hash
vector        float[1536]   ← output of text-embedding-3-small
text          chunk body    ← what GPT will see if retrieved
source_path   html/...
doc_type      html_narrative | interview_tree | bmi_table | ...
module        cancer | bmi | mortgage_process | ...
knowledge_layer  synthesis_2026 | legacy_2022_dump
title, metadata
```

This table **is** the learning data for retrieval. Rebuild it anytime with `npm run ingest` after `html/` or `data/learning/corrections/` change.

## 4. How new knowledge gets in (operators)

| Action | Effect |
|--------|--------|
| Edit a file under `html/` then ingest | Corpus learning (preferred for disease/BMI/process facts) |
| Add `data/learning/corrections/2026-08-20-dialysis.md` then ingest | Local learning without editing the HTML site |
| Call OpenAI chat | **No** weights saved locally; only the answer in the HTTP response unless you log it |

Correction file format (ingest will treat as `doc_type: correction`, `knowledge_layer: operator`):

```markdown
---
module: kidney
title: Dialysis current — SI note
---
Operator note: current dialysis is treated as Level knockout in this agency playbook.
```

## 5. Feedback API (learning data write path)

`POST /v1/feedback` appends one JSON line to `data/learning/feedback.jsonl`:

```json
{
  "ts": "2026-08-20T08:30:00Z",
  "question": "...",
  "answer_id": "...",
  "rating": "up" | "down",
  "comment": "Missed Graded path",
  "chunk_ids": ["..."]
}
```

v1 may only **write** this file. Using ratings to re-rank chunks is v1.1 (`src/learning/promoteFeedback.ts`).

## 6. Config paths

```
CORPUS_DIR=./html
VECTOR_DIR=./data/lancedb
CHUNKS_PATH=./data/chunks/chunks.jsonl
LEARNING_DIR=./data/learning
LOG_QUERIES=false
```

`src/config.ts` must expose these so it is obvious **where learning data is saved**.

## 7. Code modules (when implemented)

| File | Learning job |
|------|----------------|
| `src/ingest/embedAndStore.ts` | Writes vectors to `VECTOR_DIR` |
| `src/store/vectorStore.ts` | Read/write the learned index |
| `src/learning/logQuery.ts` | Append `queries.jsonl` |
| `src/learning/feedback.ts` | Append `feedback.jsonl` |
| `src/learning/loadCorrections.ts` | Read `corrections/*.md` during ingest |

## 8. One-sentence summary

**Models talk via OpenAI; this repo learns by saving embeddings in `data/lancedb/` and optional human files in `data/learning/`, sourced from `html/`.**
