|
| 1 | +# ENM 3800 "Learning from Data" — Curriculum Roadmap |
| 2 | + |
| 3 | +Living design doc for course structure and planned changes. Committed so it |
| 4 | +travels with the repo rather than living in chat history. Edit as decisions land. |
| 5 | + |
| 6 | +**Status legend** |
| 7 | + |
| 8 | +- ✅ **Decided** — agreed; safe to implement. |
| 9 | +- 🔧 **Proposed** — drafted, not yet agreed or implemented. |
| 10 | +- ❓ **Open** — needs a decision before work starts. |
| 11 | + |
| 12 | +> **Note:** The handoff references `docs/ENM3800_Revised_Schedule.md` (fixed, |
| 13 | +> date-bound slots; rule for changes = reallocate/re-theme, don't add). That file |
| 14 | +> is **not yet in the repo.** Lecture slots below are referenced by number (L2–L15). |
| 15 | +> Add the schedule doc so date-bound slots have a source of truth. |
| 16 | +
|
| 17 | +--- |
| 18 | + |
| 19 | +## Current module structure |
| 20 | + |
| 21 | +| Module | Title | Notebooks | Lectures | |
| 22 | +|--------|-------|-----------|----------| |
| 23 | +| M1 | Learning from Data | 1a, 1b | L2–3 | |
| 24 | +| M2 | Building & Selecting Models | 2a, 2b | L5–6 | |
| 25 | +| M3 | Probability & Inference | 3a, 3b, 3c | L8–10 | |
| 26 | +| M4 | Supervised | 4a, 4b, 4c | L7, 11–13 | |
| 27 | +| M5 | Unsupervised | 5a, 5b | L14–15 | |
| 28 | + |
| 29 | +--- |
| 30 | + |
| 31 | +## M1 split — Learning from Data ✅ IMPLEMENTED |
| 32 | + |
| 33 | +The Module 1 split was already in the repo when audited: |
| 34 | + |
| 35 | +| File | Title | Lecture | |
| 36 | +|---|---|---| |
| 37 | +| `Notebook_1a_Overview_and_Data_Representation` | Overview & Data Representation | L2 | |
| 38 | +| `Notebook_1b_Noise_and_Modeling_Data` | Noise, Variability & Modeling Data | L3 | |
| 39 | + |
| 40 | +Both are registered as separate M1 chapters in `_quarto-book.yml`. The split is |
| 41 | +clean by theme — **1a = representation** (Python, linear algebra, matrices as |
| 42 | +transformations, images, tables); **1b = variability** (noise, covariance, |
| 43 | +distributions) — and 1a hands off explicitly ("covariance… comes in the next |
| 44 | +notebook"). |
| 45 | + |
| 46 | +**Self-containment audit** surfaced two issues, both since fixed: |
| 47 | + |
| 48 | +1. **Dangling prerequisite.** 1b claimed its covariance work built on "eigenvector |
| 49 | + ideas from the previous notebook," but 1a never introduces eigenvectors. Fixed by |
| 50 | + removing the claim. |
| 51 | +2. **Cross-module duplication with the M3/M5 rework:** |
| 52 | + - *Covariance eigenvectors* — 1b (L3) duplicated the geometric-covariance content |
| 53 | + staged for M5 PCA. **Decision: cut** 1b's "Eigenvectors of the Covariance |
| 54 | + Matrix" section (too heavy for L3); keep a light covariance-as-shape teaser with |
| 55 | + a forward pointer to Module 5. `_staging_covariance_geometry.py` is now the |
| 56 | + single source of the eigenvector treatment. |
| 57 | + - *Probability distributions* — 1b (L3, intuitive) overlaps 3a's zoo (L8, |
| 58 | + rigorous). **Decision: keep as a deliberate preview** — 1b now points forward to |
| 59 | + Module 3, and 3a's zoo points back to 1b. Coordinated spiral, no cuts. |
| 60 | + |
| 61 | +--- |
| 62 | + |
| 63 | +## M3 rework — Probability & Inference ✅ IMPLEMENTED |
| 64 | + |
| 65 | +**Goal:** re-theme the three M3 notebooks along a cleaner narrative spine — |
| 66 | +**L8 Probability → L9 Statistics 1 → L10 Statistics 2** — culminating L8 in the CLT |
| 67 | +as the bridge into inference. Fits existing slots (no additions). |
| 68 | + |
| 69 | +> **Implemented** (this pass). Titles, content, **and filenames** now match the new |
| 70 | +> themes. Files were renamed (`git mv` on both the `.py` and `.ipynb` of each pair): |
| 71 | +> |
| 72 | +> | New file | Was | |
| 73 | +> |---|---| |
| 74 | +> | `Notebook_3a_Probability` | `Notebook_3a_Probability_and_Covariance` | |
| 75 | +> | `Notebook_3b_Statistics_1` | `Notebook_3b_Hypothesis_Testing` | |
| 76 | +> | `Notebook_3c_Statistics_2` | `Notebook_3c_Bootstrapping_and_Permutation_Tests` | |
| 77 | +> |
| 78 | +> `_quarto-book.yml` chapter paths and every Colab badge URL (title + `#scrollTo=` |
| 79 | +> exercise deep-links) were updated to match; exercise badges are regenerated by |
| 80 | +> `scripts/add_exercise_badges.py`, which derives the URL from the file path. |
| 81 | +> |
| 82 | +> A new generating-process exercise (`id="ex-generating-processes"`, "Name That |
| 83 | +> Process") was added to 3a's zoo. |
| 84 | +> |
| 85 | +> Geometric-covariance cells were extracted to |
| 86 | +> `notebooks/Notebook_5/_staging_covariance_geometry.py` (leading underscore → not |
| 87 | +> rendered, not in the book config). Fold them into **5a** when PCA is worked, and |
| 88 | +> update that file's badge URLs to 5a's path. |
| 89 | +
|
| 90 | +### Key findings from the current notebooks |
| 91 | + |
| 92 | +Two facts reshaped this from "write a lot of new content" into "mostly migration": |
| 93 | + |
| 94 | +1. **No Bayes content exists** in 3a or 3b. "Bayes deprioritized" is a no-op — |
| 95 | + nothing to remove. |
| 96 | +2. **The CLT/LLN capstone already exists — in 3b.** The "Sampling Distributions: |
| 97 | + Where t-tests Come From" section (`Notebook_3b:467`) *is* a CLT/LLN demo: |
| 98 | + repeated samples → distribution of the mean → it narrows → `SE = σ/√n`. |
| 99 | + Already written, tested, with embedded outputs. Migrate it into 3a rather than |
| 100 | + author from scratch. |
| 101 | + |
| 102 | +### Target: 3a → **Probability** (L8) ✅ |
| 103 | + |
| 104 | +Narrative spine: build distributions → expectation/variance → LLN (averages |
| 105 | +converge) → CLT (averages go Gaussian, `SE = σ/√n`). |
| 106 | + |
| 107 | +- **Keep:** What is probability; Random variables; Expectation & variance; |
| 108 | + Coin-flip simulation (already Bernoulli/Binomial *by generating process* — the seed). |
| 109 | +- **New (the only net-new writing):** the **distributions zoo** — Uniform, |
| 110 | + Poisson, Exponential, Gaussian, each taught *by its generating process*, with |
| 111 | + E/Var formalized per distribution. |
| 112 | +- **Migrate in from 3b:** "Sampling Distributions" + "SE shrinks with √n" → |
| 113 | + reframed as the **LLN → CLT capstone**. |
| 114 | +- **Move out:** all covariance/correlation (see split below). |
| 115 | + |
| 116 | +### Target: 3b → **Statistics 1** (L9) ✅ |
| 117 | + |
| 118 | +- **Loses** the sampling-distribution material to 3a (~95 lines freed). |
| 119 | +- **Gains** the *statistical* covariance/correlation from 3a. |
| 120 | +- **Keeps:** framing a question; test statistics & p-values; confidence intervals; |
| 121 | + t-test deep dive; two-sample t-test types; ANOVA; paired t-test. |
| 122 | +- **Opens** by cashing in 3a's CLT: "because the sample mean is ~Gaussian, the |
| 123 | + t-statistic has a known distribution" — clean handoff. |
| 124 | +- **Type I/II introduced here** (α = Type I error; power = 1−β). The existing α |
| 125 | + and "fail-to-reject ≠ accept" callouts already imply it — this is a labeling |
| 126 | + pass, not new content. |
| 127 | + |
| 128 | +### Target: 3c → **Statistics 2** (L10) ✅ |
| 129 | + |
| 130 | +- **Keeps:** bootstrap; permutation tests. |
| 131 | +- **Type I/II quantified** — the "Noise, Sample Size, and Power" demo needed for |
| 132 | + this already exists at `Notebook_3c:347`. Reframing, not writing. |
| 133 | + |
| 134 | +### Covariance split ✅ (resolves Open Decision: "covariance's home") |
| 135 | + |
| 136 | +The 3a covariance content is mostly *geometric* (shape-of-a-data-cloud, tilted |
| 137 | +clouds, interactive shape explorer) — that's PCA prep, not t-test prep. Split it: |
| 138 | + |
| 139 | +- **Statistical** part (correlation coefficient, correlation matrix, correlated- |
| 140 | + variables sim, correlation in a real dataset) → **L9 Statistics (3b).** |
| 141 | +- **Geometric** part (covariance as shape of a data cloud, tilted cloud, three- |
| 142 | + shapes comparison, interactive shape explorer) → **earmark for M5 PCA prep.** |
| 143 | + Re-invoke the covariance matrix there: "its eigenvectors are the principal |
| 144 | + components." Same spiral pattern as Type I/II. |
| 145 | + |
| 146 | +### Work summary |
| 147 | + |
| 148 | +1. **Write** the distributions zoo (3a) — the only net-new authoring. |
| 149 | +2. **Migrate** sampling-distribution/CLT section 3b → 3a; statistical covariance 3a → 3b. |
| 150 | +3. **Relabel** for the Type I/II spiral (3b introduce, 3c quantify). |
| 151 | +4. **Earmark** geometric covariance for M5 (leave a pointer; move when PCA is touched). |
| 152 | + |
| 153 | +Per repo convention: edit the `.py`, then `jupytext --sync`; re-run |
| 154 | +`scripts/execute_notebooks.sh` for fresh embedded outputs. |
| 155 | + |
| 156 | +--- |
| 157 | + |
| 158 | +## Other proposals (from handoff — not yet worked) |
| 159 | + |
| 160 | +*(M1 split — now ✅ implemented; see the M1 section below.)* |
| 161 | + |
| 162 | +🔧 **M4 deep learning** — add backprop/autograd intuition + transfer learning. |
| 163 | +Recommendation from discussion: **deepen 4c in place (L13)** rather than spin up a |
| 164 | +2-notebook DL module, which would cost an M5 slot and force autoencoders/self-sup |
| 165 | +out of 5a (weakening M5's unsupervised identity). |
| 166 | + |
| 167 | +🔧 **L7 Regularization → "Bias–Variance & Regularization"** — bias–variance is |
| 168 | +missing and lecture-worthy; reframing also fixes timing so L11 regression can apply |
| 169 | +ridge/lasso. |
| 170 | + |
| 171 | +🔧 **L4 "What is ML" → "the map of ML"** — sharpen into taxonomy / workflow / |
| 172 | +generalization. Recommendation: **keep as a structural keystone**, don't cut to a |
| 173 | +buffer or to fund DL (wrong time of year). |
| 174 | + |
| 175 | +🔧 **Type I/II spiral** (cross-module) — introduce in 3b (α, power = 1−β), quantify |
| 176 | +in 3c (power vs sample size/noise), reconnect in **4b** (confusion matrix FP/FN = |
| 177 | +Type I/II; "Error Costs" exercise = cost asymmetry). |
| 178 | + |
| 179 | +--- |
| 180 | + |
| 181 | +## Open decisions |
| 182 | + |
| 183 | +- ✅ ~~Covariance's home~~ → **split** (statistical → L9; geometric → M5 PCA prep). |
| 184 | +- ❓ **DL:** deepen 4c in place vs. 2-lecture module (costs an M5 slot). |
| 185 | + *Leaning deepen-in-place.* |
| 186 | +- ❓ **Timing of the M1 split.** *Leaning: do first — lowest risk.* |
| 187 | +- ❓ **L4:** stays "map of ML" vs. becomes a buffer. *Leaning: keep as keystone.* |
0 commit comments