Technical map of this repository: what each stage does, how data flows
between them, and what's planned next. AGENTS.md is the conventions/how-to-run
guide; this file is the "what is this and why is it shaped this way" guide.
This repository is a staged model-training workspace, split by pipeline stage rather than by model:
pre-training/ PDF corpus -> OCR text, summaries, layout, synthetic QA (CSVs)
fine-tuning/ text/summary pairs -> trained LoRA adapters
serving/ trained adapters -> inference (FastAPI)
training/ raw datasets -> from-scratch / non-LoRA trained models
Each leaf folder (pre-training/, fine-tuning/<pipeline>/,
serving/<pipeline>/) is an independent uv project: its own
pyproject.toml, uv.lock, .python-version, and pinned dependency set (in
particular, its own CUDA torch build). Nothing is shared at runtime between
folders — a pipeline can be deleted or reworked without touching its
siblings. One uv binary at the root drives all of them via
uv run --directory <folder> ... (see the root README.md's uv Commands
section for the current, verified list); there is deliberately no root-level
Python project or shared virtualenv, since the folders pin conflicting
dependency versions (e.g. different torch builds) that a single shared
resolution would fight.
Repo-wide hygiene is intentionally centralized rather than per-folder: one
.gitignore at the root (unanchored patterns match every project's drop-zone
folders at any depth), and one AGENTS.md/CLAUDE.md pair at the root
covering conventions for the whole repo. No subfolder should have its own
copy of any of these.
Every project pins .python-version to 3.12. Without it, uv run picks
the newest CPython it can find (e.g. 3.14), and some pinned dependencies —
pillow==10.4.0 in particular — have no prebuilt wheel for that new a
version yet, so uv falls back to a from-source build that fails on Windows
(missing zlib headers). Pinning 3.12 (already installed and known to have
wheels for every pinned dependency across all four projects) is what makes
uv run --directory <folder> ... work reproducibly from a clean checkout.
This was verified by reproducing the failure and fixing it during this pass —
see the Verified working list below.
Turns a PDF corpus into training data. Local, GPU-first, Surya OCR +
Gemma 3 (unsloth/gemma-3-4b-it, an ungated mirror — no HF_TOKEN needed).
Five steps (exec_1.bat … exec_5.bat, or main.bat for an interactive
menu): PDF → PNG pages → OCR CSV → summary CSV / layout CSV / synthetic-QA
CSV, all written per-run to outputs/[timestamp]_[dataset]/.
Step 1 (PDF → PNG) is not a standalone Python entry point — it's
scripts/convert_pdf_to_png.ps1, which shells out to poppler
(pdftoppm/pdfinfo on PATH) and a small compress_png_max.py helper.
Run it via exec_1.bat or the .ps1 directly, not uv run ... python scripts/convert_pdf_to_png.py — that file doesn't exist (an earlier version
of this doc incorrectly assumed it did; fixed here). Steps 2–5
(ocr_detection_png.py, summarize_ocr_gemma.py, describe_layout_gemma.py,
generate_qa_gemma.py) are genuine argparse Python scripts and do run via
uv run --directory pre-training python scripts/<name>.py.
Status: out of scope for active work — left as-is, beyond the
.gitignore/.python-version consolidation described above.
Two example pipelines, same transformers+peft pattern, both LoRA-based,
both sized for a single RTX 3090 (24GB). (A third, Axolotl-based
axolotl-ocr-summary/ pipeline existed earlier but was removed by the repo
owner — it only resolved its uv environment on Linux/WSL, never natively
on Windows, since axolotl[deepspeed] depends on triton.)
| Pipeline | Framework | Base model | Data shape |
|---|---|---|---|
fine-tuning/vicuna-7b-lora/ |
transformers + peft (manual Trainer loop) |
lmsys/vicuna-7b-v1.5 — LoRA on q_proj/v_proj, loaded directly via AutoModelForCausalLM/AutoTokenizer. No LLaVA checkpoint, no vision encoder, no multimodal projector anywhere in the dependency graph. |
JSONL with text / summary fields |
fine-tuning/qwen25-3b-lora/ |
transformers + peft (same pattern as vicuna-7b-lora/) |
Qwen/Qwen2.5-3B-Instruct — LoRA on q_proj/v_proj (same target modules as Vicuna; Qwen2ForCausalLM uses the same separate Q/K/V/O naming, confirmed via peft's own default LoRA target-module table). ChatML prompt format instead of Vicuna's USER:/ASSISTANT: (verified against the tokenizer's chat_template/eos_token). |
JSONL with text / summary fields |
Naming/loading history: vicuna-7b-lora, previously llava15-lm-lora,
originally llava15-lora. Two renames, each fixing a real overstatement:
llava15-lora→llava15-lm-lora: the pipeline only ever LoRA'd the language-model backbone (q_proj/v_proj) ofllava-hf/llava-1.5-7b-hf— never the vision encoder or multimodal projector, never an image. The-loraname alone overstated that as a VLM fine-tune.llava15-lm-lora→vicuna-7b-lora: even loading LLaVA's checkpoint at all was unnecessary once the vision half was never used — it still downloaded the full ~14 GB multimodal weights viaLlavaForConditionalGeneration/AutoProcessorto get to a submodule that is, in substance, Vicuna-7B. This pass switched to loadinglmsys/vicuna-7b-v1.5directly viaAutoModelForCausalLM+AutoTokenizer— ~13 GB instead of ~14 GB, no vision-related code path in the dependency graph at all, same LoRA config/target modules. Trade-off:lmsys/vicuna-7b-v1.5is the checkpoint LLaVA 1.5 was later visually-instruction-tuned from, not the LLaVA-tuned weights themselves — different starting point, not directly comparable to the oldllava15-lm-lorarun's results, but a cleaner, smaller, honestly-named base for a language-model-only LoRA.serving/llava15-lorawas renamed toserving/vicuna-7b-lorain step, with matching model-loading changes (functionally required: a Vicuna-7B-trained adapter's parameter names don't matchLlavaForConditionalGeneration'slanguage_model.*prefix, so serving would fail to load it otherwise) — that serving folder has since been removed, see Stage 3. A plannedllava15-full-lorasibling, trained on image+text pairs and actually exercising the vision encoder/projector, remains the natural first real VLM fine-tune in this repo — that one should load the full LLaVA checkpoint. Full reasoning infine-tuning/vicuna-7b-lora/README.md.
vicuna-7b-lora/ is a generic text-summarization LoRA, not OCR-specific —
its interface, data source, and default prompt were all cleaned up this pass
to reflect that:
- JSONL field is
text(wasocr_text);build_vicuna7b_dataset.pyandgenerate_vicuna7b_lora.py's flags are--source-csv/--text/--text-file(were--ocr-csv/--ocr-text/--ocr-text-file). build_vicuna7b_dataset.pyonly builds from a CNN/DailyMail Parquet dump now (--cnn-dailymail-dir, required) — the earlier dual-source mode that also read pre-training's image-linked OCR/SUMMARIES CSV pair (normalize_image_key/resolve_image_path/load_summaries) was removed entirely, not just renamed, since it's not needed for this pipeline's current use (generate_vicuna7b_lora.py's--source-csvbatch-eval mode still accepts any generic CSV with atextcolumn, unrelated to that removed ingestion path).train_vicuna7b_lora.py/generate_vicuna7b_lora.py'sDEFAULT_INSTRUCTIONis now the CNN/DailyMail news-article wording (was "Summarize this scanned document page... UAP-related content") — since that's the only source this pipeline builds from,--instructionno longer needs to be passed explicitly for the common case.
qwen25-3b-lora/ is a near-clone of vicuna-7b-lora/ — same dataset
builder logic, same trainer/generator structure, same CLI shape. Two
verified differences (not assumed): the ChatML prompt wrapper (see table
above), and no protobuf/sentencepiece dependency needed (Qwen2.5-3B-Instruct
ships a ready tokenizer.json, unlike Vicuna's raw SentencePiece tokenizer).
No serving/qwen25-3b-lora/ — serving/ holds no pipelines at all now (see
Stage 3), so a serving folder would have to be written from scratch if this
adapter goes to production. It could not have been shared with the removed
serving/vicuna-7b-lora/ in any case: that one was Vicuna-specific, and the
ChatML wrapper differs.
Downloaded locally to C:\Users\luisarandas\Desktop\cnn_dailymail\3.0.0\
(outside the repo, outside the root folder, gitignored regardless). Measured
against the actual files:
| Split | Rows | Size |
|---|---|---|
| train (3 shards) | 287,113 | ~772 MB |
| validation | 13,368 | ~35 MB |
| test | 11,490 | ~30 MB |
| Total | 311,971 | ~799 MB |
article (avg ~3,950 chars) → text, highlights (avg ~260 chars) →
summary. build_vicuna7b_dataset.py --cnn-dailymail-dir ... --max-samples 2000 was run end-to-end against the real files and produces valid JSONL
records; full details and commands in
fine-tuning/vicuna-7b-lora/README.md.
Before the switch to loading Vicuna-7B directly, a 2,000-sample run on the
old llava15-lm-lora pipeline (1,800 train / 200 val, 1 epoch, 450 steps,
~31 min on a single RTX 3090) showed loss dropping 1.66 → ~1.12 in the first
~50 steps then plateauing in a ~1.0–1.2 band, with eval_loss (1.11)
tracking train loss closely (no overfitting) — a normal curve for a
rank-16, 2-projection adapter on a small dataset, not evidence of a broken
run. That adapter and its data/hf_cache were deleted as part of this pass's
switch to lmsys/vicuna-7b-v1.5 (different base weights, not compatible
with the old adapter) — the numbers above are illustrative of the expected
curve shape, not a claim about the current pipeline's untrained state.
What carries forward: loss alone doesn't say whether summaries are actually
good, so generate_vicuna7b_lora.py has a --jsonl-eval <path> --num-samples N reconstruction-test mode (added the same pass as the old
run above) — it replicates the trainer's train/val split (same seed/ratio)
and prints genuinely held-out source/reference/generated triples with
token-F1, instead of requiring a manual --text string. Run against the old
adapter it produced coherent, on-topic, correctly-bulleted CNN/DailyMail
summaries (avg token-F1 0.357 across 5 samples) despite the plateaued loss —
this is the tool to use to judge the next real run on the current pipeline,
not the loss curve. See the
pipeline README.
One serving/<pipeline>/ folder per fine-tuning pipeline that has a serving
story. Currently empty. serving/vicuna-7b-lora/ — a FastAPI service
(app.py) that loaded the base Vicuna-7B model plus the trained adapter (or
a fused/merged model) once and served a JSON API plus a dataset-browser
front-end — was removed by the repo owner, the same way
fine-tuning/axolotl-ocr-summary/ was.
The design it demonstrated is still the intended shape for this stage, and
worth restating for whatever returns here: a serving folder is deliberately
decoupled from its fine-tuning counterpart, reading only that pipeline's
trained output directory (for Vicuna that was
../../fine-tuning/vicuna-7b-lora/runs/vicuna7b_lora/final_adapter) and
never importing its training code. That boundary is what makes a serving
folder independently deployable.
From-scratch / non-LoRA training of other models, as distinct from
adapting an existing checkpoint (fine-tuning/). Seven pipelines so far,
each an independent uv project and each writing out by hand whatever the
usual library would hide.
training/adult-income-logreg/ is
the first pipeline here: logistic regression on the UCI
Adult / Census Income
dataset, implemented with raw numpy rather than scikit-learn — the sigmoid,
binary cross-entropy loss, gradient derivation, and gradient-descent update
loop are all written out by hand in train_logreg.py so the math stays
visible, and build_income_dataset.py parses the raw CSV files and does
one-hot/z-score encoding without pandas. Own uv project like every other
pipeline folder, but no torch/CUDA dependency at all — just numpy.
Verified end-to-end against the real dataset: 30,162 train / 15,060 test
rows after dropping "?" rows (matches the cleaned-variant counts in
adult.names exactly), 300 epochs of batch gradient descent, 84.6% test
accuracy (in line with the 84–86% published for tree-based methods on this
same cleaned split — a from-scratch linear model landing close to that is
the expected sanity-check result, not a target to beat).
training/cifar10-vqvae/ is the next
pipeline here: a VQ-VAE
(van den Oord et al., 2017) trained
from scratch on CIFAR-10 (raw python-format pickles at
E:\datasets\cifar-10-python, parsed by hand with stdlib pickle — no
torchvision/keras, no Pillow). It is the in-place successor of the
former cifar10-vae (renamed/converted): a plain VAE's blur comes from
Gaussian-posterior averaging in the ELBO, and VQ-VAE removes that
mechanism — an 8x8 grid of D-dim encoder vectors is replaced by its
nearest neighbors in a learned 512x64 codebook (straight-through
estimator, EMA codebook updates, commitment loss), and reconstruction is
driven purely by MSE. Same hand-written philosophy as the whole training/
folder: encoder/quantizer/decoder all plain torch.nn, torch only for
tensor ops/autograd/GPU, numpy-permutation batching, no VQ-VAE library,
no DataLoader. ~0.74M params (codebook included), ~40–60 min for 100
epochs on a single RTX 3090. Reconstruction-only by design (encode ->
quantize -> decode a real image back; no learned prior over the discrete
codes, so no sampling) — the discrete code grid is the substrate for a
later learned prior (the planned cascade). Its evaluator computes the same
metric suite used to judge the predecessor VAE (MAE/PSNR/SSIM/
high-frequency retention) plus a codebook-usage check, so the VQ-VAE vs
VAE comparison is reproducible; measured numbers in the pipeline README's
"Verified runs".
training/imdb-sentiment-cnn/
is the text-classification pipeline here: a Text CNN
(Kim, 2014) trained from scratch on
the Large Movie Review Dataset
(25k train / 25k test binary sentiment; raw review .txt files at
E:\datasets\aclImdb_v1, parsed by hand — no torchtext/datasets/nltk).
Same hand-written philosophy as the whole training/ folder: a
randomly-initialized trainable embedding (no GloVe — strictly
IMDB-only data by design), three parallel 1D convs (widths 3/4/5 × 128
filters) + ReLU + 1-max-pool per filter, concat, dropout 0.5, linear → 2;
torch only for tensor ops/autograd/GPU, numpy-permutation batching, no
DataLoader. ~6.6M params (6.4M in the embedding), 20 epochs in ~31 s
on a single RTX 3090, best checkpoint by val acc (peaks at epoch 3 before
the model overfits — train acc → 100%). Measured on the held-out 25k test
split: 89.2% accuracy (neg 88.96% / pos 89.43%) — well above Kim's
published CNN-rand (82.7%), which the README attributes to full-length
reviews plus val-based early stopping. A dropout-0.7 variant scored 87.9%
on test and was discarded. Full numbers in the pipeline README's
"Verified runs".
training/flow-matching-mnist/
is the newest pipeline here, and the first generative model family in
this repo that is not an autoencoder: flow matching / rectified flow
trained from scratch on MNIST. The model learns a velocity field
v(x,t) whose ODE transports N(0,I) at t=0 into the data at t=1;
training regresses that velocity on straight-line conditional paths with
plain MSE — x_t = (1-(1-sigma_min)*t)*x0 + t*x1, target x1 - (1-sigma_min)*x0 — which is Lipman et al.'s conditional-OT path
(2210.02747) and, at the default
--sigma-min 0.0, exactly the rectified flow of Liu et al.
(2209.03003). There is no noise
schedule, no variance parameterization, and no ELBO — that absence is the
point, and it is the difference between this and a DDPM. Same hand-written
philosophy as the rest of training/: the UNet velocity field (three
resolutions, one 7x7 self-attention block, sinusoidal time embedding as a
per-channel bias), the EMA, and the Euler/Heun ODE samplers are all plain
torch.nn; no diffusers/torchcfm/torchdiffeq/torchvision, no
DataLoader, numpy-permutation batching. ~1.18M params, 40 epochs in
~5.3 min on a single RTX 3090.
It is the deliberate counterpart to training/mnist-vae:
same dataset, same data/mnist.npz contract, same hand-written zlib PNG
writer, so the two prior-sample grids are directly comparable — the flow
model's are visibly sharper, the VAE's blur being the Gaussian-posterior
averaging that cifar10-vqvae also exists to remove. That comparison is
model-family vs model-family, not a controlled ablation: 1,175,841
params of UNet against 370,945 of plain conv encoder/decoder, and the two
runs cannot separate objective from capacity. The honest cost difference
they do show: the VAE generates in one forward pass, the flow model needs
20–50 network evaluations.
Judging it needed care, and two of this pipeline's guardrails came from
getting it wrong first. (1) The evaluator's round-trip MAE/PSNR sweep
measures ODE discretization error, not sample quality — the ODE is
time-reversible, so a real digit can be integrated back to noise and
forward again, but a near-zero velocity field round-trips perfectly since
the identity is its own inverse; a 2-epoch smoke run really did post a
better round-trip PSNR than the converged model. The sweep's actual use is
choosing --num-steps, and it shows Heun winning per network evaluation,
not just per step (10 Heun steps beat 50 Euler steps by ~2 dB at the same
100 evals). (2) The nearest-neighbour memorization check reports a distance
that is meaningless without a scale, so real held-out test digits are
measured against the training set the same way as a control. There is
deliberately no FID — it would require a pretrained Inception network,
against this folder's from-scratch rule, and a substitute number would be
worse than none. Measured figures in the pipeline README's "Verified runs".
training/rvq-audio-codec/ is the
newest pipeline here, and the first audio pipeline in the repo: a
neural audio codec with residual vector quantization (the
EnCodec/SoundStream/DAC architecture) trained from scratch on LJSpeech
(13,100 wavs at E:\datasets\LJSpeech-1.1, 23.92 h, parsed by a
hand-written RIFF/WAVE chunk walker - no torchaudio, no soundfile, no
librosa, no scipy). A SEANet-style strided conv encoder maps the waveform
to 68.9 frames/s, a stack of 8 codebooks x 1,024 entries quantizes each
frame (each codebook quantizing the residual the previous one left), and a
mirrored transposed-conv decoder reconstructs it - 5.51 kbps. 7,338,658
params, plus a 2,112,582-param multi-scale STFT discriminator that exists
only during training.
It is the deliberate successor of
training/cifar10-vqvae: that folder
has one codebook of 512 entries looked up in the full 64-dim latent, this
one stacks eight looked up in an 8-dim factorized projection under cosine
distance, with EMA updates and dead-code re-initialization. 9 bits per
latent position is enough for a 32x32 thumbnail and nowhere near enough for
a waveform; RVQ is how the bit budget is bought without a K^N codebook.
It is also the layer every modern audio LM (VALL-E, MusicGen, Moshi) sits
on - those models generate codec tokens, not waveforms.
Three things are deliberate and worth not undoing. (1) LJSpeech is 22,050
Hz and is trained at that native rate - no resampler is written, so the
frame rate is 68.9 Hz and the bitrate 5.51 kbps rather than EnCodec's
published 24 kHz / 75 Hz / 6 kbps. (2) Quantizer dropout (a random
n_q in [1, N] on half of each batch) is what makes one trained model
serve the whole 1->8 codebook ladder; without it the 1->8 quality demo
would need eight separate runs. (3) The discriminator is staged behind
--adv-start-step, because a randomly-initialized generator fighting a
randomly-initialized discriminator collapses a codec in the first thousand
steps; --lambda-adv 0 turns it off entirely for a reconstruction-only
A/B.
Judging it needs the same care as flow matching's ODE sweep. SI-SDR is a
weak proxy for a GAN-trained codec - the adversarial loss trades exact
waveform/phase alignment for perceptual realism, so a model that sounds
better can post a worse SI-SDR than a reconstruction-only one. The real
evaluation is the original_NN.wav / recon_nq{8,4,2,1}_NN.wav files the
evaluator writes (hand-written 44-byte RIFF writer, the inverse of the
builder's parser), plus the per-codebook usage table that says whether the
8th codebook is doing any work. There is deliberately no ViSQOL/PESQ/
NISQA - each needs an external binary or a pretrained network, the same
rule that keeps FID out of flow-matching-mnist.
The discriminator is also where the run's cost lives, by a wide margin.
Measured on the RTX 3090 at batch 32: reconstruction-only runs at 7.75
steps/s, and turning the discriminator on in fp32 drops that to 0.95 —
8x, because its spectrograms are much larger than the waveform they
judge (173 x 257 positions at the 512-point resolution against 22,080
samples, three resolutions, three passes per step) and those conv shapes
map badly onto fp32 tensor cores. cudnn.benchmark and TF32 matmul were
both measured and change nothing (0.90-0.95 steps/s, inside the noise).
What does work is bf16 autocast on the critic only (--disc-bf16,
default on): 2.01 steps/s and 10.8 GiB peak instead of 16.8, turning a
7-hour 60-epoch run into a 3.3-hour one with no architecture change. The
generator, the codebook lookup and every EMA update deliberately stay fp32
— bf16 EMA statistics would quietly stop accumulating small updates, which
is exactly the mechanism dead-code revival exists to detect.
One guardrail came from getting it wrong first: the dead-code cutoff is a
fraction of uniform codebook usage, not the absolute 2.0 that EnCodec and
vector-quantize-pytorch use. One batch here is 32 x 69 = 2,208 vectors
over 1,024 entries, so uniform usage is only 2.16 per entry and an absolute
2.0 condemns half a healthy codebook every sweep - the first smoke run
reported 1,023 of 1,024 entries "revived" per codebook; after the fix, 0.
None of the fine-tuning pipelines ship data — DATASET/, data/, runs/,
output/ etc. are all git-ignored, drop-zone folders (via the single root
.gitignore). Both vicuna-7b-lora/ and qwen25-3b-lora/ train on
CNN/DailyMail. Example small, permissively-licensed public datasets are
listed in the root README.md's Datasets section.
training/rvq-audio-codec is the one pipeline whose prepared data is too
large for the .npz contract the others share: LJSpeech is 3.80 GB of
int16, which becomes 7.6 GB as float32. It writes a raw
data/ljspeech_audio.i16 opened with np.memmap plus a small
data/ljspeech_index.npz of offsets/lengths/ids/split, and converts crops
to float one batch at a time. Both are covered by the root .gitignore's
data/ rule like everything else.
2,000-sample JSONL (1,800 train / 200 val), 2 epochs, 900 steps, ~61 min on
a single RTX 3090. Train loss 1.76 → 0.99, eval_loss essentially flat across
epochs (1.100 → 1.094 — a mild overfitting signal in isolation, train loss
kept falling while eval_loss didn't). What matters: reconstruction-test avg
token-F1 rose from 0.357 (an earlier 1-epoch/1,800-sample run on the
predecessor llava15-lm-lora pipeline) to 0.467 on this run, and the
generated summaries reproduced exact figures from source text correctly
(e.g. "383-41" and "70-26" vote counts). Confirms the earlier lesson again:
eval_loss plateauing is not itself a stop signal — the reconstruction test
is what actually shows whether a further epoch helped.
uv run --directory <folder> python <script> --help, and further real
executions where noted, actually run, not assumed:
-
training/cifar10-vqvae(successor of the formercifar10-vae) — real runs on the RTX 3090 against the actual downloaded CIFAR-10 python pickles (E:\datasets\cifar-10-python):build_cifar10_dataset.pywrotedata/cifar10.npz(50k train / 10k test);train_vqvae.pyandevaluate_vqvae.pyrun end-to-end (100 epochs, ~10–15 min, codebook perplexity ~404/512, 512/512 codes fired on test — no collapse). Measured reconstruction on the held-out test set: PSNR 25.2 dB, SSIM 0.884, MAE 0.042, high-frequency detail kept 73.8% — vs the predecessor VAE's best (PSNR 21.5 dB, SSIM 0.742, HF 51.4%), i.e. the discrete-codebook family removes the plain-VAE blur mechanism. Full numbers in the pipeline README's "Verified runs". -
training/flow-matching-mnist— real run on the RTX 3090 against the actual MNIST IDX files atE:\datasets\mnist-dataset:build_mnist_dataset.pywrotedata/mnist.npz(60k train / 10k test);train_flow.pytrained 1,175,841 params for 40 epochs in 317.5 s (val velocity MSE 0.2263 → 0.1704, best at epoch 38 — train and val track each other throughout, no overfitting);evaluate_flow.pyscored 0.1687 velocity MSE on the held-out 10k and wrote all three PNGs. Reproduced in a second independent run of the same commands (312.1 s, best val 0.1705, test 0.1687) — expect the third decimal to move. Samples are clean, readable digits with a handful of malformed glyphs per 64. The ~0.17 loss floor is expected, not a defect:u = x1 - x0is irreducibly random given(x_t, t), so the MSE floors at that conditional variance and can never reach zero — judge the samples, the same lessonfine-tuning/vicuna-7b-lorataught about loss plateaus. Memorization check: generated samples sit farther from the training set (mean L2 4.029, min 1.858) than real unseen test digits do (3.611 / 1.169). -
training/rvq-audio-codec— real build on the RTX 3090 machine against the actual LJSpeech drop atE:\datasets\LJSpeech-1.1:build_ljspeech_dataset.pyverified all 13,100 wavs (22,050 Hz / 16-bit / mono PCM, confirmed by reading the RIFF headers, not assumed) and wrotedata/ljspeech_audio.i16— 1,898,881,532 samples, 23.92 h, 3.80 GB — plus the index (12,838 train / 262 val; shortest utterance 1.11 s, longest 10.10 s, mean 6.57 s, so nothing is dropped by the 1.0014 s crop). The model builds at 7,338,658 params (3,659,936 encoder / 3,660,162 decoder / 18,560 RVQ) with a 2,112,582-param discriminator, and the full forward/backward path, the 1→8 ladder, the odd-length padding path and the loss stack were shape- and gradient-checked before training. The hand-written WAV writer round-trips through the hand-written parser at the 16-bit quantization floor (max abs error 5.32e-05 against a floor of 3.05e-05). A 2-epoch smoke run (802 steps, discriminator from step 0) trained without NaN, val mel 7.363 → 7.224, and — the number that matters for RVQ — showed no collapse down the stack: after 2 epochs the 8th codebook's perplexity (299) was as high as the 1st's (251), unused entries fell across every quantizer, and dead-code revivals dropped from 3,087 to 788. Throughput was profiled rather than guessed (see Stage 4). Training figures in the pipeline README's "Verified runs". -
fine-tuning/qwen25-3b-lora—build_qwen3b_dataset.py,train_qwen3b_lora.py,generate_qwen3b_lora.py, plus a real 40-sample smoke train against the actual downloadedQwen/Qwen2.5-3B-Instructweights, confirmingtrainable params > 0(LoRA genuinely attached toq_proj/v_proj) rather than trustingpeft's target-module table alone. -
fine-tuning/vicuna-7b-lora— real 2,000-sample/2-epoch training run (see above), executed by the repo owner, not just a smoke test. Not executed this pass (no PDFs/poppler set up in this environment): -
pre-training/exec_1.bat/scripts/convert_pdf_to_png.ps1
fine-tuning/vicuna-7b-lorahas a verified-good real run (see above) — reasonable next moves are more samples (the eval_loss plateau suggests more epochs on this same 1,800-row set has limited further upside), judged by the reconstruction test, not loss alone.fine-tuning/qwen25-3b-lorais smoke-tested but not yet trained for real — same next step as Vicuna's first run: build a few-thousand-sample JSONL, train, then judge with--jsonl-eval.training/has a verified realcifar10-vqvaerun (see above) — the planned next rung is stage 2 of its cascade: a learned prior over the discrete code grid (PixelCNN/transformer over code indices, or a latent DDPM), latent-diffusion style.training/imdb-sentiment-cnnis verified at 89.2% test acc with random embeddings in31 s — natural next rungs: a GloVe variant (+1–3 pts expected), or the bigger from-scratch projects that use the 50k unlabeled reviews (AWD-LSTM LM-pretrain + fine-tune, or a small transformer with MLM pretraining, both ~91% territory).training/flow-matching-mnistis verified at ~5.3 min for 40 epochs and was still improving when it stopped — natural next rungs: more epochs or a wider--base-channels; class-conditioning plus classifier-free guidance (the smallest real upgrade, and what makes samples steerable); a second rectification pass (re-train on the model's own noise/sample pairs) to straighten the paths for 1–4-step sampling, which is the whole reason rectified flow is used in production; or the same objective on CIFAR-10 next tocifar10-vqvae, where the VAE-blur comparison has more room to show itself than at 28x28.training/rvq-audio-codecis the repo's first audio pipeline — natural next rungs, in rough order of value: the reconstruction-only A/B (--lambda-adv 0) to measure what the discriminator is actually worth; the collapse-mitigation ablations (--code-dim 128,--vq-l2-normalize 0,--dead-code-threshold 0,--vq-mode loss), whose findings are meant to feed back intotraining/cifar10-vqvae's single codebook; and then the obvious sequel — an autoregressive prior over the RVQ code indices, which is what turns a codec into an audio LM and is the same "learned prior over discrete codes" rung already planned forcifar10-vqvae. Multi-speaker (LibriTTS/VCTK) is the fix for the single-speaker limitation, but only if generalization becomes the goal.fine-tuning/llava15-full-lora(planned, not started): the first real VLM fine-tune in this repo — image+text pairs, vision encoder/projector actually in the training graph, unlikevicuna-7b-lora/qwen25-3b-lora.- A
phi35-mini-lorasibling (discussed, not started) would needtarget_modules=["qkv_proj"]instead of["q_proj", "v_proj"]— Phi-3 fuses Q/K/V into one linear layer (confirmed by readingPhi3Attention's source), so thevicuna-7b-lora/qwen25-3b-loratarget-module config would silently attach to nothing on that model.