a python pipeline for aligning 4-lens nimslo camera photos into smooth stereoscopic boomerang gifs. takes 4 slightly offset images and aligns them so the subject stays in place while the background shifts, creating that classic nimslo parallax effect.
- preprocessing β reduces film grain and normalizes exposure across frames
- segmentation β detects the main subject using uΒ²-net (cli default); optional fallbacks exist in
segmentation.pyfor experiments - centering β translates each frame so the mask centroid aligns to the reference
- alignment β sift matching + affine-ransac inlier rejection + translation-only warp
- render β applies transforms to the original scans (preserves film grain)
- export β boomerang gif or mp4 with crop + brightness normalization
βββββββββββββββ ββββββββββββββββ βββββββββββββββββββ
β 4 scans ββββΆβ preprocess ββββΆβ uΒ²-net segment β
β (jpg) β β denoise/exp β β + mask centroid β
βββββββββββββββ ββββββββββββββββ ββββββββββ¬βββββββββ
β
ββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββ
β per-frame pair (β ref frame 1) β
β sift (masked) β flann + lowe 0.75 β
β β affine_partial ransac (inliers only) β
β β centroid translation fit (0Β° rotation) β
ββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββ
β
ββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββ
β warp originals β boomerang β gif/mp4 β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
the final warp is always translation-only β nimslo lenses are fixed horizontally and rotation would break the stereo effect. the trick is using a looser ransac model for inlier voting (affine partial: translation + rotation + uniform scale) while discarding the rotation/scale from the estimated transform. this handles depth parallax during correspondence filtering without rotating the output.
confidence score: 0.5 Γ inlier_ratio + 0.5 Γ mask_iou
see notes.md for the aug 2026 benchmark that validated this approach.
nap/
βββ nimslo_cli.py # main cli
βββ nimslo_visualize.py # pipeline + matplotlib debug output
βββ notes.md # experiment log (tracked)
βββ benchmark_output/ # gitignored β csv/gif benchmark artifacts
βββ profile_output/ # gitignored β profiler csv output
βββ assets/ # readme media (tracked gifs; see .gitignore)
βββ .env.example # template for range / numeric interactive paths
βββ nimslo_core/
β βββ preprocessing.py # film grain reduction, exposure balancing
β βββ segmentation.py # uΒ²-net subject detection
β βββ alignment.py # sift matching, affine-ransac, translation warp
β βββ gif_generator.py # boomerang frame order + gif/mp4 encode
β βββ terminal_picker.py # ghostty/kitty interactive subject selection
β βββ rectification.py # stereo rectification utilities
βββ tests/ # unit tests (config, terminal picker helpers)
βββ notebooks/
β βββ nimslo_alignment_dashboard.py # molab walkthrough
βββ notes.md # development notes & experiment log
run the browser-oriented alignment walkthrough in molab:
open nimslo_alignment_dashboard.py in molab
the molab notebook uses the committed lightweight sample scans in notebooks/web-scans/.
keep full-resolution/private scans out of the public repo; use the local cli for those.
python nimslo_cli.py ./nimslo_raw/01/ -o output.gifpython nimslo_cli.py ./nimslo_raw/ --batch -o ./outputs/Range mode (nap START END) and numeric interactive mode (nap --interactive 132) read batch folders and default output locations from the environment.
Configure once:
cp .env.example .envThen edit .env:
NAP_INPUT_DIR=~/path/to/nimslo
NAP_GIF_OUTPUT_DIR=~/path/to/wigglegrams
NAP_MP4_OUTPUT_DIR=~/path/to/wigglegrams/output_mp4.env is gitignored; .env.example documents the portable configuration.
Existing process environment variables take precedence over values in the
file. Range mode uses the best preset and writes both formats from one
alignment run:
nap 20 137 # every batch in range β gif + mp4 under configured dirs
nap --interactive 132 # one batch, terminal picker, gif + mp4Output filenames use each batch directory name (e.g. 132 or 01). Existing
files are overwritten. Use --longer N if the MP4 should repeat its boomerang
sequence more than once.
If you use a shell alias, point it at this repoβs nimslo_cli.py (or install
the module); the alias name nap is optional.
python nimslo_visualize.py ./nimslo_raw/01/ -o output.gif --viz-dir ./viz/For a batch where automatic segmentation selects the wrong subject, use the terminal-native picker:
nap --interactive 132With a numeric batch, the CLI resolves NAP_INPUT_DIR/<batch>/ and writes
NAP_GIF_OUTPUT_DIR/<batch>.gif and NAP_MP4_OUTPUT_DIR/<batch>.mp4 (same
layout as range mode). For a custom input path or a single output file, use:
python nimslo_cli.py ./scans/132 --interactive -o 132.gif
python nimslo_cli.py ./scans/132 --interactive --format mp4 -o 132.mp4The picker uses the Kitty graphics protocol and pixel mouse reporting, both
supported by Ghostty. Click the subject in frame 1. The picker tracks a local
feature cloud through the remaining frames and shows all proposed anchors.
Click any frame to correct its anchor, press Enter to accept, or press q /
Escape to cancel.
The accepted anchors create local ROI masks for subject-specific SIFT matching; they do not add a new segmentation model. Interactive mode currently handles one batch at a time and must run in an attached compatible terminal.
python nimslo_cli.py INPUT [END] [-o OUTPUT] [OPTIONS]
positional:
INPUT path to one batch, parent dir (with --batch), or first batch number
END inclusive final batch number; enables `nap START END` range mode
options:
-o, --output PATH output path (file for single, directory for batch)
--batch process all subdirectories as batches
range mode:
`nap START END` always uses best quality and writes both a GIF and MP4 for each
existing numbered batch. Paths come from NAP_INPUT_DIR,
NAP_GIF_OUTPUT_DIR, and NAP_MP4_OUTPUT_DIR in the environment or local .env.
-q, --quality quality preset: fast, balanced, best (default: best)
--format output format: gif or mp4 (otherwise inferred from -o)
--show-masks save segmentation mask visualization
--interactive select and review subject anchors in the terminal
--preview open result after processing (single mode only)
-v, --verbose enable verbose output
--loops / --longer N mp4 only: repeat boomerang sequence N times| format | sequence | why |
|---|---|---|
| gif | 1β2β3β4β3β2 |
loops 2β1 cleanly, no duplicate hold |
| mp4 | 1β2β3β4β3β2β1 |
ends on frame 1 for seamless concatenation |
3-frame fallback (mechanical failure): 1β2β3β2
mp4 export uses ffmpeg + libx264 and is tuned to preserve grain:
- fps: 10
- loops: 1 (default). use
--longer N(or--loops N) to concatenate more loops. - tune: grain
- crf: 18
- even dimensions: enforced via 1px crop when needed (no resampling blur)
- edge crop: removes warp borders by intersecting valid regions across frames
| preset | sift features | max dimension | denoise |
|---|---|---|---|
| fast | 500 | 400px | no |
| balanced | 1000 | 600px | yes |
| best | 2000 | 800px | yes |
preprocess_image()β denoise + exposure balancenormalize_sizes()β match dimensions across frames
subject detection with fallback chain:
- uΒ²-net (primary) β rembg/onnxruntime
- depth-based β intel dpt
- grabcut β opencv refinement
exports: get_segmentation_mask() β (mask, confidence, method)
key functions:
| function | role |
|---|---|
extract_features() |
sift inside subject mask |
match_features() |
flann + lowe ratio test (0.75) |
ransac_inlier_mask() |
ransac inlier rejection (production: affine_partial) |
fit_translation_from_points() |
centroid translation on inliers |
estimate_translation_ransac() |
ransac + translation fit (single entry point) |
align_pair() / align_images() |
full per-pair / multi-frame alignment |
center_images_on_subject() |
mask-centroid pre-alignment |
make_boomerang_frames()β forward + reverse frame sequenceencode_gif()β pillow (cpu);encode_mp4()β ffmpeglibx264(multicore cpu)_crop_to_valid_region()β removes black warp borders_normalize_brightness()β prevents exposure flashing between frames
dev scripts (benchmark_alignment.py, benchmark_optimizations.py, profile_pipeline.py, smoke_test_framing.py) are gitignored β keep them locally for sweeps. outputs go to benchmark_output/ and profile_output/ (also gitignored).
compare alignment variants on real rolls:
python benchmark_alignment.py \
--input "$NAP_INPUT_DIR" \
--output benchmark_output \
--write-gifs --limit 12 --stride 7writes benchmark_output/alignment_benchmark.csv and optional gifs per variant.
quantify where wall time goes (segmentation substeps, alignment, export):
python profile_pipeline.py ./nimslo_raw/61/
python profile_pipeline.py --input "$NAP_INPUT_DIR" --limit 5
python profile_pipeline.py ./batch/ --segmentation-only --runs 3writes profile_output/pipeline_profile.csv.
key deps (pin versions in your own environment as needed):
opencv-pythonβ image processing, sift, ransacnumpy<2.0β onnxruntime compatibilityrembg+onnxruntimeβ uΒ²-net segmentationpillowβ gif encodingmatplotlibβ visualizations (optional)ffmpegβ mp4 export
rembg/onnxruntime causes jupyter kernel crashes on macos (openmp conflicts). use the cli instead:
python nimslo_cli.py ./nimslo_raw/01/ -o output.gifonnxruntime isn't fully compatible with numpy 2.x. requirements constrain to numpy<2.0.
see notes.md for experiment logs, benchmark results, and troubleshooting.
# single batch, best quality, preview
python nimslo_cli.py ./nimslo_raw/01/ -o my_photo.gif -q best --preview
# batch with mask debug images
python nimslo_cli.py ./nimslo_raw/ --batch -o ./outputs/ --show-masks
# alignment debug visualizations
python nimslo_visualize.py ./nimslo_raw/01/ -o output.gif --viz-dir ./debug_viz/the pipeline automatically handles:
- different image sizes (normalizes to smallest)
- exposure differences (brightness normalization)
- black borders from warping (auto-cropping)
- subject centering (mask-centroid alignment before sift)