AI movie generation process https://spacecruft.org/movies/process
  • Python 84.8%
  • Shell 13.3%
  • Svelte 1%
  • TypeScript 0.8%
Find a file
2026-08-08 21:15:50 -06:00
.kilo/rules Useful comments plz 2026-04-19 12:46:33 -06:00
000-project-setup bc + prices for runpod 2026-04-10 16:24:02 -06:00
010-source-material Fix credits 2026-05-14 13:53:27 -06:00
020-character-extraction more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
030-location-extraction Ruff exceptions 2026-07-24 19:51:32 -06:00
040-plot-structure Ruff exceptions 2026-07-24 19:51:32 -06:00
050-themes-and-tone Ruff exceptions 2026-07-24 19:51:32 -06:00
060-screenplay Use unbuffered python for logs 2026-05-05 19:50:39 -06:00
070-storyboard-script fixes 2026-08-08 20:57:23 -06:00
080-style-guide more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
081-character-relationships more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
082-character-visual-prompts More lint fixes, ruff 2026-07-24 20:04:52 -06:00
084-location-visual-prompts More lint fixes, ruff 2026-07-24 20:04:52 -06:00
086-prop-extraction Ruff formatting 2026-07-24 20:45:06 -06:00
090-character-design Ruff formatting 2026-07-24 20:45:06 -06:00
092-location-hero-generation more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
093-location-hero-selection Ruff formatting 2026-07-24 20:45:06 -06:00
094-location-design more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
098-prop-design more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
100-character-dataset-curation Ruff formatting 2026-07-24 20:45:06 -06:00
102-location-dataset-curation Ruff formatting 2026-07-24 20:45:06 -06:00
104-prop-dataset-curation Ruff formatting 2026-07-24 20:45:06 -06:00
110-character-image-lora-training more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
112-location-image-lora-training more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
114-prop-image-lora-training more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
115-character-identity-calibration Ruff formatting 2026-07-24 20:45:06 -06:00
116-prop-identity-calibration Ruff formatting 2026-07-24 20:45:06 -06:00
120-character-image-lora-validation Ruff formatting 2026-07-24 20:45:06 -06:00
121-character-identity-recalibration Ruff formatting 2026-07-24 20:45:06 -06:00
122-location-image-lora-validation Ruff formatting 2026-07-24 20:45:06 -06:00
124-prop-image-lora-validation Ruff formatting 2026-07-24 20:45:06 -06:00
127-character-lora-identity-reselect more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
128-prop-lora-identity-reselect more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
160-mood-boards fixes 2026-08-08 20:57:23 -06:00
170-scout-generation more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
171-scout-segmentation fixes 2026-08-08 20:57:23 -06:00
172-location-foundation-generation fixes 2026-08-08 20:57:23 -06:00
173-location-foundation-selection Ruff formatting 2026-07-24 20:45:06 -06:00
174-character-inpainting Ruff formatting 2026-07-24 20:45:06 -06:00
175-prop-inpainting Ruff formatting 2026-07-24 20:45:06 -06:00
176-harmonization Ruff formatting 2026-07-24 20:45:06 -06:00
177-face-refinement Ruff formatting 2026-07-24 20:45:06 -06:00
178-keyframe-selection Ruff formatting 2026-07-24 20:45:06 -06:00
179-keyframe-qa Ruff formatting 2026-07-24 20:45:06 -06:00
180-storyboard-image-refinement more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
189-voice-descriptions More lint fixes, ruff 2026-07-24 20:04:52 -06:00
190-voice-profiles More lint fixes, ruff 2026-07-24 20:04:52 -06:00
200-dialogue-generation Ruff exceptions 2026-07-24 19:51:32 -06:00
205-dialogue-selection More ruff fixes 2026-07-24 20:44:49 -06:00
210-narration Ruff exceptions 2026-07-24 19:51:32 -06:00
215-narration-selection More ruff fixes 2026-07-24 20:44:49 -06:00
220-video-strategy More lint fixes, ruff 2026-07-24 20:04:52 -06:00
225-video-prompt-grounding Ruff formatting 2026-07-24 20:45:06 -06:00
230-video-clip-generation more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
240-video-clip-selection fmt 2026-07-24 21:38:20 -06:00
260-video-upscaling Manual ruff fixes 2026-07-24 19:32:09 -06:00
270-video-finalization Manual ruff fixes 2026-07-24 19:32:09 -06:00
300-sfx-planning Ruff exceptions 2026-07-24 19:51:32 -06:00
310-sfx-freesound Ruff formatting 2026-07-24 20:45:06 -06:00
311-sfx-generation Ruff exceptions 2026-07-24 19:51:32 -06:00
312-sfx-selection more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
320-music-planning Ruff formatting 2026-07-24 20:45:06 -06:00
330-music-theme more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
340-music-scene-scores more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
350-music-stingers-and-credits more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
360-rough-cut more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
370-transitions Ruff formatting 2026-07-24 20:45:06 -06:00
380-audio-mixing Ruff formatting 2026-07-24 20:45:06 -06:00
390-color-grading Manual ruff fixes 2026-07-24 19:32:09 -06:00
400-title-sequence More lint fixes, ruff 2026-07-24 20:04:52 -06:00
410-credits more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
420-subtitles more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
430-final-export Ruff exceptions 2026-07-24 19:51:32 -06:00
440-quality-review Use unbuffered python for logs 2026-05-05 19:50:39 -06:00
450-issue-log Fix changed step name 2026-05-10 14:01:38 -06:00
460-revisions Use unbuffered python for logs 2026-05-05 19:50:39 -06:00
470-feedback Use unbuffered python for logs 2026-05-05 19:50:39 -06:00
480-final-delivery more code modernization, ruff cleanups 2026-07-24 20:39:16 -06:00
docs JoyAI noted 2026-08-08 21:15:50 -06:00
gui Update models used 2026-05-10 19:57:42 -06:00
img Update web screenshot 2026-04-15 13:05:39 -06:00
notes Process tracking notes 2026-03-26 15:28:18 -06:00
scripts Ruff formatting 2026-07-24 20:45:06 -06:00
.env.example Rework clip selection 2026-07-24 21:38:07 -06:00
.gitattributes git LFS tracking of large files 2026-03-26 15:27:42 -06:00
.gitignore Fix credits 2026-05-14 13:53:27 -06:00
.python-version Default python 3.13 2026-04-02 15:14:08 -06:00
0-setup-all.sh Fix download script name 2026-04-11 10:23:56 -06:00
1-run-llm-detached.sh Add scriptlets to run each step detached with logs 2026-04-03 15:10:12 -06:00
1-run-llm.sh Seperate log files for each step when running detached 2026-04-03 18:15:22 -06:00
2-run-comfyui-detached.sh Add scriptlets to run each step detached with logs 2026-04-03 15:10:12 -06:00
2-run-comfyui.sh Major renumbering, split of step 092 into multiple 2026-04-10 13:07:09 -06:00
3-run-curation-detached.sh Add scriptlets to run each step detached with logs 2026-04-03 15:10:12 -06:00
3-run-curation.sh Shutdown verda pipeline, unless env is set 2026-04-11 16:36:18 -06:00
4-run-training-detached.sh Add scriptlets to run each step detached with logs 2026-04-03 15:10:12 -06:00
4-run-training.sh Seperate log files for each step when running detached 2026-04-03 18:15:22 -06:00
cleanup.sh Add cleanup script 2026-04-05 10:20:20 -06:00
LICENSE-apache.txt Apache 2.0 2026-03-26 14:13:25 -06:00
MODELS.json Update to LongCat-Video-Avatar version 1.5 2026-05-28 11:25:15 -06:00
README.md Update to LongCat-Video-Avatar version 1.5 2026-05-28 11:25:15 -06:00
requirements-dev.txt Add some dev py deps 2026-04-11 10:20:46 -06:00
ruff.toml More ruff fixes 2026-07-24 20:44:49 -06:00

AI Movie Production Process

AI Movie Production

A complete, step-by-step pipeline for transforming short stories into fully produced AI-generated films using open source tools.

Web GUI


Quick Start

# Clone this template to start a new movie project
git clone https://spacecruft.org/movies/process my-movie-project
cd my-movie-project

# Initialize Git LFS (one-time per machine)
git lfs install

# Copy and configure environment
cp .env.example .env
nano .env   # Set API URLs, model paths, GPU config

# Start at step 000
cat 000-project-setup/README.md

Then work through each numbered directory in order. Each directory has a README.md explaining exactly what to do, what inputs you need, and what outputs to produce.


What's Inside

69 step directories organized into 13 phases with 6 GPU swap points:

Phase Steps Description
Pre-Production — Story & Script 000080 Source story, extract characters/locations, write screenplay, create storyboard, define style guide
Pre-Production — Visual Prompts 082086 LLM extracts visual descriptions for characters, locations, props
🔄 GPU Swap 1: Unload LLM → Load ComfyUI
Pre-Production — Visual Design 090098 Generate design images via ComfyUI for characters, locations (PyraCanny ControlNet), props
🔄 GPU Swap 2: Unload ComfyUI → Load Qwen3-VL
Pre-Production — Dataset Curation 100104 VLM rate (Qwen3-VL-Reranker + DreamSim/Embedding), curate, and caption training images
🔄 GPU Swap 3: Unload Qwen3-VL → Load AI Toolkit
Pre-Production — LoRA Training 110114 Train Qwen-Image LoRAs for characters, locations, props (AI Toolkit)
🔄 GPU Swap 4: Unload AI Toolkit → Load ComfyUI
Pre-Production — LoRA Validation 120124 Automated checkpoint + strength selection via DreamSim + Reranker scoring, visual validation galleries
🔄 GPU Swap 5: Unload ComfyUI → Load LLM
Production — Shot Planning 160 LLM maps shots to LoRAs, compile per-scene mood boards
🔄 GPU Swap 6: Unload LLM → Load ComfyUI
Production — Keyframe Generation 170180 Scout generation, GDINO/SAM2 segmentation, layered inpainting (location → character → prop), harmonization, face refinement, keyframe selection, outpainting
Production — Voice & Dialogue 189210 Qwen3-TTS voice design, dialogue generation, narration (audio before video — LongCat requires dialogue as input)
Production — Video Generation 220270 LongCat-Video-B200 video clips with lip sync, VLM selection, upscale, finalize
Production — Audio: Sound & Music 300350 FreeSound CC0 retrieval, MOSS-SoundEffect SFX generation, Qwen3-Omni SFX selection, ACE-Step 1.5 XL Turbo music composition
Post-Production — Assembly & Finishing 360430 Edit, mix, grade, titles, credits, subtitles, export
Review & Delivery 440480 Quality review, revisions, feedback, ship

Key Tools & Infrastructure

  • LLM — OpenAI-compatible API for all text analysis, screenplay, storyboard, prompt extraction, and shot planning
  • ComfyUI — Image generation with multi-GPU parallel dispatch (COMFYUI_API_URLS), PyraCanny ControlNet for location consistency
  • AI Toolkit — LoRA training for Qwen-Image models (characters, locations, props) with parallel GPU training
  • Qwen3-VL — Dataset curation and LoRA validation (Reranker-8B for text relevance, Embedding-8B for cluster consistency, 30B-A3B-Instruct for captioning)
  • DreamSim — Perceptual similarity scoring (CLIP + OpenCLIP + DINO ViT-B/16 ensemble) for curation, hero selection, and LoRA validation
  • Grounding DINO + SAM2.1 — Text-guided object detection and pixel-precise segmentation for keyframe entity masking
  • LongCat-Video-B200 — Video generation with audio-driven lip sync and multi-character support
  • Qwen3-TTS — Voice design (VoiceDesign model) and dialogue generation (Base model, ICL cloning)
  • FreeSound — CC0 sound effects retrieval via API with LLM-generated search queries
  • MOSS-SoundEffect — Text-to-audio SFX generation (8B, duration-controlled)
  • Qwen3-Omni — Audio understanding model (30B-A3B MoE) for SFX candidate evaluation and selection
  • ACE-Step 1.5 XL Turbo — AI music composition (4B DiT + 5Hz LM, theme, scene scores, stingers)
  • Configuration — All settings via root .env.example (API URLs, model paths, GPU IDs, training hyperparameters)

Models

  • ACE-Step 1.5 XL Turbo DiT — AI music composition diffusion transformer
  • ACE-Step 5Hz LM — AI music composition language model
  • AuraFace-v1 — face identity embedding model used for character identity calibration
  • CLIP ViT-B/16 — vision-language backbone used in DreamSim ensemble
  • DINO ViT-B/16 — self-supervised vision backbone used in DreamSim ensemble
  • DINOv2-giant — self-supervised vision backbone for body identity embeddings in character/prop calibration
  • DreamSim — perceptual similarity ensemble for image-pair distance scoring
  • GFPGAN v1.4 — face restoration for keyframe refinement
  • GLM 5.1 — text LLM for story analysis, screenplay, storyboard, prompt extraction, shot planning
  • Grounding DINO — text-guided object detection for entity segmentation
  • LibreHPS-4B v1.1 — aesthetic / human-preference scorer
  • LongCat-Video — image-to-video generation model
  • LongCat-Video-Avatar-1.5 — avatar video generation with audio-driven lip sync
  • MOSS-SoundEffect — text-to-audio sound effect generation
  • OpenCLIP ViT-B/16 — vision-language backbone used in DreamSim ensemble
  • Qwen-2.5-VL-7B — CLIP text encoder for Qwen-Image
  • Qwen-Image — image generation diffusion model
  • Qwen-Image ControlNet-Union — ControlNet for location design + character/prop inpainting
  • Qwen-Image-Edit — image editing diffusion model for outpainting
  • Qwen-Image VAE — variational autoencoder for Qwen-Image
  • Qwen3-Omni — audio understanding model for SFX evaluation and selection
  • Qwen3-TTS Base — text-to-speech base model for dialogue and narration with ICL voice cloning
  • Qwen3-TTS VoiceDesign — text-to-speech voice design model for voice profile creation
  • Qwen3-VL — vision-language instruct model for image captioning and binary classification checks
  • Qwen3-VL-Embedding-8B — vision-language embedding model for image similarity scoring
  • Qwen3-VL-Reranker-8B — vision-language reranker model for text-image relevance scoring
  • SAM2.1 — Segment Anything Model 2.1 for pixel-precise entity mask extraction
  • SigLIP2-so400m — multilingual vision-language backbone for body identity embeddings in character/prop calibration

Batch Runners

Run entire phases unattended (each has a -detached.sh variant for tmux/SSH):

Script Phase Steps
0-setup-all.sh Setup Download + setup all 69 step directories
1-run-llm.sh LLM extraction 020086 (story analysis, screenplay, visual prompts)
2-run-comfyui.sh Image generation 090098 (character/location/prop design)
3-run-curation.sh Dataset curation 100104 (rate, curate, caption)
4-run-training.sh LoRA training 110114 (character/location/prop LoRAs)

Documentation


Project Structure

my-movie-project/
├── 000-project-setup/          69 step directories (each has README.md, setup.sh, run.sh)
├── 010-source-material/        ↓
├── ...                         ↓
├── 480-final-delivery/         ↓
├── docs/                       Overview, guides, and reference docs
├── img/                        Project logo/branding
├── notes/                      Generation logs, revision notes, lessons learned
├── scripts/                    Backend tooling (runcomfy, runpod, verda)
├── .env.example                Global configuration template
├── .gitattributes              Git LFS tracking rules
├── .gitignore                  Ignore patterns for AI production
├── 0-setup-all.sh              Setup all step directories
├── 1-run-llm.sh                Batch: LLM extraction phase (020086)
├── 2-run-comfyui.sh            Batch: ComfyUI image generation (090098)
├── 3-run-curation.sh           Batch: dataset curation (100104)
├── 4-run-training.sh           Batch: LoRA training (110114)
├── cleanup.sh                  Remove all outputs, logs, completion markers
├── LICENSE-apache.txt          Apache 2.0 license
└── README.md                   This file

Per-Step Workflow

Each step from 010 onward follows this pattern:

cd NNN-step-name/
# direnv auto-activates the correct Python venv when you cd in

# First-time setup:
bash setup.sh          # Create venv, install dependencies
bash download.sh       # Download models/workflows (if present)
bash download_models.sh # Download ML models (if present)

# Run the step:
bash run.sh

# Review outputs in output/
# Commit when satisfied:
git add -A && git commit -m "NNN: Description"

Each step directory has its own isolated environment:

  • venv/ — Python virtual environment (created by setup.sh)
  • .python-version — pyenv Python version
  • .envrc — direnv config (auto-activates venv on cd)
  • pyproject.toml — Python dependencies
  • setup.sh — creates venv and installs dependencies
  • run.sh — executes all scripts in sequence
  • scripts/ — Python scripts implementing the step

License

Apache 2.0. See LICENSE-apache.txt.

Copyright © 2026 Jeff Moe.

Loveland, Colorado, USA