Model storage
Roughly 70 GB for MiniMax-H3 plus the Veda predictor and Turbo LoRA. Persistent storage avoids downloading weights after every restart.
A learned attention predictor that selects the most useful 10% of tiles per head while preserving the MiniMax-H3 backbone, motion, visual style, and native audio path.
Watch the authors' outputs and prepare a text-only, first-frame, or Foodie multi-reference request in your browser. Your files stay on this device. Live inference requires dedicated GPU hardware and persistent model storage; it is not running in this Static Space.
Dense and Veda renders use the same prompt, seed, and Turbo LoRA. The titles inside each video show the generation-time comparison.
Compose a structured prompt, choose the reference mode, then download a role-preserving JSONL manifest.
PNG, JPEG, or WebP. Selecting an image switches the manifest from text-to-video (T2VA) to first-frame-to-video (FL2VA). The file is previewed locally and is never uploaded by this page.
Foodie keeps the roles separate: the KOL image owns visible identity and appearance, the storyboard owns composition and action, and the audio owns voice timbre only. Spoken narration still comes from the prompt's dialogue tags. The downloaded role manifest targets the DBee/ComfyUI Foodie handoff; Miowtion's current encoder CLI only accepts T2VA and FL2VA manifests.
The preview is a 33B audio-video DiT. Sparse attention saves compute, not the need to host the backbone.
Roughly 70 GB for MiniMax-H3 plus the Veda predictor and Turbo LoRA. Persistent storage avoids downloading weights after every restart.
Miowtion documents about 40 GB of free pinned host RAM per process. A 24 GB GPU can run with all transformer blocks offloaded.
H100/H200 avoids block streaming. Compatible Ada and Blackwell GPUs can use offloading when sufficient host RAM is available.