Text to Audio — Single Stage
Uses LTX-2.5's audio pathway on its own to generate sound effects, ambience or short music beds from a text prompt. Because no video VAE or upscaler is loaded, it is the smallest footprint of all the example graphs.
Required models
Drop each file into the matching folder under ComfyUI/models/, or let Workflow Overview download them all.
| File | Folder | Size | Purpose | Link |
|---|---|---|---|---|
ltx-2.5-22b-distilled-transformer-bf16.safetensors | ComfyUI/models/diffusion_models/ | 42 GB | The LTX-2.5 22B distilled video model — used by every example workflow. Few sampling steps, built for fast iteration. | Hugging Face |
↳ lower-VRAM alternative ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors | ComfyUI/models/diffusion_models/ | 21.5 GB | INT8 — about half the size; the usual pick for 24 GB cards. | Hugging Face |
↳ lower-VRAM alternative ltx-2.5-22b-distilled-transformer-nvfp4.safetensors | ComfyUI/models/diffusion_models/ | 18.7 GB | NVFP4 — smallest; needs an NVIDIA GPU with FP4 support (Blackwell). | Hugging Face |
gemma4-12b-with-proj-ltx-2.5-bf16.safetensors | ComfyUI/models/text_encoders/ | 26.3 GB | Gemma 4 12B text encoder with the LTX-2.5 projection — turns your prompt into conditioning. Required by all workflows. | Hugging Face |
↳ lower-VRAM alternative gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors | ComfyUI/models/text_encoders/ | 15.4 GB | INT8 text encoder — saves ~11 GB on disk and in memory. | Hugging Face |
ltx-2.5-audio-vae-bf16.safetensors | ComfyUI/models/vae/ | 0.4 GB | Audio VAE — decodes the synchronized soundtrack LTX-2.5 generates alongside the video. | Hugging Face |
ltx-2.5-duration-head-bf16.safetensors | ComfyUI/models/model_patches/ | 0.05 GB | Small duration-prediction head used by audio workflows. | Hugging Face |
How to use
- 1
Open ComfyUI (current version, with ComfyUI-LTXVideo installed via Manager).
- 2
Workflow → Open and pick the downloaded .json, or drag the file onto the canvas.
- 3
Open Workflow Overview → Missing Models → Download all, and wait for the files to finish.
- 4
Describe the sound you want (source, texture, rhythm, space).
- 5
Set the duration.
- 6
Click Run and preview the audio in the output node.
VRAM
Which transformer build fits your card for this graph.
| VRAM | Example cards | Verdict | Transformer | Advice |
|---|---|---|---|---|
| 24 GB | RTX 3090, 4090, A5000 | OK | int8 | Workable with the INT8 transformer + INT8 text encoder and the low-VRAM loader nodes. Two-stage graphs run; keep decode tiles small. Leave bf16 files for 32 GB+ cards. |
| 32 GB | RTX 5090, V100 32G | Comfortable | bf16 | Lightricks' stated minimum. bf16 transformer with the low-VRAM loaders fits; INT8 text encoder frees headroom for larger decode tiles and longer clips. |
| 48 GB+ | RTX 6000 Ada, A6000, H100 | Comfortable | bf16 | Everything in bf16 without offloading. Increase the decode tile size (see the Decode notes in each graph) for faster runs. |
Official requirement is 32 GB+. Lower tiers are guidance for quantized builds.
Under 24 GB?
The same LTX-2 model runs in our browser generator — 4K with audio, no setup, free credits to start.
Troubleshooting
FAQ
Other workflows
Text / Image to Video — Two Stage
Start here. Base-resolution pass, 2× spatial upscale, then a 3-step refine. Video and audio together.
Text / Image to Video — Single Stage
Same inputs, one distilled pass, no upscaler. Faster and lighter; less spatial detail than two-stage.
Audio to Video — Two Stage
Generate video that follows an input soundtrack. The original waveform is kept and muxed into the output.