1-stageaudio

Text to Audio — Single Stage

Uses LTX-2.5's audio pathway on its own to generate sound effects, ambience or short music beds from a text prompt. Because no video VAE or upscaler is loaded, it is the smallest footprint of all the example graphs.

Stages
1-stage
Upscale
No
Conditioning
Text
Best for
Audio-only clips, SFX and ambience
Outputs
audio

Required models

Drop each file into the matching folder under ComfyUI/models/, or let Workflow Overview download them all.

FileFolderSizePurposeLink
ltx-2.5-22b-distilled-transformer-bf16.safetensorsComfyUI/models/diffusion_models/42 GBThe LTX-2.5 22B distilled video model — used by every example workflow. Few sampling steps, built for fast iteration.Hugging Face
lower-VRAM alternative
ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors
ComfyUI/models/diffusion_models/21.5 GBINT8 — about half the size; the usual pick for 24 GB cards.Hugging Face
lower-VRAM alternative
ltx-2.5-22b-distilled-transformer-nvfp4.safetensors
ComfyUI/models/diffusion_models/18.7 GBNVFP4 — smallest; needs an NVIDIA GPU with FP4 support (Blackwell).Hugging Face
gemma4-12b-with-proj-ltx-2.5-bf16.safetensorsComfyUI/models/text_encoders/26.3 GBGemma 4 12B text encoder with the LTX-2.5 projection — turns your prompt into conditioning. Required by all workflows.Hugging Face
lower-VRAM alternative
gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors
ComfyUI/models/text_encoders/15.4 GBINT8 text encoder — saves ~11 GB on disk and in memory.Hugging Face
ltx-2.5-audio-vae-bf16.safetensorsComfyUI/models/vae/0.4 GBAudio VAE — decodes the synchronized soundtrack LTX-2.5 generates alongside the video.Hugging Face
ltx-2.5-duration-head-bf16.safetensorsComfyUI/models/model_patches/0.05 GBSmall duration-prediction head used by audio workflows.Hugging Face
Total ≈ 68.8 GB (bf16, excluding LoRAs)See the full model list

How to use

  1. 1

    Open ComfyUI (current version, with ComfyUI-LTXVideo installed via Manager).

  2. 2

    Workflow → Open and pick the downloaded .json, or drag the file onto the canvas.

  3. 3

    Open Workflow Overview → Missing Models → Download all, and wait for the files to finish.

  4. 4

    Describe the sound you want (source, texture, rhythm, space).

  5. 5

    Set the duration.

  6. 6

    Click Run and preview the audio in the output node.

VRAM

Which transformer build fits your card for this graph.

VRAMExample cardsVerdictTransformerAdvice
24 GBRTX 3090, 4090, A5000OKint8Workable with the INT8 transformer + INT8 text encoder and the low-VRAM loader nodes. Two-stage graphs run; keep decode tiles small. Leave bf16 files for 32 GB+ cards.
32 GBRTX 5090, V100 32GComfortablebf16Lightricks' stated minimum. bf16 transformer with the low-VRAM loaders fits; INT8 text encoder frees headroom for larger decode tiles and longer clips.
48 GB+RTX 6000 Ada, A6000, H100Comfortablebf16Everything in bf16 without offloading. Increase the decode tile size (see the Decode notes in each graph) for faster runs.

Official requirement is 32 GB+. Lower tiers are guidance for quantized builds.

Under 24 GB?

The same LTX-2 model runs in our browser generator — 4K with audio, no setup, free credits to start.

Generate online free

Troubleshooting

FAQ

Workflow © Lightricks, LTX-2 Community License. Links point to the official repository: Lightricks/ComfyUI-LTXVideo