The official starting point

PartDocumented Python setup
Operating systemLinux
Python3.12
GPUNVIDIA, with BF16 support
VRAM24 GB
ConcurrencyOne request at a time
Output48 kHz stereo

This is the authors’ supported starting configuration, not a measured minimum for every implementation. A “3B” model label alone does not tell you peak memory: intermediate states and audio decoding also consume memory.

What if you have 8, 12 or 16 GB?

The native ComfyUI template uses an INT8 checkpoint. Quantization and offloading can change memory use, but the linked ComfyUI guide does not establish a universal VRAM minimum. We do not promise that every 8 GB card will finish every song.

For a lower-memory machine, begin with the native ComfyUI workflow, one short request and no other GPU jobs. Record the GPU, available VRAM, application version, precision, duration and any offloading settings when comparing a community report with your own setup.

Windows, AMD and Apple Silicon

The authors’ Python quick start is written for Linux and CUDA. The ComfyUI documentation separately mentions AMD support and recommends version 0.35.1 or later for those GPUs. Follow the instructions for the implementation you actually use.

We have not verified a native Apple Silicon setup or measured a Windows configuration. Treat community ports as separate projects with their own supported hardware. If you only want to hear a first result, the hosted demo avoids a local setup.

When generation runs out of memory

  1. Check that the selected checkpoint matches the template.
  2. Close applications and jobs using the same GPU.
  3. Reduce the requested song length and run one request.
  4. For covers, run transcription and generation sequentially.
  5. Keep the error message and environment details before changing more settings.

Do not treat a file that plays as proof of a complete run. Check the saved truncation flags as well as the ending of the song.

Sources & further reading

Documentation checked against the linked sources. We have not benchmarked generation on our own hardware.