The featured open-source model generates full songs from style prompts and lyricsAudio generation uses AI to create speech, music, sound effects or other audio from instructions and reference inputs., supports multiple languages, and can transform reference audio into a different genre or arrangement. The demonstrations cover original songs, cover-style workflows, and iterative editing through an AI agent.
The model first creates an editable score containing vocal and instrumental notes, key, and tempo, then uses that structure to synthesize audio. The ComfyUI workflow exposes practical controls for duration, seed, sampling steps, prompt adherencePrompt adherence is the degree to which a generated result accurately follows the requested subjects, actions, style, structure, constraints, and details., decoding, lyrics, and instrumental-only generation.
The installation uses separate audio-encoderAn audio encoder converts sound into a compact representation that an AI system can store, compare, or use while generating new audio. and checkpointAn AI model checkpoint is a saved version of model parameters and related state from a particular point in training or development. downloads, with a smaller quantized checkpoint available for limited VRAMModel quantization represents AI model values with fewer bits to reduce storage, memory use and often inference cost.. The code and documentation use the Apache 2.0 license, while the model weights have a separate non-commercial license that may restrict monetized use.
Watch on YouTube



