AI Music Generation, From Prompt to Production

Phlo Young54:34 · Feb 2025 · 2,807 views
Thumbnail for AI Music Generation, From Prompt to Production Watch on YouTube
TL;DR
  1. 1

    AI music includes separate workflows for text-to-music, audio-to-music, and style transfer or voice conversion.

  2. 2

    Udio works well as a songwriting partner for transforming existing ideas, while Suno tends to produce polished, radio-friendly songs from prompts.

  3. 3

    Generated songs become more controllable when you provide structured lyrics, specific musical details, and carefully chosen labels from Rate Your Music.

Summary

Phlo Young gives a practical introduction to AI music generation through Suno and Udio. He separates the field into text-to-music, audio-to-music, and style transfer or voice conversion, using examples such as AI versions of Kanye West and Drake, Randy Travis's post-voice-loss recording, and viral songs generated from text prompts. He then shows how to create songs from descriptions, custom lyrics, and instrumental settings. In his comparison, Udio is a songwriting partner that can extend, remix, inpaint, and transform uploaded audio, while Suno produces polished songs with a strong top-40 feel. Phlo also covers prompt construction, song structure, voice blending, copyright concerns, stem extraction, and the limits of his own legal knowledge. His workshop is aimed at lowering the barrier to making music, while remaining candid about moderation filters, lawsuits, and the unsettled ethics of using real artists' voices.

Key ideas
09:46

AI music describes several different generation workflows

Phlo separates AI music into text-to-music, audio-to-music, and style transfer. Text-to-music generates short samples or full songs from text prompts. Audio-to-music takes an existing sound and turns it into music. Style transfer, which he also calls voice conversion, changes a recorded voice into another voice. He says public discussion often calls voice conversion AI music without distinguishing it from a song generated from scratch. This distinction matters because the underlying process, creative input, and legal questions differ. His examples move from converted vocals over an existing beat to songs produced from a written prompt with no prior recording.

11:18

Voice conversion can change a performance while leaving the songwriting process intact

Phlo uses an AI Kanye West example to explain that the creator wrote lyrics, found a beat, recorded reference vocals, and then converted those vocals into a trained voice. He makes the same point about the viral Drake and The Weeknd track associated with Ghostwriter. The track was made by a person who handled the song apart from the vocal conversion. Phlo says the song reached about 9 million streams in its first 24 hours before it was removed after a complaint from the RIAA. He contrasts older conversions, which sounded artificial, with newer models that he says can sound almost exactly like the target voice.

18:01

An authorised voice conversion helped Randy Travis release new music after losing his voice

Phlo presents Randy Travis as an example of voice conversion used with the original artist and team. Travis's team wrote and produced a song, then used the technique so the song could sound like Travis without him singing the vocal in the usual way. Phlo says public reaction to the release was overwhelmingly positive, including comments from older listeners who were moved to hear Travis again. He is careful to frame this differently from anonymous creators imitating famous performers. For him, the example shows how the same technical method can raise different questions depending on who controls the voice and how the recording is made.

20:12

Text-to-music removes the need for a prior recording

Phlo distinguishes text-to-music from voice conversion by showing songs created from words alone. He discusses the viral BBL Drizzy song, which he attributes to comedian King Willonius using Udio. The creator entered a prompt and possibly lyrics, rather than recording vocals and converting them. Phlo calls this a significant example because a song generated in one pass became widely shared and was reused in other music. He also plays an ElevenLabs-generated song about GPUs. These examples illustrate the meme-song category, where the concept, lyrics, and musical result can come from a short written instruction instead of a conventional studio session.

25:08

Audio-to-music lets musicians turn rough sounds into arrangements

Phlo shows audio-to-music examples made with Stable Audio and Udio. In one case, an existing sound is transformed into music. In another, a person records only fingers drumming on a desk, and the system generates a fuller arrangement with guitar and other instruments. He thinks this workflow will appeal to musicians who already have musical ideas, because it preserves a starting performance while expanding it into something more developed. He encourages artists to upload their own recorded music, spoken word, or songs to Udio and compare the result with their original. In his words, Udio sometimes does him better than he does himself.

35:45

Udio and Suno encourage different working styles

Phlo describes Udio as a songwriting partner. Its default experience generates about 30 seconds at a time, which users can extend, remix, discard, or modify with inpainting. Udio also supports audio-to-audio uploads. He describes Suno as an in-house music producer that often produces clean, polished, top-40-style results. He says Suno's handling of timing and song structure had improved recently, especially when users provide custom lyrics. Earlier, vague prompts and lyrics with poorly matched syllable counts could make the beat and vocal timing go badly off. The tools are easy to start using, but getting consistent results requires more detailed prompting.

44:33

Structured lyrics give the music models better material to work with

Phlo recommends using a ChatGPT template to turn a song description into coherent lyrics before sending the result to a music model. He says this gives the user a chance to edit the lyrics instead of asking the music service to invent everything from a general topic. The music prompt can then describe the type of song, while the custom-lyrics field supplies the words and structure. He also points attendees to official Udio and Suno tips. His approach is practical: separate the lyric-writing step from the music-generation step, then adjust the text before generating another version.

52:05

Generated tracks can be opened in a digital audio workstation and split into stems

Phlo explains that adding 'daw' to the relevant song URL opens the track in WaveTool, a digital audio workstation. The service separates the generated song into stems, including vocals and individual instruments, and lets the user change their volume. This gives creators more control than the original rendered file. He also names UVR5, an open-source project that can run locally and extract stems from AI-generated songs or ordinary recordings. WaveTool accounts were free at the time of the workshop, while UVR5 required suitable computer hardware for local use. The workflow lets someone keep the arrangement while replacing or reducing a specific part, such as the vocal.

"I really have one main goal in mind and that's for this talk to be interactive."00:00
Who should watch
  • You want to try Suno or Udio but need a clear first workflow for prompts, lyrics, and song generation.
  • You are a musician with recordings, demos, or spoken-word material and want to test audio-to-audio transformation.
  • You need to understand the practical limits around voice cloning, copyright complaints, moderation filters, and stem extraction before using AI music publicly.