AI audio is entering the 2.0 era of models
Against the backdrop of the meteoric progress of large language models, generative AI systems dedicated to music sometimes seem to be treading water. ChatGPT, Gemini and Claude reason, programme and engage in dialogue with impressive fluidity, whilst music generators still struggle to offer the same level of creative control. This observation is true, but there is an explanation: those in the music industry are not… any less advanced; they are simply facing an infinitely more complex challenge. Unlike text, music is not simply a sequence of symbols. A song simultaneously combines melody, harmony, rhythm, timbres, vocal performance, mixing and emotional progression. The model must maintain this coherence for several minutes. Whereas an article consists of a few thousand tokens, a four-minute track corresponds to millions of audio samples, even when…