3 ms·
I think part of the problem is that the 'simple plan' at the foundation of the attempt was this: 1. Have a background server process generate 8-second three-pa
by codeulike 6y ago
I think part of the problem is that the 'simple plan' at the foundation of the attempt was this:
1. Have a background server process generate 8-second three-part-tune clips
2. Use some basic heuristics to guess at the key signature (“C# major”) of the clips, evaluate their intensity (lots of drum hits? loads of notes?) and save them in a huge clip database.
3. Create a search interface for the clip database.
4. Have the client request clips from the server in a specific key and intensity.
5. Weave two or three clips together, repeating them a couple of times, to make a full “song”.
I think thats just a bad plan, and if you start from that, any embellishments or improvements you try to make later just wont get anywhere.
Why is it a bad plan?
Working out what key signature a short phrase of music is in is not that simple.
e.g. you can look at the notes and say 'well that fits C major' but actually C major and A minor contain the same notes, and the difference between them is subtle, its to do with how often the root note is used, or even more abstract stuff like how often the root note is implied. There is kindof a probability distribution of notes for each key signature, and that distribution is how we recognise them. Sometimes there's ambiguity, and thats part of the art.
If the generated clip only uses (say) 6 semitones, there might be a multitude of differnt key signatures or scales that it could _potentially_ fit into to (Pentatonic, phrygian, Major, Minor etc etc with various different root notes), but the most accurate one musically would be very hard to determine out of context, because in reality the context around the clip (say the 20 bars before and after) play a big part in how a 'clip' would be experienced.
And if you're going to generate random clips to begin with (step 1), I'm hoping those werent random clips where any of the 8 tones or 12 semitones in play have an equal probability of being used, because thats just going to sound muddy.
I guess he was going for a system that had very few rules and then tried to get the system to learn from there, but I think thats just too muddy a starting point to ever get anywhere.
And obv purely generative music is hard, but what I'm talking about here I think explains why the stuff he generated was classified by him as 'awful' rather than just 'weak' or a bit boring.
TLDR: I think a misunderstanding of what a key-signature really is might be to blame. Its not just 'which notes', its a whole probability curve and its very contextual.