3 ms·
I work on music models, and this is a very cool paper! There are no papers that go into depth on how token-based AR music models (that aren't absurdly inefficie
by zaptrem 2y ago
I work on music models, and this is a very cool paper! There are no papers that go into depth on how token-based AR music models (that aren't absurdly inefficient like Yue) are trained. I'm particularly interested in your semantic tokens. I tried reproducing the CTC loss part but my curve was very spikey and didn't seem to actually figure out any characters. The semantic tokens gave great acoustic info but gibberish lyrics. What did your CTC loss curves look like and did you see anything similar at any point?
As a semi-aside, I feel like semantic tokens in general may end up being a bottleneck on how interesting model outputs can be.
- minihat 1y agoHow is it possible that text-to-score/notation is lagging text-to-audio in music generation? Generating audio seems wildly more complicated! Since you are working in this space, I wonder if you could comment on my pet theories for why this is true: 1. Not enough training data (scores not available for most songs), or 2. Difficulty with tokenization of musical notation vs. audio
- balivandi 1y agoCan you recommend any music models that can expand a composition I wrote into a full length symphony?