11 ms·
Artificial Intelligence Generates Christmas Song from Holiday Image
- pmyjavec 10y agoChilling...
- d33 10y agoIt kind of feels like a five-year-old trying to make a song. Which seems good - now they only need to improve the mechanisms and who know, maybe we'll get to the point when it's a ten-year-old?
- logicallee 10y agoInstead of five, don't you mean "two to three"? And under that comparison, isn't it scary how much it does sound that way? Like a kid who hears words together but don't know what they mean yet? (and doesn't really get the cultural things around it.) To me it sounds like a quite musical 2-3 year old stringing words together. Doesn't it strike anyone else that way? These things are going to grow up very, very soon! We know that. It's scary. You're watching a two-year old. It's interesting to think about what it will be like when it's ten, sure. But what will really blow your mind is what it'll be like when it's 23. This is happening right now, before our eyes. I cannot overemphasize that any large server farm at Google or Amazon is doing more, and much faster, processing than a human brain's neural net. Human brain is 86 billion neurons with an average of 7,000 synaptic connections each. That is a huge number. But they are firing at 15-300 herz (because they're at biological speed, instead of lightspeed like our CPU's) - which is 7 orders of magnitude slower than our silocon. Our brain is about 3 pounds (1,300-1,400 g) and uses some 20 watts. It's not a question of "if" a server farm will have as powerful neural nets. It's a question of "when". (Also although we won't be using it, the entire source code for the human brain has to be strictly less than 700 MB, because the fully sequenced human genome which obviously encodes the full human mind is less than 700 MB uncompressed.) Guys, we are at an incredible pivot point in human history. We are coming up with computerized brains with in some ways comparable architecture to humans, and they are doing human activities. Today, in 2016, there are thousands, perhaps tens of thousands, of server rooms all over the world that have more than enough computational power to do in real-time what a human adult brain does in real time - but we lack the software. when we see advances like this in artificial intelligence, this is scary. We're all but looking at the intellectual output of a two year-old in the field of music. every single day AI results are astounding. this is it.
- ythn 10y ago> These things are going to grow up very, very soon! We know that. It's scary. I'll believe it when I see it. Your post has just a bit too much futurologist science fantasy wishful thinking in it.
- monk_e_boy 10y agoYou sound like some of the kids I teach. All 17 - 18 year old kids have phones and check them at least every 5 minutes. I mentioned that it'll be weird when they are wearing some sort of augmented reality glasses and teachers won't be able to tell if students are concentrating or reading reddit. The kids all said "That'll never happen." as if technology is stuck where we are today. They were astounded when I told them that ten years ago, I never saw a phone in a classroom. So from my perspective we went from no phones, to big phones, to little dumb phones, to little connected computers... it's no stretch to think that soon they will be glasses mounted or project some sort of holograph into the eyeball. Version 2 of Microsoft HoloLens will be cool, v3 will be tiny, v4 will be mounted inside glasses for sure. The same with AI. A few years ago I couldn't talk to my computer. Today I do it all the time. Last week I had to remember passwords, today my computer recognises me and logs me in. AI is here and it's getting better and from my point of view it is getting better much faster.
- ythn 10y agoIt's one thing to say "VR will be commonplace in 10 years" or "Self driving cars will be commonplace in 10 years". That, I can believe because we already have prototypes and I've messed with them and could imagine advancing them. It is quite another thing to claim that human adult level strong AI is coming "very very soon" and could happen overnight. The latter is just science fiction wishful thinking. I have no reason to believe we will ever have truly thinking, sentient computers, let alone "very very soon." Sure, our Siris and Alexas and whatnot will get better and better at responding to our queries how we want, but that's way different from an adult level human intelligence AI. Machine Learning has limits and will not yield conscious machines anytime soon, if ever.
- dweekly 10y agoMusic for people who hate music.
- _ix 10y agoDid I just hear a new classic?
- BoringCode 10y agoNo.
- divanvisagie 10y agoOne step closer to GLaDOS.
- demolish 10y agothis was a triumph! im making a note here: huge success
- divanvisagie 10y agoIt's hard to overstate my satisfaction.
- givinguflac 10y agoThis is definitely now my favorite Christmas song. While obviously not a masterpiece, it's incredible how far this tech has come. It's almost got a Dadaist feel to it. Can't wait to see where this ends up in ten years! I can foresee music labels buying a few of these AI's, getting some pretty people with decent voices and sending them on tour.
- loganbertram 10y agoEarlier this year, an AI-written script was made into a short film. It's got the same sort of absurd vibe. http://arstechnica.com/the-multiverse/2016/06/an-ai-wrote-this-movie-and-its-strangely-moving/ http://arstechnica.com/the-multiverse/2016/06/an-ai-wrote-th...
- midgetjones 10y agoIs pop music not cheap and disposable enough for you already?
- TheOtherHobbes 10y agoPop is still recognisably human. This sounds like logic gates roasting by an open fire.
- midgetjones 10y agoThat was a brilliantly seasonal analogy.
- duke_z 10y agoit is scary, sounds like GLAdOS singing "still alive"!
- bitwize 10y agoIt sounds more like a corrupt core. "Cave here. It's Christmas time, and you know what that means: Christmas bonuses have been suspended until further notice. We've gotta pay the judgement on that pesky class-action with something. But don't let that get you out of the Christmas spirit. The lab boys have come up with a way to stay festive by hooking up the Christmas Core to the lab's PA system. So enjoy free, continuous, computer-generated Christmas music from now until January 5!" "Cave again. Apparently the Christmas music has been causing some employees severe emotional and psychological distress. We've had reports of people sticking their heads into active particle accelerators and drinking Repulsion Gel to get away from the sound. So until a full investigation has been conducted and the Christmas Core thoroughly debugged, we are discontinuing the Christmas music. We do not need another class-action on our hands, folks."
- mojuba 10y agoIs it just me or it really sounds like a randomly generated chord progression that barely makes any musical sense?
- nkozyra 10y ago"Musical sense" is a pretty subjective thing, which I suspect is one of the primary issues with artificial creativity. That said, assuming training data comes from music with progressions that could be broadly classified as 'popular music,' you would expect to find some regression to the mean with more production and deeper training data. (to abuse a phrase) One other issue that I think will come up is how insular and unevolving artificial creativity will be if it's based on present music for training data. What has historically moved creative trends is disruption; sometimes it's a slow burn and sometimes it's a few catalysts, but experimentation in artificial creativity will be hard to come by early on but quickly needed if it's to supplant human creativity.
- TheOtherHobbes 10y agoThe statistical approach is painfully naive and doesn't work - as is obvious from the example. It's like feeding a net with the complete works of Shakespeare and expecting it to produce a genius-level original play. It's simply not going to happen.
- nkozyra 10y agoThe issue is not with the statistical but with the parameters around the output and the organization of training data.
- zeveb 10y agoI think that your assumption is that the genius of Shakespeare's plays can be statistically reproduced through sufficiently-clever organisation. That is not obviously true to me. Some things are just art, capable of being truly understood only by a creature with a head and heart, arms & legs, love & hate, emotions, experiences — in short, a man.
- midgetjones 10y agoI'm not sure that Christmas songs generally use the blues scale, I wonder what made them choose that for the melody?
- dasboth 10y agoI suspect it's an easy way to get something that won't sound completely dissonant, especially because you can use the same blues scale over multiple chords.
- midgetjones 10y agoI can see the logic behind that, but without any sort of tension/release between the melody and chords, it still sounds just as dissonant to my ears.
- dasboth 10y agoGranted, it will never win any awards. Part of the choice behind it might have been the desire to "ship" it before Christmas.
- miguelrochefort 10y agoShow me the same algorithm generate songs in a different genre from different images and I'll be impressed.
- hahaker 10y agoI'm from the project team. This is a very interesting point. While it is easy to crawl many songs from the internet, it is a little harder to gather the same amount but with proper genre/style/etc labels, although it is not impossible. For now there's only one genre, which we call it "the genre of whatever is on the internet". So whatever music files on there, many of them quite "crappy", were used to train the model. Also there are many other problems on how to better structure and flavor the composition. This is just a very early-stage attempt, as a CS student's fun side project. We are working with people with real musical talent now and hoping to make better songs in the next version.
- miguelrochefort 10y agoI mean, where does the Christmas element comes from? The image alone, the music it was trained with, or is it somehow hardcoded in the algorithm?
- hahaker 10y agoThe Christmas element comes from 1. the image, and 2. a 4800-dimensional RNN sentence encoding bias generated from ~30 Christmas songs. Not sure how to hardcode this.
- midgetjones 10y agoI'm interested that you used some Christmas songs as training (which wasn't obvious from what I read of the paper). Were they pop songs, traditional, or a mix? Further to my comment up there[0] - and I don't wish to sound a grinch because this is a really cool project - but would I be right in thinking you spent more time on the image description than the music? I saw that you specify a scale for the melody, would it be either possible to use a mode to generate the accompaniment around, so that the melody can move diatonically and risk too many clashes, or to allow the melody to follow the chord sequence somehow? Again, sorry if I sound too critical. It's a really awesome thing you've done, and I'm just a guy that listens to the music instead of the lyrics. [0] https://news.ycombinator.com/item?id=13079355 https://news.ycombinator.com/item?id=13079355
- fredleblanc 10y agoThe melody is all over the place, and the rhythm is hard to tap a foot to, but one thing is certain: it's absolutely convinced that Christmas trees get decorated with flowers. Lots and lots and lots of them.
- afandian 10y agoI found the naïvité and belief that flowers are put on Christmas trees truly touching. Such a pure reaction to something that can be so cynical. For that I'll forgive the tone-deaf singing.
- klenwell 10y agoI could have mistaken it for a track off The Shaggs'[0] lost Christmas album. I like to believe Frank Zappa would have liked this, too. [0] https://en.wikipedia.org/wiki/The_Shaggs https://en.wikipedia.org/wiki/The_Shaggs
- fredleblanc 10y agoOh man, The Shaggs are something else. And it just so happens that I live two towns from Fremont, NH, where they're from.
- partycoder 10y agoSounds like "Friday" by Rebecca Black.
- eva1984 10y agoThis is not even a proper song.
- binarnosp 10y agoWhen I think that this month I'm going to hear "Last Christmas" by Wham! in every store I set foot into, this new masterpiece doesn't sound so bad.
- brudgers 10y agoIt brings this recent story to mind: https://news.ycombinator.com/item?id=13033299 https://news.ycombinator.com/item?id=13033299
- BinaryBullet 10y agoOne of the cofounders of the Echonest (acquired by Spotify) created this back in 2004: "A Singular Christmas" was composed and rendered in 2004. It is the automatic statistical distillation of hundreds of Christmas songs; the 16-song answer to the question asked of a bank of computers: "What is Christmas Music, really?" https://soundcloud.com/bwhitman/sets/a-singular-christmas https://soundcloud.com/bwhitman/sets/a-singular-christmas
- hahaker 10y agoVery interesting reference! Deep learning is statistical, so this is sorta one of the spiritual predecessors.
- x2398dh1 10y agoBy my understanding, yes, in the sense that certain parts of each algorithm are looking to minimize something...the former model used principal components analysis, which is linear in the sense that you are using transforms which pick out for the least correlated pieces of a huge chunk of data, whereas neural networks, which uses a combination of linear and non-linear layers picked by the user to minimize, "errors." What's interesting is that the former model sounds so much, "better." I wonder if anyone could chime in about how our ears and auditory nerves or perhaps auditory cognition works, and whether they are more, "principal component analysis-y" somehow than "error minimization-y" or something relating to the actual math, which may explain why this new neural network christmas song sounds like absolute crap to us, whereas the older version sounds pretty amazing. Also, whether my understanding of the underlying math is correct or not.
- diydsp 10y agoAuto-synthesis of music has been a topic of academic interest since the 1950s, when the first mainframe scribbled out code on paper tape to be translated into sheet music and performed. UToronto's work here is the latest expression of this desire. The huge gap between our cultures' actual music and these synthetic projects can to an extent be described through "receptivity" or the phenomenology of music, in other words, how it's experienced. The following fun, short talk does a great job of introducing the concept of through its analysis of "vaporwave." https://www.youtube.com/watch?v=QdVEez20X_s https://www.youtube.com/watch?v=QdVEez20X_s
- vasaulys 10y agoThat video is fantastic I watched it a week or so ago. Your explanation also explains why computer performed music is so off. It still has that uncanny valley effect. So when Sony had a computer generate a "Beatles-esque pop song", they still had a human perform and produce it. But at the point there's so much creativity and human-added value on top of it that I don't think its fair to call it computer generated imho.
- diydsp 10y agoyes. I can tell you a little more about that, too, since I used to research this stuff and think about it a lot still. One of my models of music is an external model of a regulated system that parallels and trains our own habits and responses. E.g. a song demonstrates tension and release similar to our own lives. The level of tension in a song before release occurs can inform us how much tension which should accept before performing some release activity. Music's rhythms also inform the pace of our work. E.g. verse-chorus-verse represents switching between two different activities. Even the pitch of a single note acts as a reference for the amount of intensity of a sensation we should use in our own lives. E.g. thrash metal listeners enjoy sudden shifts into massive intensity and hold it there. Dub step listeners are training themselves for unusual, but rather intense aesthetics leading up to disproportionate release. Classical music tends to be for "long-chain thinkers" tumbling ideas over from various perspectives, e.g. writers and politicans, doctors, not factory workers. With that as a background, consider that a live instrument is also a physical system with a human controlling it interactively. The live system is a bit different every time. Here's the critical part: the human must listen and provide instantaneous feedback to a varying system in order to present the piece of music as a proper response model of a regulated system. If the player fails to do this, the model communicated by the performance is different. In open-loop systems, such as a sequencer, there is no (or limited) interaction between the player and the sound, so an incidental model emerges. That incidental model represents an unintended and therefore most likely irrelevant model of how to interact with reality. e.g. it relieves tension where no relief was needed. It lingers too long on an idea, long after a human novelty-seeking circuit has starved. Some people, e.g. in discussions of unstable filters like the TB-303, chalk up the variations as being different at every performance because the instrument is random... However, they're missing the closed loop portion of the performance, in which the performer reacts to the unpredictability of the instrument in order to maintain the model. In other words, the score and notes are not the music, but the performer's response to the environment the score sets up is the music. To revivify your uncanny valley observation, the "unstable filter creates variations" crowd has a parallel in Perlin noise used to subtly animate human models to make them not look so dead. However, it's incomplete because they don't use (short-term) feedback to determine when the movement suffices to be convincing. That feedback is the essence of performance. In theory, computer scientists could implement these feedback models in performance to make the sounds more realistic. They could be used in synthesis, but the playback would still require observation of the listener! Which is possible. Personally, I just prefer playing electronic instruments live over using sequencers. It's only the sounds of electronic music I like, the zaps, peowms, zizzes, pews, and poonshes, etc. I don't care for electronics/computers to perform for me. If you like this hypothesis, you can find more references on my wiki at: http://www.diydsp.com/index.php?title=Computer_Music_Isolation http://www.diydsp.com/index.php?title=Computer_Music_Isolati...
- andrewclunn 10y agoBetter than The Christmas Shoes.
- keypulsations 10y agoI've seen more and more little A.I.-generated ditties like this recently and their reception tends to be the same: that they're interesting and funny but don't sound that great. The output would probably be more compelling if A.I. were adopted more as an instrument by individual artists/composers to automate some of their more tedious tasks by learning their own particular styles rather than a magical music box that churns out top hits.
- kingkawn 10y ago"The best Christmas present in the world is a blessing." This algorithm is throwing down some wisdom
- Florin_Andrei 10y agoI think generating fortune cookies is a really low hanging fruit for current AI. Someone could put it together in a week-end.
- hahaker 10y agoI am the main developer. You are literally right on the time spent... Our original focus was writing a research paper on hierarchical music generation. Composing a song from an image is just one of the "fun applications" that we spent a little time on, to promote the interestingness of our method. I started Saturday afternoon, and was basically done by Sunday night.
- Florin_Andrei 10y agoWell, your project is quite a bit more complex than just making fortune cookie messages. Still amazing how much can be accomplished quickly with current AI tech. Thumbs up for the cool project.
- kmill 10y agoI once developed a hierarchical Markov chain, and I decided to use it to generate fortune cookie wisdom because I thought people would be more willing to overlook grammatical mistakes or be willing to interpret it as an expression of deep truth. http://www.kylem.net/stuff/fortunes.html http://www.kylem.net/stuff/fortunes.html (You might need to increase the "sense" parameter.)
- DrPhish 10y ago"The supreme happiness in life is simply to serve as a warning to others." Oh my god, I'm in stitches. What a fun project, thanks for sharing!
- antisthenes 10y agoThis was a triumph! I'm making a note here: Huge success!
- blauditore 10y agoWhile this is hilarious, it doesn't seem like a huge achievement to me. The only thing (kind of) working well is feature/topic detection in the image (tree, christmas etc.), but that isn't really cutting edge. The core part, learning and creating music, only produced melody and lyrics that seem not much different from accumulating random sentence and chordal fragments.
- ravenstine 10y agoI'm not sure what I should be impressed by. Maybe there's some real technical feat happening here, but I feel like a basic mad-libs style algorithm could produce something better.
- hahaker 10y agoI'm not very familiar with mad-libs so correct me if I was wrong. I think generating a lyrics passage (zero hard-coded rule on content or grammar or anything) from an image would not be something you can do with mad-libs.
- ChuckMcM 10y agoI am pretty amazed at the effort that nVidia is putting into its corporate rebranding effort. I wonder if, in the not too distant future, they will be the AI company that also makes Graphics cards sometimes. The other thing I find really amazing about it, coming from IBM, is that IBM has invested a ton of money in IBM Watson but they sold off their foundry business (could have made massively parallel AI machines) and their systems business is a fraction of what it was. Looking at what can be done when you're leading versus when you are following is really sobering to me.
- debt 10y agoWe have a very long way to go it seems. That was almost nonsensical.