5 ms·
An IBM computer learned to sing in 1961
- rzzzt 3y agoThe crowdsourced version of the song is quite chilling: https://youtu.be/Gz4OTFeE5JY https://youtu.be/Gz4OTFeE5JY
- eddieroger 3y agoI don't know if paying for samples via Mechanical Turk is "crowdsourced," but you were 100% right on the being chilling. Is there an audio uncanny valley to go with the visual one we know so well? Like when something sounds close-to but not exactly like something else we know? That was so weird to listen to.
- rzzzt 3y agoI think it fits the definition if the invitation to work on a piece of the task is open for a significant chunk of the population. It doesn't have to be unpaid.
- mgdlbp 3y agoIf you listen to the individual tracks it just sounds like people "normally" making noises https://www.bicyclebuiltfortwothousand.com/ https://www.bicyclebuiltfortwothousand.com/ Maybe intent plays a part in the perceived musicality of the result, considering that even 4chan and other forums can make up more coherent virtual choirs, under equally poor recording conditions... https://www.youtube.com/watch?v=uK_SRSB9pdA&list=PLlsIiu-R8aAJLg_Flhf1-IXGaOyWOAC1x&index=159 https://www.youtube.com/watch?v=uK_SRSB9pdA&list=PLlsIiu-R8a... Edit: or perhaps the Mechanical Turk performance is more like a haka than the original song in effect - https://www.youtube.com/watch?v=BI851yJUQQw https://www.youtube.com/watch?v=BI851yJUQQw
- rzzzt 3y agoIt does feel like they wanted the choir to stay true to the speech synthesizer's sounds, not the song itself. So (rendition of (rendition of Bicycle Built for Two)).
- xaedes 3y agoEery, this song performance reminded me of the ending song of Portal.
- zabzonk 3y agothe thing i most like in the film 2010 is that HAL gets put back together and saves Helen Mirren's human crew.
- silveira 3y agoReminds me of Kapp'n in Animal Crossing.
- vidanay 3y agoDaisy, Daisy, Daisy.
- tomcam 3y agoIt didn’t “learn“ how to sing, of course. The voice was built bottom up, phoneme by phoneme, pitch by pitch.
- deleted 3y ago[deleted]
- SCAQTony 3y agoHAL did it better 40-years later in 2001: https://www.google.com/search?q=HAL+singing+D+a+bicycle+built+for+two+2001&oq=HAL+singing+D+a+bicycle+built+for+two+2001&aqs=chrome..69i57j33i160l2.13467j1j4&sourceid=chrome&ie=UTF-8#fpstate=ive&vld=cid:66e079f5,vid:E7WQ1tdxSqI https://www.google.com/search?q=HAL+singing+D+a+bicycle+buil...
- gwbas1c 3y agoThe clip is embedded in the bottom of the TFA.
- gwbas1c 3y agoOne of the things I get a kick out of in the clip of HAL singing Daisy is just how much the physical modules being removed look like hard drives being pulled out of a NAS. I could easily think of a much larger version of my Synchrony NAS looking just like that. BUT: When I first watched 2001 sometime in the early 1990s, I had only seen a 5.25" hard drive, and not on a slider / rails. I thought the inside of HAL was just tacky scifi from the 1960s. It's only later as I've seen the predictions come true that I've realized just how forward-looking 2001 is. Like the scene with watching news reports on the tablets at breakfast. It wasn't until I watched a video, on my phone, in the late 2010s, that I realized that prediction in the move was 100% spot-on. BTW, the 4k Ultra-HD bluray of 2001 is awesome.
- rthomas6 3y agoThey even nailed the headrest screens on the space "plane"! The lunar landing scenes are also mind blowing when you realize the film was released before the actual moon landing!
- chris-orgmenta 3y agoAnd more generally, a good predictor of how the free market would eventually take over the industry. All the branding, the food, etc. really does come across as well considered.
- sibeliuss 3y agoAnd the quiet, zero-point style propulsion on the carrier ship. No rocket boosters! That always sticks with me.
- datavirtue 3y agoWe built tablets based on the influence from 2001. You have it backwards.
- blincoln 3y agoAre there any detailed descriptions available of how the music and voice were synthesized? Based on the recording, the information I could find, and imagining how I'd try to do the same thing using the technology of the era, I assume the melody is based on single-cycle samples of a piano and Max Matthews playing the violin. The vocals sound like formant synthesis like the Votrax SC-01 or TI LPC series, although of course those chips didn't exist until 15+ years after the work at IBM. But I'm very curious about the details. Did the team develop a general-purpose sequencer for the melody and/or speech, or were all of the notes, slides, etc. hardcoded? Did the computer actually output all 3+ parts together, or were they separate elements mixed after the fact? I assume the output was not realtime, but it would be a neat surprise if they achieved that in the 60s. Was it all handled digitally in the computer, or was the computer controlling some add-on hardware, maybe with analogue filters? Etc.
- zebproj 3y agoThe singing synthesizer used a surprisingly sophisticated physical model of the human voice [1]. The music was mostly likely created using some variant of MUSIC-N [2], the first computer music language. The syntax and design of Csound[3] was based off of MUSIC-N, and I believe the older Csound opcodes are either ported or based off those found. Apparently the sources for MUSIC-V (the last major iteration of the MUSIC language) can be found on github [4], though I haven't tried to run it yet. 1: https://ccrma.stanford.edu/~jos/pasp/Singing_Kelly_Lochbaum_Vocal_Tract.html https://ccrma.stanford.edu/~jos/pasp/Singing_Kelly_Lochbaum_... 2: https://en.wikipedia.org/wiki/MUSIC-N https://en.wikipedia.org/wiki/MUSIC-N 3: https://en.wikipedia.org/wiki/Csound https://en.wikipedia.org/wiki/Csound 4: https://github.com/vlazzarini/MUSICV https://github.com/vlazzarini/MUSICV
- blincoln 3y agoThank you! Seems like that project was incredibly far ahead of its time. The physical-modelling aspect is super interesting. Does that mean that the similarity in sound to formant-based speech synthesis is because they're both using a sawtooth wave, noise, or other relatively simple sound as the raw input? I always imagined that a physical-modelling speech synthesizer fed by a sawtooth wave would sound more like a vocoder than Votrax or TI LPC output does, but I guess not.
- cf100clunk 3y agoA recording of it was included in some electronics hobbyist magazines of that time on shiny black flexible vinyl for playing on a phonograph at 45rpm. I seem to recall that Bell Labs was credited on the label but IBM was not. EDIT: dbarlett just posted an image of the recording's label elsewhere in this thread
- bregma 3y agoWe had a floppy vinyl 45 of that when I was a kid in the 1960s (my mother was a high school science teacher and we often had cool stuff like that around the house).
- sibeliuss 3y agoDoes anyone have any information / background on the programming behind this project? It seems incredible for so long ago and I can't quite conceive of how they were able to do it.
- zebproj 3y agoSee my other comments here for more info about the underlying technology. It is pretty incredible that sophisticated digital physical models of the human vocal tract were being done in the early 60s. This was able to be done largely due to the deep pockets of Bell Labs. A lot of R+D was put into the voice and voice transmission.
- cf100clunk 3y agoLibrary of Congress essay by Cary O'Dell on "Daisy Bell (Bicycle Built For Two)" from song origins to Bell Labs recording: https://www.loc.gov/static/programs/national-recording-preservation-board/documents/DaisyBell.pdf https://www.loc.gov/static/programs/national-recording-prese...
- dbarlett 3y agoI inherited a Southern Bell promotional card/record that includes Daisy Bell: https://dbarlett-wordpress.s3.us-east-1.amazonaws.com/wp-content/uploads/2023/05/Computer_Speaks_1-1024x762.jpg https://dbarlett-wordpress.s3.us-east-1.amazonaws.com/wp-con... https://dbarlett-wordpress.s3.us-east-1.amazonaws.com/wp-content/uploads/2023/05/Computer_Speaks_2-1024x762.jpg https://dbarlett-wordpress.s3.us-east-1.amazonaws.com/wp-con...
- zebproj 3y agoThe neat thing about this particular singing synthesizer is that it used a surprisingly sophisticated (especially for the 60s) physical model of the human vocal tract [1], and was perhaps the first use of physical modeling sound synthesis. Vowel shapes were obtained through physical measurements of an actual vocal tract via x-rays. In this case, they were Russian vowels, but were close enough for English. While this particular kind of speech synthesis[2] isn't really used anymore, it's still fun to play around with. Pink Trombone [3] is a good example of a fun toy that uses a waveguide physical model, similar to the Kelly-Lochbaum model above. I've adapted some of the DSP in Pink Trombone a few times[4][5][6], and used it in some music[7] and projects[8]of mine. For more in-depth information about specifically doing singing synthesis (as opposed to general speech synthesis) using waveguide physical models, Perry Cook's Dissertation [9] is still considered to be a seminal work. In the early 2000s, there were a handful of follow-ups to physically-based singing synthesis being done at CCRMA. Hui-Ling Lu's dissertation [10] on glottal source modelling for singing purposes comes to mind. 1: https://ccrma.stanford.edu/~jos/pasp/Singing_Kelly_Lochbaum_Vocal_Tract.html https://ccrma.stanford.edu/~jos/pasp/Singing_Kelly_Lochbaum_... 2: https://en.wikipedia.org/wiki/Articulatory_synthesis https://en.wikipedia.org/wiki/Articulatory_synthesis 3: https://dood.al/pinktrombone/ https://dood.al/pinktrombone/ 4: https://pbat.ch/proj/voc/ https://pbat.ch/proj/voc/ 5: https://pbat.ch/sndkit/tract/ https://pbat.ch/sndkit/tract/ 6: https://pbat.ch/sndkit/glottis/ https://pbat.ch/sndkit/glottis/ 7: https://soundcloud.com/patchlore/sets/looptober-2021 https://soundcloud.com/patchlore/sets/looptober-2021 8: https://pbat.ch/wiki/vocshape/ https://pbat.ch/wiki/vocshape/ 9: https://www.cs.princeton.edu/~prc/SingingSynth.html https://www.cs.princeton.edu/~prc/SingingSynth.html 10: https://web.archive.org/web/20080725195347/http://ccrma-www.stanford.edu/~vickylu/thesis/index.html https://web.archive.org/web/20080725195347/http://ccrma-www....
- vidarh 3y agoI've been fascinated by the simplicity of this since I ran into SAM (Software Automatic Mouth) on the C64, but never really taken the time to delve into it. Your links are an amazing resource...
- thrtythreeforty 3y agoAnother excellent, but quite dense, resource I've found helpful for implementing my own waveguide models is Physical Audio Signal Processing, a book available as a hard copy and online [1]. There are also an absolute ton of research papers on these topics which have failed to be summarized anywhere or cited outside the small circle of researchers, so there's a ton of institutional knowledge about physical modeling locked up in academic papers that isn't super accessible. 1: https://ccrma.stanford.edu/~jos/pasp/ https://ccrma.stanford.edu/~jos/pasp/
- krunck 3y agoNice. But why did the video creator feel the need to put in the fake film projector effects? The urge people have to add "oldness" where it is already present - though not in the form they imagine - is interesting by itself.
- jlarocco 3y agoA few years ago I picked up "Music By Computer"[1] from a used book store, and it's fascinating. Published in 1969, it's a collection of papers from the 60s about music and sound processing on the machines back then, and it goes into a lot more detail, if anybody is interested and can find a copy. It even came with recorded music on 5 paper thin flexi-discs that I've never been able to play. https://www.amazon.com/Music-Computers-Heinz-von-Foerster/dp/0471910309 https://www.amazon.com/Music-Computers-Heinz-von-Foerster/dp... https://en.wikipedia.org/wiki/Flexi_disc https://en.wikipedia.org/wiki/Flexi_disc
- jagged-chisel 3y ago“Learned” or “was instructed”?