5 ms·
Show HN: Karaoke for any song in any language
- yunusabd 7y agoPersonally I love karaoke, but looking at the repo and the website gave me no information whatsoever about this project. Maybe that's something you can work on? In the meantime I found this article, which reads quite positive: https://www.theverge.com/tldr/2020/2/19/21144452/youtube-youka-club-karaoke-lyrics https://www.theverge.com/tldr/2020/2/19/21144452/youtube-you...
- ronyfadel 7y agoI wish the readme had a description of how Youka works. Looks promising, but I’m not sure it does what I think it does.
- youka 7y agoI'll add some explanation soon. Here's the main process: Search your query in YouTube using https://github.com/youkaclub/youka-youtube https://github.com/youkaclub/youka-youtube Search lyrics using https://github.com/youkaclub/youka-lyrics https://github.com/youkaclub/youka-lyrics Split the vocals from instruments using https://github.com/deezer/spleeter https://github.com/deezer/spleeter Align text to voice (the hardest part) using some private api
- yorwba 7y ago> Align text to voice (the hardest part) using some private api That's also the part that would be most interesting to have explained. Is it language-agnostic? After all, the title says "in any language", but I can't think of any text-audio alignment algorithms that don't require a language-specific model. (Unless you just count characters and assume they map linearly to time, which I'd expect to go very badly.)
- gliese1337 7y agoHaving worked for many years in a linguistics research lab where we spent a lot of money paying people to edit and align subtitles and audio transcripts, and having largely written what was at the time the most sophisticated subtitle-and-transcript editing tool available, I can confirm: counting characters and mapping them linearly to timespan, even after isolating vocals, does indeed go very poorly. And much worse when there's singing involved.
- youka 7y agoSo let’s play, if you can guess the align method I’ll open source it :)
- ampdepolymerase 7y agoSpeech recognition?
- rnotaro 7y agoI get an error when trying to open any video : Ooops, some error occurred :( Error: [Errno 2] No such file or directory: '/tmp/tmpphtr8ehu/accompaniment.aac' When running on the official Windows 10 SandBox (https://techcommunity.microsoft.com/t5/windows-kernel-internals/windows-sandbox/ba-p/301849 https://techcommunity.microsoft.com/t5/windows-kernel-intern...) Edit: it somehow works for some songs. The concept is really nice. I love it.
- youka 7y agoLooks like a server-side bug (can't really handle more that a single split process concurrently), I'll add queue in the next version.
- rnotaro 7y agoDemo: https://peertube.co.uk/videos/watch/3c183b56-deb6-4e6b-a7a2-87cd692df483 https://peertube.co.uk/videos/watch/3c183b56-deb6-4e6b-a7a2-... edit: Swapped youtube URL to Peertube for Content ID claims issues.
- fareesh 7y agoFrom what I understand, it is software for you to align lyrics to music contained in a video, with tools to enable you to do so.
- Reubend 7y agoHey there! First of all, I want to tell you that the app is fantastic. I used the earlier version of this, when it was a website, from your previous HN post. And once again the alignment works quite well in my experience, as does the isolation. In the future, it would be great to have a "portable" version of this for Windows that doesn't install anything. It's annoying to open up an app, and have it install itself without any warning or user consent. You could just release a .zip file with the build as an option.
- youka 7y agoI’ve considered few options to install ffmpeg, and choose that way. I’m open to other suggestions
- Reubend 7y agoYou can distribute a .zip file which includes the statically linked build of FFmepeg: https://ffmpeg.zeranoe.com/builds/ https://ffmpeg.zeranoe.com/builds/ . Then just call it locally. There's no need to install it system-wide.
- youka 7y agoI don’t install it system wide, just download a single binary into youka directory.
- laurieg 7y agoI installed this on Mac OS but the program always fails with: Uncaught Exception: Error: Could not get code signature for running application at m(/Applications/Youka.app/Contents/Resources/app/.webpack/main /index.js:1:12481) at App.<anonymous> (/Applications/Youka.app/Contents/Resources/app/.webpack/main/index.js:1:14365) at App.emit (events.js:215:7)
- youka 7y agoJust reopen it and you will be fine (I don’t have free 99$/year for apple code signature)
- peterburkimsher 7y agoIs there a way to manually provide the lyrics? I have a substantial collection of songs in Chinese and Taiwanese, and it would be really helpful to use this to help me make lyrics videos for Pingtype. When I tried, I got this error: Ooops, some error occurred :( Error: name 'espeakng_supported_langs' is not defined I'll look into aeneas to see if that can give the API-level technical tools that I need - thank you for explaining that part in the other comments!
- yorwba 7y agoNote that it won't work for Taiwanese (I assume Hokkien) unless you add the necessary support to espeak-ng. If your lyrics are in Peh-oe-ji, you'll need to define how the romanization maps to phonemes. You may be able to get some inspiration for that from the definitions for Mandarin and Cantonese. Though I just looked at the "phonology" section on Wikipedia https://en.wikipedia.org/wiki/Taiwanese_Hokkien#Phonology https://en.wikipedia.org/wiki/Taiwanese_Hokkien#Phonology and the tone sandhi rules look a lot more complex than any other Sinitic language I know. If the lyrics use Chinese characters, there's the added difficulty of collecting a pronunciation dictionary, which I'd probably do by scraping https://twblg.dict.edu.tw/holodict_new/index.html https://twblg.dict.edu.tw/holodict_new/index.html , http://xiaoxue.iis.sinica.edu.tw/ccr/ http://xiaoxue.iis.sinica.edu.tw/ccr/ and Wiktionary. (If you know any other sources for pronunciation data, I'm interested.)
- peterburkimsher 7y agoYes, I know about romanisation! I wrote Pingtype, and extracted romanisation dictionaries for Taiwanese Hokkien and Hakka by parsing Bible data. https://pingtype.github.io https://pingtype.github.io Tones are difficult, so I encode those as colours. Adding code to espeak-ng sounds very difficult. Most of the songs are in Mandarin though, so I'll try those first.
- redraw 7y agooh, I had the same idea and started working here https://github.com/redraw/karaoke-machine https://github.com/redraw/karaoke-machine days after Deezer's spleeter was released, but stopped while searching for a way to sync the lyrics. thx! I'll try it out
- youka 7y agogood luck! here's the relevant code https://github.com/youkaclub/youka-api/blob/master/youka/align.py https://github.com/youkaclub/youka-api/blob/master/youka/ali...