25 ms·
Backing up Spotify
- lelouch9099 10mo agoHow legal is this with regards to copyright laws?
- phainopepla2 10mo agoNot legal
- basisword 10mo agoIt's not. It's awful people justifying awful behaviour. And it's why we can't have nice things. There are always assholes ready to exploit others.
- nemomarx 10mo agoThere's some irony here considering Spotify used pirated mp3s at the start of their operations, I suppose.
- poly2it 10mo agoSome people's urges to destroy all traces of human civilisation astonish me. What do you think Spotify is going to do with all its music when it ceases to exist in however many years? No, we must collectively feed Daniel Ek the Hungry.
- jopicornell 10mo agoMonopoly is not a nice thing. Maybe it is convenient, but not nice. People that gives money to artists are the ones going to concerts and buying music directly to artists. Spotify gives cents to artists, incetivizing awful behaviour (AI music, aggressive marketing, low effort art...).
- conception 10mo agoAre you talking about Spotify here…?
- chrneu 10mo agolol is this comedy? Cuz it's absolutely hilarious opposite humor.
- rireads 10mo agoYou must be the Spotify CEO, lol
- venturecruelty 10mo agoYou're talking about Spotify, right? Famously started by ad execs pirating music and then selling it.
- Aurornis 10mo agoNot legal. This group does not concern themselves with copyright law.
- chrneu 10mo agothey do concern themselves with it, but in a "calling it out for being shit" kind of way.
- toomuchtodo 10mo agoAdherence to the legal framework is a function of your risk appetite.
- ronsor 10mo agoVery, if we delete copyright like we're supposed to.
- luke-stanley 10mo agoCurrently it says they have released metadata and album art. Is archiving and sharing the textual track metadata alone (no images, no audio) legal in the US, or Europe? By what basis is it legal or illegal?
- layer8 10mo agoCompletely illegal.
- sneak 10mo agoThe metadata scrape might not be.
- layer8 10mo agoPretty sure any kind of scraping violates Spotify’s ToS.
- sneak 10mo agoToS is not law except in the most draconian and authoritarian interpretations of the CFAA.
- layer8 10mo agoYou are mistaken, it’s contract law.
- DannyBee 10mo agoLawyer here - A bunch of things: 1. You are all probably talking past each other - I expect the original question of legality was about criminal, and not civil, law. 2. I'm sure they did not view or sign the TOS to access this. You can't be bound to a contract you never view or intentionally assent to. At least in most countries/places. For example, in the US I can show you tons of cases in just about every state and federal court where the court decided the TOS doesn't apply because it was never viewed or assented to. IE cases like https://cases.justia.com/federal/district-courts/nevada/nvdce/3:2012cv00325/88233/114/0.pdf https://cases.justia.com/federal/district-courts/nevada/nvdc... (Ironically it works both ways, so if the contract provides you any guarantees, you can't take advantage of them to sue for breach if yuo never assented) It's different if you can prove that they knew there was a TOS they would be bound by and just never bothered to look at the terms. That is very hard to prove, and it does not suffice to prove that everybody has a TOS these days or whatever. You have to prove actual knowledge of a TOS by these particular defendants. I use the US because it tends to be on the forefront of maximal browserwrap enforcement, so if it's not going to be enforced there, it's usually not going to be enforced anywhere
- artninja1988 10mo agoWow. Anna is a godsend. Hopefully now we get some really good open source music models
- brcmthrowaway 10mo agoFirst we need good stem splitting
- artninja1988 10mo agoWhat do you think about the recent SAM audio model by meta? https://ai.meta.com/blog/sam-audio/ https://ai.meta.com/blog/sam-audio/
- brcmthrowaway 10mo agoIs it realtime?
- zoklet-enjoyer 10mo agoWow. Now I just need some hard drives and a way to download that without my ISP doing something about it. That's amazing.
- timcobb 10mo ago> and a way to download that without my ISP doing something about it. what would your ISP do?
- komali2 10mo agoWhen I left my apartment back in 2018, I was switching the Comcast account over to my housemate who was staying on there. In doing so I discovered I had a myname2342@comcast.com email account. The UI showed something like 8,000 unread emails. Bemused, I opened it to see what kind of spam it had accumulated. None at all! It was just under 8,000 DMCA / torrent warning emails from Comcast itself. "We know you torrented The.Pokemon.Movie.2001.h264.mkv, you better stop that!" A full year of these emails and nothing more than that ever happened. (if you're wondering how I hit 8000 torrents, the answer is individual album torrents)
- k12sosse 9mo agoThis is called Speculative Invoicing, some law farm submitted a c&d and requested your subscriber information, Comcast hopefully threw the request in the trash but forwarded the complaint to you.
- basisword 10mo ago[flagged]
- efilife 10mo agoWhy is this stealing? You can already listen to everything that's on Spotify with a free account. You are free to also record the audio while it's playing. I suppose grabbing the actual file should't matter? Or is this about releasing? And robbing people of plays they would otherwise get through Spotify?
- basisword 10mo agoIf you listen to something on Spotify with a free account the artists still get paid. This isn't a case where you're ripping off so mega-corp. You're ripping off thousands of artists from major label ones to tiny indies. Take the metadata and build something cool. Stealing the files and releasing them is something else entirely.
- prmoustache 10mo agoYou can record what you play from Spotify and you are already free to play the record again and again and again without the artist being paid. Most people do not because they find it less convenient than paying 20bucks a month or whatever is the current price in 2025 but that doesn't change the reality. For most people the appeal of Spotify is not the music itself but the playlists that are shared thanks to its ubiquity. This is the reason other services struggle to make a dent even if they have better quality, UI and algos. Spotify started by disrupting the market using pirated music by the way so you are pretty much endorsing and encouraging piracy when "paying" your favorite artists through Spotify.
- viraptor 10mo ago> with a free account the artists still get paid Unless they're international stars, not really. It's peanuts these days. https://www.reddit.com/r/spotify/comments/13djsl9/how_much_do_artists_make_on_spotify_case_studies/ https://www.reddit.com/r/spotify/comments/13djsl9/how_much_d...
- WD-42 10mo agoIncredible. > A while ago, we discovered a way to scrape Spotify at scale. They wont and shouldn’t divulge the details, but I imagine that would be a fun read!
- bmikaili 10mo agothey're probably just using something like https://github.com/nor-dee/spotizerr-spotify https://github.com/nor-dee/spotizerr-spotify
- WD-42 10mo agoNo way, that would take far too long.
- bigyabai 10mo agoProbably not, those tools don't actually download Spotify tracks at source quality.
- sunaookami 10mo agoThere are tools that actually download directly from Spotify (needs premium then) but yeah most of them just use the search and download from other sources like YouTube without mentioning it. I won't say which tools download directly out of fear that they get killed but they exist.
- echelon_musk 10mo agoSadly since zspotify was killed I don't know of any remaining tools.
- spatterl1ght 10mo agovotify
- DUDOS 10mo ago
- frereubu 10mo agoSite is down for me. Archive link: https://archive.is/jf3HW https://archive.is/jf3HW
- ipsum2 10mo agoIronic. But its working for me.
- mawax 10mo agoProbably not down, but blocked by your ISP. Try a VPN. Same thing happens here.
- lukan 10mo agoYes, blocked. This is what I see in germany without a VPN https://notice.cuii.info/ https://notice.cuii.info/ "Their buisness model is based on copyright infringement" Well, where to complain that Anna's Archive ain't a buisness?
- MrGilbert 10mo agoAamzingly, I don't even get this page. I just see the default "this page is not available" from my browser. I'm with Vodafone, and I wonder if it is legal to pretend a site doesn't exist without notifying me.
- croemer 10mo agoPretty sure it's DNS level block. So just using private DNS would be enough, no need for full blown VPN. It's just that VPNs also usually use their own DNS instead of the ISPs. I recommend NextDNS or similar to bypass those DNS blocks and also block ads at a very deep level that works ok mobile and even inside apps.
- nurumaik 10mo agoI'd rather complain why somebody decides for me where what websites I'm allowed to open
- xnx 10mo agoMerry Christmas!
- crazygringo 10mo agoThis is insane. I definitely was not aware Spotify DRM had been cracked to enable downloading at scale like this. The thing is, this doesn't even seem particularly useful for average consumers/listeners, since Spotify itself is so convenient, and trying to locate individual tracks in massive torrent files of presumably 10,000's of tracks each sounds horrible. But this does seem like it will be a godsend for researchers working on things like music classification and generation. The only thing is, you can't really publicly admit exactly what dataset you trained/tested on...? Definitely wondering if this was in response to desire from AI researchers/companies who wanted this stuff. Or if the major record labels already license their entire catalogs for training purposes cheaply enough, so this really is just solely intended as a preservation effort?
- Aurornis 10mo ago> The thing is, this doesn't even seem particularly useful for average consumers/listeners, since Spotify itself is so convenient, and trying to locate individual tracks in massive torrent files of presumably 10,000's of tracks each sounds horrible. I wouldn’t be so sure. There are already tools to automatically locate and stream pirated TV and movie content automatic and on demand. They’re so common that I had non-technical family members bragging at Thanksgiving about how they bought at box at their local Best Buy that has an app which plays any movie or TV show they want on demand without paying anything. They didn’t understand what was happening, but they said it worked great. > Definitely wondering if this was in response to desire from AI researchers/companies who wanted this stuff. The Anna’s archive group is ideologically motivated. They’re definitely not doing this for AI companies.
- crazygringo 10mo ago> The Anna’s archive group is ideologically motivated. Very interesting, thank you. So using this for AI will just be a side effect. And good point -- yup, can now definitely imagine apps building an interface to search and download. I guess I just wonder how seeding and bandwidth would work for the long tail of tracks rarely accessed, if people are only ever downloading tiny chunks.
- throwaway613745 10mo agoI wonder how deep the hole they're gonna put whoever runs this site into is gonna be?
- urbandw311er 10mo agoI heard they’re based in Russia so one assumes they probably will be welcomed by the current government (or even aided) rather than prosecuted.
- Etheryte 10mo agoTo put this into perspective, What.CD [0] was widely considered to be the music library of Alexandria, unparalleled in both its high quality standard and it's depth. What had in the ballpark of a few million torrents when it got raided and shut down. Anna's rip of Spotify includes roughly 186 million unique records. Granted, the tail end is a mixed bag of bot music and whatnot, but the scale is staggering. [0] https://en.wikipedia.org/wiki/What.CD https://en.wikipedia.org/wiki/What.CD
- VanTheBrand 10mo agoTrue but What.cd had a tremendous amount of notable music not available on Spotify though because it was also sourced from cds, bootlegs, vinyl, tape etc whereas Spotify only includes music explicitly licensed for streaming.
- Etheryte 10mo agoThis is true and a category of music that got hit notably hard was live recordings. What had a wide array of live recordings made by sound engineers straight from the mixer. This is something that you simply cannot find now unless you maybe know a guy.
- qingcharles 10mo agoThat's why I use YouTube Music as my streamer as they allow damned near anyone to upload any old rare record and then figure out the royalties somehow.
- alxndr 10mo agoFWIW archive.org has a lot of live music as well
- leetbulb 10mo agoYes. RIP a ton of very rare material. What.cd has a special place in my heart.
- syntaxing 10mo agoMoral and legal discussion aside, this is technically very impressive. I also wouldn’t be surprised if this somehow kickstarts open source music generative AI from China.
- robotbikes 10mo agoThis already exists and is interesting to play around with - https://github.com/ASLP-lab/DiffRhythm https://github.com/ASLP-lab/DiffRhythm
- ipsum2 10mo agoCan someone explain why C#/Db (major/minor) is the third most popular key? Very unexpected for me, since its relatively more difficult to play.
- klysm 10mo agoDifficult to play in what instrument?
- yurishimo 10mo agoC# I don’t believe was/is a common tuning for most western instruments, classical or modern. A digital piano can transpose things to make it “easier” to play. Cursory google search says that a sitar is traditionally tuned to something useful for c# I’m curious if C# is one of those notes that lines up nicely with whatever crappy consumer stereos/subs were capable of reasonable reproducing in the 90s as electronic music was taking off and it stuck around as a tribal knowledge for getting more “oomph” out of your tracks.
- klysm 10mo agoI play piano and don’t mind playing in Db at all. The chords fit nicely in the hands
- kzrdude 10mo agoElectronic dance music is the biggest genre in the data. So then easy to play shouldn't matter. It's still an interesting question. I think playing Db is pretty nice on the piano even if it's not the easiest.
- ruuda 10mo agoThere is a sweet spot for the bass. Lower is better for deep bass, but too low and it stops being a recognizable note, and consumer speakers can't reproduce it. This effect exists though I'm not sure if it is the cause of the pattern here.
- 10mo ago
- Fizzadar 10mo agoI have Spotify premium but the constant shuffle of content availability has meant I’ve stared routinely archiving my liked songs to avoid any rug pull. Zspotify and co still work a charm.
- nutjob2 10mo agoI wonder how definitive their collection is and how much ripping Google Music/YouTube would improve on this. A distributed ripping project to do that would be a fine thing.
- yegle 10mo agoNot that we should, but it's technically feasible to have a music streaming server with the torrent as the backend, and selectively download the part of the torrent in respond to on-demand streaming request from the client.
- pjerem 10mo agoYeah we shouldn’t. But we may.
- nness 10mo agoa la "Popcorn Time."
- uhfraid 10mo agospotify used to do just that (stream p2p) until 2014 or so https://www.scribd.com/document/56651812/kreitz-spotify-kth11 https://www.scribd.com/document/56651812/kreitz-spotify-kth1...
- zanderz 10mo agoThe person who wrote this Spotify p2p software also wrote uTorrent, which was bought by the company bittorrent after they struggled to make a C++ client on their own. The original bittorrent implimentation was in python, but they re-skinned uTorrent as bittorrent and shipped both for a few years. https://en.wikipedia.org/wiki/Ludvig_Strigeus https://en.wikipedia.org/wiki/Ludvig_Strigeus
- johanyc 10mo agohttps://www.csc.kth.se/~gkreitz/spotify/kreitz-spotify_kth11.pdf https://www.csc.kth.se/~gkreitz/spotify/kreitz-spotify_kth11... KTH link is better than scribd for downloading. though academic links are sometimes prone to link rot.
- willio58 10mo agoI recently got into the whole homelab *arr stack for things like movies and tv and while I know options exist for music I just don’t see the need yet price-wise. Spotify is still just cheap enough for me to not care enough. We’ll see how long this holds. That being said it’s no secret Spotify and other streaming services barely pay even popular artists. Artists make money from live shows and merch. The fact that their music is behind a paywall at all could mean they make less money from some lack of exposure. I do hope one day self-hosting music with an extremely easy setup with torrenting for sourcing is set up again. What I’m talking about exists to some extent, but it’s not trivial for most people.
- yellow_lead 10mo agoIs the music torrent not up yet? Only see the metadata one here: https://annas-archive.li/torrents/spotify https://annas-archive.li/torrents/spotify
- artninja1988 10mo agoYeah, in the article they write: The data will be released in different stages on our Torrents page: [X] Metadata (Dec 2025) [ ] Music files (releasing in order of popularity) [ ] Additional file metadata (torrent paths and checksums) [ ] Album art [ ] .zstdpatch files (to reconstruct original files before we added embedded metadata)
- yellow_lead 10mo agoOh I see, thanks! I missed that
- vlaaad 10mo agoUnrelated, but I just can't stop myself from saying that I absolutely hate Spotify even though I'm a paying customer. Fuck you Spotify. You were supposed to be a convenient way to discover and listen to music. Now you are only convenient for listening to music, and absolutely terrible for any recommendations. This is sad really. Spotify had good recommendations. It's absolutely in a position where it can provide good recommendations — it has both a vast music library and a vast amount of data on user preferences. And it chooses to push procedural/ai-generated slop instead to earn more money. I thought that maybe buying $SPOT stock will make me more at peace with its greed, but it didn't work. Spotify fucking deserves to crash and burn because it sees paying customers as idiots who might not notice they are fed garbage. Fuck you Spotify, fuck you.
- eastbound 10mo agoThis is more frequent than you would assume. I’ve neither subscribed to Apple Music nor Spotify for this exact reason: I’m a millenial who would like to discover music. Another extremely annoying effect is, being 40+, they only suggest music for my age. In “New” and “Trending”, I see Muse and Coldplay! I should make myself a fake ID just to discover new music, but that gets creepy very fast.
- layer8 10mo agoYouTube Music works pretty well for me. One great feature is that it includes not just a commercial music streaming catalog, but all user uploads of music on YouTube.
- nickthegreek 10mo agoand you can upload 100,000 of your own tracks to the service for your private use as well. It is a great service considering I am getting it as a side effect of youtube premium. Single handedly the last subscription I would cancel.
- komali2 10mo agoI had to chuck Youtube Music away when it was polluting my youtube playlists with stuff I was liking on youtube music. Me as a video viewer and me as a music listener are two completely different people.
- 827a 10mo agoHoly crap. This is going to trigger a five-alarm fire at Spotify Engineering. This has got to be among the largest proprietary datasets ever unintentionally publicized by a company.
- rightbyte 10mo agoWasn't all data available to users though?
- cm2012 10mo agoYes but very hard to scrape in bulk from user accounts
- potwinkle 10mo agoI mean... not really? Not much music is Spotify exclusive (at least from the 99.6% of what people listen to mentioned in the article), and from friends in the industry I can guarantee you all major content platforms (Netflix, Disney+, Prime Video, a large chunk of YouTube) have already been completely copied without a business agreement with the rightsholders by AI startups and big-name players.
- okokwhatever 9mo agoWho cares now, it's already downloaded and ready to be torrented... God is good
- bob1029 10mo agoI recall many interesting tracks that were very aggressively deleted from all platforms in sync. I wonder if I could find them in this archive. There is contemporary lost media being created every day because of how we distribute things now. I think in some cases, the intent of the publisher was to literally destroy every copy of the information. I understand the legal arguments for this, but from a spiritual perspective, this is one of the most offensive things I can imagine. Intentionally destroying all copies of a creative work is simply evil. I don't care how you frame it. Making media effectively lost is not much different in my mind. Is it available if it's sitting on a tape in an iron mountain bunker that no one will ever look at again?
- krick 10mo agoUh, cool, I guess? I want to applaud that, but, first off, unless you are OpenAI or Facebook, it is not exactly plausibly easy to participate in the festivities. Even if I had spare 300 TB laying around, how the fuck do I download that? But, more importantly, I cannot even say "good for you", because I don't actually think it is good for Anna's Archive. I wouldn't touch that thing, if I was them. Do we even have any solid alternatives for books, if Anna's Archive gets shot down, by the way? Don't recommend Amazon, please.
- pjerem 10mo agoBitTorrent protocol doesn’t force you to download all of the files of a torrent :) Now imagine a dedicated music client that will download and stream (and share, because we are polite) only the needed files :)
- killingtime74 10mo agoYou can download torrents selectively. I think if they adopted that cautious attitude they wouldn't exist in the first place
- Gander5739 10mo agoAnna's archive mirrors z-lib and libgen, so those are the main alternatives. But it's unlikely anna's archive would go down so easily, they take a lot of precautions.
- krick 10mo agoOh, I was somehow under impression that libgen is no more. Glad to see it's not. I guess it was just a different domain.
- chrneu 10mo agothink popcorn time for mp3s/flac instead of mp4. a client can selectively list and then stream individual files from a huge torrent. if you've ever watched illegal movies/shows on those random domain websites, you're likely streaming it from a torrent on the backend somewhere. it wouldn't surprise me if we start to see some docker images pop up in a few days to do exactly this as a sort of "quasi-self-hosted jellyfin". Where a person host a thin client on a machine that then fetches the data from the torrent, then allows the user to "select" their library. A user can just select "Top hits from the 80s" and it'll grab those files from the torrent, then stream or back them up. I don't really see why it wouldn't, from an end user perspective, be any different than a self hosted jellyfin or plexamp.
- _vqpz 10mo agoI really don't understand how focusing on source quality files is supposed to be a "major issue" with the music preservation community. It's bizarre for them to talk about these being barriers for creating a "full archive of all music that humanity has ever produced" have and their answer be scraping Spotify to end up with a music library comprised of many AI and bulk produced songs at 75/160kbps.
- zzzeek 10mo agogreat. Spotify just removes things all the time (things I actively listen to and work on for my jazz practices, one day just go "poof" because they didn't want to pay the record company anymore), and they are not as a company deserving of the role of "keeper of all the world's music". They don't give a shit and they'd vastly prefer we all listen to their AI generated royalty free crap and Joe Rogan.
- frytaped 10mo agoIt seems to be that the metadata doesn't include the lyrics, probably because they are provided by Musixmatch. It would have been nice to have a database of lyrics linked to ISRCs. AFAIK Lrclib doesn't support downloading lyrics for a given ISRC.
- siquick 10mo agoIs there a way to see the shape of the metadata?
- krackers 10mo agoNew multimodal training set just dropped.
- tjoff 10mo agoI just want to be able to backup my playlists. Maybe thats possible but last time I looked I could only find sites that wanted your login, not gonna happen.
- lelandfe 10mo agohttps://developer.spotify.com/documentation/web-api/reference/get-a-list-of-current-users-playlists https://developer.spotify.com/documentation/web-api/referenc... https://developer.spotify.com/documentation/web-api/reference/get-playlist#:~:text=The%20tracks%20of%20the%20playlist. https://developer.spotify.com/documentation/web-api/referenc... I bet you can whip up a super simple script with an LLM to do this!
- Spivak 10mo agoNot that using the Spotify API directly is all that hard but the spotipy library makes it even easier.
- hn111 10mo agoThis works nicely: https://github.com/spotDL/spotify-downloader https://github.com/spotDL/spotify-downloader
- crazygringo 10mo agoThis is where ChatGPT shines. Just ask it to write you a script, it'll give you all the instructions. I've used ChatGPT to write a whole bunch of playlist logic scripts (e.g. create a playlist that takes tracks from playlists A, B and C, but exclude tracks in playlist D.)
- emsixteen 10mo agoI worry about potential bans from scraping files through this sort of thing.
- crazygringo 10mo ago
- nighthawk454 10mo agoAmazing! I wonder if the Every Noise At Once[1] site could be updated with the metadata from this? [1] https://everynoise.com/ https://everynoise.com/
- iggldiggl 10mo agoThanks for linking that page, interesting rabbit hole that I hadn't heard about until today…
- reactordev 10mo agoOh this is going to go over real well in Nashville, TN.
- p0w3n3d 10mo agoThis is something really important, especially in the days when music and film vanishes from platforms one by one. I myself have three playlists with greyed out titles (titles are missing so there's no possibility for me to find out what was there). That's why I divide music to the one that I want to have forever - I buy it on CDs - and dance music that I can live without one day
- eightys3v3n 10mo agoI really appreciate platforms that still show the titles and metadada after something is removed. Then at least I can go find it again to maintain my collection. Tidal does this.
- tristanc 10mo agoThis is one of the greatest news I've ever heard for the digital preservation community. Just so many projects over the years could have used resources like this. Thank you for contributing to humankind!
- tolerance 10mo agoI am not enthused by this news. Let us entertain the possibility that similar institutions will eschew this catalog.
- 47282847 10mo agoHmmm I don’t like this. There are sources for music with better quality out there and all this will do is paint them a bigger target for takedowns/prosecution. I am worried about losing their ebook library. Quoting from the announcement: “Generally speaking, music is already fairly well preserved.“ They should have done this as a separate identity.
- lukan 10mo ago"and all this will do is paint them a bigger target for takedowns/prosecution" They are based in russia. And they currently do not work together so well with the west. So it is imaginable, that if some people give Trump quite some money, to make Annas takedown part of some deal to lift sanctions after a ceasefire in Ukraine, but .. it does not seem like it. I rather suspect more effort in the west to block access to unwanted sites like this. My ISP in germany is already blocking it.
- computergert 10mo agoTrump threatened the EU to tax Spotify (and others) just this week. So it doesn’t look like Trump would be happy to help Spotify out, though in exchange for money he’ll probably change his mind.
- 47282847 10mo agoYour ISP is filtering DNS records. Easily fixed by changing DNS. It may even speed up your lookups, as most ISP DNS are slower than the large ones like quad1/8/9. > They are based in russia. “Russian authorities have without any notice suspended Russia's most popular file-sharing website torrents.ru for the alleged violation of copyright laws.” (2010) https://www.petosevic.com/resources/news/2010/03/000350 https://www.petosevic.com/resources/news/2010/03/000350 “In 2016, for example, the Moscow City Court (Mosgorsud) granted more than 700 requests to protect intellectual property.” https://www.group-ib.com/blog/torrents/ https://www.group-ib.com/blog/torrents/ “The ISPs in Russia are required to block subscriber access to thepiratebay.se and thepiratebay.mn following the complaint of […]” (2015) https://www.maverickeye.de/russia-has-ordered-local-isps-to-block-the-pirate-bay/ https://www.maverickeye.de/russia-has-ordered-local-isps-to-... “Roskomnadzor, the country’s telecom and media industries regulating body wants people to pay, so in 2016 it’s going to block Russia’s 15 most popular torrent websites” https://www.inverse.com/article/9619-russia-will-crack-down-on-the-15-most-popular-torrent-websites https://www.inverse.com/article/9619-russia-will-crack-down-... etc There are plenty of Russian music labels. Big book publishers? Not so much. Some sites explicitly ban content from the hosting country to try and avoid that. Not the case here.
- jimmydoe 10mo ago[flagged]
- dmix 10mo agoI hope they get the new lossless versions
- 1dry 10mo agoYuck. Just to make it easier to train slop machines. The point of art is not to have completionist archives of EVERYthing that’s ever been made! Let it die. Death is the most natural part of life. Art is about the human experience, not “for researchers”. The point is human connection. Art is a living reflection and record of human experience. Art will persevere- the kinds of folks who prioritize what they like based on popularity were never the supporters artists (contrast with craftspeople trying to make a buck) counted on in the first place. Enjoy your derivative slop - we’ll continue on our imperfect, messy, individual, human artistic lives.
- justatdotin 10mo agoI am having a lot of trouble following you. Something has upset you: what would make you feel better? do you mean that researchers should be disallowed from accessing art? I do not see how research interferes with all the benefits you prioritise. Can't you continue to enjoy those benefits? Many people think 'real' music has electric guitars. I think they're wrong, but why argue with them? I think it's fine if you do not like music made from music, but that ship sailed last century. One detail you may be missing is that there are imperfect messy individual artistic humans who make music from music too. Computers are no more an obstacle to human connection through music than electric guitars are.
- junon 10mo ago> I am having a lot of trouble following you. Something has upset you: what would make you feel better? Don't talk to people like here, please. It's passive aggressive and unproductive. GP's comment was fine, if not a bit impassioned, regardless if you agree with it.
- justatdotin 10mo agothanks for the correction, I do not want to be aggressive. I see now I should have just asked: what do you want? to prefix my response with an admission that I'm not sure what the problem is.
- sneak 10mo ago199GB, only metadata released for now. Magnet link found here: https://annas-archive.li/torrents/spotify https://annas-archive.li/torrents/spotify Are magnet links allowed on HN?
- cranberryturkey 9mo agothat is only 199gb, the real one is 300TB
- littlecranky67 10mo agoFor some reason, the link does not work for me (spain). Works perfect at the same time in tor browser.
- mvkel 10mo agoThis work is so critical. Read an article that was published just 10 years ago, and witness the bit rot as most external links will 404, gone forever. I think it's worth questioning the value of preserving -everything-, but it seems like if we can, we should.
- larodi 9mo agoYou know, I had the (at time of writing) 600 something comments ran through Opus 4.5 and do a summary of the sentiments. It could't find a single comment that genuinely defends Spotify or expresses sympathy for the company. HN crowd is, of course, biased in the technocratic sense, but you see - everyone seems to actually rejoice the move. The closest to remorse is `linhns` and `locusofself` expressing concern about artists getting hurt (not Spotify itself), but locusofself prefaces with "I hate spotify as a company but..." (disclaimer: this text is NOT LLM generated, I wrote myself a summary of the summary. here's the Claude thread should anyone care https://claude.ai/share/cfc4ca63-2b9e-47ac-a360-202025d1a134 https://claude.ai/share/cfc4ca63-2b9e-47ac-a360-202025d1a134)
- mycall 9mo agoAre those 404 links available on web.archive.org?
- msephton 10mo agoIs this all regions? I'm assuming so but I can't be sure
- walthamstow 10mo agoVery interesting that a white noise track for babies is the 4th most popular track on Spotify.
- cluckindan 10mo agoInteresting if that is considered to be copyrightable. Any white noise track is perceptually indistinguishable from another, but none have the exact same sequence of samples except by chance, or if the noise generator happens to be deterministic as a function of time.
- zarzavat 10mo agoWhite noise isn't copyrightable.
- cluckindan 10mo agoThen how is silence copyrightable?
- al_borland 10mo agoI find it so odd that people then to streaming services for stuff like this. I have a dedicated white noise machine, and when I travel, I use the white noise (bright noise actually) built into the iPhone. Relying on an external hosted service would never cross my mind, and surely wouldn’t be something I go to on a daily basis.
- junon 10mo agoIt's not odd if you aren't the type who frequents hacker news. We are, after all, very much in a bubble here.
- komali2 10mo agoYou might find it interesting that there's an entire genre of youtube video that's designed to just be chucked one by one into slideshows for elementary school teachers to use as their lesson plan. Including videos that are just "2 minute timer for kids!" e.g. https://www.youtube.com/@Ask.the.Teacher https://www.youtube.com/@Ask.the.Teacher "Independent Reading: Count Up Timer for Classrooms": https://www.youtube.com/watch?v=AfLfJtVeME8 https://www.youtube.com/watch?v=AfLfJtVeME8 straight up just stock imagery and a timer lol
- virtualritz 10mo agoI just found out that https://annas-archive.li/ https://annas-archive.li/ is masked by my German internet provider (SIM.de/Drillisch). I usually use a VPN but I had it switched off temp. to watch Fallout (Prime Video won't let you watch through a VPN). Only when I switched Mullvad back on could I open the site. I didn't know German providers do this.
- iknowstuff 10mo agoIn that vein, I am trying to find out why searching for alextud popcorntime which should trivially yield http://github.com/alextud/PopcornTimeTV http://github.com/alextud/PopcornTimeTV results in anything but that one particular URL in every search engine: Google, Kagi, DuckDuckGo, Bing They even find a fork of that particular repo, which in turn links back to it, but refuse to show the result I want. Have't found any DMCA notices. What is going on?
- ticoombs 10mo agoThey have marked the repo as noindex (or GitHub is forcing a noindex header). Its returning a noindex flag so every serp is correctly doing what the repo has been asked. That is... except for brave! I checked on my searx instance and it still showed up in brave's results
- ZeWaka 10mo agoVery interesting. The security page does show up on kagi at #6. I wonder if GitHub flags it to not be indexed or something.
- Mythli 10mo agoTry Yandex search, trust me later. It has 0 censorship - regarding pirated content at least.
- junon 10mo agoWas also shocked to see that (Berlin, Telekom here).
- oarfish 10mo ago
- 63 10mo agoAttracting the ire of the music industry seems like a huge, unnecessary risk. I wish they had performed this as some kind of other entity to try to keep the ebook archive protected from the fallout. I fear this will not end well.
- urbandw311er 10mo agoThey can’t be touched by the music industry they’re based in Russia.
- lysace 10mo ago[flagged]
- BrokenCogs 10mo agoWhy... does Putin like music more than the next guy?
- lysace 10mo agoWhy would you want to destroy your enemies' industries, is what you're asking? Although I suppose that is predicated on seeing Russia as the enemy. Strangely not always the norm these days in the new world.
- komali2 10mo ago> Why would you want to destroy your enemies' industries, is what you're asking? Do you have any evidence that pirating is destroying industries? My guess is I can find the majority of this release by anna's archive on some combination of the pirate bay and the soulseek, or private music trackers. And yet, Spotify is still a thriving company, as is the entire music industry as a whole. There's even room for competing streaming services like Tidal and Youtube Music.
- flexagoon 10mo agoThen why would Anna's Archive also release archives of some of the largest Chinese publishers? Surely Putin wouldn't want to destroy China's industries.
- squigz 10mo agoOut of curiosity, where does Anna's Archive claim "communism" as a motivation?
- urbandw311er 10mo agoI have absolutely no idea why you’re being downvoted. This feels like exactly the sort of project that would be backed by the current Russian administration, given it serves to damage and destabilise businesses in countries that are currently hostile to Russia. — it’s not even a controversial take to say so.
- shevy-java 10mo agoHmm. This is actually not really something I need, I think; but I consider anna's archive etc... as about as important as the internet web archive. We need to preserve data, at the least important data, also historic data - how the original websites looked. Creativity of past generations. Same for games and books. It may be only ~30 years for webpages to have emerged, but there are also many young people who may not have experienced that since they are too young to have experienced it. There is always a generational change; our generation has the opportunity to store more things.
- schmuckonwheels 10mo agoI want to time-travel back to 2000 like Old Biff with the sports almanac so I can tell Shawn Fanning to use the "it's for historical preservation" defense.
- snoozebutton 10mo agois this not highly illegal?
- MuffinFlavored 10mo agoAt first I was thinking "ok maybe they only backed up artists who released under some kind of like... public open source music sharing license" then I read deeper... I had never heard of Anna's Archive before. Feels similar to ThePirateBay2.0. Surprised they are so public about their crimes?
- ZeWaka 10mo agoSince the article asks: > We're curious about the peaks at whole minutes (particularly 2:00, 3:00, 4:00). If you know why this is, please let us know! As a hobby video/audio editor, people will start with their track taking up a preset amount and fill up the time - even if it means having some dead space at the end. The other alternative is algorithmically created music.
- nemomarx 10mo agoI've heard 2:00 is some kinda sweet spot for the Spotify algorithm and payouts? You get paid per play so you don't want to it too long, but if your track is much shorter than two minutes you get penalized or something. I know they've had to remove ambient tracks that were cut into 40 second clips as part of this. So you might see a lot of anchoring just like YouTube videos kept stretching to almost exactly ten minutes?
- gyrgtyn 10mo agois there a torrent client already that is be good at partial downloads? I didn't realize how popcorn time worked until I read this thread.
- kccqzy 10mo agoAll torrent clients must necessarily support partial downloads because of the nature of torrents. The files are split into pieces which are downloaded and then assembled by the torrent client.
- flexagoon 10mo ago"Partial downloads" in the context of torrenting usually refers to downloading specific files from a torrent
- acjohnson55 10mo agoThis is incredible. I once assembled a collection of 100,000 tracks for research on exploration of large music libraries. Essentially vector search. I was limited in storage and processing power to a single machine. If I were to do it today, I could get so much farther with hyperscaler products and this dataset.
- deleted 10mo ago[deleted]
- markstos 10mo ago> ≥70% of songs are ones almost no one ever listens to (stream count < 1000). So much interesting but undiscovered music is out there!
- halperter 10mo agoIt would be interesting to find out how that has changed with the growth of the music industry over the years. I suspect that many of these <1000 streamed could be artificially generated for monetary purposes but I'm not entirely sure. That being said, there is a lot of good music with less than 1000 streams. I've been looking myslef and I've definitely found some hidden gems.
- junon 10mo agoTIL Anna's Archive is blocked in Germany (by a rather obtrusive MitM, I might add). Get redirected to a "Copyright Clearing House" or something.
- rendaw 10mo agoLooking at the analysis, I'm totally surprised opera and psytrance are so prolific. Psy-trance... I thought it was the same as any other electronic genres, but do people get high and just start shoveling psy-trance tracks out or something? Opera I thought was a very strict discipline, needing rigorous somewhat esoteric training in order to produce the right sounds. How could there be so many opera artists? I mean, I'm sure there's some misclassification, but chamber music is basically a couple people with any sort of music training on classical instruments so that doesn't surprise me nearly as much... I can easily imagine there being _lots_ of those, and you might come up with a different artist name for each unique set of people you collaborate with.
- komali2 10mo ago> Opera I thought was a very strict discipline, needing rigorous somewhat esoteric training in order to produce the right sounds. How could there be so many opera artists? My guess is just the same opera performed by a ton of different orchestras, and perhaps the same orchestra for different recordings, times however many operas there are.
- captbaritone 10mo agoFormer classical singer here. Only theory I can come up with is that opera tends to have large casts where all the singers are credited individually which would inflate the absolute numbers of "artists" relative to other generes. I still struggle to imagine this accounting for bringing such a niche genera to the top here.
- gorbachev 10mo agoMy guess is a large portion of the psytrance music is slop, whether AI or some other form of auto-generation.
- legacynl 9mo agolol. Where is all this anti-psytrance hate coming from? Are you people actually that childish that you don't understand the concept of taste, and that everyones' is different? People who have like different music than you aren't stupid. Electronic musicians aren't bad musicians. You know that nice feeling you get when you listen to music from your preferred composer/artist/genre? Other people feel exactly the same, but with different kinds of music. Some people even love the thing that you hate! wow! Who knew? Except for anybody above the age of 5. TLDR; just because you dont like Indian food, doesn't make Indian food bad. It's the same for music or other things that are dependent on taste.
- Uninen 10mo agoI hope someone builds an open API around this metadata. I'd love to have alternatives to the big player APIs.
- meysamazad 10mo agoI wonder if Spotify will pursue any legal actions to take this archive or the site down!
- shomp 10mo agoIf only Spotify paid musicians their fair share
- userbinator 10mo agoMusic files (releasing in order of popularity) Increasing or decreasing? IMHO increasing would make more sense, as the most popular music is already mirrored in countless other places. It's the rare stuff that is most in need of preservation. I wonder how much of the content there is AI-generated. Honestly, even as someone who was initially skeptical, I've found some of it to be rather good --- not knowing that it was AI-generated at first. Now if they could only reverse-engineer the prompt and only store the model, that would be an extremely efficient form of "compression".
- reassess_blind 10mo agoSame model and same prompt won’t necessarily create the same result, unless I misunderstand how these audio models work.
- squigz 10mo agoIt's possible to generate the same images and text from LMs by tweaking the settings, right? Are audio models different?
- Philpax 10mo agoYes and no - yes, in theory, but in practice, non-determinism can be introduced at different points along the stack. See Thinking Machines' post on LLM non-determinism: https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ https://thinkingmachines.ai/blog/defeating-nondeterminism-in...
- m00dy 10mo agoCongrats! I’m sure the Spotify lawyers are gonna have some sleepless nights ahead.
- djfergus 10mo agoAnna’s Archive has largely flown under the radar by focusing on books. Even perceived involvement in music piracy puts a much bigger target on their back from far more aggressive actors (RIAA, major labels)
- reassess_blind 10mo ago“Good luck, we don’t care.” is their stance, as far as I can tell.
- pmdr 10mo agoThe bulk of today's customers has no idea how to pirate music, so they're not really a threat anymore. Music streaming has been rather convenient, you pretty much get the same content across all services. Video streaming platforms have, unfortunately become fragmented and, as of late, ad-ridden.
- TheAceOfHearts 10mo agoI wonder if they'll explore other music services as well. As I understand it, Deezer, Qobuz, and Tidal can all get ripped easily enough. Although I'm not sure if they rate limit downloads past a certain point. I'm a bit sad that they chose to focus on music rather than audiobooks. Creating an archive of audiobooks seem like it would be more aligned with their mission.
- TechSquidTV 10mo agoThe metadata is gold, but I was immediately curious why why wouldnt go for Tidal first. Though what ever they have on Spotify I think is unique.
- BaudouinVH 10mo agoerror 451 https://postimg.cc/QFddnW41 https://postimg.cc/QFddnW41
- linhns 10mo agoUnlike books, which are massively overpriced, this will hurt artists a lot as they need the fees paid by Spotify to make ends meet.
- Stagnant 10mo agoI don't think so. Streaming services are used for convenience. Torrenting and managing music at this scale is inconvenient. Distributing these huge torrents is the perfect way to avoid any real damage to artists while being invaluable to preservation of culture.
- locusofself 10mo agoI hate spotify as a company but I agree, at least in my case, a large share of my wife's income comes from spotify.
- n07fr0mh34r3 9mo ago[dead]
- themusicgod1 9mo ago> this will hurt artists a lot as they need the fees paid by Spotify to make ends meet. Anyone using DRM/paracopyright to "make their ends meet" deserves what they get. This is de facto theft from the public domain.
- thih9 10mo agoThis is conspiracy theory territory but I wonder if big tech is sponsoring efforts like this as an easy way to get training data.
- dbacar 10mo agoNow, anyone with some decent info on signal processing and machine learning can build his/her own Shazam.
- verisimi 10mo agoYes, but do they have the one that goes like: to-to-to dotodoo? Hmmm? Do they?
- gorbachev 10mo agoQuoting from their page: -------------- This is by far the largest music metadata database that is publicly available. For comparison, we have 256 million tracks, while others have 50-150 million. Our data is well-annotated: MusicBrainz has 5 million unique ISRCs, while our database has 186 million. -------------- If they truly are on a mission to protect world's information from disappearing, they should work with MusicBrainz to get this data on it. Alternatively, it would be amazing, if they built a MusicBrainz like service around it. In either case, to make the data truly useful, they'd need to solve the problem on how to match the metadata to a fingerprint used to identify the music tracks, assuming that data is not part of the metadata they collected.
- 47282847 10mo ago> n either case, to make the data truly useful, they'd need to solve the problem on how to match the metadata to a fingerprint used to identify the music tracks How is that a problem? for each track in collection do extract_fingerprint
- aerozol 10mo agoIt would be reasonably trivial to set up a bot that mass-imports metadata from Spotify to MusicBrainz (note that MB rules do not allow this, community cleanup from a single user doing this with another source, years ago, is still ongoing). The value that MusicBrainz adds is the community editor who spent a few hours going through YouTube videos and wayback machine social links to figure out that Fog (Wellington, NZ, punk/post-punk) and Fog (Auckland, NZ, Post-Punk) are different bands - even if they share a Spotify profile. The editor that hunted down and listened to 5 compilations that have mixed up a radio edit and an original mix of a track, to find out which is which, and separate them in MB and make notes. [these are made up examples] That's not to imply that these two projects are 'competing', or that the ISRC figure comparison isn't useful and correct. But community database + scraped data is apples and oranges. And a mixed fruit bowl is wonderful.
- squigz 10mo agoI was wondering if MB had any rules on such things. I get the motivation, but I hope they'd be willing to work with some trusted editors to figure out if this data would be useful/could be imported without risking quality. But MB is one of the best resources out there - precisely because of what you said - so I'm not complaining too much :)
- gorbachev 10mo agoI want to peek in that metadata collection to see if it could be used to identify the AI slop that's infecting Spotify. If you could identify a track supposedly by artist X was actually AI slop not created by artist X, you could use that information to skip tracks on (web) music players, for example.
- pcblues 10mo ago[flagged]
- Kerollmops 10mo agoSo nice! That's an excellent extract and looks useful for benchmarking Meilisearch. I'll probably spend my Christmas holidays importing the tracks, albums, and artists into Meilisearch, while my CEO builds a beautiful front-end for it. I'll probably replace [the current music search demo](https://music.meilisearch.com https://music.meilisearch.com) we have with this much higher-quality dataset! That would also be a good fit for [the new delta-encoded posting lists I am working on](https://github.com/meilisearch/meilisearch/pull/5985 https://github.com/meilisearch/meilisearch/pull/5985). Let's see how good it can get. My early benchmarks showed a 50% reduction in disk usage.
- gverrilla 10mo agoGREAT DAY
- 7ero 10mo agofree the music
- ThinkBeat 10mo agoCan this last? I envision an army of lawyers and cyber security companies being prepared to unleash a scorched earth campaign that book publishers might want to be part of as well. At the end it may take down more than just this publication but most others as well.
- romanovcode 10mo ago`spotdl download "https://open.spotify.com/user/{username} https://open.spotify.com/user/{username}" --user-auth --output '{list-name}/{title} - {artists}.{output-ext}'` This is literally all you need to back up Spotify.
- Philpax 10mo agospotdl downloads from YouTube, not Spotify, afaik
- peterburkimsher 10mo agoFor a fully-legal alternative of metadata archiving, I suggest the iTunes EPF (Enterprise Partner Feed). https://performance-partners.apple.com/epf https://performance-partners.apple.com/epf The best metadata I've found, though, is the MySpace Dragon Hoard: https://archive.org/details/myspace_dragon_hoard_2010 https://archive.org/details/myspace_dragon_hoard_2010 That included the artist location, allowing me to tag songs based on their country. I then created playlists such as "NERAS" Non-English Rock Artist Sample, where the one most popular song for a particular artist was chosen, and only when the country of origin was not English-speaking, and the genre was Rock. I like listening to music while working, but English lyrics distract me because I understand what they're saying. After discovering music via the MySpace archive, I've since purchased 73 songs from 35 artists that I'd never heard of before digging into the data. I rebuilt my playlist on Spotify, but got greyed out tracks, and YouTube Music, but got "unavailable video". So I still prefer purchasing tracks via the iTunes Music Store, Qobuz, Bandcamp, and 7digital. Other data sources such as the MP3.com rescue barge, PureVolume archive, and Anna's Spotify archive lack the country-of-origin metadata, so are of less interest to me. It may be possible to use an LLM to guess the language of each track title, but someone else will have to do that. Meanwhile, if you're interested in the genre-by-country MySpace data, or have questions about the iTunes EPF, feel free to reach out and we can discuss your research.
- squigz 10mo ago> Other data sources such as the MP3.com rescue barge, PureVolume archive, and Anna's Spotify archive lack the country-of-origin metadata, so are of less interest to me. It may be possible to use an LLM to guess the language of each track title, but someone else will have to do that. I would guess that combining these sources, along with info from MusicBrainz, would help quite a bit? Still, I'm rather surprised Spotify doesn't provide more information about artists.
- o_____________o 9mo ago> Please note that Apple Music and iTunes Music data will be migrating away from the Enterprise Partner Feed (EPF). Starting July 16, 2024
- 9mo ago
- simmo9000 10mo agoWe need insane for culture to survive.
- Aldipower 10mo agoOh, just noticed my provider "Vodafone Germany" is blocking the domain annas-archive.li on DNS level.
- xandrius 10mo agoTruly amazing work. I couldn't help but being sad of the less popular songs not being currently stored, as those are definitely the ones more in risk of being lost forever. If you like the goal and you have even a few 100gb available on your server, consider "donating" some of that space to seeding the data (music or books). It's absolutely how we can fight the system, even if just a tiny bit. https://annas-archive.org/torrents https://annas-archive.org/torrents
- squigz 10mo agoGoing off the blog post, archiving the rest of Spotify (which only represents 0.4% of total listens) would bring the total size up to something like 1PB, and would likely include a huge amount of AI generated stuff, which I don't think is worth it. I'd rather see them focus resources on archiving other stuff.
- xandrius 9mo agoSure but "the other stuff" is Lady Gaga and Bunny, which we won't have issue finding a copy of. Sure, there is AI stuff but also not.
- justacrow 9mo agoYou'd hope there is room for a lot of stuff in between as well. I tried searching for some of yhe more unknown artists I follow (and have bought stuff from!) but didn't see any clear way to filter it to Spotify/music metadata only. Restricting to metadata+other cuts it down somewhat but it's still 100s of results for most topics. Will be interesting to see what's there and not once the actual music torrents come up, should make it easier to search.
- xandrius 9mo agoIt's in their roadmap, from what they said. I'd imagine they wouldn't reject someone who wanted to contribute that feature to the project.
- yoan9224 10mo agoThe metadata alone is incredibly valuable for researchers. Having 186 million ISRCs catalogued with associated genre, tempo, and popularity data is a goldmine for music analysis that doesn't even require touching the audio files. I've always found it interesting how streaming services have become the de facto music library of record, yet they can and do remove content at will. When Spotify pulled out of Russia, entire catalogs became inaccessible. Physical media and personal archives suddenly matter again in ways we thought were obsolete. The copyright discussion is complex, but from a pure preservation standpoint, I'm glad someone is doing this work.
- Yeri 10mo agowow. Blocked in Belgium. Error HTTP 451 - Unavailable For Legal Reasons https://lumendatabase.org/notices/71398835 https://lumendatabase.org/notices/71398835
- hmokiguess 10mo agoWhat an early christmas gift for humanity. Now, asking for a friend, what's the ideal setup for torrenting this? Mullvad / Tailscale?
- nmz 10mo agoThis might be the perfect time to do archiving before the entire internet gets inundated by sub-par AI generated content.
- DoctorOetker 10mo agoI'd rather see them use AI to convert all the scanned scientific articles into proper PDF or other formats. Also sort and classify the articles by binary size, vs page count, plot count, raster image count etc, in order to compress the outliers and detect when a raster image should have been a plot and convert it to vectorized images etc. How compact can we get the collective human scientific corpus?
- bguberfain 10mo agoWe can finally search for playlists with a giving song! A basic feature that Spotify is missing!
- RickyLahey 10mo agoThis will be great to train AI on.
- shmerl 10mo agoJust buy music DRM-free in the first place.
- htx80nerd 10mo ago>Over-focus on the most popular artists. There is a long tail of music which only gets preserved when a single person cares enough to share it. And such files are often poorly seeded. There is a ton of good bands with under 10k or even 1k monthly listeners.
- lawrenceFounta 10mo ago[dead]
- wartywhoa23 10mo agohttps://annas-archive.li/llm https://annas-archive.li/llm
- rldjbpin 10mo agothe metadata alone is a staggering couple hundred gb, however it contains quite handy information to play with. consider the following: > /audio-features/{id} "Get audio feature information for a single track identified by its unique Spotify ID." this combined with track metadata can finally allow those motivated enough to create their own personalized shuffle. potentially better than the slop we get nowadays. no generative ai required*.
- marstall 10mo agothe top 10,000 songs seem to be 99.9% top-40 corporate pop, which suprised me. thought a list that broad would pick up more that was outside the maintream ...
- squigz 10mo ago10,000 sounds like a lot, but it really isn't. Even my own personal music collection - which isn't all that impressive - is nearly 20,000 tracks.
- Mr_Minderbinder 10mo ago> Over-focus on the highest possible quality This is not an issue in my view. I like the fact that I can download 100 MiB ultra-high resolution TIFF files of scans of photographs from the original negative from the Library of Congress and 24-bit/96kHz FLAC files of captures of 78 RPM records from the Internet Archive. In addition to maintaining completeness and quality of information, one of the main goals of preservation is to guard against further degradation and information loss. You should try to preserve the highest quality copies available (because they contain more information) and re-encoding (deliberate degradation) should only be used to create convenient access copies. Inferior copies, in addition to being less informative, have the potential to misinform. Only the archivist will enjoy space savings. All the readers who might consult your library in the infinite future will bear the cost. > ...(e.g. lossless FLAC). This inflates the file size... This is entirely the wrong view. The file size of a raw capture compressed to FLAC should be thought of as the “true” or “correct” size. It is roughly the most efficient (balancing various trade-offs) representation of sampled audio data that we can presently achieve. In preservation we seek to preserve the item or signal itself and not simply what we might perceive thereof. This human-centric perception view is just wrong. There is data in film photographs which cannot be perceived visually yet can be of interest to researchers and be revealed with digital image analysis tools. As an example of how much information celluloid can contain see: https://vimeo.com/89784677 https://vimeo.com/89784677 (context: he is comparing a Blu-ray and a scan of a 35mm print)
- lawrenceFounta 10mo ago[dead]
- iqandjoke 10mo agoThat’s why Spotify would lose against Apple. Spotify may need to pay a fortune for this scraper behaviour while Apple Music does not.
- throw-12-16 9mo agoI love coming to these threads to read the pearl clutching of "technologists" who suddenly care about IP and copyright law.
- bekindtoartists 9mo agoI’m hugely disappointed in Anna’s archive. As much as they believed they were doing this for good, they have now allowed bad faith actors to obtain all music for AI gen. This is just horrific for all artists out there who are fighting against so many issues that impact their creativity and sustainability. Why not just digest the data and not allow the music out there. As usual artists get fucked over.
- asacrowflies 9mo agoAny serious player with ai training already had this data. This is just evening the playing field .
- puffpuff12345 9mo agoAmazing! Is there any way to search this spotify database without downloading the currently available metadata torrent?
- udoyxyz 9mo agoyo, this is insane!! why would anyone do that? I think it is for AI music generation models, like training them. Maybe ai labs people did it?? yeah that is likely
- barbari04 9mo ago[dead]
- damnitbuilds 9mo agoWell done ! Until we have reasonable copyright terms, Pirate On !
- raducanu70e 9mo ago[dead]
- thenthenthen 9mo agoFull circle! Thank you! (https://torrentfreak.com/how-the-pirate-bay-helped-spotify-become-a-success-180319/ https://torrentfreak.com/how-the-pirate-bay-helped-spotify-b...)
- MightyHousewife 9mo ago[dead]
- performative 9mo agothis is a really incredible effort. but, for the developers and analysts currently working with music metadata in a world where so much of music is being consumed thru streaming services that keep a tight hold on how their metadata and album art can be used, i am constantly yearning for a way to link streaming releases to public metadata sources that can be manipulated, embedded, and queried. i've done my best to build my own w/o a background in data science, but it's a hole that desperately needs filling to enable the new generation of scrobbling/music listening habit exploration.
- none14988 9mo agoDownloading of individual files to Anna’s Archive Please
- none14988 9mo agoDownloading of individual files to Anna’s Archive Please!
- Varaldar 9mo agoim thinking about the consolidation around minute marks. its at every minute mark below 10 minutes, albeit dropping precipitously after 4 minutes. i have 2 guesses. guess one is that people like even numbers so if a track was already going to be within so many seconds of exactly a minute mark that they are more likely to push it to that number. with people caring less above 4 minutes because you are already making a long song, i could imagine caring less at that point. but my second guess is that along with the vast increase of ai slop posted to spotify both by spotify themselves and by other people, some of the programs they use probably fix on minute increments. like how a lot of ai videos are 10 seconds long or a series of 10 second videos. just a guess, however. i have no information or facts to back this up
- haryj 9mo agowow
- eastoncrafter 9mo agoPlans to upload all this to musicbrainz soundid program?
- eastoncrafter 9mo agoPlans to upload all of this to music brainz soundid?
- machloof 9mo agoThats huge, altho as a musician myself i am kinda scared of ai just taking all this data so they could make music better then me, i dunno maybe drop in there an anti ai trap zipbomb or somthing, that way it will work for normal users but not for ai
- pmnitin 9mo ago[dead]
- provokateur 9mo ago[dead]
- provokateur 9mo ago[dead]
- aftbit 9mo agoHas anyone tried to add up the track file size from the metadata dump? In spotify_clean_track_files.sqlite3: SELECT count(*), sum(filesize_bytes) FROM track_files; 255966403|15970064861274 That's only 14.5 TiB, nowhere near 300 TiB. What makes up the other 285 TiB of content?
- squigz 9mo agoThat's curious and changes things pretty dramatically. It's a lot easier to host 15TB than 300. I wonder what's up here.
- laamla4m 9mo ago[dead]
- baxuz 9mo ago> The quality is the original OGG Vorbis at 160kbit/s. Yeah, the original quality is either a 320kbps OGG or lossless. Not 160. While this is _a_ backup, it's a pretty lossy one.
- sma3in 9mo agospotify undressed
- Motorbytes 9mo agoDoes the Spotify backup contain any so far grayed out or unavailable songs on their list? I'm a music archivist & preservationist, I've archived and found several formerly lost or on the verge of becoming lost albums, EPs, and Singles, and I've been wondering if the backup of Spotify so far, even with the available info, contain any taken down, region limited, or no longer available songs? any response is appreciated!
- ewzimm 9mo agoThe data analysis here is interesting. One thing that stood out to me is that black metal is the 6th most common musical genre for bands, right after rockabilly. I would never have expected that.
- fungonimus 9mo agoI would like a downloader! :D this is such an awesome project
- soundsgoodman 9mo agoYou need to seriously re-think this... Releasing indie music, like really low-level indie music, for free in the name of "preservation" is so misguided. Don't do this. You will only end up hurting the artists who rely on paid downloads.
- kim100 9mo ago[flagged]
- Jumpmanlives 9mo agoGood stuff Anna's Archive. The Anchormen, premium sea shanty crew from Western Australia, officially endorsed you sharing our salty tunes. https://www.facebook.com/theanchormenwa/ https://www.facebook.com/theanchormenwa/ https://open.spotify.com/album/07IyzOA9jJWPZcLDysQwpo?si=KZOt3MvXQ9SPD7yIw8S16w https://open.spotify.com/album/07IyzOA9jJWPZcLDysQwpo?si=KZO...
- haghiri75 9mo agoI guess having an API to do search on metadata may be cool. Anyone thought of that?
- pekkag 9mo agoExtremely useful statistics. However, users need to know that IRSC codes are not really unique identifiers. The code was created to identify unique digital tracks (recordings). When older analog recordings (there are millions of them) the publisher assigns it an ISRC code, which shows the year of reissue. If the recording is in public domain, anyone can reissue it and assign it a new ISRC code. Even if the recording is still in copyright, the company can assign each new rerelease a new code - all with a different year. So be careful with interpreting statistics based on these codes.
- pranavm27 9mo agoMiss anna, next time please scale down image dimensions so that us on mobile can read properly haha Jokes aside, I always thought the best way to deal with piracy was to understand or convince the demand not to do it over dealing with the supply.
- lanalanabobana 9mo agothese guys are 100% selling that data to "AI" companies for thousands of dollars so the internet and world at large can get a little more shitty. awesome -_-
- new_hair 9mo agoRookie Question, but how do i access all this metadata especially in a cleaner way, or genre-wise for my project development.
- HawkEyeSpaceMan 9mo agoNot worth the risk imo. This might backfire at some point and ruin a good thing with the book libraries.