7 ms·
I assume one reason Apple has made it more challenging to extract the dictionary resources is in order to satisfy licensing constraints with the dictionary auth
by gfaure 5y ago
I assume one reason Apple has made it more challenging to extract the dictionary resources is in order to satisfy licensing constraints with the dictionary authors. I wonder if they'd block an app like this through the App Store submission process, if submitted.
- macintux 5y agoI’d definitely assume this would be a copyright problem if used in an app.
- peterburkimsher 5y agoYes, a lot of people involved in dictionary processing are worried about copyright! After I wrote my post about Apple's dictionary files, I got a mysterious email showing up in my inbox. The email was from someone who's spent some time writing code to do the same thing, but doesn't want to post it under his own name in case he falls fowl of his country's DMCA equivalent. Crazy. He said I could post his code under the condition that I took his name off it. https://josephg.com/blog/apple-dictionaries-part-2/ https://josephg.com/blog/apple-dictionaries-part-2/
- tkgally 5y agoI have helped with the writing and editing of a number of dictionaries over the years. It's difficult, highly specialized work. For some of the jobs, the only compensation has been a share of the royalties from future sales. I doubt if many people like me are enthusiastic about the dictionary data becoming accessible for free.
- tasogare 5y agoYes, dictionary content is the revenue source of dictionary vendors, so of course they don't want anyone to use it without permission. On the other hand there are more and more open-data projects (I started one myself), often based on printed dictionaries that felt in the public domain.
- simondotau 5y agoWhat’s the state of Wiktionary like in your opinion?
- tasogare 5y agoI don't really use it often as a user nor i my projects to have a definite opinion. There is some pairs of words (about 5K) in Sino-Vietnamese that came with their chu nom writing which was very helpful to one of project. Otherwise I think it lacks structure and can't be harvested automatically easily (I don't think Wikidata integrate it all, and that website is a non-starter for me). Also every language is structured differently so Wiktionary can hardly be commented as a whole.
- thombles 5y ago> Otherwise I think it lacks structure and can't be harvested automatically easily Indeed, it depends on the language and your goals - I had a very high success rate plucking out Russian grammatical tables from English Wiktionary with a few hours of scripting the data cleaning (https://github.com/thombles/declensions https://github.com/thombles/declensions). I have a theory that you could get better results using an offline archive of the page sources but haven't tried this yet.
- Someone 5y agohttps://en.wiktionary.org/wiki/Help:FAQ https://en.wiktionary.org/wiki/Help:FAQ: Q: Is it possible to download Wiktionary? A: Yes. https://dumps.wikimedia.org/enwiktionary/ https://dumps.wikimedia.org/enwiktionary/ should have the latest copy of the main namespace. The cleanest navigation page is https://dumps.wikimedia.org/ https://dumps.wikimedia.org/. Just download a -articles.xml.bz2 file and some software to read it (for nix, for Windows). Q: Can I use data from Wiktionary in my program? A: As long as you meet the conditions of the GNU Free Documentation License or Creative Commons Attribution/Share-Alike License, certainly. Latest dump for English is from September 1. I wouldn’t know whether it has all the data or how easy it is to parse it.
- tkgally 5y ago
- dhosek 5y agoI have a vague recollection of reading somewhere that you're explicitly forbidden from creating a dictionary app using the OS dictionaries. You can do dictionary lookup within apps (so, for example, you're free to use the dictionary to look up definitions in your word processor or ePub app, but not to have an app which lets a user enter a word and get a definition back).
- avianlyric 5y agoI suspect it’s more likely they have a bunch of internal frameworks for creating, accessing, and distributing simple databases (something akin to SQLite). So the difficulty we’re seeing here isn’t deliberate obfuscation, but rather just an dev using a database structure to make app design and word lookup easier. If there really was licensing concerns I would expect there to at-least be some basic encryption, not just some basic compression.
- simonw 5y agoApple use regular SQLite for a whole bunch of other applications, such as Apple Photos. I wonder why they didn't use it here?
- spitfire 5y agoThe dictionary app (and hence its database) may go all the way back to NeXTStep/Openstep. Maybe not all of it. But I wouldn't doubt if bits of it went all the way back to the early days where a fancy dictionary was one of the star features of NeXTStep.
- simondotau 5y agoThe underlying file format and associated libraries are probably largely unchanged from the NeXTStep era. If it ain’t broke, don’t fix it. http://toastytech.com/guis/ns20mail.png http://toastytech.com/guis/ns20mail.png (Edit: beaten by 30 seconds!)
- jahewson 5y agoThat’s highly unlikely. The author is incorrect about it being a Zip file, it’s actually a simple gzip/zlib stream. Because it’s not possible to seek within such a stream, they are always chunked in any file that requires random access. Elsewhere there will be an index file which maps words to their corresponding chunks, so that the definition can be quickly loaded without having to decompress the entire file - or even load it all into memory. This is very normal stuff in the world of file formats.