10 ms·
Anki SRS Algorithm : Spaced repetition explained with code
- nmca 4y agoI'm surprised that no spaced repetition systems seem to exploit relationships between facts/cards. One might imagine that capturing a graph of similar ideas would allow for better algorithms.
- Scaevolus 4y agoAnki supports linking cards when they're different facets of the same information (e.g., each language source/target for a given vocab word): https://faqs.ankiweb.net/linking-cards-together.html https://faqs.ankiweb.net/linking-cards-together.html More precision isn't that important-- at worst, you slightly over-review some information, but that's no different from using a flash card's information in normal life outside of a review period.
- _Algernon_ 4y agoThe Anki FAQ goes into the reasoning a bit: https://faqs.ankiweb.net/linking-cards-together.html https://faqs.ankiweb.net/linking-cards-together.html >Some people want to extend this link between arbitrary cards. They want to be able to tell Anki "after showing me this card, show me that card", or "don’t show me that card until I know this card well enough". This might sound like a nice idea in theory, but in practice it is not practical. >For one, unlike the sibling card case above, you would have to define all the relations yourself. Entering new notes into Anki would become a complicated process, as you’d have to search through the rest of the deck and assign relationships between the old and new material. >Secondly, remember that Anki is using an algorithm to determine when the optimum time to show you material again is. Adding constraints to card display that cause cards to display earlier or later than they were supposed to will make the spaced repetition system less effective, leading to more work than necessary, or forgotten cards. >The most effective way to use Anki is to make each note you see independent from other notes. Instead of trying to join similar words together, you’ll be better off if you can determine the differences between them. Synonyms are rarely completely interchangeable - they tend to have nuances attached, and it’s not unusual for a sentence to become strange if one synonym is replaced with another.
- nmca 4y agoThis seems to boil down to "it would be more complicated", mixed with "the trivial thing is a bad idea", both of which seem true. But I can imagine a more complicated scheme could be much better; done right.
- _Algernon_ 4y agoI agree. It makes sense for cards which link previously learned facts to only show up once the base facts have been learned, for example. But I'm not sure how to design a good UI for it. I can easily see it as resulting in too much friction so nobody uses it.
- s_m_t 4y agoDoesn't just putting the cards in order of concepts essentially accomplish this? When I was doing Anki I adjusted my new card rate so my retention rate stayed above 90%. If you do it this way then your chances of getting stuck in a certain areas of the dependency graph (if you were to model it) is pretty low. If it does end up happening then do some extra study on those concepts outside of SRS. This way you don't have to model the actual dependency graph. You really shouldn't be doing all your study in SRS anyways because so called "flashcard blindness" (i.e. only being able to recall the concepts in the context of your SRS program) is a real thing. Furthermore, in my experience, what really made Anki effective for me was to write my new cards just in time. Meaning no pre-made decks except for something like the primitives of what you are studying (i.e. symbol names, alphabets, extreme base level concepts, etc). When you write your cards just in time you can really connect the context of what you are learning to the concept. From this perspective it really seems strange to build an entire ontology based on dependency graphs of a concept you haven't even learned yet!
- type-r 4y agoI have found this to be a huge missing piece of any SRSs I've come across. Personally, I have a pipeline that looks at what knowledge I've already learned and only then reveals successive information. For example, I'm learning numbers in Spanish. First, I ask myself to go from doce => 12; then, 12 => doce. Then, once I have a few numbers in my memory, I can start to combine those into diez + dos = ? (answer doce). The idea is that this last example requires recalling three previously-learned Spanish numbers to arrive at the answer. This is a trivial example but this underlying structure of successive layers of knowledge shows up everywhere and is horribly under-exploited by our current learning systems. Programmatically this can be done. Manually, this stuff takes way too much work.
- DennisP 4y agoHow would you do it programmatically? Seems like you'd always need someone to give it the relationships.
- trane_project 4y agoYes. That's required. I think they meant that the program uses those relationships to automatically advance the student to the next topic when it makes sense.
- trane_project 4y agoCheck out my project: https://GitHub.com/trane-project/trane https://GitHub.com/trane-project/trane It does exactly this because I also thought this was a critical limitation of current systems. I have a feature coming out soon (already merged to master but I need to make a proper release) which should make it way easier to use with Markdown files acting as the front and back of flashcards that are automatically picked up when starting up the program.
- aporetics 4y agoIf you want to look at prior art, take a look at org-drill
- kashunstva 4y ago> I'm surprised that no spaced repetition systems seem to exploit relationships between facts/cards. In fact, the Anki community largely endorses as an article of faith the complete insular atomicity of cards. I’ve always treated that orthodoxy with skepticism because I’m pretty sure that’s not how my brain works. Nor anyone else’s, probably.
- aliceryhl 4y agoAn interesting example is chessable, which is a spaced repetition system for chess. It keeps track of each move in a sequence of moves separately, but every time it asks you to re-learn a specific move, you are asked to play out the whole sequence, instead of just being shown the position and asking you to play the next move.
- optionalsquid 4y agoWanikani [1], an SRS for learning Japanese kanji and vocabulary, kinda works like that, though on a simpler level than what I think you envision. Basically their system involves first learning to recognize common parts of kanji (radicals), then learning to recognize kanji made up of those radicals, and finally learning vocabulary that uses those kanji. Later items in the graph are unlocked by getting the previous items to a certain SRS stage. Additionally, the entire syllabus is divided into 60 discreet levels of (very) roughly the same number of items, the next level being unlocked by getting 90% of the current kanji to a certain SRS stage, which helps keep it all manageable. I tried using Anki before, but found the Wanikani "method" to work much better for me since I could constantly build on things I had learned previously. [1] https://wanikani.com/ https://wanikani.com/
- trane_project 4y agoShameless plug to my own project: https://GitHub.com/trane-project/trane https://GitHub.com/trane-project/trane Dependency relationships between exercises are core to its design. I tried to use Anki but this limitation is pretty baked into Anki and made it unusable for cases where there's a clear order in which things must be learned (music in my case).
- nmca 4y agoThis looks like a very cool project! The docs include this description of the core algorithm: > The space repetition algorithm in Trane is fairly simple and relies on computing a score for a given exercise based on previous trials rather than computing the optimal time at which the exercise needs to be presented again. This will most likely result in exercises being presented more often than they would in other spaced repetition software. Trane is not focused on memorization but on the repetition of individual skills until they are mastered, so I do not believe this to be a problem. Could you say more about it?
- trane_project 4y agoSure, here's the gist of how it works. - When Trane is asked for a batch of exercises, it performs a depth-first search of every branch in the dependency graph. It picks up exercises as it goes along, and stops exploring a branch when the dependencies of the last node are not met. - Each exercise is given a 1-5 score by the student and those scores are used to produce a combined score, rather than compute a date on which the exercise should be reviewed. Two reasons: - Trane is also meant to be used to learn skills, which do not fit the assumptions made by the SuperMemo algorithm. - I believe most of the gains come from regularly being reminded to perform an exercise, so even a simple scoring algorithm should get most of the gains. Eventually, I want to replace the existing scoring with something more complicated that also makes a distinction between memorization and skill exercises to get better results. But based on my testing, this works for the most part as is and the gains to be made from more complicated algorithms are marginal. - Once you get exercises from performing the search, you reduce that number into the final batch by grouping exercises based on difficulty and selecting a percentage from each bucket. This is done to make sure not too many very difficult or easy exercises are included in each batch. There's also some elements of randomness added to prevent the same exercises from appearing all the time. For example, the branches are explored in random order and the elements from each difficulty bucket are selected at random. But that is the gist of it.
- majikaja 4y agoIn my personal system, I tag facts by sub-topic and the daily reps are ordered so that facts of the same sub-topic are shown consecutively to avoid context switching (eg. AAABBCCCCD not CABCCDBACA). In addition, there is logic so that after the mandatory daily reps are done, it will show cards that are close to the due-date (eg. within 0.25*LastInterval), again batched by sub-topic. This lets me efficiently get reviews out of the way on days when I have time to do so, without deviating too far from the basic SRS philosophy. As a result, for mature cards, I end up focusing on a small number of sub-topics on any particular day. I find that works a lot better. All of this is trivial to implement - I think it's best to create from scratch and specialize it for one's own domain instead of using a pre-existing solution. In my case the relationships are straightforward (sub-topics) but in others it might be more fuzzy. Programmatic fact-generation/deletion is also something I put a lot of effort into (again, domain-specific).
- tester457 4y agoMochi does this. It uses markdown too and it's a pretty modern app. https://mochi.cards/ https://mochi.cards/ The dev is knubie on this forum
- dandiep 4y agoI have been working on a form of this for a language learning app [1]. The basic idea is that it gets you do free form output (discussion questions, word games, etc), gives you feedback on where you have errors and then sets up SR flashcards automatically (e.g. if you don’t know a word, conjugate it wrong, etc). It uses embeddings to make inferences across exercises and flashcards. This is handy for: - detecting that you already have a flashcard for that particular word/phrase/grammar issue - you know one sense of a word but not another - you correctly used a word that you have a flashcard for and that review can be put off It can look at your vocab and grammar skills of a particular area overall using Rasch scoring. Perhaps you answer almost all the questions about “cooking” correctly but none about “parts of the body” - well then the app can focus on the second area. There is also work to be done to build out better regressions given this data. If a word I am studying is way above my skill level, I will probably forget it faster. Or maybe I forgot a conjugation which is very common. It should be easier to remember because I’ll encounter it more. Whether or not my app succeeds, I think these types of approaches hold a ton of promise to increase learning efficiency. 1. https://squidgies.app https://squidgies.app
- steve1977 4y agoSuperMemo does (in a way): https://help.supermemo.org/wiki/Neural_creativity https://help.supermemo.org/wiki/Neural_creativity
- codetrotter 4y agoThis is very nice and thorough.
- Syzygies 4y ago[flagged]
- tsumnia 4y agoThe primary hurdle is the number of "activity types" practice can come from. Simply reviewing material isn't sufficient for learning and so we need more engaging practice as well. Also, similar to how "all students are different", topics encompass a large number of subskills that need to be trained and not all activity types target those skills. Also, "learning" something is quite complex - is simple repetition of the fact sufficient, or do we need the ability to apply the skill? Its actually research we're doing at NC State. At present, we understand spaced, deliberate practice is beneficial [1] but students will largely reject AI recommendations to complete lower-level practice [2]. So our current methods are to develop a system that maintains an "instructor in the loop" to create tailored training regimens (my terminology for it), similar to how a coach builds regimens for athletes. The hope is that if we can analyze how instructors made their recommendations, we could then design AI to mirror the decisions. [1] https://dl.acm.org/doi/pdf/10.1145/3373165.3373177 https://dl.acm.org/doi/pdf/10.1145/3373165.3373177 [2] https://repository.lib.ncsu.edu/bitstream/handle/1840.20/39822/etd.pdf https://repository.lib.ncsu.edu/bitstream/handle/1840.20/398...
- Syzygies 4y agoThank you. I suppose I was downvoted, hiding your careful response, because I appeared to be pitching a product. I was attempting a precise analogy, I had nothing to gain from the promotion. I was robbed on BART recently, and my judgment is off. I've relied heavily on spaced practice for learning languages for travel. There are two phases to that experience: Study, where one takes in as much as possible, then real world practice without further study. I find myself wishing I'd memorized a play for the first phase, to access faster than a phrase book on the ground, with lots of pathway shortening while I travel. Learning models tend to assume learning is single-phased, though even more academic learning is two phased: acquire, then later apply.
- tsumnia 4y ago
- GenericDev 4y agoI'm so thrilled to have come across this. Thank you for sharing this OP!
- fallat 4y agoFor anyone interested: https://len.falken.directory/code/sm2.git/ https://len.falken.directory/code/sm2.git/ if you want to run something locally pretty easily. This post hits exactly what I found when researching months ago :) Absolutely captures everything well. This post has convinced me to even use Anki's algorithm.
- daniels11 4y agoNice article! This would've been very helpful a few years back when I was attempting to make an Anki clone while teaching myself to code. I'd still love to create a better open-source SRS algorithm at some point in the future. Mostly to use for language learning. When I was learning Chinese characters, I did a bit of a deep dive into the Anki algorithm and found that the biggest flaw (imho) was the Ease Factor knocking a card back to 0. If you mostly know a card, but miss it on one day for some reason - maybe you were tired or distracted, that card's learning progress should not be totally reset (as if you were learning it from scratch). That leads you to have too many cards to review on a daily basis. Instead, you should use some modifier to increase the interval to a reasonable level. More explanation here: https://readbroca.com/anki/ease-hell/ https://readbroca.com/anki/ease-hell/ I think it would be awesome to pair SRS with high quality images and audio, which I find most helpful for language learning. I've used Rosetta Stone and Duolingo in the past; Rosetta Stone has great audio and images but lacks a powerful SRS (it also has a number of other flaws in my mind, but I'll save that for another time). Duolingo is great for grammar and explanations, but I can't take the pronunciations and tediousness of it all.
- Anon1096 4y agoYou're describing a separate problem than ease hell, what you are looking for is fixed just by changing the lapses->new interval setting to be from 0% to something higher like 50%. I do agree that the default 0 is pretty bad though.
- imron 4y agoIf you’re interested in SRS algorithms you should visit the site of the guy who basically invented them [0]. Most SRS algorithms seen in apps today are based on one of the super memo algorithms in one shape or form. Anki was based on SM2, which was already outdated in terms of supermemo algorithms when released, but was chosen due to its simplicity. 0: https://supermemo.guru/wiki/SuperMemo_Guru#Memory https://supermemo.guru/wiki/SuperMemo_Guru#Memory
- alwayslikethis 4y agoSuperMemo is proprietary software, and only works on Windows. Anki's algorithm is deemed adequate by most, and the benefits of newer algorithms which introduce more complexity are marginal/questionable.
- jakobov 4y agoShameless plug: I developed a space repetition algorithm which has a different goal. Mine is focused on learning a large number of new things fast, rather than adding new things over time and retaining memory of them. See: https://github.com/Jakobovski/SaneMemo https://github.com/Jakobovski/SaneMemo
- googlryas 4y agoI don't quite understand how that is different? Is the idea that you have nebulous knowledge of more things instead of locked in knowledge of fewer things? What's your use case for this? Would it be good for learning say the top 5000 words in a new language or something like that?
- ATMLOTTOBEER 4y agoI think the idea is that this would be more optimized for school settings where what you actually want is an algorithm for cramming.
- _Algernon_ 4y agoIt could make sense to cram for an exam for instance
- deepu256 4y agoAwesome. I just read through the code & Blog Post and the changes make so much sense as well as the code is super easy to read and understand. Would be awesome to read part-2 of this blog post -> http://www.blueraja.com/blog/477/a-better-spaced-repetition-learning-algorithm-sm2 http://www.blueraja.com/blog/477/a-better-spaced-repetition-... Plz plz plz pen down your thoughts on reviewing(part-2 of the post) if you get some free time.
- cmehdy 4y agoThe article is great - a visual that explains SRS in one single graph near the top of the page is a great proof that the article is solid. But I also want to say: what a clean and enjoyable website!! Thank you for having shared this.
- Kelamir 4y agoI've noticed this https://github.com/open-spaced-repetition/fsrs4anki https://github.com/open-spaced-repetition/fsrs4anki set of modifications for the Anki algorithm is popular in the Anki Discord. Might be of interest to those exploring the code.
- _-____-_ 4y agoI enjoy using Anki and have found it effective in the very rare instances where I need to rapidly memorize a lot of material. But as a developer, I've never found a need for memorization. Am I missing out? Do any developers have suggestions for how to utilize Anki to make myself more professionally efficient? I suppose one starting point might be to track my most googled phrases ("trim newlines in sed," "bash printf into a variable," etc.) and turn those into cards.
- nyadesu 4y agoI think it's meant for memorization, so ideal for language learning and raw knowledge absorption (like in academic courses)
- tmtvl 4y agoWhen it comes to programming I consider SRS useful for things like the tradeoffs and characteristics of various data structures, modelling techniques, and algorithms. Things like code snippets (including shell invocations) I prefer to simply keep in a personal library because those don't require as much innate knowledge. Compare: Oh, I have to remove this user from all the services this machine provides? I've got a script for that. To: We need to sort all the users by how much they've spent in purchases? A heap sort has a fairly decent worst case but doesn't parallelize all that well and if the users are already mostly sorted it doesn't get much advantage, maybe merge sort would be better...
- lowken 4y agoI use Anki every day and it helps in the following ways... 1. It's rare that I have to use Google figure out how to do basic tasks in my Programming Languages of choice (C#, & TSQL). Don't get me wrong I, and everyone on the planet, looks up code on the internet. It's just a more rare scenario for me. 2. I write much quicker then your average developer. Of course most of software development is design and requirements gathering/clarification but when I do get down to writing code I'm very fast. 3. I spot bugs almost immediately because when I go through my Anki deck I get repeated exposure do different error scenarios. In fact I can often simply think about the scenario and it will occur to me where the bug is and what is going on. I attribute this to using Anki on a daily basis. 4. I can read code much better then before I started using Anki. I know at a glance what code is part of the language/framework vs what is custom code for the application.
- ZhangSWEFAANG 4y agoOne underrated benefit of using Anki is you become a god. In a meeting, when I can recite the Q1 revenue along with the names of everyone working on the project, everyone automatically becomes submissive to me because they think they have just observed a superior intellect.
- dyno12345 4y agowait so how do you decide what minutia to try to memorize?
- yosito 4y agoI've used Anki to learn languages with SRS for years. I keep running into the same problem, even with difficult languages: my knowledge of the language quickly outpaces the rate at which I can study new cards, and my ability to retain vocabulary doesn't really seem to fade. Once I learn words, I don't need to review them in x months or years. I just know them and, for the most part, won't forget. I learned Hungarian (one of the more difficult languages to learn) by loading up a database of ~10k cards. While it worked great in the beginning, within a year, I was just constantly telling the algorithm that I didn't need to review cards for 3-5 years. But the thing is, with a language, I'm not going to forget all those words in 5 years. There will be no point in reviewing them. So I just started deleting cards if I knew them already. And eventually I was spending so much time deleting cards that I just gave up on SRS altogether. It could be that I have some latent savant-like language learning abilities, though I doubt that. I think my experience is likely related to language learning itself. With a language, if you are immersed, you have constant daily repetition with or without the cards. All that to say, I think SRS algorithms should have a language learning mode that automatically archives vocabulary cards once you've mastered them to a certain degree. I'm not sure what would work best. Maybe once you've crossed the 6 month threshold, it could just automatically archive the card.
- TheCowboy 4y agoLanguage learning has a lot of quirks and this just isn't the case for most people who didn't learn a language early on, and if you're immersed/actively using the language then you're way less likely to forget. Your daily usage and exposure is simply your spaced repetition. If one is bilingual at an early age, then it's much easier to learn/retain additional languages. Most language learning in America starts kids off much later than it should (middle or high school). Full immersion is also a great way to learn. And as far as SRS goes, it's important to not make cards if you think they're not necessary as it's easy to just make the whole review process too much to chew on. And sometimes you can cluster information in a single card (maybe a sentence that's an idiom if using language as an example) such that you do get continued exposure, and can decide later if you need to break it up if you can't recall some parts of it.
- deleted 4y ago[deleted]
- jacquesm 4y agoThis has been up for 10 hours and I don't see Gwern's article mentioned yet, so here it is: https://www.gwern.net/Spaced-repetition https://www.gwern.net/Spaced-repetition
- siduck 4y agoBookmarked
- ukoki 4y agoNice! When I was building a spaced repetition toy a few years ago I took a different 'bottom-up' approach, where I dynamically fit the forgetting curve to the user's actual performance on any given card to come up with a unique memory half-life value for that card for that user, which was then used to schedule the next review at a given estimated recall % (typically 70-95% depending on whether the user wanted an "easy" or "hard" experience). As you need at least two data points to be able to fit the curve, you never have enough data to schedule the second review of a given card, so for this I used the average half-life accross the set of cards, or hard-coded values if it was the first item in the group. This way a set of cards (such as world flags) was always initially scheduled as though they were similarly difficult for the user, and then harder cards within the set would naturally be reviewed with a higher priority.
- adaszko 4y agoThere’s a Bayesian stats approach to spaced repetition: https://fasiha.github.io/ebisu/ https://fasiha.github.io/ebisu/ AFAIU, SM2 computes the datetime of a next review, whereas Ebisu models a probability of remembering a given flashcard. It seems it’s a more straightforward representation that’s more amenable to implementing functionalities like “show me 10 least remembered cards”.
- cyphar 4y agoIt should be noted that Ebisu currently has a severe flaw that causes it to massively underestimate how well you know a card[1]. The author has been working on a rewrite to fix the issue for a year and a half, and is currently asking for comments on the new API for Ebisu v3 which should fix this issue[2]. (The author argues this never would've caused to you to get suboptimal reviews if you use it in the "just give me a few reviews now" mode, but I'm not personally convinced.) All of that being said, I am a fan of the Bayesian approach Ebisu uses (and Ebisu v3 would allow you to eliminate the awful way SM2/Anki deal with ease factors). Thankfully the v3 Anki scheduler will be pluggable so you'll be able to use Ebisu v3 with Anki on every Anki platform. :D [1]: https://github.com/fasiha/ebisu/issues/43 https://github.com/fasiha/ebisu/issues/43 [2]: https://github.com/fasiha/ebisu/issues/58 https://github.com/fasiha/ebisu/issues/58
- adaszko 4y agoOh, interesting. Is v3 just a new API or a new algorithm altogether?
- cyphar 4y agoIt's still Anki's version of SM-2, with minor changes to its behaviour. You can see a list of the changes here[1]. [1]: https://faqs.ankiweb.net/the-2021-scheduler.html https://faqs.ankiweb.net/the-2021-scheduler.html
- c7b 4y agoReminds me of one of my backlogged projects: I'd like to have an SRS app that just draws cards at random, with probabilities proportional to how 'urgently' you need to review the cards (using the same logic of exponential growth in review intervals). I just couldn't get myself to stick with Anki so far, because I didn't like the 'X cards to review today' UX of the SM2 algorithm. I also think it's false accuracy trying to predict the exact review time down to the minute, I couldn't find much research supporting that any of the SM family algorithms would be the claimed optimum. If that already exists, ideally compatible with Anki cards, I'd be happy for recommendations. I have other projects and ideas that I feel would add more to the world than the umpteenth Anki-clone.