23 ms·
Author here, thank you so much! I really tried to make it nice to use, beyond the (quite original) modelling. A prototype I did tried to detect some grammatica
by Labo333 1mo ago
Author here, thank you so much! I really tried to make it nice to use, beyond the (quite original) modelling.
A prototype I did tried to detect some grammatical constructions, eg "it's not ..., it's ...", but I am not sure how to systematize that.
Also just a disclaimer: I am NOT tracking Claude tics, I am merely finding that a particular cluster of vocabulary increases. Tracking Claude requires labelled data IMO. I tried using model release dates in a structural model to constraint the clusters but the result was not compelling, so I ended up simplifying the model a lot!
- dmd 1mo agoI love this. I think you should be clearer about that - it's obvious to people who have experience doing this kind of data analysis, but a lot of people I've shown this were confused, as in "huh, load-bearing and genuinely makes sense but how'd he pick all these other words!?" I think it should be made clear that the sorting is sort of Texas Sharpshooter-ish - the ones on top are on top because they sort that way. The fact that load-bearing ends up on top is the proof that this works because we all know a priori that load-bearing is a Claudism.