5 ms·
We’re classifying gigabytes of intel (SOCMINT / HUMINT) per second and found semantic folding or better in classification quality vs throughput than BERT / LLMs
by 19h 3y ago
We’re classifying gigabytes of intel (SOCMINT / HUMINT) per second and found semantic folding or better in classification quality vs throughput than BERT / LLMs.
How it works — imagine you’re having these sentences:
“Acorn is a tree” and “acorn is an app”
You essentially keep record of all word to word relations internal to a sentence:
- acorn: is, a, an, app, tree
Etc.
Now you repeat this for a few gigabytes of text. You’ll end up with a huge map of “word connections”.
You now take the top X words that other words connect to (I.e. 16384). Then you create a vector of 16384 connections, where each word is encoded as 1,0,1,0,1,0,0,0, … (1 is the most connected to word, 0 the second, etc. 1 indicates “is connected” and 0 indicates “no such connection).
You’ll end up with a vector that has a lot of zeroes — you can now sparsify it (I.e. store only the positions of the ones).
You essentially have fingerprints now — what you can do now is to generate fingerprints of entire sentences, paragraphs and texts. Remove the fingerprints of the most common words like “is”, “in”, “a”, “the” etc. and you’ll have a “semantic fingerprint”. Now if you take a lot of example texts and generate fingerprints off it, you can end up with a very small amount of “indices” like maybe 10 numbers that are enough to very reliably identify texts of a specific topic.
Sorry, couldn’t be too specific as I’m on the go - if you’re interested drop me a mail.
We’re using this to categorize literally tens of gigabytes per second with 92% precision into more than 72 categories.
- lgas 3y agoNot to dogpile on all the other "isn't this just" messages, but isn't this just sparse embeddings?
- lmeyerov 3y agoYeah I'm struggling to understand at a fundamental level how this is better both in math + engineering, doubly so by the time you get to sentence embeddings . (Genuinely, it seems to use the same ideas, so curious what the specific trick is vs mature embedding packagings already doing much of this afaict.)
- mmcwilliams 3y agoYou're not wrong. This sounds curiously close to the ways I've seen word2vec used in production.
- wavemode 3y agoI'd be curious how the output of your approach compares to merely classifying based on what keywords are contained in the text (given that AFAICT you're simply categorizing rather than trying to extract precise meaning).
- mistymountains 3y agoIt’s the same as a giant one hot vector. He’s not describing anything terribly new or impressive, but if it works then god bless and good luck.
- spyckie2 3y agoJust asking, this seems very similar to the attention algorithm that powers LLMs?
- mistymountains 3y agoIt’s not similar other than that attention relates tokens.
- LewisDavidson 3y agoDo you have any code that demonstrates this? Sounds super interesting!
- mynegation 3y agoThey do but it’s probably… classified.
- 19h 3y agoUnfortunately, I can't. We have some projects bubbling around that may see the light of the day eventually but given the myriads of NDAs that stack on top of each other this is rather unlikely. That said, here's some reading material on the underlying ideas: - https://en.wikipedia.org/wiki/Semantic_folding https://en.wikipedia.org/wiki/Semantic_folding - https://arxiv.org/pdf/1511.08855.pdf https://arxiv.org/pdf/1511.08855.pdf ("Semantic Folding Theory And its Application in Semantic Fingerprinting") This is _not_ TF-IDF. Once you have built the "relation fingerprints" of each word, the fingerprint lookup complexity is o(1) as you'll essentially only load a massive LUT of type HashMap<String, Vec<u16>> (or u32 if you go above 255*255). [pro tip: our LUT has the type HashMap<Vec<String>, Vec<u16>> as our impl also considers bigrams, trigrams, quadgrams] Unfortunately I can't get extremely specific, but we're also feeding these [u8; 16348] vecs into an HTM w/ spatial pooler; feeding one word-SDR aka fingerprint into the HTM at a time allows you to leverage the HTM to make predictions for the most likely next word-SDR aka the fingerprint of the next word -- if you generalise this on a sentence level, you can use the cosine distance between the actual text-SDR aka fingerprint of the next sentence and the predicted text-SDR out of the HTM to semantically segment paragraphs in a continuous stream of text. This allows us to segment SOCMINT user2user conversations into individual semantically connected packages of text / messages that can be marked by scenario-specific heuristics to be additionally analysed by a downstream system.
- SomewhatLikely 3y agoSounds like TF-IDF vectors.
- espe 3y agovery efficient but also brittle. that must be vast amounts of relatively clean data. you have to magically set the number of top n words to in- and exclude. for most user generated content one would need to heavily normalize the text, e.g. by stemming (to keep in line with the computational austerity). 16384 is very little even if it is neatly seperated concepts. applied to that volume of data it should amount to keyword matching.. that only works if users are basically self-tagging their texts via constrained language use. edit: short version: not semantics and not a fingerprint :)
- 19h 3y agoWe also trained on all of pushshift and have an average ”unknown” word rate of less than 0.007% — the Reddit corpus is rather amazing to capture pretty much all misspellings of a word. We may only be using 16k vector values but that doesn’t mean we only have a vocab of 16k —- our vocab is more around 1.9 million words each described by a sparse fingerprint of 16k.
- espe 3y agothanks for the clarification. if your base population is that large then it's frequencies and you get a fingerprint. well done.
- foolswisdom 3y agoI'm curious though, how do you handle related forms of a word (assuming you don't use stemming)? It doesn't seem to me that this process would automatically handle that.
- mistrial9 3y agoamazing that this streaming pile of characters and its uncreative associations with three-letter-agency code names, results in exactly ninety two percent accuracy.. almost like its profoundly wrong in exactly the most important ways
- 19h 3y agoCare to elaborate? Not sure why this tone is appropriate. The 92% is an average and not the exact accuracy across all categories; the accuracy varies by category as every category is represented by its own filter.
- dr_kiszonka 3y agoIf I understand your approach correctly, you could represent relations between words as graphs and use graph/network similarity measures (of which there are tons) to possibly get over the 92%. (Or not, I have never tried it.)
- 19h 3y agoInteresting idea! Can you elaborate a bit more?