4 ms·
Do you have any code that demonstrates this? Sounds super interesting!
by LewisDavidson 3y ago
Do you have any code that demonstrates this? Sounds super interesting!
- mynegation 3y agoThey do but it’s probably… classified.
- 19h 3y agoUnfortunately, I can't. We have some projects bubbling around that may see the light of the day eventually but given the myriads of NDAs that stack on top of each other this is rather unlikely. That said, here's some reading material on the underlying ideas: - https://en.wikipedia.org/wiki/Semantic_folding https://en.wikipedia.org/wiki/Semantic_folding - https://arxiv.org/pdf/1511.08855.pdf https://arxiv.org/pdf/1511.08855.pdf ("Semantic Folding Theory And its Application in Semantic Fingerprinting") This is _not_ TF-IDF. Once you have built the "relation fingerprints" of each word, the fingerprint lookup complexity is o(1) as you'll essentially only load a massive LUT of type HashMap<String, Vec<u16>> (or u32 if you go above 255*255). [pro tip: our LUT has the type HashMap<Vec<String>, Vec<u16>> as our impl also considers bigrams, trigrams, quadgrams] Unfortunately I can't get extremely specific, but we're also feeding these [u8; 16348] vecs into an HTM w/ spatial pooler; feeding one word-SDR aka fingerprint into the HTM at a time allows you to leverage the HTM to make predictions for the most likely next word-SDR aka the fingerprint of the next word -- if you generalise this on a sentence level, you can use the cosine distance between the actual text-SDR aka fingerprint of the next sentence and the predicted text-SDR out of the HTM to semantically segment paragraphs in a continuous stream of text. This allows us to segment SOCMINT user2user conversations into individual semantically connected packages of text / messages that can be marked by scenario-specific heuristics to be additionally analysed by a downstream system.