7 ms·
A new semantic chunking approach for RAG
- refulgentis 2y agoIt'd be interesting to see examples involving the "G" in RAG: I'm wary that having a fine-grained breakdown of, say, 2 vs. 5 sentences, would affect things significantly. If I'm understanding correctly, you end up with a "turtles all the way down" problem: you need to chunk using some size = K, in order to get the inputs you need to decide bespoke chunk sizes X, Y, and Z.
- bbor 2y agoYeah but this becomes very, very relevant for larger corpuses. It’s not just about pulling the right sized snippet in the end (really that could be done by an intermediate LLM call…), it’s also about supporting effective vector similarity searches to find the right snippet in the first place. If you’re embedding all five pages into a single vector, and only two of them are relevant to your query, that makes the search much less effective. AFAIU.
- refulgentis 2y agoExactly: you have to do all the small embeddings to get the larger embeddings. Also, in practice, embedding are much smaller than that, the popular MiniLM v2 is 256 words. Leaving me unsure how this helps
- lmeyerov 2y agoThis is fun. At a fractal level, it's what's going on in graph RAG, where you're doing hierarchical similarity modeling and using that structure in different ways for smarter summarization, indexing, & retrieval.
- kewp 2y agoWhat is RAG?
- Sai_Praneeth 2y agoRetrieval Augmented Generation. Fancy way of saying, retrieve chunks from your document corpus similar to your input using a similarity (mostly cosine) of embedding vectors of the chunks and input vectors, then inject those relevant chunks into your prompt to the LLM. Useful for Document Intelligence.
- kewp 2y agoThank you for the explanation, I remember hearing about this now
- daemonologist 2y agoRetrieval Augmented Generation - using search (usually with some kind of semantic component) to find relevant context and provide it to the language model to help it respond, give it knowledge about a specific document, etc.
- kewp 2y agoOh right, I remember now. Thank you for explaining it
- pamelafox 2y agoHow does your chunker actually work? The gist only shows an API call. Or is that secret?
- nutanc 2y agoWill be sharing the code next week. The basic idea is to find the embeddings of sentences and then finding the distance in the latent space to see if there is too big a jump in context.
- magicalhippo 2y agoA single sentence can have a lot of different meanings. I wonder if it would be more fruitful to use rolling pairs and triplets? Ie was thinking about how you use two moving averages with different window sizes to detect trend shifts.
- nutanc 2y agoYes. We are using rolling sentences. Not just pairs. But the whole story. I have written a little about this in my previous post, shape of stories, https://gpt3experiments.substack.com/p/the-shape-of-stories-or-how-ai-sees https://gpt3experiments.substack.com/p/the-shape-of-stories-... So we can see how as we are appending sentences and getting the embedding we get the movement of the story in the latent space.
- ravishar313 2y agoHow does it work? I want to know how it works!!!
- bbor 2y agoWow, this is absolutely deviously simple, and drop-dead obvious in hindsight. Unless this person is copying someone else without knowing it, this little blog post is the most important LLM paper to come out in months! Assuming I understood it properly, ofc. Some random musings off the top of my head on quantitative analysis of tokenized text: 1. Could you enhance this by generating piecemeal summaries and comparing those to find chunks? Presumably it would decrease noise. 2. If we’re dealing with graphable data, can’t we just apply some old-school clustering algorithms to it? Say, KNN? 3. Idk if this is implied or not already these days, but can you filter out incidental words before tokenizing to get a cleaner analysis? I.e. take out the prepositions to better highlight the unique words for each chunk?
- nutanc 2y agoHey thanks. Yeah, even we felt it was very obvious and why no one has done this before :) Will share more details soon.
- Ey7NFZ3P0nzAe 2y agoBtw, i'm pretty sure we can use this on spoken output of schizophrenics to categorize optimal first line treatment.
- yawnxyz 2y agoI wonder if this only works for stories with narratives or will it break down for other kinds of reports? I guess even research papers have "narratives"
- extr 2y agoVery cool! I'm wondering how to apply it to the type of RAG problems I usually encounter which are more technical in nature. Specification A refers to Specification C which has references to X, Y, and Z, but only X is really important context. You can traverse the tree, but how to select the important nodes? Semantic similarity is not appropriate since the specifications refer to different components of a system and are not semantically similar at all.
- throwaway314155 2y agoedit: Some of this language was a tad harsh and cynical. Apologies, was in a bit of a mood. I'll leave it as-is for posterity. > We have experimented with this and are releasing an API that you can explore to see if this chunking strategy works for you. Call me when you release literally any details at all. Even better - call when you've open sourced your code. It's kind of hard to just take your word for it when it comes to a alleged superior method. Demo's are interesting but again, I don't really trust that even those results do what you claim (semantic chunking). The best way to do that is with a rigorous evaluation and probably good old fashioned "read and compare" by a human being. Somehow, I doubt that this UMAP technique always maps directly to semantically important aspects of a narrative, at least not in a way that is substantially more accurate and useful than previous methods. I'd be happy to be proven wrong, but again this blog post is _extremely_ light on details. Hell, even saying "we don't want to release our algorithm because of the competitive landscape" as OpenAI does would be more useful than this.
- nutanc 2y agoHey, no problem on the language. Always happy for any thoughts :) We will open source next week. A couple of corrections: I never meant to say our method was superior. We have not done any benchmarks etc enough to write a paper. We have just been using this approach in our RAG pipeline and we have been happy with it. Sorry, will add a disclaimer in the post. Agreed. For semantic chunking, only read and compare works best. Not sure how to scale that to a bench mark. Guess for the PG essay, only PG can answer if the topics align with his thoughts :)[https://x.com/nutanc/status/1838813258972549549 https://x.com/nutanc/status/1838813258972549549] We will release the algorithm. I just put this out on Hackernews thinking this will also not be noticed as my all other posts :) But looks like some interest is there. Will post a follow up update soon. Sorry for the lack of details on the original post.
- jejwusud 2y ago[flagged]
- postalcoder 2y agoDo I smell astroturfing here? Ton of comments, lots of fascination and compliments glazing a post that has a few hundred words saying nothing. On top of that, I don't see how using a 2D reduction of embeddings is any more rigorous than the current "semantic chunking" approach (which I am not a fan of).
- nutanc 2y agoNo astroturfing here. Atleast I am not doing it. If somebody is doing it on my behalf, why? 2d reduction is only to visualize the flow of an article. When we chunk, we chunk on the full dimensions. Sorry, the article does not have more details. My bad. Will add more details along with the code doing this also.
- ofou 2y agoThis is an output example from the raw transcription of "10 Programmer Stereotypes" (https://www.youtube.com/watch?v=_k-F-MMvQV4 https://www.youtube.com/watch?v=_k-F-MMvQV4) [ "the programmer an offshoot of the great ape family closely related to chimps and gorillas distinguished by its minimal bipedal movement and ability to stare at a computer screen for the majority of its lifetime there's an estimated 30 million specimens alive in the world today normal humans use stereotypes to help understand and generalize this unusual variant which experts estimate are about 99 accurate about 12 of the time in today's video we'll take a look at 10 different programmer stereotypes to find out which one you fall into first up we have the gear head this variant owns the bleeding edge version of everything like the latest m1 mac a big ass curved monitor mechanical keyboard tesla in the garage ai generated synthetic meat in the fridge and a smart lock on the house to keep it all safe when programming he goes wherever the hype train takes him in 96 it was java in o6 it was jquery in 2016 it was graphql and in 2026 he'll be first in line at neural link to get a chip that can help him write blazingly fast code it doesn't matter what the tech does if it's trendy it belongs in the stack this stereotype may be true sometimes but programming can actually push many people in the opposite direction the guy who works in tech but hates text stereotype knows exactly how unreliable and dangerous code can be like that the rac25 incident where a little software bug accidentally killed some people by giving them a massive overdose of radiation this guy would never buy a car that can be remotely summoned back to elon when you stop paying the bill and he would definitely never put a smart lock on his house because the nsa probably has backdoor access or at the very least there's an undiscovered exploit in its code if you broke into his farmhouse you'd find a single monitor linux machine a flip phone some gold bullion and a shotgun barrel pointed in your face the most stereotypical programmer though has to be the introvert he's a savant who still sleeps in a car bed and his vision of the ideal lifestyle is what the rest of society calls quarantine he's super good at math and can actually program stuff without using google and stack overflow but couldn't hold a conversation to save his life extroverts like jobs use these nerds like woz to get super rich this stereotype used to be 100 true back when programming was hard like pre-1990s but as programming has become more mainstream it's led to a new paradigm the programmer this guy got a computer science degree while mostly partying with his frat in college his name is usually chad and he has more mating opportunities than the introvert but it comes at a cost of reduced code quality which he refuses to test because test driven development is for losers nah bro however he has better communication skills than the introverts which is annoying because i wish this guy would stop talking to me eventually he evolves into your manager where he can torment you with code reviews and team building exercises now so far in this video i've been using a lot of masculine pronouns that's because 95 of my audience is male which is actually pretty close to the real world distribution today what you may not realize though is that back in the day women used to dominate the programming space kathleen booth created the first assembly language grace hopper created the first compiler and margaret hamilton led the team who wrote the code for the apollo moonlander code that was so flawless and perfectly executed that some people think it's proof we didn't actually go to the moon that was the apex of code quality since that time everything's gone downhill the next specimen we'll look at is the influencer or code fluencer his natural habitat is not a code editor but rather a social media platform most commonly twitter after figuring out how to print hello world in php he immediately rose to the top of the dominance hierarchy in his own mind now he makes the world a better place by regurgitating code tips and hot takes all day long and he just landed a better paying job than you because he mastered the art of virtue signaling and that's what we call a good culture fit another popular stereotype is the hacker this guy's able to open up a terminal connect to some remote mainframe and break all of its security protocols one by one with awesome fancy animations between each step this stereotype is what most people think programmers do but is 100 manufactured by hollywood real hacking is extremely tedious and boring and is done primarily by the people who have all the guns now a stereotype that is actually real is the 10x developer this guy is an extremely rare unicorn that can do the work of 10 other developers combined some say they're a myth but i've seen developers first hand who write code like durant plays basketball or kasparov plays chess there are people out there with a natural problem-solving ability that just goes far beyond the rest of the population you'll know a 10x developer when you see one because you'll feel very incompetent and also very jealous now i think the ideal stereotype for most of us to fall into is the lazy programmer to the outside world it doesn't look like this guy does much he sits at a computer all day hitting the keyboard and if you glance at a screen it looks like he's just copying and pasting things from the internet what he's actually doing though is building a million dollar side hustle so he can retire in his 30s he also has a remote job with a 400k salary but he eats ramen for dinner while sharing a crappy apartment with four other dudes his wardrobe is 50 swag from tech conferences and 50 thinks his mom bought him he leverages code to work smarter and not harder now on the other end of the spectrum we have the old jaded guy he has long silver hair and a big white beard he only codes in c not c plus plus and definitely not any of the hipster garbage that you're using in fact he probably wrote the compiler for the silly toy language that you're trying to learn his depth of knowledge transcends the normal apes idea of reality when he discovered through psychedelics that we're all just one entity that found a hack in the universe to experience itself in parallel with primate bodies and computers are the tool that will ultimately make us one again and that concludes our presentation on programmer stereotypes let me know which one you fall into in the comments below thanks for watching and i will see you in the next one" ] Only one chunk. Not too semantic if you ask me.