3 ms·
Thanks for the concrete numbers — 4 minutes for 63k lines on a 3060 is pretty reasonable for a full index pass, especially with incremental after that. I work
by muin_kr 7mo ago
Thanks for the concrete numbers — 4 minutes for 63k lines on a 3060 is pretty reasonable for a full index pass, especially with incremental after that.
I work with a few TypeScript/Python monorepos in the 5-8k file range. Most are mixed codebases — application code, configs, generated types, test fixtures. Happy to run rawq against one and report back with timings and chunk counts. The interesting test would be whether search quality degrades gracefully as you push past the 50k chunk threshold, or if it falls off a cliff.
Curious about your thinking on the vector store side — is HNSW (or something like FAISS IVF) on the roadmap, or are you leaning toward a different approach for scaling? The flat search has the advantage of exact results, so there's a real tradeoff there. For code search specifically, I'd imagine recall matters more than in typical document search since missing a relevant function can break an agent's whole approach.
- Yerzhigit 7mo agoThe search quality should not degrade as rawq chunks code by its structure into similar sizes across the codebase, the only thing that can get worse is the full indexing time for large codebases, but it is a one-time action and depends on the hardware capabilities. I wanted to implement HNSW to make searching faster for 50K+ chunks, but there are difficulties with that right now, I couldn’t get it to work properly, but it is on the roadmap.