6 ms·
Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
- punnerud 2mo agoWhat’s in that “Search tool”, could actually be the same solution presented in a different way
- kumama 2mo agocastform founder here. it uses lakebases's native bm25 and vector search and fuses the results using rrf (https://medium.com/@devalshah1619/mathematical-intuition-behind-reciprocal-rank-fusion-rrf-explained-in-2-mins-002df0cc5e2a https://medium.com/@devalshah1619/mathematical-intuition-beh...)
- softwaredoug 2mo agoUsually in these cases you don't need to do much tuning to the retriever. So you just give it BM25 or somesuch. I'm hesitant to say absolutely zero tuning, because there are cases where you do want to say, bias towards trustworthy results or recent results etc to help the model avoid wasting tokens. But probably not much beyond that. You can also just create a param in the tool for the agent that selects for "recent" or "popular" or "trustworthy" in ranking.
- mrinterweb 2mo agoThere is so much opportunity for purpose built models like this. Ideally a harness should spin up a subagent to offload to targeted models for specific tasks like this. I know this is not a novel idea. Claude code does some of this by handing off the "explore" agent work to haiku. I just love seeing that specialized LLMs are being developed.
- foota 2mo agoI feel like the future is people building applications with tightly integrated LLMs that work hand in hand with the application's own lifecycle and code. I also didn't realize that people were using agentic harnesses for search, it's an interesting idea. If the context length is short enough it should be fairly cheap compared to running "normal" agentic coding workloads where you have O(100k) context length for doing almost anything.
- kumama 2mo agocastform founder here. that's a future we are really excited about too :) ideally, you can post-train the llm within the application itself, as it's being used. both interesting infrastructure & algorithmic challenges here
- Malp 2mo agoThere are! Chroma has Context1, SID has SID-1, and you'd actually be surprised at how easy it is to post-train your own with pretty good pass@ recall@ ndcg@ etc. There's also Hornet who have shared some interesting talks & blogs lately. I don't know that I'd exclusively use agents for retrieval the way Neon outlines here as well. I think distillation similar to what ZeroEntropy has done for bespoke retrieval & reranking with _some_ agent manipulation on top-k results works better (IME).
- fennecfoxy 2mo agoBy post-train I presume you mean a finetune? Unless that's wrong (please correct me if so). I haven't looked into model architecture people are working with for this stuff too deeply yet but I presume the core idea is fine-tuning a lightweight reasoning-enabled LLM specifically using search as a metric for training?
- Malp 2mo agoThat or providing a concrete RL env for $your_search_corpus_etc_here
- devolving-dev 2mo agoModels keep on improving though, so doesn't fine tuning become an ongoing task with ongoing maintenance burden?
- kumama 2mo ago(one of the blog post authors here) -> once you set up a finetuning pipeline, it's often trivial to rerun it on top of a new open weights model. so, it's orthogonal to base model improvements
- Razengan 2mo ago> There is so much opportunity for purpose built models like this. OpenAI etc could themselves do this, and maybe they already do? Where the public-facing interface delegates to multiple little goblins behinds the scenes
- mrinterweb 2mo agoExactly. There could be a lot of value for inference companies to do this. Could save a lot of money being able to hand off highly repetitive known tasks to far smaller specialized models.
- kumama 2mo agocastform founder here. openai actually deprecated their finetuning apis a few months back weirdly.
- phainopepla2 2mo ago> Claude code does some of this by handing off the "explore" agent work to haiku This is no longer necessarily true. As of 2.1.198 [0] (released July 1st): "The built-in Explore agent now inherits the main session’s model (capped at opus) instead of running on haiku" [0] https://code.claude.com/docs/en/changelog#2-1-198 https://code.claude.com/docs/en/changelog#2-1-198
- benjiro29 2mo ago> Claude code does some of this by handing off the "explore" agent work to haiku. That is not handing off to a specialized model, its just handing off to a lighter and interior model (compared to the parent model). That by itself can create issues like the lighter model not capturing all the data that the parent needs. The idea is that we get specialized models that are better then general purpose models. But its rare for a specialized model to beat a strong general model. There is a reason why we hear less about this idea of smaller expert models, because large strong models to the tasks just as good. And if the tasks is repetitive to the point that specialization is useful, you can get into a situation that your better off having a program written for that reputative nature, then delegating to other models. And then have the main strong model, deal with the (semi)cleaned up data.
- BikiniPrince 2mo agoYou can register models with mcp. I think it’s an expensive solution, but it is available in the framework. I use a light weight bus protocol that lets agents interact and pass short messages with pointers. It’s very efficient.
- tyre 2mo ago> There is a reason why we hear less about this idea of smaller expert models, because large strong models to the tasks just as good. Smaller models are cheaper, sometimes faster. I agree that the “we’re an LLM fine-tuned for X” hasn’t worked out because you can just train Claude to do X (and Anthropic will), but not burning Opus/Fable tokens on dumb-but-token-heavy tasks is good sense. As we move from “integrate AI into Y” to “optimize the ROI on Y”, we’ll see more of this.
- kumama 2mo agocastform founder here. the roi optimization makes sense. i think there are lots of usecases for which even a 2% gain in accuracy can be quite useful. off the top of my head - high volume customer support. higher accuracy means fewer escalation, reducing labor costs - fraud detection. catching even one extra fraud attempt could mean a lot in savings - and ofc the classic ads use-case where at scale bps in improvement could mean millions in revenue :)
- oliver236 2mo agothis is exactly what leopold talks about in situational awareness
- try-working 2mo agoyes, and this is why we need model routing
- kumama 2mo agocastform founder here: totally! we also think model routing is also a post-training problem i.e. getting a model to predict the difficulty of a task and match it to the right model -> we're gonna be sharing more on that soon :)
- dd8601fn 2mo agoI’ve been trying to do this with a pet project and admit that I’m getting terrible results. My small llm as a classifier/router stage just isn’t getting the job done.
- nikcub 2mo agoThere has been an over-obsession with frontier models and benchmarks. Most of the work will be done by task specific models. You don't put Phds on the factory floor.
- bizzletk 2mo agoIf the PhDs don't cost very much more than your equivalent of factory-technicians but still get the job done, why wouldn't you do that, at least in the blunt case before cost control rears up?
- polotics 2mo agobecause they will take initiatives for localized improvements you don't want them to take
- dwaltrip 2mo agoThis only works if the tasks are actually specific and don’t benefit from broad competency. IMO, this doesn’t match most things that people use LLMs for.
- nullbio 2mo agoYep. It's a real shame that the labs are incentivized not to go in this direction. They all want to try and suck us into the cloud and take away full control and local processing, but there's far more opportunity by building small AI systems and tools into the harness itself to make the models more intelligent. They naturally don't like this direction, because it draws the intelligence away from their systems, and onto the local machine, where idea moats cannot be protected and hidden, and costs can be dramatically cut. Imagine though, how powerful our harnesses could be if the best researchers were thinking about how to utilize the power of the gaming GPUs that most PC users have (or can get), to supplement the frontier model processing. Instead of trying to have the frontier model do everything, the frontier model can serve as the orchestrator over all of the smaller dedicated harness models. Right now my rtx4090 sits there unused for most of the day while I'm paying for inference in the cloud... It's such a waste of parallel intelligence bandwidth. I'm not just talking about LLMs either, most people seem unaware that there are a plethora of dedicated AI models for all sorts of conceivable pipeline usecases, from all sorts of classification tasks all the way down to things like code duplication detection. Right now the LLMs completely suck at cleaning up code and architecture, and a big part of that is because the frontier LLM cannot fit your entire codebase + all of its long-chain reasoning into the context window. But using small local models and tools bypasses this problem because small fast models can iterate over an entire codebase quickly. A harness that creates a big model bundle + routing system + DAG-based memory/execution management over all of these has the potential to be incredibly powerful. Even better, building a framework around this concept and having the frontier model dynamically and adaptively generate the ideal execution system for any given task/domain. We're working on coding today? Okay, here's a recipe we can use: ..., and it generates a local model pipeline execution system that it feeds all of your prompts through in real time by using pre-defined or shared recipe building blocks, etc... Lots of interesting possibilities.
- kinnth 2mo agoisn't this also the threat to frontier AI houses? As in they want you to expend tokens in their ecosystem, but the optimization at 100x is their profit?
- kumama 2mo agocastform founder here. i'd say it's a threat but the frontier ai labs' argument would basically be that the market opportunity for intelligence is infinite so it doesn't matter. at the same time, i do believe there will continue to be a massive, growing market for big labs, esp for super-frontier use-cases. today that's longer-horizon coding tasks but in the future it can be things like scientific discovery, etc.
- boonzeet 2mo agoThis is how Sakana's Fugu model works, achieving similar performance to Opus/Fable with a mix of open source & mainstream LLMs.
- drob518 2mo agoIMO, purpose built or “adapted” models are The Next Big Thing. If I’m using a model to write Python code, for instance, I really want the 8B or 27B expert model for exactly that, which would also be runnable locally. I don’t care about the 1.8T model that can answer every query under the sun and that only runs in a remote data center.
- fennecfoxy 2mo agoIt is nice to have a model that can "do it all", though. And surely that's still the end goal? Like how MoE is still somewhat popular in certain areas even after its heyday. I am wondering if models will end up being some sort of evolution of MoE where it has something internally like the model the author refers to that gets surfaced when it needs to search in some way. I guess it makes sense; our own brains have so many distinct task-specific regions.
- drob518 2mo agoYes, agreed, but I think it’s going to be difficult to have a high quality model that “does it all” and have it be local. For a quality “does it all” model, you’re going to need a lot of parameters and that means you’re almost always going to be running in the cloud. But it seems like you could probably get a smaller, focused model that runs locally and is also high quality. In other words, I want the programming expert portion of a 2T model that is maybe 35B or 70B parameters or whatever but it’s running locally (and yes, I know you don’t just carve out an expert from a larger model, but conceptually that’s what I’m after).
- fennecfoxy 2mo agoYeah, I always thought the future of this stuff would be hot-pluggable MoE modules or LORAs that are able to be downloaded and applied/used at will like how skills have become a thing. Like, atm most architectures seem limited by a single context and fixed architecture with no hot loading. Especially for robotics, being able to load/unload various specialised skills on limited mobile hardware will (I hope) definitely become a thing.
- deleted 2mo ago[deleted]
- ramon156 2mo agoBit unrelated, I realized that z.ai gives you access to deepseek 4 flash. It's incredible how well it performs when given a detailed spec. I'm not sure I've seen a model one-shot like that, and I was already impressed by gemma 4's speed and efficiency.
- swiftcoder 2mo agoDeepseek flash (especially after the recent update) has to be one of the most slept-on models. Price-performance is ridiculous, and its available on a number of cheap coding subscriptions
- try-working 2mo agoFlash is the most used model in the world since last week
- esafak 2mo ago> I realized that z.ai gives you access to deepseek 4 flash. How? Can you give details?
- aliljet 2mo agoThere is a more serious question in here that's not being answered. How effective is the retrieval in finding buried needles in larger and larger haystacks. And there's a correlary question, how effective could you be in finding paired needles in that haystack where you need to hold a needle to unlock finding another needle.
- Foobar8568 2mo agoConsidering the state of the field ( RAG/retrieval/evaluation) I have 0 trust in it, even more if it's closed source with bullshit claim like that. Everything is vibe sloped to death, and dead after a few months to a couple of years (and not hard to be 100 cheaper than GPT-5.6 sol ... DS is basically free and I guess already 100 times cheaper or more, and here another slope ).
- kumama 2mo agocastform founder here. we should have made it more prominent on the blog but here's the full code example: https://github.com/castform-ai/benchmax/tree/main/examples/neon_rag https://github.com/castform-ai/benchmax/tree/main/examples/n...
- sreekanth850 2mo agoAre you using blind chunking or section aware chunking?
- kumama 2mo agofor the example here the chunking is section aware -> but the general training data synthesis pipeline is agnostic to type of chunking
- sreekanth850 2mo agoif chunking is section aware, how do you manage large section embeddings?
- richwater 2mo agoOne thing that plagues [insert current FAANG] is the large amount of corpus knowledge that is outdated/misleading or just plain wrong. I'm curious how this addresses that if it's deriving the reward function from the corpus itself.
- seahyinghang8 2mo ago(founder of castform here!) - having worked at FAANG / big tech, i totally get this. our example was on gitlab's open source company handbook but i think a real company's corpus is way more messy and has many sources of truth. a few ideas i have yet to validate are: - prioritize recently updated docs when generating the training questions (assumption those docs are more correct than others) - actually including contradicting documents that talks about the exact same topic might be a good training example - ideally the model should surface all the relevant info it can find, and explain what it has found. (usually contradiction comes from the fact that the later document is the updated stance) - you could also mine high quality Q&A from public slack / communication channels where questions were asked and someone else in the team linked some docs / answer. those are strongly validated "ground truth" answers
- richwater 2mo agoThanks! I think those could all be good heuristics.
- JCharante 2mo agoI have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead.
- andrenotgiant 2mo agoAny data or public links you can share? That surprises me
- barake 2mo agoAnecdotally, it feels like Opus, Fable, and Sol "get distracted" when you use them for writing code. Great at reasoning and coordination but they will go off on a tangent and refactor half the code base. I only use them for reasoning (of course) and coordinating subagents.
- CoolCold 2mo agomind sharing hints/links on your harness/flow setup? I did several attempts with naive prompting, but spent more time babysitting than actual flow
- jorl17 2mo agoHave been feeling the same. There's a sweet spot that threads the needle between "too dumb to search the right thing / relay the correct results" and "too smart to just stop overthinking and just report the damn thing"
- hankbond 2mo agoJust an anecdote but thats why Deepseek v4 flash 0731 is my current favorite model. It's really not very "eager" and stays on the task at hand.
- seahyinghang8 2mo agowe actually have the test benchmark against luna too! it's just not in our title but you can see it in the first diagram below the title. luna does pretty well tbh but sol is just a tad bit better. but luna is way cheaper. if you want to dive down into the various traces of the benchmark, you can check this out: https://app.castform.com/train/a7a898f6-d802-4908-b044-acb812f14a48?tab=comp https://app.castform.com/train/a7a898f6-d802-4908-b044-acb81... - founder of castform
- breadislove 2mo agoOn what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report?
- krm01 2mo agoKeeping track of any AI progress is becoming harder by the day, because there's ambiguity around common/clear/consistent benchmarks. Everything is constantly skewed into favourable directions.
- alansaber 2mo agoTBF gaming benchmarks is not something new to AI
- seahyinghang8 2mo ago(founder of castform here) - we didn't get to dive too deep into the dataset we were using for the retrieval in the blogpost for brevity, but we did link the training run (which shows the dataset) here: https://app.castform.com/train/a7a898f6-d802-4908-b044-acb812f14a48?tab=comp https://app.castform.com/train/a7a898f6-d802-4908-b044-acb81... the page shows the exact trace of all the models we are comparing against and the aggregate scores we generated the question & answer pair from gitlab product handbook (https://handbook.gitlab.com/ https://handbook.gitlab.com/) since the point is to show that you can generate training questions from raw data corpus (something a company already has today)
- BedVibe_Studios 2mo agoThis feels like the database equivalent of "use the right data structure." We've spent two years assuming the biggest general-purpose model should do everything. It makes more sense for retrieval, reranking, reasoning, and generation to each have their own optimized model if the routing cost is negligible.
- seahyinghang8 2mo agofounder of castform here, we believe that as agent deployment moves from experimentation phase where cost doesn't matter as much to deployment (what's the margin of serving the request), there will be a rise in interest in optimized models.
- nullbio 2mo agoYes, and it's only logical things move in this direction because there's massive hardware incentive to do so. If frontier models can be broken down into small models, networks of smaller-GPUs can be utilized. Right now the smaller GPUs are basically paper weights for frontier intelligence.
- dev_l1x_be 2mo agoI am not sure about GPT-5.6. It usually 10x more verbose for no apparent reason than GPT-5.5. Maybe it is only me.
- jr3592 2mo agoHave you tried Claude? 5.6 feels less verbose, and less messy to me.
- dev_l1x_be 2mo agoYeah Opus 4.8 / GPT 5.5 what I use. Fable is okish, the coding experience is a bit weird with it.
- alansaber 2mo agoit is definitely more verbose.
- wahnfrieden 2mo agotry adjusting model_verbosity. it defaults to verbose. and of course, use agents.md.
- dbbk 2mo agoYou can just tell it not to be verbose...
- skybrian 2mo agoMaybe, but a specific example showing how to do it would have been a more compelling argument.
- kumama 2mo ago(castform founder here) we should have made it more prominent on the blogpost but here's the github repo: https://github.com/castform-ai/benchmax/tree/main/examples/neon_rag https://github.com/castform-ai/benchmax/tree/main/examples/n...
- andai 2mo agoNice, but there's no mention of how Luna or DSFlash perform on the same task? (Being 25x and 50x cheaper respectively.) Nor of how much faster their custom model performs?
- deleted 2mo ago[deleted]
- seahyinghang8 2mo agowe actually have the test benchmark against luna but no deepseek flash (we haven't added DSFlash into our benchmarking model pipeline) you can check out the full comparison against all the other models here: https://app.castform.com/train/a7a898f6-d802-4908-b044-acb812f14a48?tab=comp https://app.castform.com/train/a7a898f6-d802-4908-b044-acb81... - founder of castform
- cmiles8 2mo agoThe big lab models are academically interesting but business wise they seem toast long term. There’s no way for these model companies to compete when the models are becoming a pure commodity and others offering options that are orders of magnitude cheaper. It’s not that the big labs couldn’t theoretically just also put out 100x cheaper options but their business model requires them to generate huge revenues from higher priced tokens or they’ll implode.
- kumama 2mo agocastform founder here. despite our bet on fine-tuned smaller open-source models, i'm still quite bullish on the big labs. i think scaled closed models will continue to dominate for more general purpose use-case like codegen, search, etc. but intelligence has lots of long-tail applications and i think for these longer-tail applications, finetuned custom models will rule
- sroussey 2mo agoProduct search model: https://www.linkedin.com/posts/introducing-ontology-1-ugcPost-7488272626277294080-XxvU/ https://www.linkedin.com/posts/introducing-ontology-1-ugcPos... Edit: more direct links, sorry: https://onton.com/research/ontology-1 https://onton.com/research/ontology-1 https://onton.com/research/ontology-1-benchmarks https://onton.com/research/ontology-1-benchmarks
- liesliy 2mo ago[flagged]
- abratabia 2mo ago[flagged]
- sreekanth850 2mo agoIMHO, what is broken is retrieval, the whole blind chunking which was first generation is still the default standard in RAG, this has to be changed. I'm saying this by seeing the results when we used richer parent candidates model for LLM and child segments as search probes. Even without reranking we got solid results.
- modgate 2mo ago[flagged]
- linux_devil 2mo agoWhy do we need to train the model to solve for retrieval within the org, so we have to keep training it whenever new dataset is introduced , or am I missing something here ?
- jmalicki 2mo agoIt's a matter of cost. Did you see the 100x cheaper? If you have a workload that is going to be very heavy, incurring a large training cost to make a cheaper model work well with the dataset will be dramatic cost reduction. Most large AI workloads can't afford, or truly need, the expense or capability of GPT 5.6 Sol when cheaper models can do. Of course you could skip that and just use GPT-5.6 Sol everywhere instead. If you're running a fast food restaurant you could hire Michelin star chefs to make your burger and fries without further training. Or you could have a training program for teenagers, a sourcing program, etc. to scale up to your chain to still get consistent quality without needing that level of cost in each store, but replacing it with a centralized repeatable process.
- kumama 2mo ago(founder of castform here) the model you post-train should ideally learn general patterns & search strategies over your dataset that should transfer to new docs you add to the search corpus (unless its super out of distribution)
- srvraw 2mo agocould you share more about what you mean by "general patterns & search strategies"? I can think of it being along the lines of searching over specific tables or databases for queries in certain context. It's an exciting line of work and I'm interested because I need something like this for the problem I'm solving atm. So, I'd like to understand how the training generalizes
- kumama 2mo agoyup! it’s mostly about getting better at using the right search keywords. for more complex multi-hop question, it's also about knowing which sections of a document to look up and in what order.
- mukundzzha 2mo ago[dead]
- nc55g3g 2mo ago[dead]
- nullbio 2mo agoI (and I imagine many others) would love to use something like this, but can't, because my data is too sensitive to be uploaded to a cloud of which I have no gaurantees of privacy/security. Is there any way we do this using rented GPUs and open-source software stacks? Paying for the service isn't the issue, I don't care if it's free or if a cut is taken in some capacity, I just don't want the provider to have access to my data.
- i5heu 2mo agoIn OpenRouter there are Zero Retention options. And if you use a EU provider you can be somewhat sure that your privacy is given. Other than that there is not really a difference to renting a GPU since the GPU provider can also just steal your data. Local GPU(s) are always an option if you have the possibility. It is also not that difficult to run with stuff like “LocalAI”
- kumama 2mo agocastform founder here. unfortunately, we are cloud-hosted at this point. but some easy options on the open-source side include huggingface's trl & unsloth. you can run our data-generation scripts here: https://github.com/castform-ai/benchmax https://github.com/castform-ai/benchmax and then hook it up trl/unsloth for training. should be able to do all of this on your own compute.
- softwaredoug 2mo agoPeople are building agentic search one of three ways: 1. Actually good retireval. There’s been a lot of progress on serving the kinds of queries agents tend to serve, from places like Hornet, MoxedBread, LightOn. Particularly in late interaction 2. Smarter harnesses with models/judges validating the result. This is now just seen as the generator/ evaluator pattern. Here’s where people try to just use grep or some other naive retrieval system. Let the agent figure it out. But it’ll consume a lot of tokens to get good results as it iterates and loops. 3. A model trained for retrieval. Give it dumb retriever like in (2) but it is fine tuned on the task as in (1). This article is 3. But we’ve been seeing this all year with SID.ai, Gleans Waldo model etc. if this interests you I’d check those out, particularly SID. I wrote about these 3 approaches here https://softwaredoug.com/blog/2026/06/08/three-kinds-of-agentic-search https://softwaredoug.com/blog/2026/06/08/three-kinds-of-agen...
- kumama 2mo ago+1 on SID-1. we were definitely inspired by that paper
- softwaredoug 2mo ago*MixedBread https://www.mixedbread.com/ https://www.mixedbread.com/
- fennecfoxy 2mo agoThank you for the link! Super interesting and I appreciate how it was written. It seems like a lot of the problems I have been running into with RAG on large/complex documents with generally low contrast in the information is not one that has been perfectly solved yet - here I am thinking I'd been a bit behind. It's just unfortunate that none of the cloud providers are flexible enough to deal with the pace of change. Probably going to have to shove one of those 8b~ models into an instance to use when needed. It's interesting you mention late interaction (retrieval), I had recently been using ChatGPT as a mirror to throw ideas back at me on this issue and had been musing about how nice it would be to have some sort of hierarchical embeddings that capture a whole chunk, then sentences and then sentence fragments or individual word and it seems that that fits the bill!
- zkmon 2mo agoNice things like this, get "blue-washed" by the parent org's tech. For general adoption, kakebase should have been swapped out for a more generic non-databricks tech. 'Blue-wash' is a reference to an IBM's practice.
- genshro 2mo ago[flagged]
- oedemis 2mo agowhats about self-consistency like in grpo with majority voiting?
- kumama 2mo agocastform founder here. i'm personally a little against techniques like self-consistency/majority voting during rl training because they tend to result in the model's output distribution "sharpening" a lot. this means the model will lose it's exploration ability and probably won't be able to explore/discover new solution strategies, which can be harmful for both rl training + generalization to unseen cases
- jillesvangurp 2mo agoModels without tools and harnesses are not really that useful. My observation is that the tool ux is driving most progress at this point. There are of course open source tools and harnesses but they require more effort to setup properly. The key challenge is to pick the right model for the right task or sub task and doing that automatically rather than manually. A big part of the problem here is that everybody is picking the most expensive and resource intensive models by default just in case they hit something that is a bit more difficult to get right. It's overkill. Most work people actually do is completely routine and would not have been a challenge for most of the mainstream OSS models. I'm starting to suffer a bit from model fatigue. There are announcements almost on a daily basis about this or that new model. I can't keep up with that and I don't have time to try them out or evaluate them. I don't want to waste brain cycles on which one to use. I just want to get shit done without micromanaging AI models. All this marketing BS and confusing naming isn't helping either. It seems a lot of that is just about tricking people into picking the expensive model so they'll burn through more tokens.
- drob518 2mo agoI agree with this. Interestingly, I’ve been using Deepseek v4 Flash a lot these days but I definitely have to constrain it a lot with tests. Fortunately, I can have it write the tests. It’s extremely cost effective. Still trying to figure out whether the latest update last week that made it smarter actually translates into something I can see in the output. It’s not dumber, but it’s still an open question as to whether it’s “real world smarter.”
- apparent 2mo agoWhat does "100x cheaper" mean? Costs are 1 / 100? If so, is there a more appropriate term to use?
- adityazero 2mo ago[flagged]
- fennecfoxy 2mo agoHmmmm, this is a problem I have been facing recently. Basic embeddings give decent-ish results (in the top say 20 chunks). Basic agentic retrieval gives slightly better results so long as the agent part of it doesn't go down the wrong track. I like the idea of what's discussed in the link, however atm we are on Bedrock KBs and so locked in to a very basic implementation of RAG, because Amazon doesn't have the foresight to make things flexible enough - including making it an absolute pita to use their hybrid search. But, I guess they "work" reliably. One of our core issues centers around a 1300 page document all about the same overall topic but with minor various for specific procedures/situations. Typical embeddings waters this down so that each chunk really just represents the common theme and therefore lacks a lot of contrast. But now that luna's (and others) price has been cut, perhaps I'll start experimenting with giving it free rein to explore the data a little in the same way that I do a web search. One thing that definitely helped was providing a separate index of each section where I had another model summarise the primary unique topics in each section to act as a guide for the agent. I think either we should be chucking the entire doc at a model (400k tokens...so not really ideal at this time) or improving RAG accuracy. For the latter I think even with embeddings, meaning of words and semantic connections are not enough at all - attention is KV so it is 2 dimensional and once I started getting into it I've kind of realised that 2 dimensions aren't really enough to represent the logic that exists between tokens (i.e. sections of documents that refer to a sequence of actions dependent on some logic that references "variables" from another section, i.e. "if x, y has happened then refer to z sequence). There's much deeper meaning to human language than I think basic embeddings covers. I think it's becoming clear to me that in the same way that embeddings encode the web of semantic meaning of a chunk of text, I need something similar to a hybrid of the author's model + reranker + super-embeddings that encodes as much of the entire meaning of a text as possible and not just semantic.