5 ms·
SubQ 1.1 Small
- EDM115 4mo agohttps://subq.ai/docs/subq-1-1-small-model-card.pdf https://subq.ai/docs/subq-1-1-small-model-card.pdf
- giancarlostoro 4mo agoThis one's interesting, and I think the next frontier for LLMs should really just be, how can we get something like Opus 4.6 to cost drastically less, for the same output? I say 4.6 because from 4.6 onwards it's been pretty darn good, at least for me, always feels like every model upgrade someone hates it, heck even 4.5 was fine.
- robmccoll 4mo agoYes - I want that and dramatically faster. Newer models don't seem to need any more or less guidance and iteration, so let's make the time-to-wrong-answer as short as possible.
- giancarlostoro 4mo agoI'm not as crazy about speed as long as it's reasonably as "quick" as Opus. Which is faster than most developers can spit out code. I do get annoyed with Claude Code because it looks like it chooses to be as slow as possible, but maybe that's by design so its not pounding their backend every milisecond? Would probably be bad. Local inference is insanely fast on my M4 Pro MBP though, so I can understand where you're coming from, but I don't need it too much faster. I still need time to review, test, review and provide feedback to the model. Fast is okay I guess for true vibe coding.
- robmccoll 4mo agoI just don't want to have to have a pipeline going in order to fully occupy my time. I don't want to wait on the model to review the prompt, read the parts of the codebase indicated, do its own research in the codebase and documentation, plan, run agents ... actually write the code and NOW I can start reading it and reviewing it. That means I either need to run a lot of operations in parallel so that I always have something to do and the agent(s) are highly utilized or I'm writing something on my own that I keep getting that keeps getting interrupted. It's the constant context switching that kills me. I want to work on one problem at a time and really focus on it - even if I'm not writing every line myself.
- mritchie712 4mo agoI agree on opus 4.5-4.8, but Fable 5 was a noticeable upgrade.
- mstkllah 4mo agoDid not feel as an upgrade to me at all, felt way slower at the same quality level as 4.8 to me.
- NetOpWibby 4mo agoMan I miss it
- aesthesia 4mo agoDisappointing they don't actually say how their sparse attention mechanism works.
- cmogni1 4mo agoI don’t understand why this lab is allergic to providing details on what they actually made, especially when Chinese labs are more than willing to share architectural specs/code/kernels (eg NSA/FSA, RAMBa, HISA, DSA LightningIndexer, etc). I don’t doubt that they’ve done something here, but the lack of details makes me default not trust this, particularly when this is the second time that they’ve released a “technical report” that just waxes poetic about the concept.
- famouswaffles 4mo agoBusiness wise, it would make sense to hold off on details till they're at least ready to serve. Look at what happened with Open AI and reasoning models. Everyone struggled with getting RL to work with LLMs for a good while. Open AI figured it out, and a few months later everyone had their prototypes out in short order. Don't forget who these labs employ. They're some of the brightest people around. Sub-q aren't really in a position for that lol. If they'd shared details at the first announcement for instance, the big labs might have had something out by now while they're still pulling resources to scale and then what ?
- cmogni1 4mo agoI don't think it makes sense from a business perspective to hold off on details as a new lab. OpenAI will not implement new architectural changes unless they've tested the changes themselves internally. Even if someone claims some great innovation, they'd need to do scaling experiments to somewhere between the size of GPT-4 to GPT-5 before they'd decide it is worth it to implement themselves. Plenty of mechanisms that seem to work at one scale do not translate to the next. Because the cost to OpenAI to make an architectural shift is far greater than the cost to a new lab to try something different, providing details is usually a net benefit for recruiting, building trust, getting acquired, etc. The lack of details is a poor business decision because it makes them seem untrustworthy. I'm not advocating that they should open source their model, but there is already so much noise in the space and many bad papers that being cagey is a poor strategy for winning over talent, developers, etc.
- 4mo ago
- embedding-shape 4mo ago> SubQ 1.1 Small scores near-perfect at 1M, 2M, 6M, and 12M tokens. The model was trained predominantly at 1M tokens yet the retrieval held near perfectly at 12x that length, despite compressing attention to just 0.13% of relationships. This generalization is a direct consequence of SSA routing attention based on content relevance rather than fixed positional patterns. If the results persists from 1M to 12M, why not 24M or 48M? Sounds almost too good to be true. With back of the napkin math from inside my head, that'd be like 0.5/1 million LOC, depending on language/code density, could just fold the entire codebase into one prompt if it's a small one, that'd be neat :)
- monster_truck 4mo agoIt likely falls off very steeply after that. 8 to 1 (which I am assuming based on the 0.13% figure) is a pretty common ratio for sparse matrix stuff.
- chrsw 4mo agoThere was, let's say, significant skepticism the last time they announced something. What's changed?
- supern0va 4mo agoI have no idea if the evaluator themselves is trustworthy, but it was supposedly independently evaluated by Appen: https://www.appen.com/whitepapers/benchmarking-subquadratics-latest-model-ssa-kernel https://www.appen.com/whitepapers/benchmarking-subquadratics...
- wxw 4mo ago> SSA replaces the O(n²) dense attention pass with a learned sparse formulation that scales linearly with context length. > At 1M tokens, SubQ 1.1 Small requires 64.5x less compute than dense attention and runs 56x faster than FlashAttention-2. Awesome stuff. Solving context at the model architecture layer rather than trying to bolt on extra memory is the right direction IMO.
- satyarohith 4mo agoIt's been all talk and no action ever since their first announcement.
- maz1b 4mo agoThey've done multiple "evaluations" by third parties, but still, it seems that they aren't being fully transparent. I think the approach is quite interesting and novel, but this feels like deja vu. I get why they aren't disclosing all the details, but it seems more hype-train-esque to me for this moment. I don't disagree that this could be big.
- Depurator 4mo agoWhat kind of hardware would be needed to serve an instance with the full 12m context? And what kind of speeds can one expwct at those extremes at 10m+?
- samber 4mo agoAccording to Subquadratic, Needle in a Haystack is strong up to 12m tokens, but RULER has not been tested above 128k tokens ??
- samber 4mo agoComparing compute cost versus FlashAttention-2 is not very honest to me. FlashAttention-2 is not used anymore for at least 2y. This architecture would have been a massive improvement 3 years ago, but it is a ~solved~ problem IMO.
- ballon_monkey 4mo agoIts funny that some people on HN think this whole thing is legit. The company is started by a bunch of no-bodies with 0 experience in AI in general let alone ML/Data. Edit: Typical HN "I can downvote but I cannot dispute facts"
- phantasmat 4mo ago[flagged]
- kristjansson 4mo agoIt's easy[1] to promise, it's hard to deliver. I hope the best for them. [1]: https://magic.dev/blog/ltm-1 https://magic.dev/blog/ltm-1 (note the date)
- bthornbury 4mo agowe need some better standard long-context benchmarks. needle in a haystack is not good for this, yes it proves the model can attend to its context, but in its usual form, somewhat trivializes the query-key relationship. something like long-form Q&A would be more ideal. Like reading a book and answering questions that require synthesizing information derived from either the whole thing or disparate portions of it. Like describing an entire character arc in a 1000 page novel with examples and evidential moments.
- mark_l_watson 4mo agoInteresting idea but until I get my grubby little fingers in it, to try it - difficult to have an opinion. I am hopefully expectant that we will see all sorts of optimizations in the next few years that will enable even more local model use and slash commercial API costs. I get excited by the results when I enjoy one or two short coding sessions a week with Claude Opus but it is even more exciting to get a major task done and see that I only used $0.05 for DeepSeek v4 Flash or perhaps $0.15 for DeepSeek v4 Pro. It was exciting in even a different way when I two shotted a complete TypeScript/Tauri app using gemma-12b-qat with little-coder on a cheap laptop a few days ago.
- dundunUp 4mo agoWhat is this?
- projektfu 4mo ago"At 1M tokens, SubQ 1.1 Small requires 64.5x less compute than dense attention and runs 56x faster than FlashAttention-2." 6450% less compute? Is Trump working there?
- freeaitools 4mo agoEverything in south Florida is a scam, this is just one of many