8 ms·
Bamba: An open-source LLM that crosses a transformer with an SSM
- mh- 1y agoSSM = state-space model, for the unfamiliar. https://en.wikipedia.org/wiki/State-space_representation https://en.wikipedia.org/wiki/State-space_representation
- samanator 1y agoYummy
- antirez 1y agoDear IBM name pickers: "Bamba", in Italian, means cocaine.
- deleted 1y ago[deleted]
- rzzzt 1y agoPara bailar La Bamba / Se necesita una poca de gracia
- amitport 1y agoMaybe? https://en.m.wikipedia.org/wiki/Bamba_(snack) https://en.m.wikipedia.org/wiki/Bamba_(snack) ;)
- akovaski 1y agoOr https://en.wikipedia.org/wiki/La_Bamba_(song) https://en.wikipedia.org/wiki/La_Bamba_(song)
- dantastic 1y agoOr (where I'm from) a school cafeteria: https://www.thelocal.se/20221125/swedish-word-of-the-day-bamba https://www.thelocal.se/20221125/swedish-word-of-the-day-bam...
- ofrzeta 1y agoSpot on. From the linked blog post "The refrain of La Bamba, the Mexican folk song that Ritchie Valens made famous, goes: Para bailar La Bamba/Se necesita una poca de Gracia. "
- francasso 1y agoSSMs never stop
- alex7o 1y agoIt's just a mamba (https://github.com/state-spaces/mamba https://github.com/state-spaces/mamba) but with a transformer. Idk where the B comes from.
- iddan 1y agoAnd in Heberw it's the name of a snack made of peanut-butter-flavored puffed maize https://en.wikipedia.org/wiki/Bamba_(snack) https://en.wikipedia.org/wiki/Bamba_(snack)
- bonzini 1y agoAs an Italian who has tried (only) the Israeli Bamba, I can certify that it is pretty addictive.
- kridsdale1 1y agoI imported these to America to feed my infant. Data shows the prevalence of peanut allergies lines up with when AAP guidelines started recommending that babies do NOT eat peanut. Israel never went along with this and thus has the lowest rates of allergies in the world.
- arijun 1y agoI think the difference in allergy rates between UK and Israeli Ashkenazi Jews (10x higher in UK Jews!) [1] is strong evidence for that. Also, they sell Bamba at Trader Joe’s now. [1] https://www.jacionline.org/article/S0091-6749(08)01698-9/fulltext#:~:text=Conclusions,genetic%20background%2C%20or%20peanut%20allergenicity. https://www.jacionline.org/article/S0091-6749(08)01698-9/ful...
- cycomanic 1y agoLatest research does strongly suggest that introducing small amounts of common allergens (peanuts, shellfish,milk products...) as early as possible does significantly reduce risk for allergies later. Many early childhood organisations already recommend this. Official health recommendations are often slow to catch up (often for good reasons, but introducing peanuts etc. early is already officially recommended in quite a few countries (Australia, NZ, Sweden for example AFAIK). Not all health professionals are always up to date either though.
- itayd 1y agoYou actually don't need to self import these. Usually Safeway (is it only a west coast thing?) always have these stocked in the Kosher section.
- beanjuiceII 1y agoi mean that sounds good to me
- rdtsc 1y agoSo someone can get fired for picking IBM after all! Or get a bonus, depending on the organization...
- vienzo 1y agoAnd in Lithuanian it's a navel
- _davide_ 1y agoWhen I read the title 'IBM crossed a transformer with an SSM and got ‘Bamba’' I laughed so hard I woke up my kid
- folgoris 1y agoA very funny and friendly way to say "cocaine" among italians. I'm struggling to read it seriously.
- lenerdenator 1y agoabout time they did something to liven things up at big blue
- dismalaf 1y agoSeems like a good fit.
- fb03 1y agoand in Portuguese, it means "flimsy". What a great name.
- deleted 1y ago[deleted]
- joshjob42 1y agoFor some reason this link isn't loading, but it's on https://archive.ph/Ks0xt https://archive.ph/Ks0xt
- jmward01 1y agoThis type of architecture is definitely the future. Unlimited attn is a dead end. As a human you don't need to scan an entire book just to guess what the next word will be and LLMs shouldn't need that either.
- quantadev 1y agoNot be contrarian, but if the next word prediction happens to be someone's name or a place or something discussed multiple places in the book then often, yes, a knowledge of the full plot of the book is "required" just to predict the next word, as you get to the middle or end of a book. For example you could never fill in the last chapter of any good book without having knowledge of every previous chapter. Not highly detailed knowledge, but still knowledge.
- parrit 1y agoWhat an LLM does is stuff it all into short term memory. Humans dump the first pages into long term memory and "make sense" of it. Humans have a massive context window because of this (and sheer brain size and efficiency).
- boroboro4 1y agoWe don’t put things into long term memory after we read it. We usually put it after night of sleep. I personally think that context (and kv cache correspondingly) in the models are akin to our short term memory, while training process (and actual weights) are to our long term memory. And we can’t be sure our short term memory doesn’t work in a way of matching the current context towards currently stored short term memory. From this perspective transformers are enough and just fine.
- parrit 1y agoSo if you now hide my original comment and try to recall what I said, do you know it word for word (and are thinking if every word, e.g. did I use one or 2 spaces somewhere as that would change tokens) or do you have a rough concept of what I said? OTOH if you had to remember a phone number to write it down, how does that differ?
- aantix 1y agoWhere's the code?
- beklein 1y agoI could find these two resources: Hugging Face: https://huggingface.co/collections/ibm-ai-platform/bamba-674f1388b9bbc98b413c7bab https://huggingface.co/collections/ibm-ai-platform/bamba-674... GitHub: https://github.com/foundation-model-stack/bamba https://github.com/foundation-model-stack/bamba
- jwilber 1y agoLLM/state space models have been popular for some years now, see: https://arxiv.org/abs/2212.14052 https://arxiv.org/abs/2212.14052 More recently, hybrid architectures that utilize attention plus other operators are gaining traction. See https://arxiv.org/abs/2503.01868 https://arxiv.org/abs/2503.01868
- adt 1y agohttps://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/ Love those GPQA scores hovering around 5% when chance (on 4-way multi-choice) would have got them 25%!
- gryfft 1y agoA stopped clock is right twice a day, but a running clock set to the wrong time is always wrong.
- parrit 1y agoThe RMS of wrongness of the running clock is probably lower.
- cwt137 1y agoNot always true! Your statement is only true when the running clock's speed is the same as time. Thus, regular time and the clock's time will never meet. If the clock is running faster than regular time, it will at point catch up to regular time and thus be correct for a split second. If the clock is slower than regular time, regular time will catch up to the clock and the clock will be right for a split second.
- actionfromafar 1y agoIf we are being pedantic, running clocks never run exactly the same as time. So they'll be right (very) much more seldom than the stopped clock, which is right twice a day.
- nathan_douglas 1y agoIf the clock is running backwards at very high speed, it would be right infinitely many times but the proportion of the time that it is right would approach some finite constant.
- k__ 1y agoMy girlfriend's microwave-clock runs faster than normal. Somehow this thing manages to accumulate an error of ~15 minutes in a month.
- mentalgear 1y ago> chose to make just about everything associated with Bamba open-source — the training recipes, the data, the data loader IBM designed for largescale distributed training, and a quantization framework aimed at shaving storage and inferencing costs.
- cubefox 1y agoAnother recent transformer/SSM hybrid is "M1", with a more than 3x claimed inference speed-up compared to equivalent transformers: https://arxiv.org/pdf/2504.10449 https://arxiv.org/pdf/2504.10449 IBM is claiming at least a 2x inference speed-up with Bamba. Both groups say that future SSM optimizations to vLLM would lead to further inference speed improvement.
- roger_ 1y agoNever got how mamba models work in multiple dimensions and non-causally.
- gitroom 1y agothe name bamba is killing me lol, all i can see is the snack now
- bushbaba 1y agoWonder if the name is inspired by my favorite snack, bamba. The best are the hazelnut bamba. Btw bamba if given to kids at a young age can drastically reduce the chance of peanut allergies
- visarga 1y agoLet me show you the etymology of Bamba: SSM (state space model) -> SSSM (structured state space model) -> (it's like a snake ssss...) Mamba -> Bamba
- flaviolivolsi 1y agoBamba means cocaine in Italian. Better not to give it to kids
- ericol 1y agoWell, have you ever heard of the Mitsubishi Pajero? [1] https://en.wikipedia.org/wiki/Mitsubishi_Pajero https://en.wikipedia.org/wiki/Mitsubishi_Pajero
- anentropic 1y ago> they added another trillion tokens and shrank the model from 18 GB to 9 GB through quantization, reducing its bit width from Mamba2’s 16-bit floating-point precision to 8-bits. This sounds like what they call "Bamba-9B" is actually an 18B model quantised to 8 bits. I thought generally we were naming models "nB" by their number of params and treating quantisation as a separate concern. Are there any other models that instead treat the name as an indicative memory requirement? Is this an attempt to hide that it fares poorly vs other ~18B parameter models? EDIT: no, I just misunderstood
- tmalsburg2 1y agoYeah, that's confusing, but the HuggingFace page says it has 9.78 B parameters. https://huggingface.co/ibm-ai-platform/Bamba-9B-fp8 https://huggingface.co/ibm-ai-platform/Bamba-9B-fp8
- cubefox 1y ago> This sounds like what they call "Bamba-9B" is actually an 18B model quantised to 8 bits. No it doesn't? The fact that it is 18 GB with 16 bit per parameter before quantization means that it is a 9B parameter model.
- anentropic 1y agoAh thanks, I see where I got confused now.
- OldSystemsFart 1y agoBamba in italian slang is cocaine, just to tell you