6 ms·
Here's a glossary to understand this post: - mixtral-8x7 or 8x7: Open source model by Mistral AI. - Dolphin: An uncensored version of the mistral model - 3.5
by EmilStenstrom 3y ago
Here's a glossary to understand this post:
- mixtral-8x7 or 8x7: Open source model by Mistral AI.
- Dolphin: An uncensored version of the mistral model
- 3.5-turbo: GPT-3.5 Turbo, the cheapest API from OpenAI
- 4-series preview OR "4.5 preview": GPT-4 Turbo, the most capable API from OpenAI
- mistral-medium: A new model by Mistral AI that they are only serving through AI. It's in private beta and there's a waiting list to access it.
- Perplexity: A new search engine that is challenging Google by applying LLM to search
- Sama: Sam Altman, CEO of OpenAI
- RenTech: Renaissance Technologies, a secretive hedge fund known for delivering impressive returns improving on the work of others
- DPO: Direct Preference Optimization. It is a technique that leverages AI feedback to optimize the performance of smaller, open-source models like Zephyr-7B1.
- Alibi: a Python library that provides tools for machine learning model inspection and interpretation2. It can be used to explain the predictions of any black-box model, including LLMs.
- Sliding window: a type of attention mechanism introduced by Mistral-7B3. It is used to support longer sequences in LLMs.
- Modern mixtures: The process of using multiple models together, like "mixtral" is a mixture of several mistral models.
- TheBloke: Open source developer that is very quick at quantizing all new models that come out
- Quantize: Decreasing memory requirements of a new model by decreasing the precision of weights, typically with just minor performance degradation.
- 4070 Super: NVIDIA 4070 Super, new graphics card announced just a week ago
- MSFT: Microsoft
- azeirah 3y agoThat's an impressive list of jargon whaha Love how deep the rabbithole has gone in just a year. I am unfortunately in the camp of understanding the post without needing a glossary. I should go outside more :|
- rrr_oh_man 3y agoI love you, Emil
- benreesman 3y agoI'm clearly spending far too much time tuning/training/using these things if a glossary to make my post comprehensible to HN is longer than my remark: thank you for correcting my error in dragging this sub-sub-sub-field into a thread of general interest.
- neals 3y agoCrazy, your post feels like downloading martial arts in the Matrix. I read the parent, didn't get a thing and though the guy was on substances. Read yours. Read the parent again. I speak AI now! I'm going to use this new power to raise billions!
- Smerity 3y agoI think you've done a great explanation expansion except I believe it's ALiBi ("Attention with Linear Biases Enables Input Length Extrapolation"), a method of positional encoding (i.e. telling the Transformer model how much to weight a distant token when computing the current output token). This has been used on various other LLMs[2]. [1]: https://arxiv.org/abs/2108.12409 https://arxiv.org/abs/2108.12409 [2]: n.b. Ofir Press is co-creator of ALiBi https://twitter.com/OfirPress/status/1654538361447522305 https://twitter.com/OfirPress/status/1654538361447522305
- benreesman 3y agoThis is indeed what I was referring to and along with RoPE and related techniques is a sort of "meta-attention" in which a cost-effective scalar pointwise calculation can hint the heavyweight attention mechanism with super-linear returns in practical use cases. In more intuitive terms, your bog-standard transformer overdoes it in terms of considering all context equally in the final prediction, and we historically used rather blunt-force instruments like causally masking everything to zero. These techniques are still heuristic and I imagine every serious shop has tweaks and tricks that go with their particular training setup, but the Rope shit in general is kind of a happy medium and exploits locality at a much cheaper place in the overall computation.
- Kerbonut 3y agoimo mistral-medium is worse than mixtral. Do you have API access?
- lhl 3y agoMy understanding is that Mistral uses a regular 4K RoPE that is "extends" the window size with SWA. This is based on looking at the results of Nous Research's Yarn-Mistral extension: https://huggingface.co/NousResearch/Yarn-Mistral-7b-128k https://huggingface.co/NousResearch/Yarn-Mistral-7b-128k and Self-Extend, both of which only apply to RoPE models. There are quite a few recent attention extension techniques recently published: * Activation Beacons - up to 100X context length extension in as little as 72 A800 hours https://huggingface.co/papers/2401.03462 https://huggingface.co/papers/2401.03462 * Self-Extend - a no-training RoPE modification that can give "free" context extension with 100% passkey retrieval (works w/ SWA as well) https://huggingface.co/papers/2401.01325 https://huggingface.co/papers/2401.01325 * DistAttention/DistKV-LLM - KV cache segmentation for 2-19X context length at runtime https://huggingface.co/papers/2401.02669 https://huggingface.co/papers/2401.02669 * YaRN - aforementioned efficient RoPE extension https://huggingface.co/papers/2309.00071 https://huggingface.co/papers/2309.00071 You could imagine combining a few of these together to basically "solve" the context issue while largely training for shorter context length. There are of course some exciting new alternative architectures, notably Mamba https://huggingface.co/papers/2312.00752 https://huggingface.co/papers/2312.00752 and Megabyte https://huggingface.co/papers/2305.07185 https://huggingface.co/papers/2305.07185 that can efficiently process up to 1M tokens...
- pandemic_region 3y agoDid you just paste that into an LLM and asked it to create a glossary? :-P (but seriously: Thanks !)
- coldtea 3y agoEmil didn't, but I did (and yeah, it's useless): Mixtral-8x7: This appears to be a technical term, possibly referring to a software, framework, or technology. Its exact nature is unclear without additional context. Dolphin locally: "Dolphin" could refer to a software tool or framework. The term "locally" implies it is being run on a local machine or server rather than a remote or cloud-based environment. 3.5-turbo: This could be a version name or a type of technology. "Turbo" often implies enhanced or accelerated performance. 4-series preview: Likely a version or iteration of a software or technology that is still in a preview or beta stage, indicating it's not the final release. Emacs: A popular text editor used often by programmers and developers. Known for its extensibility and customization. Mistral Medium: This might be a product or service, possibly in the realm of technology or AI. The specific nature is not clear from the text alone. Perplexity: Likely a company or service provider, possibly in the field of AI or technology. They seem to have a partnership offering involving Mistral Medium. RenTech of AI: RenTech, or Renaissance Technologies, is a well-known quantitative hedge fund. The term here is used metaphorically to suggest a pioneering or leading position in the AI field. DPO, Alibi, and sliding window: These are likely technical concepts or tools in the field being discussed. Without additional context, their exact meanings are unclear. Modern mixtures: This could refer to modern algorithms, techniques, or technologies in the field of AI or data science. TheBloke: This could be a reference to an individual, a role within a community, or a specific entity known for certain expertise or actions. 4070 Super: This seems like a model name, possibly of a computer hardware component like a GPU (Graphics Processing Unit). MSFT: An abbreviation for Microsoft Corporation. On-premise: Refers to software or services that are operated from the physical premises of the organization, as opposed to being hosted on the cloud.
- aftoprokrustes 3y agoThis is actually hilarious. It looks like a student who did not learn for the exam but still tries their best to scratch a point or two by filling the page with as many reasonnable sounding statements (a.k.a. "bullshit") as they can. Not that I expect more of a language model, no matter how "large".
- deleted 3y ago[deleted]
- vincentrolfs 3y agoI asked ChatGPT to rewrite the original post using your glossary, which worked well: I've set up my system to use several AI models: the open-source Mixtral-8x7, Dolphin (an uncensored version of Mixtral), GPT-3.5 Turbo (a cost-effective option from OpenAI), and the latest GPT-4 Turbo from OpenAI. I can easily compare their performances in Emacs. Lately, I've noticed that GPT-4 Turbo is starting to outperform Mixtral-8x7, which wasn't the case until recently. However, I'm still waiting for access to Mistral-Medium, a new, more exclusive AI model by Mistral AI. I just found out that Perplexity, a new search engine competing with Google, is offering free access to Mistral Medium through their partnership. This makes me question Sam Altman, the CEO of OpenAI, and his claims about their technology. Mistral Medium seems superior to GPT-4 Turbo, and if it were expensive to run, Perplexity wouldn't be giving it away. I'm guessing that Mistral AI could become the next Renaissance Technologies (a hedge fund known for its innovative strategies) of the AI world. Techniques like Direct Preference Optimization, which improves smaller models, along with other advancements like the Alibi Python library for understanding AI models, sliding windows for longer text sequences, and combining multiple models, are now well understood. The real opportunity lies in quickly adapting these new technologies before they become mainstream and affordable. Big companies are cautious about adopting these new structures, remembering their dependence on Microsoft in the past. They're willing to experiment with AI until it becomes both affordable and easy to use in-house. It's sad to see the old technology go, but exciting to see the new advancements take its place.
- benreesman 3y agoThe GP did a great job summarizing the original post and defining a lot of cryptic jargon that I didn't anticipate would generate so much conversation, and I'd wager did it without a blind LLM shot (though these days even that is possible). I endorse that summary without reservation. And the above is substantially what I said, and undoubtedly would find a better reception with a larger audience. I'm troubled though, because I already sanitize what I write and say by passing it through a GPT-style "alignment" filter in almost every interaction precisely because I know my authentic self is brash/abrasive/neuro-atypical/etc. and it's more advantageous to talk like ChatGPT than to talk like Ben. Hacker News is one of a few places real or digital where I just talk like Ben. Maybe I'm an outlier in how different I am and it'll just be me that is sad to start talking like GPT, and maybe the net change in society will just be a little drift towards brighter and more diplomatic. But either way it's kind of a drag: either passing me and people like me through a filter is net positive, which would suck but I guess I'd get on board, or it actually edits out contrarian originality in toto, in which case the world goes all Huxley really fast. Door #3 where we net people out on accomplishment and optics with a strong tilt towards accomplishment doesn't seem to be on the menu.
- spuz 3y agoAs someone who follows AI pretty closely, this was unbelievably helpful in understanding the parent post. It's crazy how much there is to keep on top of if you don't want to fall behind everything that is going on in AI at the moment.
- hmottestad 3y agoThanks for this. I was initially wondering what this new GPT 4.5 model was and if I had somehow missed out on something big.