8 ms·
MK-1
- Philpax 3y ago...isn't this just quantization?
- atlas_hugged 3y agoExactly what I was thinking. Everyone already does this. Unless they’re doing something else, they’ll have to show why it’s better than just quickly quantizing to 8 bits or 4 bits or whatever.
- bhouston 3y agoWhatever it is, it will likely be copied into the open source tooling like llama.cop soonish or something similar will arrive in llama.cpp. It doesn’t seem defensive advantage. It seems like a feature and fighting against fast moving open source alternatives.
- amelius 3y agoIf you look at the demo video, the output is exactly the same for both cases, so I doubt it uses quantization.
- pestatije 3y ago> Today, we’re announcing our first product, MKML. MKML is a software package that can reduce LLM inference costs on GPUs by 2x with just a few lines of Python code. And it is plug and play with popular ecosystems like Hugging Face and PyTorch
- cududa 3y agoNo judgement, but I’m genuinely curious why you saw the need to comment with a random sentence in their post?
- qup 3y agoIt's not a random sentence, it's the main sentence everyone wants to read. They posted it to be helpful.
- pestatije 3y agoNo judgement taken...i try to go through hn article headers as quickly as possible. If theres an idiot that thinks "MK-1" is an appropriate title id prefer they dont bother to be honest. Missing that, i go through some comments to find out what is it about. If i have to waste minutes to find out what it is about then ill go and comment a summary
- metadat 3y agoToo bad it's not an open source effort. I'm not a fan of proprietary dependencies in my stack, full stop.
- lolinder 3y agoI seriously doubt this will go anywhere. The open source community has already achieved basically the same performance improvements via quantization. This feels like someone has repackaged those libraries and is going to try to sell them to unwary and uninformed AI startups.
- lolinder 3y agoIt's weird that not once do they mention or compare their results to the already-available quantization methods. I normally try to give benefit of the doubt, but there's really no way they're not aware that there are already widely used techniques for accomplishing this same thing, so the comparison benchmarks really should be there. To fill in the gap, here's llama.cpp's comparison chart[0] for the different quantizations available for Llama 1. We can't compare directly with their Llama 2 metrics, but just comparing the percent change in speed and perplexity, MK-1 looks very similar to Q5_1. There's a small but not insignificant hit to perplexity, and a just over 2x speedup. If these numbers are accurate, you can download pre-quantized Llama 2 models from Hugging Face that will perform essentially the same as what MK-1 is offering, with the Q5 files here: https://huggingface.co/TheBloke/Llama-2-13B-GGML/tree/main https://huggingface.co/TheBloke/Llama-2-13B-GGML/tree/main [0] https://github.com/ggerganov/llama.cpp#quantization https://github.com/ggerganov/llama.cpp#quantization
- moffkalast 3y agoQ5_1 is already old news too, K quants are faster and more space efficient for the same perplexity loss. https://old.reddit.com/r/LocalLLaMA/comments/142q5k5/updated_relative_comparison_of_ggml_quantization/ https://old.reddit.com/r/LocalLLaMA/comments/142q5k5/updated...
- lolinder 3y agoFor sure, but I couldn't find numbers for the K quants that included inference speeds, so I settled on the older one. If MK-1 were trying to be honest they'd definitely want to benchmark against the newest methods!
- andy_xor_andrew 3y agoAlso, using the word "codecs" kind of puts a bad taste in my mouth. It's like they're trying to sound like they invented an entirely new paradigm, with their own fancy name that reminds people of video compression.
- 3y ago
- ipsum2 3y agoIsn't FasterTransformer (NVidia, OSS) and text-generation-inference (Huggingface, not OSS) are faster than this?
- xianshou 3y agoNot a single mention of existing quantization techniques? Ten bucks says this is just a wrapper around bitsandbytes or ggml.
- drtournier 3y agoMKML == abstractions and wrappers for GGML?
- Scene_Cast2 3y agoI've worked on ML model quantization. The open source 4-bit or 8-bit quantization isn't as good as one can get - there are much fancier techniques to keep predictive performance while squeezing size. Some techniques (like quantization-aware training) involve changes to training.
- lolinder 3y agoI'm sure there are better methods! But in this case, MKML's numbers just don't look impressive when placed alongside the prominent quantization techniques already in use. According to this chart [0] it's most similar in size to a Q6_K quantization, and if anything has slightly worse perplexity. If their technique were better, I imagine that the company would acknowledge the existence of the open source techniques and show them in their comparisons, instead of pretending the only other option is the raw fp16 model. [0] https://old.reddit.com/r/LocalLLaMA/comments/142q5k5/updated_relative_comparison_of_ggml_quantization/ https://old.reddit.com/r/LocalLLaMA/comments/142q5k5/updated...
- Scene_Cast2 3y agoFrom what I remember, non-power-of-2 compression schemes tank inference speed (assuming Q6_k is 6-bit; I haven't actually verified if ggml q6_K llama is slow). Meanwhile, the site claims a speed-up. But I do actually agree with you - they should really be benchmarking against popular competitors. In my experience, fancier quantization is a _lot_ of work for fairly little gain (at least for neural nets). I also think that ML techniques such as quantization (or fancy param sweeps, feature pruning, that kind of stuff) tend to either get in-housed (i.e. the model will come quantized from the source) or get open-sourced. In-housing of ML techniques tends to happen more often if there's a money-making model where the hardware running the model costs money, but running the model brings in money.
- KRAKRISMOTT 3y agoWhat about Unum's quantization methods? https://github.com/unum-cloud/usearch https://github.com/unum-cloud/usearch
- rvz 3y agoAnother AI startup grift, using GGML and closing it up to beg for VC cash. Yet another AI wrapper company doing the same thing and jumping on the LLM hype train before it dries up. If it is not open source and it is closed, it is immediately dead in the water.
- modeless 3y agoHow does this compare to mlc-llm with 4 bit quantization? It runs llama2 13B incredibly fast on my 4090. Multiples of the speed of llama.cpp even on GPU with the same 4 bit quantization.
- deleted 3y ago[deleted]
- brucethemoose2 3y agoYeah, that TVM Vulkan autotuning is incredible. And its not even using the matmul Vulkan extension, I dont think. MLC's 4 bit quantization is "dumb" compared to llama.cpp, which reduces perplexity (and also explains some of the speed difference), but the biggest missing feature is CPU offloading (which would allow you to run 70B reasonably well on a 4090). I think the holy grail of local llm inference is llama 70B, run in TVM, split between the GPU and IGP. It feels like we are inches away... All the pieces are there, but there are no front end devs connecting those dots.
- modeless 3y agoWow, using the IGP for the parts that don't fit on the discrete GPU is a great idea.
- brucethemoose2 3y agoYeah. It might be possible in llama.cpp soon, but the vulkan implementation may or may not be fast on the IGP.
- gsuuon 3y agoInterestingly, mlc's web-llm runs much better on my iGPU than dGPU when the model size goes over available dGPU vram. Llama 2 7B runs faster on dGPU, but when I switch to Llama 2 13B suddenly my iGPU outperforms. I think because the iGPU effectively utilizes shared memory?
- 3y ago
- lyapunova 3y agoI don't think I can use this if it's not open source... sorry. The field moves too fast and the convenience is just not there otherwise. edit: also the branding makes me think of MK-ultra which is probably something to avoid
- dheera 3y ago[flagged]
- haswell 3y agoThis is completely unrelated
- mugivarra69 3y agowe have no moat and so does not openai is getting better by time.
- hardwaresofton 3y agoIs this the true effect of Ultra Instinct^H^H Llama2? Facebook is effectively supercharging the ecosystems and tool builders and smaller inference services. This company had access to a credible, popular model (with an actual OSS license), and the relevant weights so they could optimize on it and sell the optimization without worrying about the license/restrictions on the weights themselves.
- radicaldreamer 3y agoYou can do this stuff on a MacBook Pro these days... not sure why you'd want to be locked into another vendor here. Either use the best (OpenAI, Anthropic) or just roll your own.
- ushakov 3y agoThis seems more like a VC Pitchdeck rather than a technical paper explaining why their approach is better