20 ms·
Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
- deleted 2y ago[deleted]
- TheAceOfHearts 2y agoI still can't access the hosted model at meta.ai from Puerto Rico, despite us being U.S. citizens. I don't know what Meta has against us. Could someone try giving the 90b model this word search problem [0] and tell me how it performs? So far with every model I've tried, none has ever managed to find a single word correctly. [0] https://imgur.com/i9Ps1v6 https://imgur.com/i9Ps1v6
- deleted 2y ago[deleted]
- Workaccount2 2y agoThis is likely because the models use OCR on images with text, and once parsed the word search doesn't make sense anymore. Would be interesting to see a model just working on raw input though.
- simonw 2y agoImage models such as Llama 3.2 11B and 90B (and the Claude 3 series, and Microsoft Phi-3.5-vision-instruct, and PaliGemma, and GPT-4o) don't run OCR as a separate step. Everything they do is from that raw vision model.
- paxys 2y agoNon US citizens can access the model just fine, if that's what you are implying.
- TheAceOfHearts 2y agoI'm not implying anything. It's just frustrating that despite being a US territory with US citizens, PR isn't allowed to use this service without any explanation.
- deleted 2y ago[deleted]
- paxys 2y agoJust because you cannot access the model doesn't mean all of Puerto Rico is blocked.
- TheAceOfHearts 2y agoWhen I visit meta.ai it says: > Meta AI isn't available yet in your country Maybe it's just my ISP, I'll ask some friends if they can access the service.
- paxys 2y agometa.ai is their AI service (similar to ChatGPT). The model source itself is hosted on llama.com.
- TheAceOfHearts 2y agoI'm aware. I wanted to try out their hosted version of the model because I'm GPU poor.
- elcomet 2y agoYou can try it on hugging face
- daemonologist 2y agoBoth Llama 3.2 90B and Claude 3.5 Sonnet can find "turkey" and "spoon", probably because they're left-to-right. Llama gave approximate locations for each and Claude gave precise but slightly incorrect locations. Further prompting to look for diagonal and right-to-left words returned plausible but incorrect responses, slightly more plausible from Claude than Llama. (In this test I cropped the word search to just the letter grid, and asked the model to find any English words related to soup.) Anyways, I think there just isn't a lot of non-right-to-left English in the training data. A word search is pretty different from the usual completion, chat, and QA tasks these models are oriented towards; you might be able to get somewhere with fine-tuning though.
- gunalx 2y agoTry and find where the words are in this word puzzle undefined ''' There are two words in this word puzzle: "soup" and "mix". The word "soup" is located in the top row, and the word "mix" is located in the bottom row. ''' Edit: Tried a bit more probing like asking it to find spoon or any other word. It just makes up a row and column.
- nmwnmw 2y ago- Llama 3.2 introduces small vision LLMs (11B and 90B parameters) and lightweight text-only models (1B and 3B) for edge/mobile devices, with the smaller models supporting 128K token context. - The 11B and 90B vision models are competitive with leading closed models like Claude 3 Haiku on image understanding tasks, while being open and customizable. - Llama 3.2 comes with official Llama Stack distributions to simplify deployment across environments (cloud, on-prem, edge), including support for RAG and safety features. - The lightweight 1B and 3B models are optimized for on-device use cases like summarization and instruction following.
- deleted 2y ago[deleted]
- opdahl 2y agoI'm blown away with just how open the Llama team at Meta is. It is nice to see that they are not only giving access to the models, but they at the same time are open about how they built them. I don't know how the future is going to go in the terms of models, but I sure am grateful that Meta has taken this position, and are pushing more openness.
- nickpsecurity 2y agoDo they tell you what training data they use for alignment? As in, what biases they intentionally put in the system they’re widely deploying?
- warkdarrior 2y agoDo you have some concrete example of biases in their models? Or are you just fishing for something to complain about?
- ericjmorey 2y agoEven without intentionally biasing the model, without knowing the biases that exist in the training data, they're just biased black boxes that come with the overhead of figuring out how it's biased. All data is biased, there's no avoiding that fact.
- slt2021 2y agobias is some normative lens that some people came up with, but it is purely subjective and is a social construct, that has roots in the area of social justice and has nothing to do with the LLM. the proof is that all critics of AI/LLM have never ever produced a single "unbiased" model. If unbiased model does not exist (at least I never seen an AI/LLM sceptics community produce one), then the concept of bias is useless. Just a fluffy word that does not mean anything
- semi-extrinsic 2y ago
- resters 2y agoThis is great! Does anyone know if the llama models are trained to do function calling like openAI models are? And/or are there any function calling training datasets?
- refulgentis 2y agoYes (rationale: 3.1 was, would be strange to rollback.) In general, you'll do a ton of damage by constraining token generation to valid JSON - I've seen models as small as 800M handle JSON with that. It's ~impossible to train constraining into it with remotely the same reliability -- you have to erase a ton of conversational training that makes it say ex. "Sure! Here's the JSON you requested:"
- Closi 2y agoWhat about OpenAI Structured Outputs? This seems to do exactly this.
- refulgentis 2y agoCorrect, I think so too, seemed that update must be doing exactly this. tl;dr: in the context of Llama fn calling reliability, you don't need to reach for training, in fact, you'll do it and still have the same problem.
- zackangelo 2y agoI'm building this type of functionality on top of Llama models if you're interested: https://docs.mixlayer.com/examples/json-output https://docs.mixlayer.com/examples/json-output
- refulgentis 2y agoI'm writing a Flutter AI client app, integrates with llama.cpp. I used a PoC of llama.cpp running in WASM, I'm desperate to signal the app is agnostic to AI provider, but it was horrifically slow, ended up backing out to WebMLC. What are you doing underneath, here? If thats secret sauce, I'm curious what you're seeing in tokens/sec on ex. a phone vs. MacBook M-series. Or are you deploying on servers?
- moffkalast 2y agoI've just tested the 1B and 3B at Q8, some interesting bits: - The 1B is extremely coherent (feels something like maybe Mistral 7B at 4 bits), and with flash attention and 4 bit KV cache it only uses about 4.2 GB of VRAM for 128k context - A Pi 5 runs the 1B at 8.4 tok/s, haven't tested the 3B yet but it might need a lower quant to fit it and with 9T training tokens it'll probably degrade pretty badly - The 3B is a certified Gemma-2-2B killer Given that llama.cpp doesn't support any multimodality (they removed the old implementation), it might be a while before the 11B and 90B become runnable. Doesn't seem like they outperform Qwen-2-VL at vision benchmarks though.
- Patrick_Devine 2y agoHoping to get this out soon w/ Ollama. Just working out a couple of last kinks. The 11b model is legit good though, particularly for tasks like OCR. It can actually read my cursive handwriting.
- jsarv 2y agoNaah, Qwen2-VL-7b still is much much better than 11b model for handwritten OCR from what i have tested. The 11b model hallucinates in case of handwritten OCR.
- sumedh 2y agoWhere can I try it out. The playground on their homepage is very slow. I am willing to pay for it as well if the OCR is good.
- Ey7NFZ3P0nzAe 2y agoOpenrouter.ai
- gdiamos 2y agoLlama 3.2 includes a 1B parameter model. This should be 8x higher throughput for data pipelines. In our experience, smaller models are just fine for simple tasks like reading paragraphs from PDF documents.
- deleted 2y ago[deleted]
- gdiamos 2y agoDo inference frameworks like vllm support vision?
- woodson 2y agoYes, vLLM does (though marked experimental): https://docs.vllm.ai/en/latest/models/vlm.html https://docs.vllm.ai/en/latest/models/vlm.html
- theaniketmaurya 2y agoYou can run with LitServe. here is the code - https://lightning.ai/lightning-ai/studios/deploy-llama-3-2-vision-with-litserve https://lightning.ai/lightning-ai/studios/deploy-llama-3-2-v...
- minimaxir 2y agoOff topic/meta, but the Llama 3.2 news topic received many, many HN submissions and upvotes but never made it to the front page: the fact that it's on the front page now indicates that moderators intervened to rescue it: https://news.ycombinator.com/from?site=meta.com https://news.ycombinator.com/from?site=meta.com (showdead on) If there's an algorithmic penalty against the news for whatever reason, that may be a flaw in the HN ranking algorithm.
- makin 2y agoThe main issue was that Meta quickly took down the first announcement, and the only remaining working submission was the information-sparse HuggingFace link. By the time the other links were back up, it was too late. Perfect opportunity for a rescue.
- deleted 2y ago[deleted]
- senko 2y agoYeah I submitted what turned out to be a dupe but I could never find the original, probably was buried at the time. Then a few hours later it miraculously (re?)appeared. AIUI exact dupes just get counted as upvotes, which hasn’t happened in my case.
- dhbradshaw 2y agoTried out 3B on ollama, asking questions in optics, bio, and rust. It's super fast with a lot of knowledge, a large context and great understanding. Really impressive model.
- tomComb 2y agoI question whether a 3B model can have “a lot of knowledge”.
- foxhop 2y agoMy guess is it uses the same vocabulary size as llama 3.1 which is 128,000 different tokens (words) to support many languages. Parameter count is less of an indicator of fitness than previously thought.
- lolinder 2y agoThat doesn't address the thing they're skeptical about, which is how much knowledge can be encoded in 3B parameters. 3B models are great for text manipulation, but I've found them to be pretty bad at having a broad understanding of pragmatics or any given subject. The larger models encode a lot more than just language in those 70B+ parameters.
- cyanydeez 2y agoOk, but what we are probably debating is knowledge versus wisdom. Like, if I know 1+1 = 2, and I know the numbers 1 through 10, my knowledge is just 11, but my wisdom is infinite in the scope of integer addition. I can find any number, given enough time. I'm pretty sure the AI guys are well aware of which types of models they want to produce. Models that can intake knowledge and intelligently manipulate it would mean general intelligence. Models that can intake knowledge and only produce subsets of it's training data have a use but wouldn't be general intelligence.
- BoorishBears 2y ago
- sva_ 2y agoCurious about the multimodal model's architecture. But alas, when I try to request access > Llama 3.2 Multimodal is not available in your region. It sounds like they input the continuous output of an image encoder into a transformer, similar to transfusion[0]? Does someone know where to find more details? Edit: > Regarding the licensing terms, Llama 3.2 comes with a very similar license to Llama 3.1, with one key difference in the acceptable use policy: any individual domiciled in, or a company with a principal place of business in, the European Union is not being granted the license rights to use multimodal models included in Llama 3.2. [1] What a bummer. 0. https://www.arxiv.org/abs/2408.11039 https://www.arxiv.org/abs/2408.11039 1. https://huggingface.co/blog/llama32#llama-32-license-changes-sorry-eu- https://huggingface.co/blog/llama32#llama-32-license-changes...
- _ink_ 2y agoOh. That's sad indeed. What might be the reason for excluding Europe?
- paxys 2y agoPunishment. "Your government passes laws we don't like, so we aren't going to let you have our latest toys".
- Arubis 2y agoGlibly, Europe has the gall to even consider writing regulations without asking the regulated parties for permission.
- pocketarc 2y agoBetween this and Apple's policies, big tech corporations really seem to be putting the screws to the EU as much as they can. "See, consumers? Look at how bad your regulation is, that you're missing out on all these cool things we're working on. Talk to your politicians!" Regardless of your political opinion on the subject, you've got to admit, at the very least, it will be educational to see how this develops over the next 5-10 years of tech progress, as the EU gets excluded from more and more things.
- getcrunk 2y agoStill no 14/30b parameter models since llama 2. Seriously killing real usability for power users/diy. The 7/8B models are great for poc and moving to edge for minor use cases … but there’s a big and empty gap till 70b that most people can’t run. The tin foil hat in me is saying this is the compromise the powers that be have agreed too. Basically being “open” but practically gimped for average joe techie. Basically arms control
- swader999 2y agoYou don't need an F-15 to play at least, a decent sniper rifle will do. You can still practise even with a pellet gun. I'm running 70b models on my M2 max with 96 ram. Even larger models sort of work, although I haven't really put much time into anything above 70b.
- int_19h 2y agoWith a 128Gb Mac, you can even run 405b at 1-bit quantization - it's large enough that even with the considerable quality drop that entails, it still appears to be smarter than 70b.
- ComputerGuru 2y agoJust to clarify, you are saying 1b-quantized 405b is smarter than 70b unquantized?
- int_19h 2y agoYou need to quantize 70b to run it on that kind of hardware as well, since even float16 wouldn't fit. But 405b:IQ1_M seems to be smarter than 70b:Q4_K_M in my experiments (admittedly very limited because it's so slow). Note that IQ1_M quants are not really "1-bit" despite the name. It's somewhere around 1.8bpw, which just happens to be enough to fit the model into 128Gb with some room for inference.
- foxhop 2y ago4090 has 24G So we really need ~40B or G model (two cards) or like a ~20B with some room for context window. 5090 has ??G - still unreleased
- kingkongjaffa 2y agollama3.2:3b-instruct-q8_0 is performing better than 3.1 8b-q4 on my macbookpro M1. It's faster and the results are better. It answered a few riddles and thought experiments better despite being 3b vs 8b. I just removed my install of 3.1-8b. my ollama list is currently: $ ollama list NAME ID SIZE MODIFIED llama3.2:3b-instruct-q8_0 e410b836fe61 3.4 GB 2 hours ago gemma2:9b-instruct-q4_1 5bfc4cf059e2 6.0 GB 3 days ago phi3.5:3.8b-mini-instruct-q8_0 8b50e8e1e216 4.1 GB 3 days ago mxbai-embed-large:latest 468836162de7 669 MB 3 months ago
- taneq 2y agoFor a second I read that as “it just removed my install of 3.1-8b” :D
- fragmede 2y agohttps://github.com/KillianLucas/open-interpreter/ https://github.com/KillianLucas/open-interpreter/
- PhilippGille 2y agoAren't the _0 quantizations considered deprecated and _K_S or _K_M preferable? https://github.com/ollama/ollama/issues/5425 https://github.com/ollama/ollama/issues/5425
- Patrick_Devine 2y agoFor _K_S definitely not. We quantized 3b with q4_K_M since we were getting good results out of it. Officially Meta has only talked about quantization for 405b and hasn't given any actual guidance for what the "best" quantization should be for the smaller models. With The 1b model we didn't see good results with any of the 4b quantizations and went with q8_0 as the default.
- aryehof 2y agoOn what basis do you use these different models?
- sk11001 2y agoCan one of thse models be run on a single machine? What specs do you need?
- Y_Y 2y agoAbsolutely! They have a billion-parameter model that will run on my first computer if we quantize it to 1.5 bits. But realistically yes, if you can fit in total ram you can run it slowly, if you can fit it in gpu ram you can probably run it fast enough to chat.
- sumedh 2y agoThe 8B models run fine on a M1 pro 16GB.
- GaggiX 2y agoThe 90B seem to perform pretty weak on visual tasks compare to Qwen2-VL-72B: https://huggingface.co/Qwen/Qwen2-VL-72B-Instruct https://huggingface.co/Qwen/Qwen2-VL-72B-Instruct, or am I missing something?
- kombine 2y agoAre these models suitable for Code assistance - as an alternative to Cursor or Copilot?
- bboygravity 2y agoI use Continue on VScode, works well with Ollama and llama3.1 (but obviously not as good as Claude).
- a_wild_dandan 2y ago"The Llama jumped over the ______!" (Fence? River? Wall? Synagogue?) With 1-hot encoding, the answer is "wall", with 100% probability. Oh, you gave plausibility to "fence" too? WRONG! ENJOY MORE PENALTY, SCRUB! I believe this unforgiving dynamic is why model distillation works well. The original teacher model had to learn via the "hot or cold" game on text answers. But when the child instead imitates the teacher's predictions, it learns semantically rich answers. That strikes me as vastly more compute-efficient. So to me, it makes sense why these Llama 3.2 edge models punch so far above their weight(s). But it still blows my mind thinking how far models have advanced from a year or two ago. Kudos to Meta for these releases.
- adtac 2y ago>WRONG! ENJOY MORE PENALTY, SCRUB! Is that true tho? During training, the model predicts {"wall": 0.65, "fence": 0.25, "river": 0.03}. Then backprop modifies the weights such that it produces {"wall": 0.67, "fence": 0.24, "river": 0.02} next time. But it does that with a much richer feedback than WRONG! because we're also telling the model how much more likely "fence" is than "wall" in an indirect way. It's likely most of the neurons that supported "wall" also supported "fence", so the average neuron that supported "river" gets penalised much more than a neuron that supported "fence". I agree that distillation is more efficient for exactly the same reason, but I think even models as old as GPT-3 use this trick to work as well as they do.
- snovv_crash 2y agoYou are in violent agreement with GP.
- refulgentis 2y agoThey don't, they're playing "hide the #s" a bit. Llama 3.2 3B is definitively worse than Phi-3 from May, both on any given metric and in an hour of playing with the 2, trying to justify moving to Llama 3.2 at 3B, given I'm adding Llama 3.2 at 1B.
- whimsicalism 2y agoyeah i mean that is exactly why distillation works. if you just were one hotting it would be the same as training on same dataset
- bottlepalm 2y agoWhat mobile devices can the smaller models run on? iPhone, Android?
- jillion 2y agoapparently so, but im trying to find a working example / some details on what specific iOS / android devices are capable of running this
- simonw 2y agoI'm absolutely amazed at how capable the new 1B model is, considering it's just a 1.3GB download (for the Ollama GGUF version). I tried running a full codebase through it (since it can handle 128,000 tokens) and asking it to summarize the code - it did a surprisingly decent job, incomplete but still unbelievable for a model that tiny: https://gist.github.com/simonw/64c5f5b111fe473999144932bef4218b https://gist.github.com/simonw/64c5f5b111fe473999144932bef42... More of my notes here: https://simonwillison.net/2024/Sep/25/llama-32/ https://simonwillison.net/2024/Sep/25/llama-32/ I've been trying out the larger image models to using the versions hosted on https://lmarena.ai/ https://lmarena.ai/ - navigate to "Direct Chat" and you can select them from the dropdown and upload images to run prompts.
- GaggiX 2y agoLlama 3.2 vision models don't seem that great if they have to compare them to Claude 3 Haiku or GPT4o-mini. For an open alternative I would use Qwen-2-72B model, it's smaller than the 90B and seems to perform quite better. Also Qwen2-VL-7B as an alternative to Llama-3.2-11B, smaller, better in visual benchmarks and also Apache 2.0. Molmo models: https://huggingface.co/collections/allenai/molmo-66f379e6fe3b8ef090a8ca19 https://huggingface.co/collections/allenai/molmo-66f379e6fe3..., also seem to perform better than Llama-3.2 models while being smaller and Apache 2.0.
- dannyobrien 2y agoWhat interface do you use for a locally-run Qwen2-VL-7B? Inspired by Simon Willison's research[1], I have tried it out on Hugging Face[2]. Its handwriting recognition seems fantastic, but I haven't figured out how to run it locally yet. [1] https://simonwillison.net/2024/Sep/4/qwen2-vl/ https://simonwillison.net/2024/Sep/4/qwen2-vl/ [2] https://huggingface.co/spaces/GanymedeNil/Qwen2-VL-7B https://huggingface.co/spaces/GanymedeNil/Qwen2-VL-7B
- Eisenstein 2y agoMiniCPM-V 2.6 is based on Qwen 2 and is also great at handwriting. It works locally with KoboldCPP. Here are the results I got with a test I just did. Image: * https://imgur.com/wg0kdQK https://imgur.com/wg0kdQK Output: * https://pastebin.com/RKvYQasi https://pastebin.com/RKvYQasi OCR script used: * https://github.com/jabberjabberjabber/LLMOCR/blob/main/llmocr.py https://github.com/jabberjabberjabber/LLMOCR/blob/main/llmoc... Model weights: MiniCPM-V-2_6-Q6_K_L.gguf, mmproj-MiniCPM-V-2_6-f16.gguf Inference: * https://github.com/LostRuins/koboldcpp/releases/tag/v1.75.2 https://github.com/LostRuins/koboldcpp/releases/tag/v1.75.2
- JohnHammersley 2y agoOllama post: https://ollama.com/blog/llama3.2 https://ollama.com/blog/llama3.2
- gunalx 2y ago3b was pretty good at multimodal (Norwegian) still a lot of gibberish at times, and way more sensitive than 8b but more usable than Gemma 2 2b at multi modal, fine at my python list sorter with args standard question. But 90b vision just refuses all my actually useful tasks like helping recreate the images in html or do anything useful with the image data other than describing it. Have not gotten as stuck with 70b or openai before. Insane amount of refusals all the time.
- thimabi 2y agoDoes anyone know how these models fare in terms of multilingual real-world usage? I’ve used previous iterations of llama models and they all seemed to be lacking in that regard.
- dharma1 2y agoare these better than qwen at codegen?
- oulipo 2y agoCan the 3B run on a M1 macbook? It seems that it hogs all the memory. The 1B runs fine
- arnaudsm 2y agoIs there an up-to-date leaderboard with multiple LLM benchmarks? Livebench and Lmsys are weeks behind and sometimes refuse to add some major models. And press releases like this cherry pick their benchmarks and ignore better models like qwen2.5. If it doesn't exist I'm willing to create it
- threatripper 2y agohttps://artificialanalysis.ai/leaderboards/models https://artificialanalysis.ai/leaderboards/models "LLM Leaderboard - Comparison of GPT-4o, Llama 3, Mistral, Gemini and over 30 models Comparison and ranking the performance of over 30 AI models (LLMs) across key metrics including quality, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others. For more details including relating to our methodology, see our FAQs."
- aussieguy1234 2y agoWhen using meta.ai, its able to generate images as well as understand them. Has this also been open sourced or just a GPT4o style ability to see images?
- 404mm 2y agoCan anyone recommend a webUI client for ollama?
- iKlsR 2y agoopenwebui
- papascrubs 2y agohttps://get.big-agi.com/ https://get.big-agi.com/
- fungi 2y agoive been using https://github.com/valiantlynx/ollama-docker https://github.com/valiantlynx/ollama-docker which comes with https://github.com/open-webui/open-webui https://github.com/open-webui/open-webui
- Ey7NFZ3P0nzAe 2y agoOpen webui has promising aspects, the same authors are pushing for "pipelines" which are a standard for how inputs and outputs are modified on the fly for different purposes.
- 404mm 2y agoNewbie question, what size model would be needed to have a 10x software engineer skills and no knowledge of the human kind (ie, no need to know how to make a pizza or sequence your DNA). Is there such a model?
- keyle 2y agoNo, not yet. And such LLM wouldn't speak back in English or French without some "knowledge of the human kind" as you put it.
- _lvbh 2y ago10x relative to what? I’ve seen bad developers use AI to 10x their productivity but they still couldn’t come anywhere close to a good developer without AI (granted, this was at a hackathon on pretty advanced optimization research. Maybe there’s more impact on lower skilled tasks)
- pants2 2y ago
- freedomben 2y agoIf anyone else is looking for the bigger models on ollama and wondering where they are, the Ollama blog post answered that for me. The are "coming soon" so they just aren't ready quite yet[1]. I was a little worried when I couldn't find them but sounds like we just need to be patient. [1]: https://ollama.com/blog/llama3.2 https://ollama.com/blog/llama3.2
- xena 2y agoAs a rule of thumb with AI stuff: it either works instantly, or wait a day or two.
- refulgentis 2y agoollama is "just" llama.cpp underneath, I recommend switching to LM Studio or Jan, they don't have this issue of proprietary wrapper that obfuscates, you can just use any ol GGUF
- lolinder 2y agoWhat proprietary wrapper? Isn't Ollama entirely open source?
- calgoo 2y agoI use gguf in ollama on a daily basis, so not sure what the issue is? Just wrap it in a modelfile and done!
- vorticalbox 2y agoI think because the larger models support images.
- Patrick_Devine 2y agoWe're working on it. There are already draft PRs up in the GH repo. We're still working out some kinks though.
- notpublic 2y agoLlama-3.2-11B-Vision-Instruct does an excellent job extracting/answering questions from screenshots. It is even able to answer questions based on information buried inside a flowchart. How is this even possible??
- bboygravity 2y agomagic
- vintermann 2y agoOh, this is promising. It's not surprising to me: image models have been very oriented towards photography and scene understanding rather than understanding symbolic information in images (like text or diagrams), but I always thought that it should be possible to make the model better at the latter, for instance by training it more on historical handwritten documents.
- Ey7NFZ3P0nzAe 2y agoBecause they trained the text model. Then froze the weights. Then trained a vision model on text image pairs of progressively higher quality. Then trained an adapter to align their latent spaces. So it became smart on text then gain a new input sense magically without changing its weights
- ComputerGuru 2y agoIs this - at a reasonable guess - what most believe OpenAI did with 4o?
- faangguyindia 2y agoHow good it is at comic reading?
- xrd 2y agoI'm currently fighting with a fastapi python app deployed to render. It's interesting because I'm struggling to see how I encode the image and send it using curl. Their example sends directly from the browser and uses a data uri. But, this is relevant because I'm curious how this new model allows image inputs. Do you paste a base64 image into the prompt? It feels like these models can start not only providing the text generation backend, but start to replace the infrastructure for the API as well. Can you input images without something in front of it like openwebui?
- bombi 2y agoIs Termux enough to run the 1B model on Android?
- brrrrrm 2y agodepends on your phone, but try a couple of these variants with ollama https://ollama.com/library/llama3.2/tags https://ollama.com/library/llama3.2/tags e.g. `ollama run llama3.2:1b-instruct-q4_0`
- alexcpn 2y agoIn KungfuPanda there is this line that the Panda says "I love KungFuuuuuuuu", well I normally don't tell like this, but when I saw this and (starting to use this), I feel like yelling"I like Metaaaaa or is it LLAMMMAAA or is it Open source.. or is it this cool ecosystem which gives such value for free...
- alanzhuly 2y agoLlama3.2 3B feels a lot better than other models with same size (e.g. Gemma2, Phi3.5-mini models). For anyone looking for a simple way to test Llama3.2 3B locally with UI, Install nexa-sdk(https://github.com/NexaAI/nexa-sdk https://github.com/NexaAI/nexa-sdk) and type in terminal: nexa run llama3.2 --streamlit Disclaimer: I am from Nexa AI and nexa-sdk is an open-sourced. We'd love your feedback.
- alfredgg 2y agoIt's a great tool. Thanks! I had to test it with Llama3.1 and was really easy. At a first glance Llama3.2 didn't seem available. The command you provided did not work, raising "An error occurred while pulling the model: not enough values to unpack (expected 2, got 1)".
- alanzhuly 2y agoThanks for reporting. We are investigating this issue. Could you help submit an issue to our GitHub and provide a screenshot of the terminal (with pip show nexaai)? This could help us reproduce this issue faster. Much appreciated!
- grahamj 2y agoor grab lmstudio
- Zuiii 2y agoFor people who really care about open source, this is not.
- mikestaub 2y agoor https://chat.webllm.ai/ https://chat.webllm.ai/
- desireco42 2y agoI have to say that running this model locally I was pleasantly suprised how well it ran, it doesn't use as much resources and produce decent output, comparable to ChatGPT, it is not quite as OpenAI but for a lot of tasks, since it doesn't burden the computer, it can be used with local model. Next I want to try to use Aider with it and see how this would work.
- Ey7NFZ3P0nzAe 2y agoInteresting that its scores are somewhat helow Pixtral 12B https://mistral.ai/news/pixtral-12b/ https://mistral.ai/news/pixtral-12b/
- deleted 2y ago[deleted]
- stogot 2y agoSurprised no mention of audio?
- edude03 2y agowas surprised by this as well
- josephernest 2y agoCan it run with llama-cpp-python? If so, where can we find and download the gguf files? Are they distributed directly by meta, or are they converted to gguf format by third parties?
- kgeist 2y agoTried the 1B model with the "think step by step" prompt. It gets "which is larger: 9.11 or 9.9?" right if it manages to mention that decimals need to be compared first in its step-by-step thinking. If it skips mentioning decimals, then it says 9.11 is larger. It gets the strawberry question wrong even after enumerating all the letters correctly, probably because it can't properly count.
- khafra 2y agoOf course, in many contexts, it is correct to put 9.11 after 9.9--software versioning does it that way, for example.
- KeplerBoy 2y agoThat's why it's an interesting question and why it struggles so hard. A good answer would explain that and state both results if the context is not hundred percent clear.
- khafra 2y agoI'm not sure how useful that question is in exploring capabilities--"ask the user clarifying questions if the answer is ambiguous" is more of a rlhf or fine tune thing than a base model thing.
- vergessenmir 2y agoWhat is the "think step by step" prompt? An example would be great, Is this part of the system prompt?
- potatoman22 2y agoIt's appending "think step-by-step" to the end of the prompt to elicit a chain-of-thought response. See: https://arxiv.org/abs/2205.11916 https://arxiv.org/abs/2205.11916
- bick_nyers 2y ago
- sgt 2y agoAnyone on HN running models on their own local machines, like smaller Llama models or such? Or something else?
- karpatic 2y agoFor sure dude! Top comment thread is all about using ollama and other ways to get that done.
- grahamj 2y agoDoesn’t everyone? X) it’s super easy now with ollama + openwebui or an all in 1 like mlstudio
- sgt 2y agoWas just concerned I don't have enough RAM. I have 16GB (M2 Pro). Got amazing mem bandwidth though (800GB/s)
- kristianp 2y agoM2 Pro has 200GB/s
- sgt 2y agoYou're right. M1 and M2 Pro: 200GB/s M3 Max: 300GB/s M1/M2 Max: 400GB/s M1/M2 Ultra: 800GB/s Seems to be the case. An Ultra .. wow. But 200GB is also still good, so not complaining.
- grahamj 2y agoIt's on the low side but plenty for something like this new 3b model. Anything up to 8GB and you've still got as much as a base model Air left over.
- troupo 2y agoIt's hallucinating so badly, it's kinda hilarious Literally everything about the quote below is wrong. The problem with these "amazing ground-breaking models" is that you never know which knowledge they actually possess. --- start quote --- >>> what is the last version of Ash Framework for Elixir you know My training data goes up to December 2023, and according to the Ash Framework releases, the last stable version that I'm familiar with is: * *Ash 2.14*: Released on April 8, 2021. However, it's worth noting that Ash has been merged into Elixir itself, and most of its features have been integrated into the Elixir language and standard libraries. In fact, Ash was one of the first frameworks to be included in Elixir, and it played a significant role in shaping the language. Many of Ash's core concepts, such as `defstruct/1` and `defmodule/1`, were later adopted by Elixir itself. ... Keep in mind that Ash is no longer actively maintained as a separate framework, but its legacy continues to influence Elixir and its community. --- end quote ---
- borromakot 2y agoWildly incorrect
- l5870uoo9y 2y ago> These models are enabled on day one for Qualcomm and MediaTek hardware and optimized for Arm processors. Do they require GPU or can they be deployed on VPS with dedicated CPU?
- KeplerBoy 2y agoDoesn't require a GPU, it will just be faster with a GPU.
- taytus 2y agometa.ai still running on 3.1
- chriskanan 2y agoThe assessments of visual capability really need to be more robust. They are still using datasets like VQAv2, which while providing some insight, have many issues. There are many newer datasets that serve as much more robust tests and that are less prone to being affected by linguistic bias. I'd like to see more head-to-head comparisons with community created multi-modal LLMs as done in these papers: https://arxiv.org/abs/2408.05334 https://arxiv.org/abs/2408.05334 https://arxiv.org/abs/2408.03326 https://arxiv.org/abs/2408.03326 I look forward to reading the technical report, once its available. I couldn't find a link to one, yet.
- Jackson__ 2y agoLooking at their benchmark results and my own experience with their 11B vision model, I think while not perfect they represent the model well. Meaning it's doing impressively bad compared to other models I've tried in similar sizes(for vision).
- monkfish328 2y agoZuckerberg has never liked having Android/iOs as gatekeepers i.e. "platforms" for his apps. He's hoping to control AI as the next platform through which users interact with apps. Free AI is then fine if the surplus value created by not having a gatekeeper to his apps exceeds the cost of the free AI. That's the strategy. No values here - just strategy folks.
- jsemrau 2y agoAgents are the new Apps
- acedTrex 2y agoI mean, just because he is not doing this as a perfectly altruistic gesture does not mean the broader ecosystem does not benefit from him doing it
- monkfish328 2y agoFor sure
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- 84adam 2y agoexcited for this
- ofermend 2y agoGreat release. Models just added to Hallucination Leaderboard: https://github.com/vectara/hallucination-leaderboard https://github.com/vectara/hallucination-leaderboard. TL;DR: * 90B-Vision: 4.3% hallucination rate * 11B-Vision: 5.5% hallucination rate