8 ms·
Tencent Hunyuan-Large
- helloericsf 2y ago- 389 billion parameters and 52 billion activation parameters, capable of handling up to 256K tokens. - outperforms LLama3.1-70B and exhibits comparable performance when compared to the significantly larger LLama3.1-405B model.
- Etheryte 2y agoIt's a bit funny to call the 405B reference "significantly larger" than their 389B, while highlighting the fact that their 389B outperforms the 70B.
- klipt 2y agoIt's a whole 4% smaller!
- rose_ann_ 2y agoMoE model with 52 billion activated parameters means its more comparable to a (dense) 70b model and not a dense 405b model
- HPsquared 2y agoDoes this mean it runs faster or better on multiple GPUs?
- chessgecko 2y agoFor decode steps it depends on the number of inputs you run at a time. If your batch size is 1 then it runs in line with active params, then as you get to like batch size 8 it runs in line with all params, then as you increase to 128ish it runs like the active params again. For the context encode it’s always close to as fast as a model with a similar number of active params. For running on your own the issue is going to be fitting all the params on your gpu. If you’re loading off disk anyways this will be faster but if this forces you to put stuff on disk it will be much slower.
- phkahler 2y ago>> MoE model with 52 billion activated parameters means its more comparable to a (dense) 70b model and not a dense 405b model Only when talking about how fast it can produce output. From a capability point of view it makes sense to compare the larger number of parameters. I suppose there's also a "total storage" comparison too, since didn't they say this is 8bit model weights, where llama is 16bit?
- trump2026 2y ago[flagged]
- helloericsf 2y agoCall Jensen and Lisa!lol
- trump2026 2y ago[dead]
- deleted 2y ago[deleted]
- eptcyka 2y agoDefinitely not trained on Nvidia or AMD GPUs.
- rb2k_ 2y agoThe readme mentioned H20 GPUs. Nvidia's "China compatible" card (41% Fewer Cores & 28% Lower Performance Versus Top Hopper H100 Config)
- 1R053 2y agoyou can get a long way on something with 41% less performance than your favorite supercar...
- acchow 2y agoHow do you know this? Apparently 20% of Nvidia's quarterly revenue is booked in Singapore where shell companies divert product to China: https://news.ycombinator.com/item?id=42048065 https://news.ycombinator.com/item?id=42048065
- mrob 2y agoNot open source. Even if we accept model weights as source code, which is highly dubious, this clearly violates clauses 5 and 6 of the Open Source Definition. It discriminates between users (clause 5) by refusing to grant any rights to users in the European Union, and it discriminates between uses (clause 6) by requiring agreement to an Acceptable Use Policy. EDIT: The HN title was changed, which previously made the claim. But as HN user swyx pointed out, Tencent is also claiming this is open source, e.g.: "The currently unveiled Hunyuan-Large (Hunyuan-MoE-A52B) model is the largest open-source Transformer-based MoE model in the industry".
- vanguardanon 2y agoWhat is the reason for restrictions in the EU? Is it due to some EU regulations?
- ronsor 2y agoMost likely yes. I don't think companies can be blamed for not wanting to subject themselves to EU regulations or uncertainty. Edit: Also, if you don't want to follow or deal with EU law, you don't do business in the EU. People here regularly say if you do business in a country, you have to follow its laws. The opposite also applies.
- troupo 2y ago[flagged]
- ronsor 2y agoI will address both points: 1. No one is training on users' bank details, but if you're training on the whole Internet, it's hard to be sure if you've filtered out all PII, or even who is in there. 2. This isn't happening because no one has time for more time-wasting lawsuits.
- troupo 2y ago
- 1R053 2y agothe paper with details: https://arxiv.org/pdf/2411.02265 https://arxiv.org/pdf/2411.02265 They use - 16 experts, of which one is activated per token - 1 shared expert that is always active in summary that makes around 52B active parameters per token instead of the 405B of LLama3.1.
- the_duke 2y ago> Territory” shall mean the worldwide territory, excluding the territory of the European Union. Anyone have some background on this?
- jmole 2y agoI believe the EU has (or is drafting) laws about LLMs of a certain size which this release would not comply with.
- troupo 2y agoAlso existing privacy laws (GDPR) and AI Act (foundational models have to disclose and document their training data)
- mattlutze 2y agohttps://artificialintelligenceact.eu/high-level-summary/ https://artificialintelligenceact.eu/high-level-summary/ There's many places where the model might be used which could count as high-risk scenarios and require lots of controls. Also, we have: GPAI models present systemic risks when the cumulative amount of compute used for its training is greater than 10^25 floating point operations (FLOPs). Providers must notify the Commission if their model meets this criterion within 2 weeks. The provider may present arguments that, despite meeting the criteria, their model does not present systemic risks. The Commission may decide on its own, or via a qualified alert from the scientific panel of independent experts, that a model has high impact capabilities, rendering it systemic. In addition to the four obligations above, providers of GPAI models with systemic risk must also: - Perform model evaluations, including conducting and documenting adversarial testing to identify and mitigate systemic risk. - Assess and mitigate possible systemic risks, including their sources. - Track, document and report serious incidents and possible corrective measures to the AI Office and relevant national competent authorities without undue delay. - Ensure an adequate level of cybersecurity protection." They may not want to meet these requirements.
- lcnPylGDnU4H9OF 2y ago
- 2OEH8eoCRo0 2y ago[flagged]
- azinman2 2y agoI just did, and it tells me it has no information on that issue. It also responded back in Chinese to that English query, which either suggests to me that the censorship instruction tuning is heavily weighted towards Chinese, or the model has a hard time staying in English (which I believe has been the case for other Chinese LLMs in the past)
- the5avage 2y agoI once triggered the ChatGPT censorship (by trying to manipulate an image of my face) and it also responded in english to a german query.
- dyauspitr 2y agoTry testing it on some of the US’ taboo topics like LGBT, feminism, racism etc.
- azinman2 2y agohttps://huggingface.co/spaces/tencent/Hunyuan-Large https://huggingface.co/spaces/tencent/Hunyuan-Large
- a_wild_dandan 2y agoThe model meets/beats Llama despite having an order-of-magnitude fewer active parameters (52B vs 405B). Absolutely bonkers. AI is moving so fast with these breakthroughs -- synthetic data, distillation, alt. architectures (e.g. MoE/SSM), LoRA, RAG, curriculum learning, etc. We've come so astonishingly far in like two years. I have no idea what AI will do in another year, and it's thrilling.
- csomar 2y agoIt is insane because 52B can run on my current 3 years old laptop. 3B LLMA 3.2 from Facebook can already autocomplete for me. I didn't try this model but if the scores are to be believed, this can give useful and actionable insights into a project source code. Probably not as good as Claude 3.5 but I can run it locally. This is a game changer.
- z3ncyberpunk 2y agoMoving fast or just completely inefficient
- geenkeuse 2y ago[flagged]
- adt 2y agohttps://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- Tepix 2y agoI'm no expert on these MoE models with "a total of 389 billion parameters and 52 billion active parameters". Do hobbyists stand a chance of running this model (quantized) at home? For example on something like a PC with 128GB (or 512GB) RAM and one or two RTX 3090 24GB VRAM GPUs?
- lanceflt 2y agoRAM for 4-bit is 1GB per 2 billion parameters. So you will want 256GB RAM and at least one GPU. If you only have one server and one user, it's the full parameter count. (If you have multiple GPUs/servers and many users in parallel, you can shard and route it so you only need the active parameter count per GPU/server. So it's cheaper at scale.)
- zamadatix 2y agoDo the inactive parameters need to be loaded into RAM to run an MoE model decently enough?
- bick_nyers 2y agoYou would need to fit the 389B parameters in VRAM to have a speed that is usable. Different experts are activated on a per token basis, so you would need to load/unload a large chunk of the 52B active parameters every token if you were trying to offload parameters to system RAM or SSD. PCIE 4.0 x16 speed is 64GB/s, so you can load those active parameters maybe 1 or 2 times per second, yielding an output speed of 1-2 tokens per second, which most would consider "unusable".
- o11c 2y agoDoes that have to be same-node VRAM? Or can you fit 52B each on several nodes, and only copy the transient state around?
- bick_nyers 2y agoGenerally speaking this works well, pending your definition of node and the interconnect between them. If by node you mean GPU, and you have multiple of them on the same system (interconnect is PCIE, doesn't need to be full speed however for inference), you're good. If you mean multiple computers connected by 1 Gigabit Ethernet? More challenging. When splitting models layer by layer, users in r/LocalLLaMA have reported good results with as low as PCIE 3.0 x4 as the interconnect (4GB/s). For tensor parallelism, the interconnect requirements are higher but the upside can be faster speeds in accordance to number of GPUs split across (whereas layer by layer operated like a pipeline, so isn't necessarily faster than what a single GPU can provide, even if splitting across 8 GPUs).
- iqandjoke 2y agoHow does it compare with LLama3.2?
- Tepix 2y agoLlama 3.2 has the same performance for text as Llama 3.1 and the largest model hasn't been released.