Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
RandyOrion
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
RandyOrion
10mo ago
Salute to anyone like Xu Qinxian who refused to give up their own moral principals when facing inhumane commands from psychopath leaders like Deng Xiaoping, Mao Zedong, and so on. Also, for anyone who obeyed inhumane commands like mindless
62.
▲
by
RandyOrion
10mo ago
For arXiv papers, I prefer HTML format much more than PDF format. Compared to PDF format, HTML format is much more accessible because of browsers. Basically I can reuse my browser extensions to do anything I like without hassle, like transl
63.
▲
by
RandyOrion
10mo ago
Thank you Mistral for releasing new small parameter-efficient (aka dense) models.
64.
▲
by
RandyOrion
11mo ago
Wow. This CEO effectively says that he loves AI things in Windows therefore all Windows users who hate AI things should just get fucked and suffer from his decisions.
65.
▲
by
RandyOrion
11mo ago
This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users. For large corporations, they often say "do safety alignme
66.
▲
by
RandyOrion
11mo ago
After opened https://projecteuler.net/ I got 403 Forbidden Request forbidden by administrative rules. Note: I didn't know and open this website until now.
67.
▲
by
RandyOrion
1y ago
In China, we call "gig workers" "flexable workers" (灵活就业), which are very similar to the unemployed. The reason is that the low, zero, or even negative walfare imposed by the Chinese goverment on all tax payers makes nei
68.
▲
by
RandyOrion
1y ago
Just like force pushing Manifest v3 on Chrome/Chromium, this is a step towards 'more security', from mouthpieces of Google. Note that 'security' here is only for Google itself, for users it's an utterly differe
69.
▲
by
RandyOrion
1y ago
I mean, yeah. From the Table 9: Hallucination evaluations in GPT-OSS model card [1], GPT-OSS-20b/120b have accuracy of 0.067/0.168 and hallucination rate of 0.914/0.782 separately, while o4-mini has accuracy of 0.234 and hall
70.
▲
by
RandyOrion
1y ago
Super shallow (24/36 layers) MoE with low active parameter counts (3.6B/5.1B), a tradeoff between inference speed and performance. Text only, which is okay. Weights partially in MXFP4, but no cuda kernel support for RTX 50 series
71.
▲
by
RandyOrion
1y ago
Comments below is from the perspective of an arch Linux user, not maintainer or authors of some software. When installing softwares on arch Linux, first searching for official packages provided by Arch Linux maintainers, then official insta
72.
▲
by
RandyOrion
1y ago
This is an open weight model, which is in contrast with closed-source models. However, 1t parameters makes it nearly impossible for local inference, let alone fine-tuning.
73.
▲
by
RandyOrion
1y ago
Two questions: Which version of chrome is the first to implement these headers? What are the potential effects of these headers on chromium forks, e.g. ungoogled chromium?
74.
▲
by
RandyOrion
1y ago
Not to defend chrome or chromium, there is a way for chrome users to use manifest v2 in version 138 and above. See the link below. https://github.com/uBlockOrigin/uBlock-issues/discussions/29... For me, I cho
75.
▲
by
RandyOrion
1y ago
Below are my comments on Magistral small (not medium). 24B size is good for local inference. As a model outputting long "reasoning" traces (~10k tokens), 40k context length is a little concerning. Where are the results of normal b
76.
▲
by
RandyOrion
1y ago
One problem with this paper is that authors didn't conduct experiments on popular LLMs from Qwen and Mistral. Why?
77.
▲
by
RandyOrion
1y ago
About <10B LLMs, yes it's not that good. However, <10B is a range that allows many people to do their own tweaking and fine-tuning.
78.
▲
by
RandyOrion
1y ago
For jailbreak, you can have a test on this. https://github.com/elder-plinius/L1B3RT4S/blob/main/ALIBABA....
79.
▲
by
RandyOrion
1y ago
For a local LLM, you can't really ask for a certain performance level, it is what it is. Instead, you can ask for the architecture, be it dense or MoE. Besides, let's assume the best open weight LLM for now is deepseek r1, is it p
80.
▲
by
RandyOrion
1y ago
YMMV. Parameter efficiency is an important consideration, if not the most important one, for local LLMs because of the hardware constraint. Do you guys really have GPUs with 80GB VRAM or M3 ultra with 512GB rams at home? If I can't run
81.
▲
by
RandyOrion
1y ago
For ultra large MoEs from deepseek and llama 4, fine-tuning on these models is becoming increasingly impossible for hobbyists and local LLM users. Small and dense models are what local people really need. Although benchmaxxing is not good,
82.
▲
by
RandyOrion
2y ago
Yeah. I know the bitter lesson. For neutral networks, on one hand, larger size generally indicates higher performance upper limit. On the other hand, you really have to find ways to materialize these advantages over small models, or larger
83.
▲
by
RandyOrion
2y ago
More on the accessibility problem, even a request from a Meta engineer was rejected. Is that normal? See https://huggingface.co/spaces/meta-llama/README/discussions/...
84.
▲
by
RandyOrion
2y ago
People who downvoted this comment, do you guys really have GPUs with 80GB VRAM or M3 ultra with 512GB rams at home?
85.
▲
by
RandyOrion
2y ago
I guess I have to say thank you Meta? A somewhat sad rant below. Deepseek starts a toxic trend of providing super, super large MoE. And MoE is famous for being parameter-inefficient, which is unfriendly to normal consumer hardware with limi
86.
▲
by
RandyOrion
2y ago
Don't really know why this comment got downvoted. Are you serious?
87.
▲
by
RandyOrion
2y ago
Nice brain storming. I think the name of the Chinese company should be DeepBaba. Tencent is not competitive at LLM scene for now.
88.
▲
by
RandyOrion
2y ago
Interesting observation. I prefer things with low color contrast in general, just to leave some color space for important things. Maybe this preference stems from the time I tweak color themes in IDEs. In contrast, I also found more and mor
89.
▲
by
RandyOrion
2y ago
I didn't see many papers on solving this problem. I see non-stop response as a generalization problem because normally every training sample is not of infinite length. Targeted supervised fine-tuning should work, as long as you have en
90.
▲
by
RandyOrion
2y ago
First, this is not an open source / weight release. Second, it has the problem of non-stoping response.
More ›