Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
omneity
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
omneity
9mo ago
Unless using docker, if vllm is not provided and built against ROCm dependencies it’s going to be time consuming. It took me quite some time to figure the magic combination of versions and commits, and to build each dependency successfully
32.
▲
by
omneity
9mo ago
https://omarkamali.com - I write here sometimes https://wikilangs.org - A bunch of cool language playgrounds I'm working on
33.
▲
by
omneity
9mo ago
That's just an implementation artifact and not a fundamental fact of life. https://docs.vllm.ai/en/latest/features/batch_invariance/
34.
▲
by
omneity
9mo ago
Do you have any evals on how good LLMs are at generating Glyphlang? I’m curious if you optimized for the ability to generate functioning code or just tokenization compression rate, which LLMs you tokenized for, and what was your optimizatio
35.
▲
by
omneity
9mo ago
Sounds like most of this is simply taking shortcuts instead of properly parsing[0]. 0: https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
36.
▲
by
omneity
9mo ago
Thinking from first principles, a large part of the content on stack overflow comes from the practical experience and battle scars worn by developers sharing them with others and cross-curating approaches. Privacy concerns notwithstanding,
37.
▲
by
omneity
9mo ago
It's likely to either be an approach like this [0] or something even less involved. 0: https://github.com/apple/ml-clara
38.
▲
by
omneity
9mo ago
Related, I heard about curriculum learning for LLMs quite often but I couldn’t find a library to order training data by an arbitrary measure like difficulty, so I made one[0]. What you get is an iterator over the dataset that samples based
39.
▲
by
omneity
9mo ago
I tried to implement something similar to optimize sampling semi-random documents from (very) large datasets on Huggingface, unfortunately their API doesn't support range requests well.
40.
▲
by
omneity
10mo ago
Given its age, I thought this language had a chance to predate the eponymous computer mouse, but nope. It was created a good decade after the first documented reference to computer mice [0]. 0: https://en.wikipedia.org/wiki&
41.
▲
by
omneity
10mo ago
> With a typical LLM SDK What is a typical LLM SDK?
42.
▲
Picomon 0.2.0: From AMD Crash Fix to GPU Monitoring That Doesn't Suck
(omarkama.li)
2 points
by
omneity
10mo ago
|
0 comments
43.
▲
by
omneity
10mo ago
Thanks, this was helpful! Reading the seminal paper[0] on Universal Transformers also gave some insights: > UTs combine the parallelizability and global receptive field of feed-forward sequence models like the Transformer with the recurr
44.
▲
by
omneity
10mo ago
It tells you about their ambitions..
45.
▲
by
omneity
10mo ago
Isn’t this in a sense an RNN built out of a slice of an LLM? Which if true means it might have the same drawbacks, namely slowness to train but also benefits such as an endless context window (in theory)
46.
▲
by
omneity
10mo ago
If performant FPGAs were more accessible we’d be able to download models directly into custom silicon, locally, and unlock innovation in inference hardware optimizations. The highest grade FPGAs also have HBM memory and are competitive (on
47.
▲
by
omneity
10mo ago
Not until it gets tensor parallelism.
48.
▲
by
omneity
10mo ago
You're right. Let me correct myself: a hobbyist-friendly hardware solution. Dolphin's PCIe switches cost more than 8 RTX 3090 on a Threadripper machine.
49.
▲
by
omneity
10mo ago
I wish for a hardware + software solution to enable direct PCIe interconnect using lanes independent from the chipset/CPU. A PCIe mesh of sorts. With the right software support from say pytorch this could suddenly make old GPUs and und
50.
▲
by
omneity
10mo ago
To my knowledge it's not as much hurting the parent domain as having two separate "worlds". Your docs which are likely to receive higher traffic will stop contributing any SEO juice to your main website.
51.
▲
by
omneity
10mo ago
The way I understood OP’s point is that because LLMs have been trained on the entirety of humanity’s knowledge (exemplified by the internet), then surely they know as much as the entirety of humanity. A cursory use of an LLM shows this is o
52.
▲
by
omneity
10mo ago
While an LLM is trained on trillions of tokens to acquire its capabilities, it does not actively retain or recall the vast majority of it, and often enough is not able to make deductive reasoning either (e.g. X owns Y does not necessarily t
53.
▲
by
omneity
10mo ago
Nemotron now works on LM Studio if you update the runtime (from the settings > Runtime screen). The default chat template is incorrect though and will fail but I published a corrected one you can replace it with: https://gist.
54.
▲
by
omneity
10mo ago
https://www.reddit.com/r/LocalLLaMA/comments/1plpc6h/mistral...
55.
▲
by
omneity
10mo ago
What a strange naming choice, mixing two things (vLLM and LoRA) while being related to neither..
56.
▲
by
omneity
10mo ago
Mistral Large 3 is reportedly using Deepseek V3.2 architecture with larger experts and fewer of them, and a 2B params vision module.
57.
▲
by
omneity
10mo ago
You might actually get that desired behavior through reasoning, or if the model was reinforced for coding workflows involving COM, or at least enough stack diversity for the model to encounter the need to develop this capability. In the cas
58.
▲
by
omneity
10mo ago
While not solving everything (such as a single bill), I implemented borgllm[0,1] as a way to get some of that convenience and development speed with no lock-in or even infra. Perhaps you’ll find it interesting. 0: https://github.
59.
▲
by
omneity
10mo ago
Slightly related, I made a monitor[0] for AMD gpus with a nifty chart. I had many issues with nvtop, it is a bit too strict for some situations and ends up crashing too often. 0: https://github.com/omarkamali/picomon
60.
▲
by
omneity
10mo ago
Hello HN! I've been trying out some AMD gpus for Gen AI workloads lately and found the state of monitoring quite unsatisfying. The best option out there, `nvtop`, has some very strict correctness asserts crashing it too often. With `pi
More ›