Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
om8
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
om8
1y ago
I'm currently in Russia trying to get a US visa for my CS PhD. Because I do CS, I got into a thing called administrative processing. For 95% of people, it takes days -- weeks, tops. Because of the colour of my passport, it has already
32.
▲
by
om8
1y ago
I believe he knows
33.
▲
by
om8
1y ago
What would you do if some of the deps started to have conflicts in them? Also, what are your plans for migration when you'll need to move from one os version to another? Implicit solutions like yours have lower cost of entrance, but la
34.
▲
by
om8
1y ago
I have a demo that runs llama3-{1,3,8}B in browser on cpu. It can be integrated with this thing in the future to be fully local https://galqiwi.github.io/aqlm-rs
35.
▲
by
om8
1y ago
> a local LLM is going to be the way to go Non-technicals don't know how LLMs work, and, more importantly, don't care about their privacy. For a technology to be widely used, by definition, you need to make it appealing to the
36.
▲
by
om8
1y ago
> 50 tokens is not really very much Yes! And also llama3.1’s tokens are different from Qwen and llama1 tokens. That’s the first model where meta started to use very large vocab_size.
37.
▲
by
om8
1y ago
I trust cloudflare more than my ISP, since I live in a place where internet is very state controlled. Some of the websites just don't open without DoH.
38.
▲
by
om8
1y ago
I thing you are just arguing about words, not about meanings. I’d call what you are referring to “secure llm infrastructure ”, not “secure llm”. But the thing is that we both agree about what’s going on, just with different words
39.
▲
by
om8
1y ago
Even human level intelligence (whatever that means) is not enough. Social engineering works fine on our meat brains, it will most probably work on llms for foreseeable non-weird non-2027-takeoff-timeline future. Based on “bug level of intel
40.
▲
by
om8
1y ago
Here is an example of Rust implementation being faster than C. https://trifectatech.org/blog/zlib-rs-is-faster-than-c/
41.
▲
by
om8
1y ago
Channels are useful when they are really (rarely) needed. IMO Channel API should've been as ugly as reflect API to be considered only in extra cases.
42.
▲
Show HN: HIGGS – new sota data-free LLM quantization
(huggingface.co)
3 points
by
om8
1y ago
|
0 comments
43.
▲
by
om8
2y ago
Burden of proof lies on you, since you mentioned corporations first
44.
▲
by
om8
2y ago
Not surprising, llama.cpp code is a mess. It's sad that hacked things that emerge first are way more popular than properly done projects that come later.
45.
▲
by
om8
2y ago
You can’t build fault tolerant system without network lag. Your hardware can fail you any second, and every backup will be inconsistent with your local db. To solve it, you need to have a system with proper control plane and consensus algor
46.
▲
by
om8
2y ago
Sure, what's the issue? Do you think that there is an algorithm that a) does not do this, b) can't be implemented in rust, but can be implemented in C or other language?
47.
▲
by
om8
2y ago
Bitwarden is the place where I store stuff safely ><. This update is just awful
48.
▲
by
om8
2y ago
Same here. I'm very sad about this 2FA thing. Bitwarden was so easy to use, I could always get an access to my accounts with just my secure master password. Does anybody know good alternative?
49.
▲
by
om8
2y ago
Not really a WebNN expert, but looks like it doesn't support CPU inference yet and only works in Chrome. It also lacks support for custom kernels, which we need for running AQLM-quantized models. When you ask about spec change, do you
50.
▲
by
om8
2y ago
This is a demo of what's possible to run on edge devices using SOTA quantization. Other similar projects that try to run 8B models in browser are either using webgpu or 2 bit quantization that breaks the model. I implemented inference
51.
▲
Show HN: Llama 3.1 8B CPU Inference in a Browser via WebAssembly
(galqiwi.github.io)
4 points
by
om8
2y ago
|
4 comments
52.
▲
by
om8
2y ago
The cool thing about rust is that it forces you to use thread safety primitives. It's called fearless concurrency
53.
▲
by
om8
2y ago
Looks like it is a drop in replacement for attention, but models will need to be retrained for this one, yes.
54.
▲
by
om8
2y ago
Younger me would've said that WYSIWYG editors were a mistake and that researchers should've used LaTeX. Now I think these errors are a small price to pay for convenience. One could waste a lifetime fighting small things like this
55.
▲
by
om8
2y ago
How to you define memorization and reasoning? There is a large grey area in between them. Some say that if you can memorize facts and algorithms and apply them to new data, it is a memorization. Some say that it is reasoning. More than that
56.
▲
by
om8
2y ago
TLDR: use shitty allocators, win shitty memory leaks
57.
▲
by
om8
2y ago
> to minimize the amount of computation IMO backprop is the most trivial implementation of differentiation in neural networks. Do you know an easier way to compute gradients with larger overhead? If so, please share it.
58.
▲
by
om8
2y ago
Single responsibility principle in its finest.
59.
▲
by
om8
2y ago
Not if you use a language that bans implicit type conversions (like go)
60.
▲
by
om8
2y ago
> What i say to people for at least 2 years now, is that "Remember when governments were not just some cryptographic algorithms?" Yeah, that's gonna change. Cryptography is here to stay, it is not as dead as people think a
More ›