Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
espadrine
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
31.
▲
by
espadrine
1y ago
At least it is not unprecedented. Palantir raised a series I in 2020 after 17 years of operation.
32.
▲
by
espadrine
1y ago
It would be interesting to have two generations per model without cherry picking, so that the Elo estimation can include an easy-to-compute standard deviation estimation.
33.
▲
by
espadrine
1y ago
The best model there is 2.5B parameters. I can believe that a model 10x bigger is somewhat better. One element of comparison is OpenAI Whisper v3, which achieves 7.44 WER on the ASR leaderboard, and shows up as ~8.3 WER on FLEURS in the Vox
34.
▲
by
espadrine
1y ago
I agree that there are some robotic designs that unnecessarily mimic human limbs. I have in mind heads, and feet (instead of wheels). A hand however, is useful because so many manufactured objects have been constructed for their purpose.
35.
▲
by
espadrine
1y ago
Here is someone that had significant corruption until they stopped: https://www.xda-developers.com/why-not-to-spin-down-nas-hard... There are many similar articles.
36.
▲
by
espadrine
1y ago
Easy yes. Even VPS providers need to maintain the IP, since your DNS typically points to that IP. You can also typically move the IP to another machine from the same provider. But as a resut, VPS often have a different price for public IPs
37.
▲
by
espadrine
2y ago
> I rather define freedom by the government not deciding what's good for me Does that mean you are against this bill? Before the bill, your community could either use your own water system, without fluoride, or use the wider syste
38.
▲
by
espadrine
2y ago
> I still haven't seen any good reasoning for why NASA would delay the return flight So many comments spread unsourced assertions on the topic on both sides. Let me change that. On 24 Aug 2024, NASA stated in a conference publishe
39.
▲
by
espadrine
2y ago
In RAM, yes. But if you compute an activation, you need to load the weights from RAM to the GPU core.
40.
▲
by
espadrine
2y ago
There is no question that quantization degrades quality. The GGUF R1 uses Q4_K_M, which, on Llama-3-8B, increases the perplexity by 0.18[0]. Many plots show increasing degradation as you quantize more[1]. That said, it is possible to train
41.
▲
by
espadrine
2y ago
> if a llm will run with usable performance at that scale? Yes. The reason: MoE. They are able to run at a good speed because they don't load all of the weights into the GPU cores. For instance, DeepSeek R1 uses 404 GB in Q4 quant
42.
▲
by
espadrine
2y ago
The blind test at lmarena.ai does give it a higher Elo than GPT-4o (API), Claude, and Gemini 1.5 Pro. It seems that people do enter real-life scenarios in the arena.
43.
▲
by
espadrine
2y ago
Hernan Moraldo is from Argentina. That may be all there is to it.
44.
▲
by
espadrine
2y ago
> I'm quite nervous for the future. Videos like these were already achievable through VFX. The only difference here is a reduction in costs. That does mean that more people will produce misinformation, but the problem is one that
45.
▲
by
espadrine
2y ago
Philosophically, it always bugged me that distros were so centrally in control of packaging: app developers are not allowed to package them in the way they see fit, which has caused some friction in the past[0]. It seems healthier and more
46.
▲
by
espadrine
2y ago
When Meta prevented the EU from using meta.ai or even downloading its vision models, I sunk my head in the AI legistation. Here, I am honestly not sure which part they rely on, to say that what they made might be unlawful. The closest thing
47.
▲
by
espadrine
2y ago
> Furthermore, by rotating the vector, we have absolutely zero impact on the norm of the vector, which encodes the semantic information of our token. Doesn’t the angle encode semantic information? Cosine similarity works for embeddings
48.
▲
by
espadrine
2y ago
You’re right. I understood it to require taking the top 2^30 tokens, but instead they sample 2^30 times with replacement. Too bad they only formulate the detection positive rate empirically. I am curious what the exact probability would be
49.
▲
by
espadrine
2y ago
The academic paper: https://www.nature.com/articles/s41586-024-08025-4 They use the last N prefix tokens, hash them (with a keyed hash), and use the random value to sample the next token by doing an 8-wise tournament,
50.
▲
by
espadrine
2y ago
That makes me wonder though what the best loss function was. I assume you used MSE on the logscore. I wonder if a sigmoid on which of two articles has the higher score would yield better results for the downstream RLHF task.
51.
▲
by
espadrine
2y ago
It takes no time at all to find other major mistakes. For instance, the Mixtral diagram § 6.6.1 shows a single router that selects separate 32-layer transformers. Instead, Mixtral has one router per layer (inside of each block), and it does
52.
▲
by
espadrine
2y ago
Terminology. Since they targeted very low risk, they did a geographically-segmented rollout, starting with Phoenix, which is one of the easiest places to drive: a lot of photons for visibility, very little rain, wide roads.
53.
▲
by
espadrine
2y ago
At this week’s dotAI conference, Ines Montani (who works on the SpaCy project) highlighted this ex-job as a warning to AI builders, so that they do not work on systems that have no future, because better and cheaper alarm clocks (for knocke
54.
▲
by
espadrine
2y ago
KataGo has a special model weights release with human-like play at various Elo: https://github.com/lightvector/KataGo/blob/master/docs/Analy... You can see in the release notes a few screenshot exa
55.
▲
by
espadrine
2y ago
> Hard to see how can Mistral compete with Meta One significant edge: Meta does not dare even distribute their latest models (the 3.2 series) to EU citizens. Mistral does.
56.
▲
by
espadrine
2y ago
Five years to build each SMR is faster than what it takes to build a typical EPR (from 10 to 20 years), but it is still longer than what I expected. Besides, the 5-year figure might have overruns as they work through regulatory hoops toward
57.
▲
Google purchases nuclear energy from Kairos Power SMRs
(blog.google)
1 points
by
espadrine
2y ago
|
2 comments
58.
▲
by
espadrine
2y ago
This is a topic I love to study. The mathematical analysis is reasonable; the policy gradient is a classic approach; I love Sutton’s RL book on it: http://incompleteideas.net/book/RLbook2020.pdf Even though nowadays ma
59.
▲
by
espadrine
2y ago
The Pixtral report[0] compares positively to Molmo. (Also, beware, molmo.org is an AI-generated website to absorb through SEO Allen AI’s efforts; the real website is molmo.allenai.org. Note for instance that all tweets listed here are from
60.
▲
by
espadrine
2y ago
That is a concern that is shared with ReLU. But since the weights are shared across the context/minibatch, perhaps that would not be an issue, similar to ReLU.
More ›