Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Manabu-eo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
Manabu-eo
2y ago
That Dragon came with it's own crew, not empty.
32.
▲
by
Manabu-eo
2y ago
>> No way it can not be real. Any technology unleashed on the wider internet is going to be used for this, obligatory xkcd: https://xkcd.com/1289/
33.
▲
by
Manabu-eo
2y ago
> you sent some medical pics of your son's crotch and permanently lose access to your google acount, even after lots of stress and the police declaring you innocent[1]. Replace google acount by some other important thing. https:&#x
34.
▲
by
Manabu-eo
2y ago
Dark Silicon. Having proportionally less transistors firing at any given time.
35.
▲
by
Manabu-eo
2y ago
You are overlooking the easier explanation: they used "velocity" in the more colloquial sense meaning speed, not in the vector sense. Suddenly everything is logically consistent. And everything matches what they said, what they fi
36.
▲
by
Manabu-eo
2y ago
> SpaceX's claim was that IFT3 would enter an orbit You are claiming that, SpaceX never claimed that, only "orbital velocity" meaning "orbital speed" in this context. They wanted to test if the Starship is capabl
37.
▲
by
Manabu-eo
2y ago
No need for NVLink just for inference, not even with tensor parallelism. And you can get used 3090 much cheaper than that.
38.
▲
by
Manabu-eo
2y ago
He is a founder at Tesla, as when he entered the round A financing Tesla was just three guys, some networking, and nothing more. Even the "Tesla" trademark and logo registration was made by SpaceX people. He didn't found his
39.
▲
by
Manabu-eo
2y ago
Even if it is contained within the model _somewhere_, it might be encoded in such a way that it's impractical to extract. Might need an exponential time algorithm, for example. Or this proxy method of hundreds of deceiving attempts. An
40.
▲
by
Manabu-eo
2y ago
There are at least two reasons for transformers poor performance on that prompt: - Transformers view the word as tokens, not words or characters. - The positional encoding might be holding them back. See this recent paper discussed here: Tr
41.
▲
by
Manabu-eo
2y ago
>but it sounds like their current design cannot tolerate the loss of more than a few tiles. What he said is more concerning than that: > Right now, we are not resilient to loss of a single tile in most places, as the secondary contain
42.
▲
by
Manabu-eo
2y ago
How much % of the theoretical FLOPs are you getting with those 7900 XTX on training?
43.
▲
by
Manabu-eo
2y ago
Just put it inside a zip and then you can send it. Maybe some plugin will be made to automatically unzip files and show contents.
44.
▲
by
Manabu-eo
2y ago
Only since crypto boom I see people (crypto aficionados) thinking of money as an investment. And that makes no sense, as you explain yourself. Dollars are much less volatile and thus less risky than any crypto currency I know. A perfect int
45.
▲
by
Manabu-eo
2y ago
Just anecdotally, I personally didn't know about the existence of that movie before this whole drama began. Sam tweeting that probably knew however.
46.
▲
by
Manabu-eo
2y ago
tl;dr Hubble is not dead, but solar storms make its orbit decay faster.
47.
▲
by
Manabu-eo
2y ago
Submarines are encircled by water, that is very good at absorbing heat. Nuclear power stations are also invariably close to a large source of water. There is no such thing on Mars, and that significantly increases the mass and size of your
48.
▲
by
Manabu-eo
2y ago
MXM was problematic because the inflexibility of the form factor to upgrade a given system. If your laptop size, power and cooling was designed for a gtx1030 you couldn't replace it with a gtx1080 module. In framework's case, the
49.
▲
by
Manabu-eo
2y ago
Your japanese phrases are quite easy for machine translation, as they have clear subjects stated and not much ambiguity. The main problem is the implicit subject due to omission of pronoms in regular japanese, among other ambiguous things t
50.
▲
by
Manabu-eo
2y ago
Japanese text in uft-8 is frequently rendered with the Chinese version of the kanji due to han unification, not "representing their writing method exactly". Shift-JIS encoding comunicates that the text is in Japanese via the encod
51.
▲
by
Manabu-eo
2y ago
They gathered images from pinterest.
52.
▲
by
Manabu-eo
2y ago
Neptune is still fully imbued of planet-hood, regardless if you prefer to call Ceres a planet or dwarf-planet.
53.
▲
by
Manabu-eo
2y ago
For decent performance, you need to keep all the parameters on memory for both. Well, with a raid-0 of two PCIe 5 SSDs (or 4 PCIe 4) you might get 1 t/s loading experts from disk on snowflake-artic... but that is slooow.
54.
▲
by
Manabu-eo
2y ago
The old google's Switch-C transformer [1] had 2048 experts, 1.6T parameters, with only one activated for each layer, so much more sparse. But also severely undertrained as the other models of that era, and thus useless now. 1. https:&
55.
▲
by
Manabu-eo
2y ago
Wrong. MoE models like this one usually chose a different and unpredictable mix of experts for each token, and as such you need all parameters at memory at once. It lessens the number of parameters that need to be moved from memory to compu
56.
▲
by
Manabu-eo
2y ago
But those people usually have more system RAM than VRAM. At those scales, most people become bandwidth and compute constrained using CPU inference instead of multiple GPUs. In those cases, an MOE with a low number of active parameters is th
57.
▲
by
Manabu-eo
2y ago
> Might be time to invest in a Mac studio! The highest end Mac Studio with 196GB of ram won't even be enough to run a Q4 quant of the 400B+ (don't forget the +) model. At this point, one have to consider an Epyc for CPU infere
58.
▲
by
Manabu-eo
2y ago
Nope, but this guy has a similar build: https://www.reddit.com/r/LocalLLaMA/comments/1bt8kc9/compari... It seems to reach only a little above half the theoretical speed, and scale only up to 32 threads f
59.
▲
by
Manabu-eo
2y ago
But one hypermiler reporting a long range under perfect conditions is exactly what he described at length while making that claim. It was pretty clear to me. And that was exactly what we got in 2017. If one wants to lie by twisting what he
60.
▲
by
Manabu-eo
2y ago
SpaceX was the company he was most involved, and still is pretty involved. Just watch any factory tour he made, be it the 2005 ones or the new Texas ones. Or hear what people who worked with him talk. And I can't trust that site. An ex
More ›