Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tsurba
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
31.
▲
by
tsurba
2y ago
Proper initialization of layers keeps gradient magnitudes from vanishing/exploding in deep networks. If you make sure the output of each layer has mean 0, std 1, the gradients will be reasonable as well, for example. I recommend e.g. t
32.
▲
by
tsurba
2y ago
That’s the difference between truly new approaches to modelling an existing problem, or coming up with a new problem. No set of a bit different results or missing exact hyperparameter settings really invalidates the value of the aforementio
33.
▲
by
tsurba
2y ago
If you are not even going to bother writing them up properly, no one is going to care. Seems fair to me. You don’t have to make a ”paper” out of it, maybe make blog post or whatever if that is more your style. Maybe upload a pdf to arxiv. H
34.
▲
by
tsurba
2y ago
My favorite quote in this topic: ”If intelligence lies in the process of acquiring new skills, there is no task X that solving X proves intelligence” IMO it especially applies to things like solving a new IQ puzzle, especially when the mode
35.
▲
by
tsurba
2y ago
The article kinda sucks as it does not really answer the question it poses. Why ”noise tends to be a an unwanted amplitude modulation, not a frequency modulation”? Is it due to naturally occurring background noise being low frequency high a
36.
▲
by
tsurba
2y ago
Shadertoy also has nice ones https://www.shadertoy.com/view/Ms2SD1
37.
▲
by
tsurba
2y ago
Beating all other models in latest benchmarks by a wide margin with a completely new approach, while having the only properly working audio chat with transformer isn’t enough? Not to even mention SORA if actually comes out someday. I’m not
38.
▲
by
tsurba
2y ago
GPL doesn’t, but as others have already said, this discussion is not about that. If they want to keep using someone else’s free server capacity, maybe they should be giving something in return.
39.
▲
by
tsurba
2y ago
The writer should stick with nontechnical points since they clearly don’t understand why the models currently used are not deterministic. It’s trivial to make them deterministic.
40.
▲
by
tsurba
2y ago
>Tennis court size is in feet Lol.
41.
▲
by
tsurba
2y ago
Yet understanding is necessary but not sufficient when you read university math, especially advanced courses. Proofs assume you have the elusive thing referred to as ”mathematical maturity”, which means many algebraic manipulation steps are
42.
▲
by
tsurba
2y ago
In a sense even this is not true, as in any sufficiently complex (which turns out to be quite simple) formal system you can create proofs that are true and untrue at the same time creating a contradiction. In other words, mathematics work
43.
▲
by
tsurba
2y ago
Unfortunately also the majority of scientific papers for eg. image generation have had completely cherry-picked examples for a long time now.
44.
▲
by
tsurba
3y ago
Nice! I would like to someday finish writing an OS for my IRC bot that is still running. Maybe the most useless comment but: that non-linear mouse movement (aka acceleration) is the very first thing I turn off when I boot up a new OS. It li
45.
▲
by
tsurba
3y ago
Very cool that the dataset and model weights are open right away! This paper also doesn't have a bunch of weird architectural choices pulled out of nowhere like the other TS foundation models recently. Looks like it will actually be us
46.
▲
by
tsurba
3y ago
Getting a job to make money and solve customer problems as fast as possible was not a stated goal by the OP. Besides, you are also wrong, having good fundamentals in the maths will help you pick up new methods much faster as they pop up. An
47.
▲
by
tsurba
3y ago
Very nice read from today’s perspective. Many points still hold, while a few have clearly progressed by leaps from back then. In hindsight language is a pretty nice generalizer as so much of it is available.
48.
▲
by
tsurba
3y ago
It could just be that improving the models does not parallelize that well, so you need to wait for the massive months long training runs to finish to see what worked and what to try next. OpenAI got started earlier going full on with scalin
49.
▲
by
tsurba
3y ago
What Brockman tweets is from technical standpoint the most mundane, boring, and obvious stuff I’ve read from a programmer. My reads of this guy have been he’s not working on any problems that are technically difficult (or interesting). It’s
50.
▲
by
tsurba
3y ago
Whisper is generally better than the one in youtube.
51.
▲
by
tsurba
3y ago
Like others said, subscriptions suck, and especially since there are apps like ”Habit List” (which I use) that have a one-time payment of only 6€, I’ll rather use that. It should not be a complicated app to make and maintain. The yearly App
52.
▲
by
tsurba
3y ago
I’m pretty sure what communities you are in are not actual research but some hype alarmist bullshit communities, since as a ML researcher absolutely zero of my peers think the things you say.
53.
▲
by
tsurba
3y ago
Essentially you are advocating against information being more efficiently available. Come on. It’s true we are fucked if bioweapons become easy to make, but that is not a question of ”AI”.
54.
▲
by
tsurba
3y ago
In the text they say you need to cram all information needed to predict the next token into a single 6KB word embedding, but isn’t that wrong? Rather, isn’t the autoregressively predicted single next token a combination (based on attention
55.
▲
by
tsurba
3y ago
The Theory That Would Not Die, by McGrayne was fun, about the history of Bayesian statistics. Also not exactly what was asked, but so good: James Gleick’s ”Information”. It is about the history of information theory or how to statistically
56.
▲
by
tsurba
3y ago
The Finnish capital area gov bought the health system from Epic systems made with Mumps for ~500 million € around 5 years ago, despite every software professional saying they shouldn’t. And just like predicted it’s utter dogpoo now that it’
57.
▲
by
tsurba
4y ago
Today I spent 2 hours cursing at GPT-4 for not being able to fix a stupid indexing mistake in the code it wrote. Just like code before. It’s helpful and I wouldn’t have the energy to work on this hobby project without GPT. But for now at le
58.
▲
by
tsurba
4y ago
Mj earlier versions were around before SD came out. Before dall-e 2 too, but after 1 IIRC. So I assume they have their own custom setup. Perhaps based on dall-e 1 paper originally (not weights as they were never published) and improved from