Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
joefourier
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
joefourier
1y ago
Vibe coding large projects isn’t feasible yet, but as a developer here’s how I use AI to great effect, to the point where losing the tool greatly decreases my productivity: - Autocomplete in Cursor. People think of AI agents first when they
92.
▲
by
joefourier
1y ago
I’m someone with ADHD who takes prescribed stimulants and they don’t make me work faster or smarter, they just make me work . Without them I’ll languish in an unfocused haze for hours, or zone in on irrelevant details until I realise I hav
93.
▲
by
joefourier
1y ago
I’ve tried it in the latest plug-in I’ve worked on - it’s just a webview embedded in the window. It’s really great for faster development of more complex plug-ins but there’s some downsides when it comes to performance and integration with
94.
▲
by
joefourier
1y ago
This type of anti-AI article is as vacuous and insipid as the superficial hype pieces peddled by pro-AI influencers. Saying generative AI is inherently shit, that there is 0 future in it, that it’s not good at anything, that it hasn’t impro
95.
▲
by
joefourier
1y ago
Autism is a spectrum disorder and I don’t think it should be controversial to cure low functioning autism. However, high functioning autists can be argued to be more of a personality variant than a disability, with different strengths and w
96.
▲
by
joefourier
1y ago
While true in a theoretical sense (an MLP of sufficient size can theoretically represent any differentiable function), in practice it’s often the case that it’s impossible for a certain architecture to learn a specific task no matter how mu
97.
▲
by
joefourier
2y ago
That’s not at all how multi-modal LLMs work - their visual input is not words generated by a classifier. Instead the image is divided into patches and tokenised by a visual encoder (essentially, it is compressed), and then fed directly as a
98.
▲
by
joefourier
3y ago
Not all vision transformers have weak priors. Shifted-window transformers and neighborhood attention have priors well suited to images; the latter is showing extremely strong performance in image generation (such as the recent hourglass dif
99.
▲
by
joefourier
3y ago
If you look at the supplementary material, they do a test where they train the Lora with a randomly initialised Unet and it is largely incapable of extracting any surface normals as opposed to using the pretrained Stable Diffusion Unet - cl
100.
▲
by
joefourier
3y ago
I think you must have misunderstood me, I didn’t say the SD-XL VAE had the same issue as in OP. What I said was that it didn’t take into account some of my points that came up during my research: - Bounding the outputs to -1, 1 and optimisi
101.
▲
by
joefourier
3y ago
Not necessarily if the model is trained with an appropriate adversarial loss. The reason that VAEs are blurry isn’t directly because of the KL divergence loss term but because of the L1/L2 loss. Since VAEs sample from a Gaussian distri
102.
▲
by
joefourier
3y ago
From the SD-XL paper: > To this end, we train the same autoencoder architecture used for the original Stable Diffusion at a larger batch-size (256 vs 9) and additionally track the weights with an exponential moving average. The resulting
103.
▲
by
joefourier
3y ago
The SD-XL VAE doesn’t take into account any of those insights, it’s the exact same as the SD1/2 one, just trained from scratch with a batch size of 256 instead of 9 and with EMA.
104.
▲
by
joefourier
3y ago
I’ve done a lot of experiments with latent diffusion and also discovered a few flaws in the SD VAE’s training and architecture, which have hardly no attention brought to them. This is concerning as the VAE is a crucial competent when it com
105.
▲
by
joefourier
3y ago
Well… a car crash is absolutely not binary. There’s a huge spectrum between a fender bender at low speed and a deadly multi-vehicle pile up on the motorway. And there’s a huge spectrum of injuries from whiplash, fractures, loss of limbs, an
106.
▲
by
joefourier
3y ago
So from my understanding, the Dunning-Kruger Effect paper doesn’t show the distribution of the perceived test scores nor the standard deviation, only an average, which rises with actual test score level. If they showed the spread bar in eac
107.
▲
by
joefourier
3y ago
This is why it’s a good thing it’s staying private, public investors would focus only on the high profile explosions of a test vehicle rather than on the fact that SpaceX has a 99.3% mission success rate, the record for the most launches in
108.
▲
by
joefourier
3y ago
AWS is extremely overpriced for nearly every service. I don’t know why anyone else outside of startups with VC money to burn or bigcos that need the “no one ever got fired for buying IBM” guarantee would use them. You’re better off with Lam
109.
▲
by
joefourier
3y ago
I wouldn’t compare GenAI to cryptocurrency because there’s no equivalent pyramid scheme mentality of “invest early in this asset and sell it later 100x” - it’s much closer to a traditional tech hype bubble. The average user isn’t roped in b
110.
▲
by
joefourier
3y ago
A100s aren’t $2/h at MJ scale tho. Firstly they likely own quite a few and they’re relatively cheap now at $4k a card, and secondly you can get much better deals than $2/h if you rent a large amount and pre-pay.
111.
▲
by
joefourier
3y ago
Those surely wouldn’t count because they’re entire CPUs on one chip, the point is making a CPU from parts that aren’t, which is sometimes dodgy sure.
112.
▲
by
joefourier
3y ago
Not impossible at all - classifier networks are much, much easier to train than generative networks. However you can’t directly integrate the logic into the generator, you’d have to train the generator against the discriminator network. Thi
113.
▲
by
joefourier
3y ago
You don’t have to run Llama 70B on a rented 2xA100 80GB which is of course going to be quite pricy. Quantising it to 4-bit as brucethemoose2 mentioned allows you to run it on far cheaper hardware - it’ll fit on a single A6000 which can be r
114.
▲
by
joefourier
3y ago
Can I ask what’s the point of testing students in a scenario where they don’t have access to AI tools? As a software developer I use these tools daily, and if I or my colleagues stopped our productivity would suffer - it sounds like the sam
115.
▲
by
joefourier
3y ago
How on earth do GCP and Azure lose money when they (just like AWS) are ridiculously expensive? Like frequently >3x the competition for VMs and for bandwidth egress, I’ve sometimes seen 10x, or charging for things that competitors don’t.
116.
▲
by
joefourier
4y ago
Not exactly the point but current smartphones are comparable to 90s supercomputers in terms of FLOPs - an iPhone 14 can hit 2 TFLOPs whereas ASCI Red achieved 2.38 TFLOPs in 1999. Now the comparison is muddied by the fact that supercomputer
117.
▲
by
joefourier
4y ago
LLMs certainly can “will” and “do things” when provided with the right interface like LangChain: https://github.com/hwchase17/langchain See also the ARC paper where the model was capable of recruiting and convincing a
118.
▲
by
joefourier
4y ago
The Jevons paradox was noted in the 19th century when the world was far less wealthy than today. Recent American monetary policy has little relevance. There’s many more examples if you look up induced demand - a similar effect is the Downs–
119.
▲
by
joefourier
4y ago
Increased efficiency leads to more demand, not less. This is known as the Jevons paradox in economics (see: https://en.m.wikipedia.org/wiki/Jevons_paradox ). The historical precedent would be for increased productivity
120.
▲
by
joefourier
4y ago
The Chinese Room is a fallacious bait-and-switch argument that should stop being brought up. It, along with the Russian game from the TFA, don’t actually provide any insight - all they do is have the computer program be executed by humans i
More ›