Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ollin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
ollin
1y ago
This is very encouraging progress, and probably what Demis was teasing [1] last month. A few speculations on technical details based on staring at the released clips: 1. You can see fine textures "jump" every 4 frames - which mean
32.
▲
by
ollin
1y ago
I think the most likely explanation is that they trained a diffusion WM (like DIAMOND) on video rollouts recorded from within a 3D scene representation (like NeRF/GS), with some collision detection enabled. This would explain: 1. How c
33.
▲
by
ollin
1y ago
Got it, that makes sense! In terms of raw compute capability, a Snapdragon 888's GPU should have more than enough power to run this demo smoothly. I think I just need to optimize the inference setup better (maybe switch to WebGPU if th
34.
▲
by
ollin
1y ago
Curious, which device/OS/browser? I did all my testing on 4-year old hardware (iPhone 13 Pro, M1 Pro MBP), and the model itself is extremely tiny (~1GFLOP) so I'm optimistic that performance issues would be solvable with a be
35.
▲
by
ollin
1y ago
I think https://diamond-wm.github.io is a reasonable place to start (they have public world-model training code, and people have successfully adapted their codebase to other games e.g. https://derewah.dev/project
36.
▲
by
ollin
1y ago
Yup, definitely similar! There are a lot of video-game-emulation World Models floating around now, https://worldarcade.gg had a list. In the self-driving & robotics literature there have also been many WMs created for policy
37.
▲
by
ollin
1y ago
Mostly 1xA10 (though I switched to 1xGH200 briefly at the end, lambda has a sale going). The network used in the post is very tiny, but I had to train a really long time w/ large batch to get somewhat-stable results.
38.
▲
by
ollin
1y ago
Thanks! My favorite failure mode (not mentioned in the post - I think it was during the first round of upgrades?) was a "dry" form of soupification where the texture detail didn't fully disappear https://imgur.com&
39.
▲
by
ollin
1y ago
Yes! This was a solo project done in my free time :) to learn about WMs and get more practice training GANs. The special aspect of NNs (in the context of simulating worlds) is that NNs can mimic entire worlds from videos alone, without acce
40.
▲
by
ollin
2y ago
yup, george also commented on the cruise news here https://x.com/realGeorgeHotz/status/1866617393436651688
41.
▲
by
ollin
2y ago
We can't assess the quality of gameplay ourselves of course (since the model wasn't released), but one author said "It's playable, the videos on our project page are actual game play." ( https://x.com
42.
▲
by
ollin
2y ago
does anyone have links to a canonical implementation of modern-style first person controls? I’ve been looking into this recently (both the WASD+mouselook PC version and the corresponding dual-stick console/touch version) and the closes
43.
▲
Gen-3 Alpha: A New Frontier for Video Generation
(runwayml.com)
22 points
by
ollin
2y ago
|
2 comments
44.
▲
by
ollin
2y ago
karpathy gave a good high-level history of the transformer architecture in this Stanford lecture https://youtu.be/XfpMkf4rD6E?si=MDICNzZ_Mq9uzRo9&t=618
45.
▲
by
ollin
3y ago
more q&as from the author on twitter https://twitter.com/mov_axbx/status/1772548497566294345 --- why the crt monitor? > What I had handy with a VGA input I am wondering how you went over the p2p problem wit
46.
▲
Building WOPR: A 7x4090 AI Server
(mov-axbx.com)
28 points
by
ollin
3y ago
|
11 comments
47.
▲
by
ollin
3y ago
They use three text encoders to encode the caption: 1. CLIP-G/14 (OpenCLIP) 2. CLIP-L/14 (OpenAI) 3. T5-v1.1-XXL (Google) They randomly disable encoders during training, so that when generating images SD3 can use any subset of the
48.
▲
by
ollin
3y ago
it may turn out more like the imagen timeline 2022-05 - google imagen research paper posted https://news.ycombinator.com/item?id=31484562 2022-12 - imagen developers leave google to form ideogram 2023-08 - ideogram ships a
49.
▲
by
ollin
3y ago
some points that stood out to me: 1. they made a lot of careful tweaks to the unet network architecture - it seems like they ran many different ablations here ("In total, our endeavor consumes approximately 512 TPUs spanning 30 days&qu
50.
▲
by
ollin
3y ago
"Character-Aware Models Improve Visual Text Rendering" https://arxiv.org/abs/2212.10562 for anyone curious
51.
▲
by
ollin
3y ago
Yes - the document is covering their entire risk-mitigation strategy. I've extracted the sections that seemed relevant to me below. The purpose of the prompt transformation system: > we share the work done to prepare DALL·E 3 for de
52.
▲
by
ollin
3y ago
For more context on why this system prompt exists, see https://cdn.openai.com/papers/DALL_E_3_System_Card.pdf
53.
▲
by
ollin
3y ago
There's a second blog post here with substantially more technical detail: https://developer.nvidia.com/blog/breaking-mlperf-training-r... Additionally, code for the actual submission is available here https:/
54.
▲
by
ollin
3y ago
It generates 64x64 for the first stage, but there's a button to upscale your favorite 64x64 image to usable resolution.
55.
▲
Training Stable Diffusion from Scratch for <$50k
(mosaicml.com)
6 points
by
ollin
3y ago
|
2 comments
56.
▲
Bare-Bones Diffusion Models
(madebyoll.in)
4 points
by
ollin
4y ago
|
0 comments
57.
▲
by
ollin
4y ago
yeah, running the full decoder takes a while. though, since the "latent" is just 4 channels and pretty close to representing RGB, you can use a linear combination of latent channels and get a basic (grainy, low-res) preview image
58.
▲
by
ollin
4y ago
here's a direct app store link, if anyone wants to try the iPhone app immediately: https://apps.apple.com/us/app/draw-things-ai-generation/id64... congratulations to liuliu on the launch!
59.
▲
Stable Diffusion inference on iOS / macOS using MPSGraph
(github.com)
2 points
by
ollin
4y ago
|
0 comments
60.
▲
by
ollin
4y ago
Author here - thanks for the thoughtful points! I think we agree on a lot, but do disagree on the conclusion :) My mental model (mostly == Karpathy's "software 2.0" one) is: source code :: training data (and scori
More ›