Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
liuliu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
121.
▲
Metal FlashAttention v2.5 with Neural Accelerators on the Apple M5 Chip
(releases.drawthings.ai)
3 points
by
liuliu
11mo ago
|
0 comments
122.
▲
by
liuliu
11mo ago
Ah, I get what you mean now. I am mixing up the nn module and the tensor execution bits. (to be fair, the PyTorch nn module carries over many these quirks!).
123.
▲
by
liuliu
11mo ago
LuaTorch is eager-execution. The problem with LuaTorch is the GC. You cannot rely on traditional GC for good work, since each tensor is megabytes (at the time), now gigabytes large, you need to collect them aggressively rather than at inter
124.
▲
by
liuliu
11mo ago
That's wrong. Llama.cpp / Candle doesn't offer anything on the table that PyTorch cannot do (design wise). What they offer is smaller deployment footprint. What's modern about LLM is the training infrastructure and singl
125.
▲
by
liuliu
11mo ago
Note that busy_timeout is not applicable to SQLite in this case (the SQLITE_BUSY issued immediately, no wait in this case). Also this is because WAL mode (and I believe only for WAL mode, since there is really no concurrent reads in the oth
126.
▲
by
liuliu
1y ago
Developers can still choose to enable sandbox for apps delivered outside of App Store. Some of them simply choose to not do so: https://developer.apple.com/documentation/security/hardened-...
127.
▲
by
liuliu
1y ago
How to do site-to-site traffic over Tailscale / WG encryption? From preliminary testing, it seems have difficulty to saturate a 10Gbps connection while plain HTTP (nginx) traffic does that fine. Of course it should vary from CPU to CPU
128.
▲
by
liuliu
1y ago
Yeah, luckily, you can unit tests these and fix them. They are not concurrency bugs (again, luckily). BTW, numeric differentiation can only be tested very limitedly (due to algorithmic complexity when you doing big matrix). It is much easie
129.
▲
by
liuliu
1y ago
And it is always felt to me that has lineage from neural Turing machine line of work as prior. The transformative part was 1. find a good task (machine translation) and a reasonable way to stack (encoder-decoder architecture); 2. run the ex
130.
▲
by
liuliu
1y ago
mmap is a good crutch when you 1. don't have busy polling / async IO API available and want to do some quick & dirty preloading tricks; 2. don't want to manage the complexity of in-memory cache, especially cross-processes
131.
▲
by
liuliu
1y ago
Is this just https://developer.apple.com/documentation/uikit/uiglasseffec... ?
132.
▲
by
liuliu
1y ago
I am not an expert on ANE, but I think it is related to the size of register files and how that is smaller than what we need for GEMM on modern transformers (especially these fat ones with MoE).
133.
▲
by
liuliu
1y ago
Faster compute helps, for things like vision language model that requires bigger context to be filled. My understanding is that ANE is still optimized for convolution load, and compute efficiency while the new neural accelerators optimized
134.
▲
by
liuliu
1y ago
Feels like a side-effect of forever 0.x version symptom (I am guilty of as well). Even though semi-ver says 0.x can do whatever, people don't associate enough disruptive changes to it, whereas 0.4.x if it is 1.x, then it is much cleare
135.
▲
by
liuliu
1y ago
There is nothing for us to take away in this discussion. So let me be the first to tune down: all I want to say is: don't take that 300ms as given, it sits in this uncomfortable region too short to be an async op and too long to be not
136.
▲
by
liuliu
1y ago
To put in perspective, 300ms is about looping over 30GiB data from RAM, loading 800MiB data from SSD, or doing 1TFLOPS on a single core computer. 300ms to generate a report would be able to go through ~100M rows at least (on a single core).
137.
▲
by
liuliu
1y ago
And 300ms for a DB call is slow, in any case. We really shouldn't accept that as normal cost of doing business. 300ms is only acceptable if we are doing scrypt type of things.
138.
▲
by
liuliu
1y ago
https://releases.drawthings.ai/p/iphone-17-pro-doubles-ai-pe...
139.
▲
by
liuliu
1y ago
No, iPad Pro won't be faster than 4090s or 4070s (or even 5% of the speed of 4090). But newer chips might contain Neural Accelerator to close the gap a little bit (i.e. 10%??). (I maintain https://apps.apple.com/us/
140.
▲
by
liuliu
1y ago
They were acquisition target since 2017 (from the OpenAI internal emails). So lacking of acquisition is not because lacking of interests. Let you wonder what happened in these due-diligence.
141.
▲
by
liuliu
1y ago
Video generation is extremely exciting a.k.a. https://video-zero-shot.github.io/ However, personalization (teleporting yourself into a video scene) is boring to me. At its core, it doesn't generate new experience to me
142.
▲
by
liuliu
1y ago
Very interesting read. I first learned this method from a random reddit post a while ago and very happy to see a systematic study on this (wish I would save the original post somewhere to reference to!).
143.
▲
by
liuliu
1y ago
- 454 instances of "Rc<" in Servo: https://github.com/search?q=repo%3Aservo%2Fservo+Rc%3C&type=... - 6 instances of "Rc<" in AWS SDK for Rust: https://github.com/search?q=repo%3
144.
▲
by
liuliu
1y ago
My understanding is that you cannot talk about warp specialization without talking about the alternative: multi-stage pipelining. And the final example code given is multi-stage pipeline with double buffers. And here is my understanding whe
145.
▲
iPhone 17 Pro Doubles Qwen Image Generation On-Device
(releases.drawthings.ai)
2 points
by
liuliu
1y ago
|
0 comments
146.
▲
by
liuliu
1y ago
Or the GluonCV by mxnet guys (ancient! https://github.com/dmlc/gluon-cv )
147.
▲
by
liuliu
1y ago
The reason why people go distances to package PyTorch is because the skill of translating models between different frameworks manually is "easy" but not well dispensed in developer community. That's why people will go stupid
148.
▲
Optimizing Qwen Image for Edge Devices
(engineering.drawthings.ai)
4 points
by
liuliu
1y ago
|
0 comments
149.
▲
by
liuliu
1y ago
That's right, but it is much easier to just use blob without application logic to worry about chunking. It is the same reason why we use SQLite in the first place, a lot of transaction / rollback logic now is on SQLite layer, not
150.
▲
by
liuliu
1y ago
One thing I would call out, if you use SQLite as an application format: BLOB type is limited to 2GiB in size (int32). Depending on your use cases, that might seem high, or not. People would argue that if you store that much of binary data i
More ›