Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thegeomaster
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
91.
▲
by
thegeomaster
2y ago
This universally fails, on anything from frontier models to Gemini 2.0 Flash in its custom fine-tuned bounding box extraction mode.
92.
▲
by
thegeomaster
2y ago
Viber is alive and well nowadays and is the dominant messaging app in quite a few geographies. Given that Facebook Messenger seems to also have about the same MAU as WhatsApp (and seems to be dominant in the US), I don't think you can
93.
▲
by
thegeomaster
2y ago
It's quite a bit better than coding --- they hint that it can tie o1's performance for coding, which already benchmarks higher than 4o. And it's significantly cheaper, and presumably faster. I believe API costs account for th
94.
▲
by
thegeomaster
2y ago
Love the Bleach reference in the name.
95.
▲
by
thegeomaster
2y ago
The parent comment was talking about coding specifically, not the average score. I see o1 at 69.69, and Claude 3.5 Sonnet at 67.13.
96.
▲
by
thegeomaster
2y ago
LiveBench (which I like because it tries very hard to avoid contamination) ranks Sonnet 3.5 second only to o1 (which is totally expected).
97.
▲
by
thegeomaster
2y ago
I simplified in my comment. It was a much better story for tooling, since you could reuse large parts of existing backends/codegen, optimization passes, and debugging. The mental model of execution would remain too, rather than being a
98.
▲
by
thegeomaster
2y ago
WASM nowadays has become quite the monstrosity compared to NaCl/PNaCl. Just look at this WASM GC spaghetti, trying to compile a GC'd language but hooking it up V8/JavaScriptCore's GC, while upholding a strict security mo
99.
▲
by
thegeomaster
2y ago
> And once we start training these larger systems from the ground up via full reinforcement learning (rather than composing existing models), Agree with this totally. I wouldn't call what the CoT models are doing exactly being able
100.
▲
by
thegeomaster
2y ago
Exactly. LLMs are gullible . They will believe anything you tell them, including incorrect things they have told themselves. This amplifies errors greatly, because they don't have the capacity to step back and try a different approach
101.
▲
by
thegeomaster
2y ago
On your object detection point, Gemini 2.0 Flash has bounding box detection: https://ai.google.dev/gemini-api/docs/models/gemini-v2#bound... . I haven't found it to work particularly well for some more do
102.
▲
by
thegeomaster
2y ago
Is it actively worse, though? My impression is that all of the other, classical-based in-painting methods are still alive and well in Adobe products. And I think their in-painting works well, when it does work. To me, this honestly sounds l
103.
▲
by
thegeomaster
2y ago
For isolated coding tasks I've been using o1-preview instead of Sonnet for a while now, I just didn't mention it. Haven't had a chance to test o1 proper, but I assume it's also a jump in performance. However, for more &q
104.
▲
by
thegeomaster
2y ago
Yeah, I would like to know this. From my perspective, even frontier models by the big players (4o, 3.5 Sonnet) can be unreliable at times, and are at best just walking the line of usefulness for a lot of "exact" tasks (for me: pro
105.
▲
by
thegeomaster
2y ago
This is literally chain-of-thought! Even better than generic chain-of-thought prompting ("Think step by step and write down your thought process."), you're doing a domain-specific CoT, where you use some of your human intuiti
106.
▲
by
thegeomaster
2y ago
Reading between the lines, it sounds like they are creating an AI product for more than just their own codebase. If this is the case, they'd probably be keeping a lot of the secret sauce hidden. More broadly, it's nowadays almost
107.
▲
by
thegeomaster
2y ago
This is not true. CNNs perform 2D convolution, conceptually "sliding" a 2 dimensional kernel with learnable weights over the input image across two dimensions. Perhaps it wasn't a convolutional network after all, but a simple
108.
▲
by
thegeomaster
2y ago
The classifier was likely a convolutional network, so the assumption of the image being a 2D grid was baked into the architecture itself - it didn't have to be represented via the shape of the input for the network to use it.
109.
▲
by
thegeomaster
2y ago
I think the parent commenter was making a joke.
110.
▲
by
thegeomaster
2y ago
In one of my jobs, a 1% perf regression (on a more stable/reproducible system, not PCs) was a reason for a customer raising a ticket, and we'd have to look into it. For dynamically dispatched but short functions, the overhead is e
111.
▲
by
thegeomaster
2y ago
Look, it's likely we just come from different backgrounds. Most of my perf-sensitive work was optimizing inner loops with SIMD, allowing the compiler to inline hot functions, creating better data structures to make use of the CPU cache
112.
▲
by
thegeomaster
2y ago
I get it. This frustrated me to no end . But still I did what I had to do --- recompiled random software throughout the stack, enabled random flags, etc. It was doable and now I can do it much faster. I don't think it's fair for
113.
▲
by
thegeomaster
2y ago
There are no performance winners if you include them by default. There will be an additional >0% overhead when you are executing additional code in the prologue and epilogue, and increasing the register pressure by removing rbp from bein
114.
▲
by
thegeomaster
2y ago
I'd shy away from storing any non-volatile state on Hetzner. As I said, I'd mostly consider it for stateless compute-bound applications. If I was looking to scale up an existing operation considerably and minimize costs as much as
115.
▲
by
thegeomaster
2y ago
I would probably host even some business-critical services on Hetzner's infra. I'm thinking of "worker"-type workloads, where each machine is 100% stateless and just serves to do some compute-intensive work. With that co
116.
▲
by
thegeomaster
2y ago
One important approach is missing: extrapolation. Usually, your entities all have velocities, which you can use to extrapolate from the last simulated state to the current one (after <dt time has passed). For things like visual effects,
117.
▲
by
thegeomaster
2y ago
I suspect that plus vs minus is arbitrary in this case (as you said, due to being able to learn a simple negation during training), but they are presenting it in this way because it is more intuitive. Indeed, adding two sources that are noi
118.
▲
by
thegeomaster
2y ago
A bubble doesn't necessarily imply no underlying worth. The dot-com bubble hit legendary proportions, and the same underlying technology (the Internet) now underpins the whole civilization. There is clearly something there, but a bub
119.
▲
by
thegeomaster
2y ago
Joining the praise in this thread. Extremely reliable, fast, versatile, plays anything. What I haven't seen mentioned is that it has 1- or 2-key keyboard shortcuts for almost everything, down to adjusting audio/video delay, subtit
120.
▲
by
thegeomaster
2y ago
mpv does it! Period key (.) for next frame, comma (,) for previous. Super handy.
More ›