Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
benreesman
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
benreesman
8mo ago
The thing you want has a kind of academic jargon name (coeffects algebra with graded/indexed monads for discharge) but is very intuitive, and it can do useful and complete attestation without compromising anyone credentials (in the lim
32.
▲
by
benreesman
8mo ago
I compile nix derivations to well-posed effect/coeffect/graded monad algebra so I can do real bill of materials and build on an action cache engine maintained by professionals, but that's mostly for long-tail stuff. These day
33.
▲
by
benreesman
8mo ago
They can all write lean4 now, don't accept numbers that don't carry proofs. The CAS I use for builds has a coeffect discharge cert in the attestation header, couple lines of code. Graded monads are a snap in CIC.
34.
▲
by
benreesman
9mo ago
As with many things the secret sauce of AI-assisted software engineering is mathematics, skepticism, and hard work. The first thing anyone should do is immediately understand that they are in a finite-sum adversarial relationship with all o
35.
▲
by
benreesman
9mo ago
I mostly write lean4 now and emit proof-carrying System F Omega via rfl. It's the right level of abstraction when the droids have been pinned to theory laden symbolisms. It's also just pleasant to use.
36.
▲
by
benreesman
9mo ago
I mean I'm running TensorRT-LLM on a basket of spot vendors at NVFP4 with auction convexity math and Clickhouse Keeper and custom passthrough. I need more tokens not less because the available weight models aren't quite as strong,
37.
▲
by
benreesman
9mo ago
Anthropic might be the first gigantic company to destroy itself by bootstrapping a capability race it definitionally cannot win. They've been leading in AI coding outcomes (not exactly the Olympics) via being first on a few things, not
38.
▲
A zero-overhead bridge between C++23 std:mdspan and CUTLASS cute layouts
(github.com)
2 points
by
benreesman
9mo ago
|
1 comments
39.
▲
by
benreesman
9mo ago
This library is a small thing. A few hundred lines of C++. A template adapter. But it stands on the shoulders of two decades of work by people who thought deeply about the shape of computation itself.
40.
▲
by
benreesman
9mo ago
It's a test of polyhedral layout algebra, what NVIDIA calls CuTe and the forthcoming C++ standard calls std::mdspan. This is the general framework for reasoning about correct memory addressing in the presence of arbitrary constraints
41.
▲
by
benreesman
9mo ago
It would need to implement a few dozen ioctls, correctly stub the kernel module in guests, do a probably memory-safe assignment of GPU memory to guest, and then ultimately map that info to BAR/MSI-X semantics of a real kernel module. Y
42.
▲
by
benreesman
9mo ago
firecracker vm: https://gist.github.com/b7r6/26b3e5c48a00d903ef617f1b073eb98...
43.
▲
by
benreesman
9mo ago
Systems programming in the large is hard, owning the category for decades harder still. Even languages that have tried to fast-follow and disrupt C++ end up looking a lot like C++. There is an irreducible complexity.
44.
▲
by
benreesman
9mo ago
I'm telling your it works now. It's just not called `tcgen05`. Put this in nsight compute: https://github.com/NVIDIA/cutlass/blob/main/examples/79_blac... (I said 83, it's 79). If you
45.
▲
by
benreesman
9mo ago
sm_120 (aka 1CTA) supports tensor cores and TMEM just fine: example 83 shows block-scaled NVFP4 (I've gotten 1850 ish dense TFLOPs at 600W, the 300W part caps out more like 1150). sage3 (which is no way in hell from China, myelin knows
46.
▲
by
benreesman
9mo ago
NVFP4 (and to a lesser extent, MXFP8) work, in general. In terms of usable FLOPS the DGX Spark and the GMTek EVO-X2 both lose to the 5090, with NCCL and OpenMPI set up the DGX is still the nicest way to dev for our SBSA future. Working on t
47.
▲
by
benreesman
9mo ago
Fun fact: if you say the right prayers to the Myelin Gods it will fuse straight through sage3 at D/DQ like it's seen it before, which of course it has. https://gist.github.com/b7r6/94f738f4e5d1a67d4632a8fbd18d
48.
▲
by
benreesman
9mo ago
I have been a big Astral and uv booster for a long time. But specifications like this one: https://gist.github.com/b7r6/47fea3c139e901cd512e15f42355f26... have me re-evaluating everything. That's TensorRT-LLM in i
49.
▲
by
benreesman
10mo ago
NVFP4 is the thing no one saw coming. I wasn't watching the MX process really, so I cast no judgements, but it's exactly what it sounds like, a serious compromise in resource constrained settings. And it's in the silicon pipe
50.
▲
by
benreesman
1y ago
clang++ $(pkg-config --cflags --libs libtorch) qwen-3-nvfp4.cpp -o ./qwen-3-infer Your move.
51.
▲
by
benreesman
1y ago
I'm a bit later in my career and I've been involved with modern machine learning for a long time which probably affects my views on this, but I can definitely relate to aspects of it. I think there are a couple of good signals in
52.
▲
by
benreesman
1y ago
There's nontrivial historical precedent for this exact playbook: when a new paradigm (Lisp machines and GOFAI search, GPU backprop, softmax self-attention) is scaling fast, a lot of promises get made, a lot of national security money g
53.
▲
by
benreesman
1y ago
Now imagine if someone combined Jia Tan patience with swiss-cheese security like all of our editor plugins and nifty shell user land stuff and all that. Developer stuff is arguably the least scrutinized thing that routinely runs as mega r
54.
▲
by
benreesman
1y ago
I personally regard posterior/score-gradient/flow-match style models as the most interesting thing going on right now, ranging from rich media diffusers (the extended `SDXL` family tree which is now MMDiT and other heavy transform
55.
▲
by
benreesman
1y ago
I'd like to "reclaim" both AI and machine learning as relatively emotionally neutral terms of art for useful software we have today or see a clearly articulated path towards. Trying to get the most out of tools that sit somew
56.
▲
by
benreesman
1y ago
Yeah, you and I are both entitled to our guesses about Rust and Zig respectively in the future. Both have serious production projects (bun and TigerBeetle are best-in-class to head off any "hobby project" stuff), neither is making
57.
▲
by
benreesman
1y ago
If you don't have good tests with coverage information then your Python and Rust code is buggy too. If you want "as low of defects as I can get from the compiler alone" then your options are things like Haskell and theory-lad
58.
▲
by
benreesman
1y ago
The fast interconnect between nodes has aaplications in inference at scale (big KV caches and other semi-durable state, multi-node tensor parallelism on mega models). But this article in particular is emphasizing extreme performance ambitio
59.
▲
by
benreesman
1y ago
Nah, distributing rootkits under false pretenses is a dick move. That's not even a little controversaial. You put a thing on the web that says "Just a harmless XYZ" and it roots TLS forever? Malware. Black and white.
60.
▲
by
benreesman
1y ago
I dramatically prefer modern C++ to either of Python or Rust in domains where it's a toss up. It's really nice these days. Like any language that lasts (including Python and Rust) you subset it over time: you end up with linters a
More ›