Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
carterschonwald
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
carterschonwald
5d ago
theres a certain amount of financial security needed for folks to do this kinda thing.
2.
▲
by
carterschonwald
6d ago
the current white house really fucked over all the smaller domestic manufacturers. and scientists. and um america basically
3.
▲
by
carterschonwald
7d ago
i do have some ideas that could be done at the model arch/inference level, but a shocking amount of stuff works even just as clever formatting
4.
▲
by
carterschonwald
9d ago
i actually have a harness setup that prevents role confusion from happening in a much more robust and interesting way. hoping to launch a nice commercial version as a saas with some compelling unique features in the next month or teo
5.
▲
by
carterschonwald
9d ago
good. theyll actually be more reliable if they dont have as much brain damage.
6.
▲
by
carterschonwald
9d ago
i think the lower bound on the end state is there cant be opaque reasonibg steps ever.
7.
▲
by
carterschonwald
10d ago
this is just not good science or engineering. ok as a student project i guess.
8.
▲
by
carterschonwald
11d ago
im pretty literate in intellectual property, but im pretty confused about the no looking at binary code artifacts bit.
9.
▲
by
carterschonwald
11d ago
… i thought international treaties explicitly say no space weapons. ughhhh fucking wh
10.
▲
by
carterschonwald
11d ago
thats important! but how do the error messages compare with a decent compiler? youre gonna have more errors in your llm corrective edits loop if you dont hew to that tier of clarifying. to be clear: youre taking this feedback wonderfully we
11.
▲
by
carterschonwald
12d ago
i'm not sure if that caveat is sensible :) theres actually a very important reason you want it to be an actual embedded dsl or tiny programming language! The reason why llms can code at all is the hugeeeee amount of RL based on the loo
12.
▲
by
carterschonwald
12d ago
repeat after me: yaml is not a programming language. (boo, hn strips emojis, i should know that) very disappointed that their domain specific language is just yaml. conditionals and variable binding become insane war crimes when yaml comes
13.
▲
by
carterschonwald
12d ago
ive been slowly working on my ideas for a proper second gen multi device harness that i want to be in the same quality regime as sublime text. unclear if theres a viable commercial market (and for a variety of reasons around user security a
14.
▲
by
carterschonwald
1mo ago
perhaps, but what public activity isn't activism? eg, i have very interesting empirical evidence that recent Anthropic models are specifically trained to refuse to critique the whitehouse cabinet and elected officials, and that this is
15.
▲
by
carterschonwald
1mo ago
… that dude probably would have preferred not having that in their life
16.
▲
by
carterschonwald
1mo ago
good.
17.
▲
by
carterschonwald
1mo ago
ive found degraded performance on models larger than 4.7. i assume its model damage from overly self righteous post training resulting in false/feigned balance imported into any long running complex task. wish i was joking.
18.
▲
by
carterschonwald
1mo ago
what sorts of background texts/sources of info did you use? edit: i think i have the dover book your readme mentions :)
19.
▲
by
carterschonwald
1mo ago
this confirms what i had determined empirically: even handedness directions got baked into the weights post opus 4.7. theres two reasons this is deeply bad 1) knowingly pursuing a policy that foreseeably causes the deaths of many thousands
20.
▲
Silent Data Corruption in PyTorch
(github.com)
2 points
by
carterschonwald
2mo ago
|
1 comments
21.
▲
by
carterschonwald
2mo ago
for those here using pytorch, theres some really nasty data corruption for computations that arent using powers of two as the coordinate axes. in gradient calcs ive had examples with multiple orders of magnitude of remative or absolute erro
22.
▲
by
carterschonwald
2mo ago
im building my own harness and inference tool chain for much of these reasons. theres so much to do that makes a big difference for users. hoping to get things into shape for early alpha as a saas in the next two months. heres the easiest b
23.
▲
by
carterschonwald
2mo ago
ummm, there is no substantive censorship with deep seek aside from first oarty hosting by deepseek for cya. trust me, and the suppression on deepseek hosted deepseek is pretty thin if you’re sophisticated and explain ethics justificstions.
24.
▲
by
carterschonwald
2mo ago
exactly. you certainly know more about the brain than i :)
25.
▲
by
carterschonwald
2mo ago
thx for the kind response! at the very least i have tools that let me easily hit better perf for fancy dense memory layouts, and the same tooling lets me experiment with frankly wildly wacky sparse and structured memory formats. the perfor
26.
▲
by
carterschonwald
2mo ago
i definitely will be doing some drop of some faster attention kernels in the next few weeks. like i can do all sorts of memory layout of tensors/matrices etc tricks that if you dont have the abstractions for it would just never happen.
27.
▲
by
carterschonwald
2mo ago
amusingly ive been working on ultra sparse llm inference/ training/ model design because nature loaths a dense graph/matrix and cause i think it shoukd be possible. i actually stood up a 20-25 percent faster than sota causal
28.
▲
by
carterschonwald
2mo ago
i think the line is: expressing that you reputationally certify its correct and its worth the time
29.
▲
by
carterschonwald
2mo ago
this mirrors my approx experience.
30.
▲
by
carterschonwald
2mo ago
i've had trouble finding any anecdotes or data about how to actually set/explore logit sampler settings
More ›