Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
carterschonwald
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
carterschonwald
4mo ago
some of the sandboxing ive been playing with gives me the best of both yolo and like logic programming tier perms on llm actions in env. still not ready for prime time though ;)
62.
▲
by
carterschonwald
4mo ago
its about overuse of rhetoric in a dilutive way. sitting with something is for way more personal or emotionally intense shock than one ceo saying other ceos are lying sacks of shit about layoffs, albeit in ceo speak. a republican party die
63.
▲
by
carterschonwald
4mo ago
“sit with it a moment” or similar phrases is one of the more common anthropic tells ive seen. ughh
64.
▲
by
carterschonwald
5mo ago
i cant find anything substantiated in the code that actually differentiates it from any other harness. my fork of oh my pi that i have a lot of experiments in, is lterally designed to only work well with models that have decent reasoning le
65.
▲
by
carterschonwald
5mo ago
just use the same context cache and a separate completion assuming elastic resources
66.
▲
by
carterschonwald
5mo ago
in my llm agent harness code i force all streamed reasoning to be turned into specially formatted visible to everyoneb msg blocks. it really does a nice job of fixing that while not caring about which provider for what edit: pretty sure ful
67.
▲
by
carterschonwald
5mo ago
this sounds like a nice non prescriptive direction. jose is a pretty cool dude. pre covid we theoretically were gonna poke a defining a memory model of ghc haskell c— / hs / core , but life intervened
68.
▲
by
carterschonwald
5mo ago
oh my, i see what youre saying. at this point youd hope everyone has realized that the best way to keep models more reliable is to force them to stay honest via very very string static typing as a feedback loop. bags of text with hyperlinks
69.
▲
by
carterschonwald
5mo ago
i do kinda appreciate that memetic corruption is now a thing thats real and mechanical. wizardry!
70.
▲
by
carterschonwald
5mo ago
it seems like with some care and disabling sip, that some pretty good work arounds using llm assisted kext hackery would get pretty far
71.
▲
by
carterschonwald
5mo ago
yeah shared “did you this weeks X” is lame, but it was social glue for a long time.
72.
▲
by
carterschonwald
5mo ago
good one
73.
▲
by
carterschonwald
5mo ago
this is literally just “leave a child at the work computer with a real doc open playing office”. otoh it is good to design benchmarks tonground these things. on the flip side if you’re literally just using a bare bones harness on top of a
74.
▲
by
carterschonwald
5mo ago
i mean of course. ive been working on this the past few months and ive a bunch of tech towards this in flight, including some harness forks to layer my ideas in. eg my oh punkin pi test bed on my github.com/cartazio page , theres some
75.
▲
by
carterschonwald
5mo ago
16weeks plus week or so per year of service is pretty good
76.
▲
by
carterschonwald
5mo ago
yup thats mine. :) i actually had some stuff layered into mono pi, and i frankly hit my limit in terms of architecture issues in monopi, omp aka oh my pi is frankly better architectured. if you pared back the fearure set to be minimal, yo
77.
▲
by
carterschonwald
5mo ago
the funny thing is once the llms got mostly good enough in november 2025 for me, it was mind boggling how much it helped me get stuff out of my head with ease. its easier for me to code now, because its like i have a 24/7 insane intern
78.
▲
by
carterschonwald
5mo ago
im def working on benchmarks for how my own general harness improves task performance vs same model in a commodity setup. its hard to do! i will say that my current harness: https://github.com/cartazio/oh-punkin-pi is
79.
▲
by
carterschonwald
5mo ago
i might borrow the skills etc for good ideas sometime. thats a lot of integration surface
80.
▲
by
carterschonwald
5mo ago
check out my pi forks.
81.
▲
by
carterschonwald
5mo ago
more than that, its pretty clear that there is an insane underinvestment in the harness layer. ive been iterating on my own ideas in that area through the lens of increasing reliability. and holy crap is there so much low hanging fruit. i
82.
▲
by
carterschonwald
5mo ago
says the org thats no longer non profit.
83.
▲
by
carterschonwald
6mo ago
while i cant speak regarding arbitrary prompt injections, ive been using a simple approach i add to any llm harness i use, that seems to solve turn or role confusion being remotely viable. i really need to test my toolkit (carterkit) augme
84.
▲
by
carterschonwald
6mo ago
im pretty stoked about the llm harness theyre using. cause I wrote all the code thats not monopi code in that fork! despite it’s paucity of features, the changes i landed in it from my design notes actually have been so smooth in terms of
85.
▲
by
carterschonwald
6mo ago
yeah its honestly full of vibe fixes to vibe hacks with no overarching desig. . some great little empirical observations though!i think the only clever bit relative to my own designs is just tracking time since last cache ht to check ttl.
86.
▲
by
carterschonwald
7mo ago
this is black bar grade great. give us black bar
87.
▲
by
carterschonwald
7mo ago
the main thing ive been hacking on recently is what i consider to be the first next gen llm harness, ive a demonstrator that implements like 40percent of what ive pretty complete specs for on top of mono pi. theres some pretty big differen
88.
▲
by
carterschonwald
7mo ago
oh?! what do they handle well? how do they fail? the 3.5 9b model on my laptop at full fp8 is outlandish in its seeming reasoning capacity, though i haven’t really stress tested it
89.
▲
by
carterschonwald
7mo ago
omg this is so cool. because im writing my own harness and i need some cognitive benchmarks. i have a bunch of harness level infra around llm interactions that seems to help with reasoning, but i dont have a structured way evaluate things
90.
▲
by
carterschonwald
7mo ago
they just released the first small models that i would consider even vaguely articulate for edge inference involving a human. maybe they want to do a mistral and raise a kajillion and work from their home town?
More ›