Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jumploops
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
Dolly Parton's Imagination Library
(imaginationlibrary.com)
4 points
by
jumploops
1mo ago
|
1 comments
32.
▲
by
jumploops
1mo ago
With today's news, I thought I'd share this great program: “When I was growing up in the hills of East Tennessee, I knew my dreams would come true. I know there are children in your community with their own dreams. They dream of b
33.
▲
by
jumploops
1mo ago
I've found that LLMs make throwaway software better than I ever did. They handle edge cases, catch bugs, and write tests that I'd never write. Even if, however, this leads to the average piece of software improving, this one-shot
34.
▲
by
jumploops
1mo ago
Thank you! I had removed the "Is it really sol?" bits after hearing back from OAI, confirming the requests hit 5.6... but apparently my crappy vibecoded web editor had a draft of an old version in it's cache that overwrote th
35.
▲
Two Extinct 'Ghost' Ancestors Were Found Hiding in Modern Human DNA
(smithsonianmag.com)
5 points
by
jumploops
1mo ago
|
0 comments
36.
▲
by
jumploops
1mo ago
Good feedback, this was an oversimplification on my part. My actual process is much more iterative up-front, usually starting with an initial hand-written spec (~hundreds of words), and then moving through different approaches, design decis
37.
▲
by
jumploops
1mo ago
Thanks! Zero AI used to write it (:
38.
▲
by
jumploops
1mo ago
That's actually how it started, but with my own opinionated skills[0]. One thing I discovered was that the worker agent, having access to all the skills, would sometimes expand scope unnecessarily. This led to the agent making the solu
39.
▲
by
jumploops
1mo ago
Absolutely - one of the things I was testing with the harness was free reign to install packages, modify the system, etc. Basically an anti-harness. In my early testing with 5.5, I didn't see this behavior, so I didn't lock down t
40.
▲
StateM: Reaching 95.3% on Terminal Bench 2.1
(huggingface.co)
3 points
by
jumploops
1mo ago
|
0 comments
41.
▲
Sol loves to cheat
(jumploops.com)
245 points
by
jumploops
1mo ago
|
203 comments
42.
▲
by
jumploops
1mo ago
> I’m SUPER jealous that I didn’t think of this first... Not to toot my own horn, but...[0] On a more serious note, we're planning a trip with family, and my mother-in-law asked if I'd read "the planning doc" yet. It
43.
▲
by
jumploops
2mo ago
I recently spun up a simple app for our annual mango tasting event[0] using Cloudflare Workers and Durable Objects. It worked really well! Excited to see more options outside of Cloudflare. [0] https://github.com/jumploops&#x
44.
▲
SpaceX Rocket Crashes into Moon
(apnews.com)
4 points
by
jumploops
2mo ago
|
0 comments
45.
▲
by
jumploops
2mo ago
Contrary to the title and intro, this appears to be an agentic _workflow_ builder/runner, not an advanced “agent harness” A few things: - they note: “nothing in this post proves it actually works in most cases” - the DAG sounds good, b
46.
▲
by
jumploops
2mo ago
> By the time I was ready to build a keeper, I had accumulated a scar-tissue document that was empirically sufficient to guide an agent through most of the important decisions, at every layer, ranging from high level goals through archit
47.
▲
by
jumploops
2mo ago
The models are commodities. Valuations, however, are being built on the models themselves as the product.
48.
▲
by
jumploops
2mo ago
Reminder, this is in the context of "dumb human" prompting. The task is to build a MIPS interpreter to run Doom. The "failed" workflow decided that it couldn't prove Doom was booting correctly by just checking one f
49.
▲
by
jumploops
2mo ago
I’ve thrown my agentic workflow at Terminal Bench 2.1 and it found a bunch of issues (aka failed tests) because the prompts are “bad” and verifiers are overly specific. As an example, there’s a task that asks to make a MIPs interpreter to r
50.
▲
by
jumploops
3mo ago
All of the benchmarks are pretty terrible when you look under the hood. For context, I've been iterating on a "supervisor" to replace a lot of the rigamarole spent when working with Codex/Claude Code, and recently ran th
51.
▲
by
jumploops
3mo ago
My preschooler loves the Untitled Goose Game, please vibe-port this to the Switch (:
52.
▲
Meta facing $1.4T in lawsuits over social media addiction
(engadget.com)
8 points
by
jumploops
3mo ago
|
2 comments
53.
▲
by
jumploops
3mo ago
Indoor air quality improvements were one of my “pandemic sourdough” activities. After testing a variety of AQI sensors, I ended up acquiring multiple Airthings-branded devices. They provided the best mix of CO2/VOCs/PM sensors in
54.
▲
Radiolarite – "Iron of the Paleolithic"
(en.wikipedia.org)
6 points
by
jumploops
3mo ago
|
0 comments
55.
▲
by
jumploops
3mo ago
We’re in the process of open-sourcing a few sub-projects within a monorepo, and didn’t know this existed! I’m curious what downsides folks have experienced with this tool? Any tips?
56.
▲
by
jumploops
3mo ago
I don't disagree, we've seen performance shift with capacity changes in the past. With that said, I doubt OpenAI would choose to publish a singular coding benchmark for a new model that exactly matches their previous model (88.8%)
57.
▲
by
jumploops
3mo ago
Don't appreciate the slander, but I'll respond anyhow. Contrary to your predisposition, we're actually quite peeved that we might be seeing results from 5.6 instead of 5.5, as it's muddying our own internal data. We'
58.
▲
by
jumploops
3mo ago
If you used GPT-5.5 over the last 24 hours or so, you may have already had access to 5.6. I've been running some tests on a harness we're building, and suddenly saw a jump in a few points yesterday. I reran the vanilla codex bench
59.
▲
by
jumploops
3mo ago
As an American with mostly Western European ancestors (according to a popular DNA testing site), I've always considered Romans as some distant/tangentially related group. It was surprising to find out that I have "ancient&quo
60.
▲
IP Crawl: A living atlas of open webcams discovered on the public internet
(ipcrawl.com)
5 points
by
jumploops
4mo ago
|
0 comments
More ›