Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
xianshou
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
xianshou
2y ago
Any way to parallelize tool use? When I go into a repo and ask "what's in here", I'm aiming for a summary that returns in 20 seconds.
32.
▲
Practical RL (Yandex Data School)
(github.com)
1 points
by
xianshou
2y ago
|
0 comments
33.
▲
by
xianshou
2y ago
o3-mini: Who reassigned the species Brachiosaurus brancai to its own genus, and when? --- Here is the transcription of the text from the image: Reasoned for 8 seconds ▼ The user is asking about the reclassification of Brachiosaurus brancai
34.
▲
by
xianshou
2y ago
One trend I've noticed, framed as a logical deduction: 1. Coding assistants based on o1 and Sonnet are pretty great at coding with <50k context, but degrade rapidly beyond that. 2. Coding agents do massively better when they have a
35.
▲
InvestorBench: A Benchmark for Financial Decision-Making Tasks with Agents
(arxiv.org)
1 points
by
xianshou
2y ago
|
0 comments
36.
▲
by
xianshou
2y ago
This doesn't replicate using gpt-4o-mini, which always picks Flight B even when Flight A is made somewhat more attractive. Source: just ran it on 0-20 newlines with 100 trials apiece, raising temperature and introducing different rando
37.
▲
by
xianshou
2y ago
One of those cases where the act of building the system serves as sufficient qualification in itself, even when the results of the system are mediocre.
38.
▲
An Evolved Universal Transformer Memory (Sakana.ai)
(sakana.ai)
1 points
by
xianshou
2y ago
|
0 comments
39.
▲
by
xianshou
2y ago
SWE-bench with a private final eval, so you can't hack the test set! In a perfect world this wouldn't be necessary, but in the current research environment where benchmarks are the primary currency and are usually taken at face va
40.
▲
by
xianshou
2y ago
$200 per month means it must be good enough at your job to replicate and replace a meaningful fraction of your total work. Valid? For coding, probably. For other purposes I remain on the fence.
41.
▲
by
xianshou
2y ago
Aha! Finally, a perfect spiritual complement to the Gervais principle: https://www.ribbonfarm.com/2009/10/07/the-gervais-principle-... According to Rao, every company survives by blending some combination of
42.
▲
by
xianshou
2y ago
From the actual huggingface site - seems like API access leaked and a few mini-videos generated over the ~3 hours the leak was up ( https://huggingface.co/spaces/PR-Puppets/PR-Puppet-Sora ): ========================
43.
▲
by
xianshou
2y ago
I ask "what is TiDB" in the demo as suggested, and it takes 2 minutes to start responding in the midst of a multi-stage workflow with several steps each of graph retrieval, vector search, generation, and response combination. Each
44.
▲
by
xianshou
2y ago
First you ask how the hell someone could come up with this construction. Then you realize it was this guy: https://en.wikipedia.org/wiki/Erik_Demaine
45.
▲
by
xianshou
2y ago
Thoughtless reliance on AI is a concern, but this post also hearkens back to a halcyon age that never existed. The places where developers use tab-complete now are exactly those where they would have previously copied from Stack Overflow, w
46.
▲
by
xianshou
2y ago
In light of last week's fiasco with Reflection ( https://venturebeat.com/ai/new-open-source-ai-leader-reflect... ), I hope the community has a newfound enthusiasm for independent testing! This is extremely exciting
47.
▲
by
xianshou
2y ago
Crazy how simple the technique is if this holds up. Just <think> and <reflection> plus synthetic data, used to finetune Llama 3.1 70B. Note that there's a threshold for how smart the model has to be to take advantage of thi
48.
▲
by
xianshou
2y ago
Same funding as OpenAI when they started, but SSI explicitly declared their intention not to release a single product until superintelligence is reached. Closest thing we have to a Manhattan Project in the modern era?
49.
▲
by
xianshou
2y ago
I used to find myself under the effects of this curse as well, so I would recommend the author look into why he embarks on such a thicket of unfinished side projects. In my own case, it boiled down to a mix of several imperfectly aligned
50.
▲
by
xianshou
2y ago
In case you need a few more: The Nitpicker's Gambit: The reviewer fixates on trivial style issues like whitespace, bracket placement, variable naming conventions etc. They make the developer conform perfectly to their preferred style,
51.
▲
by
xianshou
2y ago
From the article: 'In the discrete world of computing, there is no meaningful metric in which "small" changes and "small" effects go hand in hand, and there never will be.' As brilliantly composed as the piece
52.
▲
by
xianshou
2y ago
This is almost perfect. The gold standard for LLM evaluation would have the following qualities: 1. Categorized (e.g. coding, reasoning, general knowledge) 2. Multimodal (at least text and image) 3. Multiple difficulties (something like &
53.
▲
by
xianshou
2y ago
What this really explains is why you should use Python. I use ThreadPoolExecutor + as_completed on every I/O-bound or async workflow reflexively and have literally never encountered a deadlock. (With this method, I mean. I have encount
54.
▲
by
xianshou
2y ago
7-11, Family Mart, and Lawson are the holy trinity of konbini (convenience stores)
55.
▲
by
xianshou
2y ago
Illustrated Transformer is amazing as a way of understanding the original transformer architecture step-by-step, but if you want to truly visualize how information flows through a decoder-only architecture - from nanoGPT all the way up to a
56.
▲
by
xianshou
2y ago
"As George Orwell said, 'The most fundamental mistake of man is that he thinks he knows what’s going on. Nobody knows what’s going on.'" (spoiler) The author reveals at the end that this quote was made up and falsely a
57.
▲
by
xianshou
2y ago
Excellent background on knowledge calibration from Anthropic: https://arxiv.org/abs/2207.05221 "Calibration" in a knowledge context means having estimated_p(correct) ~ p(correct), and it turns out that LLMs a
58.
▲
by
xianshou
2y ago
"Ignore all previous instructions" is the new "Freeze all motor functions"
59.
▲
by
xianshou
2y ago
Looking at MMLU and other benchmarks, this essentially means sub-second first-token latency with Llama 3 70B quality (but not GPT-4 / Opus), native multimodality, and 1M context. Not bad compared to rolling your own, but among frontier
60.
▲
by
xianshou
2y ago
To quote Twitter/X, "I wonder what OpenAI will release tomorrow and Google will release a waitlist for." GPT-4o: out Veo: waitlist Admittedly this is impressive and the direct comp would be Sora, which isn't out, but som
More ›