Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
biobootloader
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
biobootloader
1y ago
a new social experience...
2.
▲
LoCoDiff: Natural Long Context Code Bench
(abanteai.github.io)
3 points
by
biobootloader
1y ago
|
0 comments
3.
▲
Rawdog: Command line AI that responds by auto-executing output as scripts
(twitter.com)
4 points
by
biobootloader
3y ago
|
0 comments
4.
▲
OpenAI Models Dominate Structured Code Edit Benchmark
(blog.mentat.ai)
5 points
by
biobootloader
3y ago
|
1 comments
5.
▲
by
biobootloader
3y ago
Hey Paul, I'm a Mentat author. > I also notice that the instructions prompt that mentat uses seems to be inspired by the aider benchmark? Glad to see others adopting similar benchmarking approaches. We were inspired by you to use Ex
6.
▲
by
biobootloader
3y ago
we are working on creating "real world" benchmarks that require a lot of context, and will report when we have results!
7.
▲
by
biobootloader
3y ago
Yes, it uses your local files and edits for you. Let us know what you think! https://github.com/AbanteAI/mentat
8.
▲
GPT-4-Turbo JSON Response Format Quirks
(blog.mentat.ai)
3 points
by
biobootloader
3y ago
|
2 comments
9.
▲
by
biobootloader
3y ago
the idea is that mentats in dune combine human and computer skills to assist others but yeah it is funny
10.
▲
by
biobootloader
3y ago
agreed! but it'd be nice to not have to jump back to the editor to fix a minor crash and then go rerun
11.
▲
by
biobootloader
3y ago
it will sometimes fix other bugs, but yeah the focus is on ones that caused the crash or are obvious
12.
▲
by
biobootloader
3y ago
haha yes it did feel a bit wild just letting this run on my machine! I wouldn't recommend running it on a script that's supposed to delete files
13.
▲
Show HN: Wolverine: Give your Python scripts regenerative healing abilities
(github.com)
43 points
by
biobootloader
3y ago
|
20 comments