Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tylermarques
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
tylermarques
10d ago
Yes they definitely do - not claiming it's a perfect solution, but as a V1 product it shows a lot of promise.
2.
▲
by
tylermarques
11d ago
We had early access and found it to be pretty useful. Having a second form of verification, where you can ask multiple questions (in the form of Nouls) raised our confidence in the outputs of other models. [0] IMHO This type of model works
3.
▲
by
tylermarques
1mo ago
I've had a lot of success combating this by adding "All summaries need to adhere to ASD-STE100 Simplified Technical English standards" [0] which I discovered from another HN thread [1] [0] https://www.asd-ste100.or
4.
▲
by
tylermarques
4mo ago
There is some evidence.[1] The best reviewer is a different model with fresh context, worst is same model with same context. 1. https://arxiv.org/pdf/2603.04582
5.
▲
by
tylermarques
8mo ago
In the same vein, we recently released a version v0.1 of our humor benchmark. [1] We use human answers from a cards against humanity style game call Bad Cards [2] as ground truth for what is funny. The models get to choose a card from a han
6.
▲
IPv6 Based Canvas
(canvas.openbased.org)
91 points
by
tylermarques
1y ago
|
25 comments
7.
▲
by
tylermarques
1y ago
Sorry about that! We've toned down the music a bit, trying to put more emphasis on the narrator. Thanks for the feedback :) We're hoping to create more like this! :)
8.
▲
Show HN: AI Playing Diplomacy
(twitch.tv)
2 points
by
tylermarques
1y ago
|
2 comments