Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
st-at-picnic
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
st-at-picnic
2y ago
A few thoughts -- with some color on 'the why' because we'd love to get your input on how best to get the story and data across. And thoughts you have would be great. So for method: We did NOT force single token responses. Ou
2.
▲
by
st-at-picnic
2y ago
A few thoughts. (Apologies, having trouble editing my response, will post a new message)
3.
▲
by
st-at-picnic
2y ago
Thanks for reading! We'll definitely include our Sonnet results in the next revision. It's worth pointing out that we're comparing accuracy on text responses and not log probability based scoring, which I think is the number
4.
▲
by
st-at-picnic
2y ago
One other interesting comment in there -- the note about how people think the worst records to deal with are the old handwritten notes. But actually, content-wise they tend to be very to-the-point. Clean printouts from EHR software have so
5.
▲
by
st-at-picnic
2y ago
I think that's very true -- and it felt like one of the real opportunities we had in the paper: that we have real production tasks whose results we need to stand behind, and so we can try to explain and show examples of what matters in
6.
▲
by
st-at-picnic
2y ago
Steve here, one of the co-authors. Totally valid on OpenBio. I will say that comparison numbers for this paper were such a challenge, in part because we found that a lot of the LLMs on the Medical LLM leaderboard struggled to follow even sl