Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fudoshin2596
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
LLM workflows: Why not just run an experiment?
(docs.parea.ai)
2 points
by
fudoshin2596
2y ago
|
1 comments
2.
▲
by
fudoshin2596
2y ago
In so many channels I hear the question, "anyone know if this works" or "what's the impact of X" and I always think to myself, why not just run an experiment? Figured i'd outline a quick workflow that I think s
3.
▲
by
fudoshin2596
3y ago
No fine tuning. Looks like he's do raw model capabilities with simple prompt. repo: https://github.com/parea-ai/tool-use-benchmark
4.
▲
by
fudoshin2596
3y ago
Yea, new tool calling now in public beta is way better. And avoids all of that XML they had before.
5.
▲
LLM evaluation metrics for RAG, chatbots and summarization
(docs.parea.ai)
3 points
by
fudoshin2596
3y ago
|
0 comments
6.
▲
by
fudoshin2596
3y ago
The cross region stats are actually quite helpful, this is cool!
7.
▲
by
fudoshin2596
3y ago
Interesting. I want to see the GPT 4 results also! But the impact of system message or not was an interesting discovery. Are system messages becoming more or less supported with the new models?
8.
▲
by
fudoshin2596
3y ago
Productionizing LLMs is a Multi headed chimera. From getting the vector db retrieval in order, to proper classification, monitoring, and reliable generations. Still searching for a solution to many of these problems, but https://