Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jipster
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
jipster
3y ago
R2R uses deepeval for their evaluation :) https://github.com/confident-ai/deepeval
2.
▲
by
jipster
3y ago
Hey, cofounder here. As far as I know, LangSmith focuses on observability (token usage, chain of thought insights, LLM cost), is vendor locked into LangChain, and uses LLMs to evaluate themselves. Here are the problems with what LangSmith o
3.
▲
by
jipster
3y ago
Hey HN! We've built this platform that allows you to evaluate how well your LLM implementation is performing, despite that be using open source tools such as LangChain, lLamaIndex, or even your own internal framework. The idea is you w
4.
▲
Be confident about your LLM stack
(confident-ai.com)
1 points
by
jipster
3y ago
|
0 comments
5.
▲
Evaluate LLMs Rigorously
(twilix.io)
1 points
by
jipster
3y ago
|
0 comments
6.
▲
How we evaluated LLMs in production
(old.reddit.com)
2 points
by
jipster
3y ago
|
0 comments
7.
▲
(Discussion) What method are yall using to evaluate LLM outputs?
1 points
by
jipster
3y ago
|
2 comments
8.
▲
(Discussion) What evaluation methods is best for LLMs?
1 points
by
jipster
3y ago
|
0 comments