Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
3d27
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
How to evaluate multi-turn LLM chatbots
(confident-ai.com)
3 points
by
3d27
2y ago
|
0 comments
2.
▲
We wrote a comprehensive guide on LLM security
(confident-ai.com)
1 points
by
3d27
2y ago
|
0 comments
3.
▲
How to generate synthetic data using SOTA data evolution methods
(confident-ai.com)
1 points
by
3d27
2y ago
|
0 comments
4.
▲
How to build your own LLM evaluation framework
(confident-ai.com)
2 points
by
3d27
2y ago
|
0 comments
5.
▲
Overview of All Major LLM Benchmarks
(confident-ai.com)
1 points
by
3d27
3y ago
|
0 comments
6.
▲
by
3d27
3y ago
Checkout this instead: https://github.com/confident-ai/deepeval Also has native ragas implementation but supports all models.
7.
▲
I wrote an article about everything I know about LLM metrics
(confident-ai.com)
2 points
by
3d27
3y ago
|
1 comments
8.
▲
by
3d27
3y ago
This is great. I'm also building an LLM evaluation framework with all these benchmarks integrated in one place so anyone can go benchmark these new models on their local setup in under 10 lines of code. Hope someone finds this useful:
9.
▲
Best practices I learnt from helping health tech enterprise test LLMs
(confident-ai.com)
1 points
by
3d27
3y ago
|
0 comments
10.
▲
Am I too needy? From a data science perspective
(medium.com)
1 points
by
3d27
3y ago
|
1 comments
11.
▲
by
3d27
3y ago
(found this interesting post on medium, this is not my original work)
12.
▲
by
3d27
3y ago
There's a lot more in the evaluation space, including this one: https://github.com/confident-ai/deepeval
13.
▲
Best Practices for Unit Testing RAG Systems in Prod
(confident-ai.com)
4 points
by
3d27
3y ago
|
0 comments
14.
▲
Tried Apple's Vision Pros, would not recommend it
(theverge.com)
2 points
by
3d27
3y ago
|
0 comments
15.
▲
Everything I know about LLM evaluation metrics
(confident-ai.com)
7 points
by
3d27
3y ago
|
0 comments
16.
▲
Google 2024 Layoffs on a rolling-basis
1 points
by
3d27
3y ago
|
0 comments
17.
▲
by
3d27
3y ago
I'm just imagining Jensen Huang laughing in his sleep right now...
18.
▲
Meta Going All in on GenAI
(datacenterdynamics.com)
3 points
by
3d27
3y ago
|
2 comments
19.
▲
by
3d27
3y ago
How did you calculate accuracy and bias?
20.
▲
I used QAG to implement an LLM text summarization evals
(confident-ai.com)
3 points
by
3d27
3y ago
|
0 comments
21.
▲
I found a way to code like Shakespear
(shakespearelang.com)
1 points
by
3d27
3y ago
|
1 comments
22.
▲
I implemented 12+ LLM evaluation metrics so you don't have to
(old.reddit.com)
4 points
by
3d27
3y ago
|
1 comments
23.
▲
AI Makes Commercial Masterpiece [video]
(youtube.com)
2 points
by
3d27
3y ago
|
0 comments
24.
▲
by
3d27
3y ago
The package I built is like a provider for 10+ different evaluation metrics that run both locally on your machine using models from hugging-face but also on the cloud IF you want more functionality. If you want to evaluate a fine-tuned mode
25.
▲
Show HN: I implemented evals metrics for LLMs that runs locally on your machine
(github.com)
22 points
by
3d27
3y ago
|
3 comments
26.
▲
Overcoming the biggest barrier to practical quantum computers
(breakingdefense.com)
1 points
by
3d27
3y ago
|
0 comments
27.
▲
Google's new model is good but the demo's not reproducible in Bard
(boingboing.net)
1 points
by
3d27
3y ago
|
0 comments
28.
▲
What Is RAG? (With Examples)
(confident-ai.com)
1 points
by
3d27
3y ago
|
0 comments
29.
▲
Found this weird programming language
(en.wikipedia.org)
2 points
by
3d27
3y ago
|
0 comments