Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sourabh03agr
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
sourabh03agr
3y ago
I have experienced such issues with GPT-4 returning the wrong syntax, especially when asking questions about complex operations in Polars. We are actually running performance monitoring for GPT-3.5/4 and Claude-2 to systematically trac
2.
▲
Tracking Prompt Drift for GPT-4 and Claude
(demo.uptrain.ai)
9 points
by
sourabh03agr
3y ago
|
1 comments
3.
▲
by
sourabh03agr
3y ago
Hello HN! I am happy to share the Monitoring reports we have been running for the past few months to identify regression in popular LLMs like GPT-4-turbo, Claude-2, etc. There have been numerous informal observations about prompt drifts in
4.
▲
Curiosity-Driven Red-Teaming for Large Language Models
(github.com)
1 points
by
sourabh03agr
3y ago
|
0 comments
5.
▲
by
sourabh03agr
3y ago
Congrats on the launch! Do you need Github permissions to answer questions on open-source repos as well?
6.
▲
India asks tech firms to seek approval before releasing 'unreliable' AI tools
(reuters.com)
2 points
by
sourabh03agr
3y ago
|
0 comments
7.
▲
AI could make the four-day workweek inevitable
(bbc.com)
2 points
by
sourabh03agr
3y ago
|
1 comments
8.
▲
Is Gemini a case of a biased tuning dataset or model overfitting?
1 points
by
sourabh03agr
3y ago
|
0 comments
9.
▲
by
sourabh03agr
3y ago
Great read! On being visible on social media efforts, how do you measure the ROI of your efforts, like how much time do you spend vs how much revenue does it bring?
10.
▲
by
sourabh03agr
3y ago
Nice idea! From the perspective of someone who is in need of such a solution, I have explored similar apps before, and one of the biggest issues was that I don't look for events regularly (frequency is once a week or two). As a result,
11.
▲
by
sourabh03agr
3y ago
That's fair, but a lot of use cases require strict information retrieval and don't want the LLM to get creative. I am of the opinion that having an LLM which is always factually correct, is an almost impossible task and we would a
12.
▲
How does one detect hallucinations?
5 points
by
sourabh03agr
3y ago
|
2 comments
13.
▲
Climate change: The 1.5C threshold explained
(bbc.com)
4 points
by
sourabh03agr
3y ago
|
0 comments
14.
▲
by
sourabh03agr
3y ago
Nice work! I see you have an evaluation module - what all are you evaluating for? Primarily Question-answer accuracy via Exact Match?
15.
▲
Review: Tuxedo InfinityBook Pro 14 Linux Laptop
(wired.com)
2 points
by
sourabh03agr
3y ago
|
0 comments
16.
▲
Have we lost faith in technology?
(bbc.com)
2 points
by
sourabh03agr
3y ago
|
2 comments
17.
▲
Ask HN: Including irrelevant documents improve RAG accuracy by 30%?
3 points
by
sourabh03agr
3y ago
|
0 comments
18.
▲
A collection of different LLM jailbreak techniques and research papers
(blog.uptrain.ai)
3 points
by
sourabh03agr
3y ago
|
1 comments
19.
▲
Canadian tar sands pollution is up to 6,300% higher than reported
(theguardian.com)
7 points
by
sourabh03agr
3y ago
|
0 comments
20.
▲
AI news presenters: Can they be trusted?
(bbc.com)
1 points
by
sourabh03agr
3y ago
|
0 comments
21.
▲
by
sourabh03agr
3y ago
Guess Indian is far better than US and Japan when it comes to producing plastic waste. A lot of Indian states have plastic bans, drastically reducing the amount of plastic wasted via packaging or grocery shopping
22.
▲
The Japanese philosophy for a no-waste world
(bbc.com)
45 points
by
sourabh03agr
3y ago
|
43 comments
23.
▲
by
sourabh03agr
3y ago
Love the product, a happy user here!
24.
▲
Can AI Get an Oscar?
(bbc.com)
2 points
by
sourabh03agr
3y ago
|
0 comments
25.
▲
by
sourabh03agr
3y ago
The authors claim that it works well for a wide variety of cases. They have defined a novel categorisation of these evaluations which help guide LLMs to generate relevant assertions
26.
▲
by
sourabh03agr
3y ago
Operationalizing large language models (LLMs) is challenging, mainly due to their unpredictable behaviors and potentially catastrophic failures. A few of the essential requirements for productionizing LLM applications are evaluating them fo
27.
▲
Integrating Spade: Synthesizing Assertions for LLMs into My OSS Project
(github.com)
6 points
by
sourabh03agr
3y ago
|
5 comments
28.
▲
Roblox: A student beat Gucci and Karlie Kloss to app award
(bbc.com)
2 points
by
sourabh03agr
3y ago
|
1 comments
29.
▲
Clocks and watches have shaped our civilisation
(bbc.com)
1 points
by
sourabh03agr
3y ago
|
0 comments
30.
▲
How OpenAI CEO Sam Altman Was Fired by Rival Board Members
(sfstandard.com)
60 points
by
sourabh03agr
3y ago
|
28 comments
More ›