Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jneagu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Show HN: HalluMix – A Benchmark for Real-World LLM Hallucination Detection
(huggingface.co)
4 points
by
jneagu
1y ago
|
0 comments
2.
▲
by
jneagu
2y ago
Hey @echollama! Can you share more about what kind of integrations you are looking for? Also, when you say "enable customers to configure their usage" - what types of configurations would you want? How are you doing all of this to
3.
▲
by
jneagu
2y ago
"trained on licensed content" - this is a bit misleading. The correct framing is in the article: "made with content owner permission." Most (all?) LLMs out there are trained on licensed content, they just try to catch an
4.
▲
Cheating Automatic LLM Benchmarks
(arxiv.org)
3 points
by
jneagu
2y ago
|
0 comments
5.
▲
by
jneagu
2y ago
I can save you a bit of market research and tell you that’s unfortunately not the case yet in the market today. There are a few reasons for it - the main one in my opinion being that it’s hard to measure the value vs cost of switching to an
6.
▲
by
jneagu
2y ago
The challenge with A/B experiments is how you design them to have sufficient power and draw a meaningful conclusion out of them. So, you either need a big % difference between the test and the control, or you need a big number of sampl
7.
▲
The promise and perils of synthetic data
(techcrunch.com)
3 points
by
jneagu
2y ago
|
0 comments
8.
▲
by
jneagu
2y ago
Nice! Overall, frameworks that re-prompt an LLM with feedback or failure modes of the original outputs do very well.
9.
▲
by
jneagu
2y ago
To your latter point - that’s where I think most of the value of LLMs in education is. They can explain code beyond the educational content that’s already available out there. They are pretty decent at finding and explaining code errors. So
10.
▲
by
jneagu
2y ago
I am very curious to see how this is going to impact STEM education. Such a big part of an engineer's education happens informally by asking peers, teachers, and strangers questions. Different groups are more or less likely to do that
11.
▲
by
jneagu
2y ago
Anecdotally, synthetic data can get good if the generation involves a nugget of human labels/feedback that gets scaled up w/ a generative process.
12.
▲
by
jneagu
2y ago
Fair point - I actually had parsed OP's sentence differently. I'll edit my comment. I agree, LLMs performance for coding tasks is super biased in favor of well-represented languages. I think this is what GitHub is trying to solve
13.
▲
by
jneagu
2y ago
Yeah, There was a reference in a paywalled article a year ago ( https://www.theinformation.com/articles/openai-made-an-ai-br... ): "Sutskever's breakthrough allowed OpenAI to overcome limitations on obtaining h
14.
▲
by
jneagu
2y ago
Edit: OP had actually qualified their statement to refer to only underrepresented coding languages. That's 100% true - LLM coding performance is super biased in favor of well-represented languages, esp. in public repos. Interesting - I
15.
▲
Show HN: Smell – A framework for aligning LLM evaluators to human feedback
(quotientai.co)
5 points
by
jneagu
2y ago
|
0 comments
16.
▲
Show HN: Quotient AI – Build better AI products fast
(quotientai.co)
1 points
by
jneagu
2y ago
|
1 comments
17.
▲
by
jneagu
2y ago
Hey HN, We are excited to reveal Quotient, a platform enabling developers to evaluate, improve and ship high-quality AI products through fast, real-world, data-backed experimentation. [1] Every developer understands that for a change in the