Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fatso784
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
fatso784
3y ago
Thanks!
32.
▲
by
fatso784
3y ago
There is a long term vision of supporting fine-tuning through an existing evaluation flow. We originally created this because we were worried about how to evaluate ‘what changed’ between a fine-tuned LLM and its base model. I wonder if Vert
33.
▲
by
fatso784
3y ago
Hey Eric! Thank you! As an aside, we are looking to interview some people who’ve used ChainForge (you see, we are academics who must justify our creations through publications… crazy, I know). Would you or anyone on your team be interested
34.
▲
by
fatso784
3y ago
Thank you for the kind words! Looking at the photo, I think you wouldn’t need the last prompt node there. As far as evaluating functions go, that’s unfortunately a ways off. But, we generally prioritize things based on how many people poste
35.
▲
by
fatso784
3y ago
We don’t view it as a competition here. These other tools are for LLM app building, which is great, but isn’t the focus of ChainForge. Instead, we’re focused on helping devs find the right prompt, and evaluate and inspect LLM outputs. So wh
36.
▲
by
fatso784
3y ago
ChainForge looks similar visually, but is very different in practice. We target evaluating and inspecting LLM outputs, rather than building LLM applications. So, some things are certainly easier to do in CF, while others will certainly be e
37.
▲
Apple’s ML model and dataset introspection API
(apple.github.io)
1 points
by
fatso784
3y ago
|
0 comments
38.
▲
Show HN: ChainForge, a visual tool for prompt engineering and LLM evaluation
(chainforge.ai)
177 points
by
fatso784
3y ago
|
29 comments
39.
▲
by
fatso784
3y ago
This is a pretty cool idea. One comment —the print representation should be entirely akin to a dict, just braces. No DotDict wrapper. That’d be cool.
40.
▲
by
fatso784
3y ago
Thanks for the clarification! Yes, I see now that auto-evals here is more AI agent-ish, than a one-shot approach. Still has the trust issue. For suggestions, one thing I'm curious about is how we can have out-of-the-box benchmark datas
41.
▲
by
fatso784
3y ago
What does 'top 50%' responses mean here, though? You'd need to have a ground truth of how 'good' each score was to calculate that --and if you had ground truth, no need to use an LLM evaluator to begin with. If you
42.
▲
by
fatso784
3y ago
I like the support for Vector DBs and LLaMa-2. I'm curious as to whether and what influences compelled PromptTools, and how it differs from other tools in this space. For context, we've also released a prompt engineering IDE, Chai
43.
▲
Continue multiple conversations simultaneously across multiple LLMs
(github.com)
2 points
by
fatso784
3y ago
|
0 comments
44.
▲
ChainForge now supports chat evaluation
(github.com)
2 points
by
fatso784
3y ago
|
0 comments
45.
▲
by
fatso784
3y ago
No problem! I guess I will make a plug myself --we've been working on a similar 'prompt engineering' tool, ChainForge ( https://github.com/ianarawjo/ChainForge ). It's targeted towards slightly differ
46.
▲
by
fatso784
3y ago
You’re missing the point here. It’s not even getting the LLM’s opinion on evaluating the responses to the prompts (which itself is fraught for some tasks, and benchmarks are known to be limited —even OpenAI admits this, it’s why they made e
47.
▲
by
fatso784
3y ago
This tool doesn’t benchmark based on how a model actually responds to the generated prompts. Instead, it trusts GPT4 to rank prompts simply in terms of how well it imagines they will perform head-to-head. Thus, there’s no way to tell if t
48.
▲
ChainForge: A visual tool for prompt engineering in the browser
(chainforge.ai)
1 points
by
fatso784
3y ago
|
0 comments
49.
▲
Show HN: Evaluate LLMs, right in the browser. Share your experiments as links
(ianarawjo.medium.com)
1 points
by
fatso784
3y ago
|
0 comments
50.
▲
You can now run OpenAI evals in ChainForge
(ianarawjo.medium.com)
3 points
by
fatso784
3y ago
|
0 comments
51.
▲
Pen-Based Computing: Still Looking for the Write App?
(cacm.acm.org)
2 points
by
fatso784
3y ago
|
0 comments
52.
▲
ChainForge: A visual programming environment for prompt engineering
(ianarawjo.medium.com)
2 points
by
fatso784
3y ago
|
0 comments
53.
▲
Show HN: ChainForge, a visual tool for evaluating LLM responses
(github.com)
1 points
by
fatso784
3y ago
|
0 comments
54.
▲
Early demo of ChainForge, a data flow environment for prompt engineering
(youtube.com)
1 points
by
fatso784
3y ago
|
0 comments
55.
▲
by
fatso784
3y ago
Nth time this week someone repackages a call to GPT4 as a way to improve/evaluate LLM outputs. Guys, just stop.
56.
▲
by
fatso784
3y ago
Looks a bit like snakeoil to me. A lot of companies now spinning up simple demos with opaque backends, making huge claims they’ve solved X hard problem for/with AI, then saying “trust us” and “join our waitlist” without hard details or
57.
▲
Shorten a paragraph of text using ChatGPT, from the command line
(github.com)
2 points
by
fatso784
4y ago
|
0 comments
58.
▲
by
fatso784
4y ago
Honestly, I wouldn’t consider these domains “foundational” for undergrads —data structures and algorithms and programming and complexity, but database SQL queries? Maybe the issue is what you’re calling foundational here.
59.
▲
Ask HN: API developers, what do you think of LLMs?
1 points
by
fatso784
4y ago
|
2 comments
60.
▲
Ask HN: Does anyone review code on their iPad/tablet?
9 points
by
fatso784
4y ago
|
10 comments
More ›