Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
yelmahallawy
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
yelmahallawy
7mo ago
Gotchu. Yeah that's pretty quick, awesome thanks!
2.
▲
by
yelmahallawy
7mo ago
Makes sense, simpler=better. Thanks!
3.
▲
by
yelmahallawy
7mo ago
The guy who reviews all of this, is his role in the company fully dedicated to reviewing these eval pipelines?
4.
▲
by
yelmahallawy
7mo ago
Yeah, it feels like an unsolved problem still. I've also seen many teams spend hours on human review in eval pipelines (and this accumulates with each new model that gets released).
5.
▲
by
yelmahallawy
7mo ago
Have some good takeaways / feedback on this? First time I hear about Braintrust (the eval platform) so I'll look into it but I'm curious on your experience with it so far.
6.
▲
by
yelmahallawy
7mo ago
And I think this is a common problem actually — figuring out what to measure and how to measure it – it's not black and white. What I do is have a few dimensions to measure it against (this may or may not fit your use case): relevance,
7.
▲
by
yelmahallawy
7mo ago
I'd love to hear more about what you're working on (if you're open to sharing!). I like to play with knowledge base powered chatbots but what's most useful to me (and probably my primary use case) is coding agents since
8.
▲
by
yelmahallawy
7mo ago
Any takeaways on Promptfoo?
9.
▲
by
yelmahallawy
7mo ago
Any takeaways? Has it been helpful? OpenAI just acquired them so it's probably useful but I was curious to hear more from people who've actually used it.
10.
▲
by
yelmahallawy
7mo ago
Yeah it's a super tedious process and I was hoping that _maybe_ there is a tool out there that can help with this.
11.
▲
by
yelmahallawy
7mo ago
Ah good read, thanks for sharing!
12.
▲
by
yelmahallawy
7mo ago
Yeah that's essentially what I'm looking for. Since now that AI has become such a core part of most businesses, it's pretty critical to use the _best_ models + prompts for whatever your use case is.
13.
▲
by
yelmahallawy
7mo ago
Ah, interesting – yeah only swapping out the model isn't super insightful since models perform differently given different prompts. I'm going to look into GEPA, thanks!
14.
▲
by
yelmahallawy
7mo ago
This is interesting approach, thanks for the insight! If I may ask, _approximately_ how long does it take to test a newly-released model with the current strategy?
15.
▲
by
yelmahallawy
7mo ago
Do you play with the temperature/top k parameters at all?
16.
▲
by
yelmahallawy
7mo ago
This makes sense. I am particularly interested in your invoice processing app example because the accuracy of those outputs can be quantitatively measured from 0%-100% accuracy. I'm curious as to what is _good enough_ and how many iter
17.
▲
by
yelmahallawy
7mo ago
What I've noticed is that it's hard to measure outputs that aren't binary right or wrong, and that's where most human intervention is needed. The biggest examples of this are chatbots and coding agents – basically any ou
18.
▲
Ask HN: How are people doing AI evals these days?
30 points
by
yelmahallawy
7mo ago
|
43 comments
19.
▲
Show HN: An interactive platform for fine-tuning OpenAI models
(modeltunerai.com)
1 points
by
yelmahallawy
3y ago
|
0 comments
20.
▲
by
yelmahallawy
3y ago
Location: Toronto/San Francisco Bay area Remote: Yes Willing to relocate: Yes Technologies: React, Next.js, Node.js, Typescript, PostgreSQL, GraphQL, Resume: https://drive.google.com/file/d/1BTNnLB11txX5_Idboe
21.
▲
Show HN: Google's Unofficial Bard (Palm-2) API
(bardapi.dev)
3 points
by
yelmahallawy
3y ago
|
1 comments
22.
▲
Show HN: Track and bill users' AI token usage through an API
(tiktokenizer.dev)
2 points
by
yelmahallawy
3y ago
|
0 comments
23.
▲
by
yelmahallawy
3y ago
By having users pay for credits. There's 3 different types of credits, prompt credits, chat credits, and image generation credits. Prompt credits are used when generate prompts, chat credits are used when testing ChatGPT on your improv
24.
▲
Show HN: Generate and improve ChatGPT, DALLE, or Midjourney prompts using GPT-4
(repromptify.com)
1 points
by
yelmahallawy
3y ago
|
3 comments
25.
▲
by
yelmahallawy
3y ago
I just built a website that uses GPT-4 to create prompts for ChatGPT, DALLE•2, and Midjourney where you can test out the prompts directly on the website. I'm looking for some feedback on the website + some initial users to validate the
26.
▲
by
yelmahallawy
4y ago
Thank you! I will definitely test and iterate.
27.
▲
by
yelmahallawy
4y ago
How do I position myself to get luckier?
28.
▲
by
yelmahallawy
4y ago
How does one get their projects from 0 to 1 (or even get beta users) if all of their socials have <10 karma. All of my previous HN threads were basically me trying to shill my projects, but no one viewed it or commented (along with my po
29.
▲
by
yelmahallawy
4y ago
They also sometimes say opposite things, so in case you were going to say "just listen to them all", it's a little bit harder than that.
30.
▲
What's better for showing off your SaaS products: Hacker News or ProductHunt?
3 points
by
yelmahallawy
4y ago
|
9 comments
More ›