Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
maxrmk
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
maxrmk
2y ago
Can't wait for all the spam calls I'll get.
62.
▲
by
maxrmk
2y ago
This reads like something generated by AI. Or at least heavily using it in the writing process.
63.
▲
by
maxrmk
2y ago
I’d consider it a web browser but that’s a vague enough term that I can understand seeing it differently. I’d be disappointed if it became common to block clients like this though. To me this feels like blocking google chrome because you do
64.
▲
by
maxrmk
2y ago
The author has misunderstood when the perplexity user agent applies. Web site owners shouldn’t dictate what browser users can access their site with - whether that’s chrome, firefox, or something totally different like perplexity. When retr
65.
▲
by
maxrmk
2y ago
Thanks for writing this up. It’s one of the best explanations of this problem I’ve seen.
66.
▲
by
maxrmk
2y ago
On one hand: an action virtually guaranteed to get you fired. On the other: $150 I used to work at FB and they have a team that tries to catch employees selling access like this. I can’t imagine risking that for what is essentially an hours
67.
▲
by
maxrmk
2y ago
I'd love to try the demo but there's no way I'm putting my openai key into that site.
68.
▲
by
maxrmk
3y ago
> By the Song dynasty, since all literate people could be assumed to have memorized the text, the order of its characters was used to put documents in sequence in the same way that alphabetical order is used in alphabetic languages. I st
69.
▲
by
maxrmk
3y ago
Are you me? This exact same thing happened to me when I graduated back in 2017. I wonder if was the same year, or if this is a recurring thing the IoT team does.
70.
▲
by
maxrmk
3y ago
Any feedback like this helps -- shoot me an email at max@talc.ai with the name of the topic you saw incorrect labels on. We didn't expect this much traction on the demo, or I'd have built this functionality in!
71.
▲
by
maxrmk
3y ago
We're in a similar space, and their work on FinanceBench is great for the whole community, so I appreciate that. Otherwise there's not much out their about their product so I can't directly compare.
72.
▲
by
maxrmk
3y ago
Good question! We aren't really focusing on this area, but I'm willing to speculate. I'd expect broaded constraints than just substring matching. For example, if the user requests that a certain plot point in the story occur
73.
▲
by
maxrmk
3y ago
Yep, in real use cases the latency for generating questions doesn't really matter. But in the demo I was really worried about it.
74.
▲
by
maxrmk
3y ago
I was worried people would run into this quirk in the demo. We have several 'advanced' question generation strategies. You correctly guessed the one we're using in the demo; forming complex questions by finding another page i
75.
▲
by
maxrmk
3y ago
Good idea! There's no limitation in the generation or grading, but we didn't set up the search to support this. I'll see if it's possible to enable this in the wikipedia search component.
76.
▲
by
maxrmk
3y ago
who tests the testing tool? Thanks though -- and let us know if you hit any issues while playing around with the demo!
77.
▲
by
maxrmk
3y ago
Thanks! We aren't hiring right now, but if you shoot me an email at max@talc.ai I'll follow up in a few months.
78.
▲
by
maxrmk
3y ago
Yeah, someone is going to build this. We considered quizzing the user on the topic instead of chatgpt for our demo. It's a lot of fun to test your knowledge on any topic, but it was a worse demo because it was way less related to our c
79.
▲
by
maxrmk
3y ago
Thanks! I've been following promptfoo, so I'm glad to see you here. In addition to automatic evals I think every engineer and PM using LLMs should be looking at as many real responses as they can _every day_, and promptfoo is a gr
80.
▲
by
maxrmk
3y ago
Thanks for flagging!
81.
▲
by
maxrmk
3y ago
Will take a look, thanks!
82.
▲
by
maxrmk
3y ago
Ah! Think of this more like software testing that goes in CI/CD rather than an ML test or validation set. We're providing this testing for applications built on top of language models. For example if you're a SWE working on b
83.
▲
by
maxrmk
3y ago
Totally - certain types of failures are much harder to test than others. We have a couple of different test generation strategies. As you can see in the demo and examples, the most basic one is "ask about a fact". Two of our other
84.
▲
by
maxrmk
3y ago
You're exactly right about the chevy tahoe reference. I wasn't sure if anyone would get it. I liked that post a lot, because as much as I think LLMs are going to be useful they have limitations that we haven't solved yet.
85.
▲
by
maxrmk
3y ago
Great questions here! > How do you rate the correctness? Some complex LLM answers seemed to be correct but not in as much detail as the expected answer. We support two different modes: a strict pass/fail where an answer has to have
86.
▲
Launch HN: Talc AI (YC S23) – Test Sets for AI
132 points
by
maxrmk
3y ago
|
47 comments
87.
▲
by
maxrmk
3y ago
From the press release: "Before Inflection-2 is released on Pi, it will undergo a series of alignment steps to become a helpful and safe personal AI." I wonder how the post-alignment will perform compared to Claude-2 (which is pre
88.
▲
OpenAI gets a C+ in high school English
(talcai.substack.com)
2 points
by
maxrmk
3y ago
|
1 comments
89.
▲
by
maxrmk
3y ago
Ahh good suggestion, I should clarify this in the article. I tried to compensate with volume -- I used a set of 200 questions for the testing. I was using temperature 0, so I'd get the same answer if I ran a single question multiple ti
90.
▲
by
maxrmk
3y ago
I've always been skeptical of benchmarking because of the memorization problem. I recently made up my own (simple) date reasoning benchmark to test this, and found that GPT-4 Turbo actually outperformed GPT-4: https://open.s
More ›