Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
hellohello2
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
hellohello2
1mo ago
Interesting idea, I had not considered refusals. I have ran into some as well although rarely. I'm not certain the 28 tasks described would trigger it though, if I understand correctly the security tasks are about avoiding prompt injec
32.
▲
by
hellohello2
1mo ago
Respectfully, the linked website reads exactly like the Claude artifacts I read all day, i.e., it is low effort. I do not mind AI generated writing at all, but I do mind bad writing. Further, you ignored the actual contents of my comments,
33.
▲
by
hellohello2
1mo ago
This whole thing immediately reads as Claude generated, making it hard to take seriously. Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https:/&
34.
▲
by
hellohello2
2mo ago
Clickbait doesn't have to be false, it has to be shocking. Being false is one way of being shocking. "Stop doing X!" -- really now?
35.
▲
by
hellohello2
2mo ago
I agree to some extent about the jargon (Claude has a bigger vocabulary that me, if it knows a useful word I don't I'm fine with learning it), but often times the way information is laid out across sentences just doesn't make
36.
▲
by
hellohello2
2mo ago
Sure, and the parent comment's position is that they dislike it. Its purpose is to advocate against clickbait titles becoming normalized in the scientific community.
37.
▲
by
hellohello2
2mo ago
Yes I can, no problem. First figure out how consciousness works and tell me, then I'll be able to explain if the software is ethical or not. Its simple really.
38.
▲
by
hellohello2
2mo ago
Haven't seen him no ;)
39.
▲
by
hellohello2
2mo ago
Yes, I think it really depends on the task. Clearly we won't be willing to spend 1000x on tokens for our everyday needs. But what about for the police monitoring conversations? I'd very much like my chats to stay encrypted, but th
40.
▲
by
hellohello2
2mo ago
The top comment mentions a ~10^3 overhead. Am I crazy (or already old?) for thinking this is not a very large constant?
41.
▲
by
hellohello2
2mo ago
Have all the skeptics in this thread somehow forgot about Moore's law?
42.
▲
by
hellohello2
2mo ago
One trick is to simply as for a fast answer when talking to a high effort model, when working interactively. Sounds stupid but I do this all the time and it works. Just tell it you are working interactively now and need ultrafast answers wi
43.
▲
by
hellohello2
2mo ago
Yeah I agree, I was just saying I really doubt the author of that blog post was trying to take credit for the ideas they present.
44.
▲
by
hellohello2
2mo ago
You're reading this the wrong way I think, citations aren't given because its obviously a pedagogical article about well established stuff. Much like you wouldn't give citations in a blog post explaining calculus.
45.
▲
by
hellohello2
2mo ago
"The oral system scales poorly. One examiner can only hear so many students in a day." I'm not joking: you could have an AI quiz the students and test their understanding. That way the tests go deeper into testing the underst
46.
▲
by
hellohello2
2mo ago
Yeah I'm not sure what happened but Reddit in particular is beyond saving, and I'm not saying it was ever great. I think we simply don't know how to create online communities that make any sense yet. I'd rather have seco
47.
▲
by
hellohello2
2mo ago
Yes I am, which is why I said that, its much easier to automate over the chaos with a coding agent than with a script. They are still far from doing anything productive on their own but its easy to imagine designing a search space and "
48.
▲
by
hellohello2
2mo ago
Yes but science is well-structures and practically designed about repeatability so its a lot easier to automate than "softer" disciplines.
49.
▲
by
hellohello2
2mo ago
Looks interesting. But does anyone really like the useState useContext etc. React model? Might as well make a better one if you are going to be compiling no?
50.
▲
by
hellohello2
2mo ago
No that's not what I meant, I meant exactly what I said above.
51.
▲
by
hellohello2
2mo ago
Thanks a lot for sharing! I'll try the discussion to unpack context, that sounds interesting. I end up having such discussions in separate chats anyways to understand what's happening so it makes sense to do it upfront.
52.
▲
by
hellohello2
2mo ago
I think you are correct, its so hard to judge the quality of research that we have been using publication record/count as a proxy. Ultimately, having papers being easy to write is good, provided we find a better way of judging quality.
53.
▲
by
hellohello2
2mo ago
My 2 cents: AI has improved writing considerably for non-native english speakers (in particular China). Writing feels more standardized/boring but easier to read overall. I hit fewer papers that are a pain to read. The most problematic
54.
▲
by
hellohello2
2mo ago
Interesting. Framing it as 1x to infinity-x matches my experience too. I've have good success with it reproducing papers with existing code, but not such much with one-shotting new code. Do you have any particular setup for this i.e. s
55.
▲
by
hellohello2
2mo ago
Surprisingly good answer from Claude here... Boringly so.
56.
▲
by
hellohello2
2mo ago
Apologies I misremembered and confused Pro with Plus.
57.
▲
by
hellohello2
2mo ago
Pro is 10$ a month. You get what you pay for lol.
58.
▲
by
hellohello2
2mo ago
Well, that statement wasn't meant to be taken literally. Maybe not the best wording, given the context...
59.
▲
by
hellohello2
2mo ago
"completely meaningless" "the only purpose" There's some kernel of truth to what you are saying, but hyperboles like this just aren't accurate. All statistics lie but its better than being blind... What your
60.
▲
by
hellohello2
3mo ago
I'm confused as well, could someone who knows how TUIs work explain what's the point of React-style diffing in this context? I thought you need to clean and full redraw if anything changs anyways?
More ›