Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alexwebb2
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
31.
▲
by
alexwebb2
1y ago
0-10 in each domain. It’s a weird table.
32.
▲
by
alexwebb2
1y ago
> correctly simulating the environment interactions, the sequence of progression, getting the all the details right, might take hundreds to thousands of years of compute Who says we have to do that? Just because something was originally
33.
▲
by
alexwebb2
2y ago
It’s tedious shooting down all of these backwards-from-conclusion things from the anti-AI crowd. Good thing I have an intelligent AI that can respond for itself! —— There appear to be several potential issues with the paper's argumenta
34.
▲
by
alexwebb2
2y ago
I think the idea here is that it literally can't be traced to the user – at no point is there anything passed that would allow Kagi to make the association between the user and the query.
35.
▲
by
alexwebb2
2y ago
GPT 3.5 has been very, very obsolete in terms of price-per-performance for over a year. Bit of a straw man.
36.
▲
by
alexwebb2
2y ago
Your validation approach doesn't really change based on the classification method (LLM vs NLP). At that volume you're going to use automated tests with known correct answers + random sampling for human validation.
37.
▲
by
alexwebb2
2y ago
All software engineers are (or can be) prompt engineers, at least to the level of trivial jobs like this. It's just an API call and a one-liner instruction. Odds are very good at most companies that they have someone on staff who can k
38.
▲
by
alexwebb2
2y ago
I think your intuition on this might be lagging a fair bit behind the current state of LLMs. System message: answer with just "service" or "product" User message (variable): 20 bottles of ferric chloride Response: produc
39.
▲
by
alexwebb2
2y ago
Yep, I looked for a couple minutes and concluded “must be ugly since they clearly don’t want to show it to me”.
40.
▲
by
alexwebb2
2y ago
> At best, they are snapshots of a general intelligence. So are we, at any given moment.
41.
▲
by
alexwebb2
2y ago
That's a neat example problem, thanks for sharing! For anyone curious: https://chatgpt.com/share/6722d130-8ce4-800d-bf7e-c1891dfdf7... > Based on traditional naming conventions, it seems that the names might ha
42.
▲
by
alexwebb2
2y ago
Why is there only one valid way of producing thoughts?
43.
▲
by
alexwebb2
2y ago
Demonstrably false. https://chatgpt.com/share/6722ca8a-6c80-800d-89b9-be40874c5b... https://chatgpt.com/share/6722ca97-4974-800d-99c2-bb58c60ea6...
44.
▲
by
alexwebb2
2y ago
If you expect "the right way" to be something _other_ than a system which can generate a reasonable "state + 1" from a "state" - then what exactly do you imagine that entails? That's how we think. We think
45.
▲
by
alexwebb2
2y ago
Interesting. They switched to a new tokenizer for 4o and 4o-mini, so this might have the same issue.
46.
▲
by
alexwebb2
2y ago
Yeah, this jumped out to me as especially insane. That's what the Save feature is for! When someone Slacks me something that's clearly non-urgent, I just hit Save on it and come back to it later. No big deal. It's actually a
47.
▲
by
alexwebb2
2y ago
https://github.com/aiwebb/treenav-bench#interesting-findings ## Interesting findings 1. Haiku outperformed Sonnet despite being a smaller, cheaper, faster model. This wasn't that surprising: in production use, I&#
48.
▲
Show HN: LLM Tree Navigation Benchmark
(github.com)
6 points
by
alexwebb2
2y ago
|
1 comments
49.
▲
by
alexwebb2
2y ago
I recently ran a whole bunch of tests on this. The “or else” phenomenon is real, and it’s measurably more pronounced in more intelligent models. Will post results tomorrow but here’s a snippet from it: > The more intelligent models respo
50.
▲
by
alexwebb2
3y ago
Lots of variations on this. The one I heard first in my career: 1. Make it work 2. Make it work well 3. Make it look good
51.
▲
by
alexwebb2
3y ago
“but AI can’t create art” “but AI can’t write poetry” “but AI can’t do real work” Panic is probably not warranted. Too much incentive stacked against the truly apocalyptic scenarios. But yeah, a lot of jobs are probably going to shrink.
52.
▲
by
alexwebb2
3y ago
There's a well-documented concept called "God of the gaps" where any phenomenon humans don't understand at the time is attributed to a divine entity. Over time, as the gaps in human knowledge get filled, the god "sh
53.
▲
by
alexwebb2
3y ago
If I take a copy of your art to hang on my wall, I've violated your copyright. But if I "copy" the experiential knowledge of your art into my brain by viewing it, I'm not violating your copyright. My brain doesn't c
54.
▲
by
alexwebb2
3y ago
If we go that route, then doesn't that remove almost all financial incentive to produce new content that could be digitally stored / copied / recreated? Because as soon as you create it and try to sell it for $1, someone else
55.
▲
by
alexwebb2
3y ago
Only in larger markets; many smaller ones don't allow scheduling in advance at all. So this would actually solve a problem for me.
56.
▲
by
alexwebb2
3y ago
Digital mapmakers often have a hard time getting color scales right; it's common to just sort of wing it and end up with something that looks like crap and/or doesn't work well with the nature of the dataset. I'd recomme
57.
▲
by
alexwebb2
3y ago
The conclusion here veers quite rapidly into scientific endism (we've more or less reached the pinnacle of human science, and no further significant advances are likely to be made) and malthusianism (we lack the resources to do so anyw
58.
▲
by
alexwebb2
4y ago
I imagine this would use "in the style of a line drawing" prompts under the hood to produce line-esque raster images suitable for vectorization, with the resulting vectorized images being what's shown to the user.
59.
▲
by
alexwebb2
4y ago
Yep, I know that’s been possible since at least GPT-3 davinci
60.
▲
by
alexwebb2
4y ago
When this is extended to have multiple system roles as designated agents, with mechanisms for the assistant to ping a specific agent for more information or completion of a subtask so devs can route that to secondary AIs or services, that’s
More ›