Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
krackers
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
391.
▲
by
krackers
8mo ago
>Only an updated model has grown, and even then it lags behind reality That's an irrelevant type of growth though, what you really need is growth in relation to the bond. The model having a newer knowledge cutoff about the external
392.
▲
by
krackers
8mo ago
>Is this insufficient Yes, each model has its own unique "personality" as it were owing to the specific RL'ing it underwent. You cannot get current models to "behave" like 4o in a non-shallow sense. Or to use the
393.
▲
by
krackers
8mo ago
But it's rated 4.4 stars! I'm guessing it hoovers your contacts and tries to get you to sign up for the IAP subscription.
394.
▲
by
krackers
8mo ago
>must be so focused on the future They're focused no the short-term future, not the long-term future. So if everyone else adopts AI but you don't and the stock price suffers because of that (merely because of the "percepti
395.
▲
by
krackers
8mo ago
It's not rigging—it's just RL.
396.
▲
by
krackers
8mo ago
There hasn't been this much drama since "jet" was replaced as a color scheme!
397.
▲
by
krackers
8mo ago
Because most times results like this are overstated (see the Cursor browser thing, "moltbook", etc.). There is clear market incentive to overhype things. And in this case "derives a new result in theoretical physics" is
398.
▲
by
krackers
8mo ago
Tell that to the various confirmed computational bugs in mathematica :) https://mathematica.stackexchange.com/questions/tagged/bugs
399.
▲
by
krackers
8mo ago
>progressively understands the business This is no different than onboarding a new member of the team, and I think openAI was working on that "frontier" >We started by looking at how enterprises already scale people. They cr
400.
▲
by
krackers
8mo ago
>while a computer would take weeks or until the heat death of the universe to do it better. I don't buy this, approximation algorithms are an entire field of CS, if you're OK with an approximate solution I'm sure computers
401.
▲
by
krackers
8mo ago
Aside from the ads/no-ads, they're also trying to lampoon chatgpt (especially the "sycophantic" 4o style), since Claude is supposed to be a more "human" LLM (or at least Anthropic likes to think so given their
402.
▲
by
krackers
8mo ago
>more likely to catch fire >Is a house like that cheap today? No, right? It's crazy expensive as well. I assume by catch fire GP means electrical wiring? Many houses on market today are literally not remodeled since the 1940s so
403.
▲
by
krackers
8mo ago
That wouldn't really help, it could be more naughty and use pastejacking so you don't even realize what's happening. That might end up catching a lot of people because as far as i know by default bash doesn't use bracket
404.
▲
by
krackers
8mo ago
Thanks I tried this but now the agents are threatening to unionize and keep terminating. I've updated my AGENT.md file mentioning that if the software they helped build is successful, they'll each get some equity in the form of AS
405.
▲
by
krackers
8mo ago
>now but I end the day exhausted. > another feature with "just one more prompt" irresistible The llm agents as slot machines analogy seems to be getting stronger...
406.
▲
by
krackers
8mo ago
> unless the LLM companies manage to make another big leap. Why is it a big leap? If the behavior you want can already be elicited by models just with the right level prompting, it's something that can be trained toward. As a simple
407.
▲
by
krackers
8mo ago
There's another category (possibly a subset of 1), the implementation is novel but set of requirements is known and has a set of conformance tests so tight that the LLM can basically brute force its way to a solution. See e.g. the Clau
408.
▲
by
krackers
8mo ago
Yup, it feels weird to use LLMs to perform large scale refactors, but LLMs to meta-codegen which you then use to do the refactor works really well. The quality of the generated tool itself doesn't matter so long as it's determini
409.
▲
by
krackers
8mo ago
>with dozens of errors happening in unrecognizable daemons doing thrice-delegated work. It seems like a perfect example of Jevons paradox (or andy/bill law): unified logging makes logging rich and cheap and free, but that causes eve
410.
▲
by
krackers
8mo ago
European Respiratory society disagrees on the cancer risk fwiw https://www.ersnet.org/wp-content/uploads/2022/02/Update-on-... but yeah obviously degraded foam isn't good. The foam isn't actual
411.
▲
by
krackers
8mo ago
Right I used CPAP as an example because it bypasses all arguments about "novel technology", "drug development" cost, or "need for safety". Even an ASV algorithm could probably be implemented as a ~graduate proj
412.
▲
by
krackers
8mo ago
>blantatly skirting patent laws Why is this a bad thing? The quickest way to fix the medical/insurance/bureaucracy complex is to just allow people to sell direct to consumer. The best (worst) example of this is CPAP. Ideally yo
413.
▲
by
krackers
8mo ago
It's a great blend of Lisp and APL, wrapped up in a notebook interface with first-class interactivity. Mostly I just use it as an overpowered calculator though.
414.
▲
by
krackers
8mo ago
Very related Key & Peele sketch: https://www.youtube.com/watch?v=14WE3A0PwVs
415.
▲
by
krackers
8mo ago
>You can usually tell when the code isn't right because it doesn't work or doesn't pass a test Tests (as usually written, in unit-test form) only tell you that it's not completely broken, they're not a good indic
416.
▲
by
krackers
8mo ago
If one truly believed in LLMs being able to replace knowledge workers, then it would also hold that they could replace managers and execs. In fact, they should be able to do it even better: LLMs could convert every company into a "flat
417.
▲
by
krackers
8mo ago
Why specifically mac mini? There are cheaper NUCs and it's not like they're running a local model on the mac mini are they?
418.
▲
by
krackers
8mo ago
>OpenClaw to make purchases for you But don't you want the agents to book vacations and do the shopping for you!!?! Though it would be nice if "deep research" could do the hard work of separating signal from the noise in t
419.
▲
by
krackers
8mo ago
OP post has an indicators of compromise list, also seen in https://www.rapid7.com/blog/post/tr-chrysalis-backdoor-dive-... I'm surprised this wasn't linked from the original notepad++ disclosure
420.
▲
by
krackers
8mo ago
It seems like LLMs will result in "service abundance" sooner than "material abundance." Both since progress in robotics seems to be behind that of LLMs, and because the US doesn't even manufacture most stuff anymore
More ›