Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alach11
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
alach11
6mo ago
On my private internal oil and gas benchmark, I found a counterintuitive result. Opus 4.7 scores 80%, outperforming Opus 4.6 (64%) and GPT-5.4 (76%). But it's the cheapest of the three models by 2x. This is mainly driven by reduced rea
32.
▲
by
alach11
8mo ago
A significant part of Anthropic's cachet as an employer is the ethical stance they profess to take. This is no doubt a tough spot to be in, but it's hard to see Dario making any other decision here. What I don't understand is
33.
▲
by
alach11
9mo ago
I believe Nat Friedman said "pessimists sound smart, optimists make money." It's certainly much easier to give a snarky/negative take and shoot an idea down than think creatively about how to make it work. Also, negative
34.
▲
by
alach11
9mo ago
I have to imagine governments are closely monitoring prediction markets as part of their intelligence apparatus. But then you just add another layer of subterfuge. Imagine a D-Day prediction market... "Will the Allies Land in Normandy,
35.
▲
by
alach11
10mo ago
I really wish these models were available via AWS or Azure. I understand strategically that this might not make sense for Google, but at a non-software-focused F500 company it would sure make it a lot easier to use Gemini.
36.
▲
by
alach11
10mo ago
Usually the first day or two are readily solvable in Excel with just regular spreadsheet formulas.
37.
▲
by
alach11
11mo ago
This is the biggest news of the announcement. Prior Opus models were strong, but the cost was a big limiter of usage. This price point still makes it a "premium" option, but isn't prohibitive. Also increasingly it's beco
38.
▲
Evals drive the next chapter in AI for businesses
(openai.com)
2 points
by
alach11
11mo ago
|
0 comments
39.
▲
by
alach11
11mo ago
Just curious - are you using Open WebUI or Librechat as a local frontend or are all your workflows just calling the models directly without UI?
40.
▲
by
alach11
11mo ago
This is a really impressive release. It's probably the biggest lead we've seen from a model since the release of GPT-4. Seems likely that OpenAI rushed out GPT-5.1 to beat the Gemini 3 release, knowing that their model would under
41.
▲
by
alach11
11mo ago
"Quantity has a quality all its own". It's categorically different to be able to do harm cheaply at scale vs. doing it at great cost/effort.
42.
▲
by
alach11
11mo ago
Can you cite your source for inference being at a loss? This disagrees with most of what I've read.
43.
▲
by
alach11
11mo ago
This is almost certainly the issue. It's very unintuitive for users, but LLMs behave much better when you clear the context often. I run /clear every third message or so with Claude Code to avoid context rot. Anthropic describes t
44.
▲
by
alach11
1y ago
Computer use is the most important AI benchmark to watch if you're trying to forecast labor-market impact. You're right, there are much more effective ways for ML/AI systems to accomplish tasks on the computer. But they all h
45.
▲
by
alach11
1y ago
I'm really interested in the progress on computer use. These are the benchmarks to watch if you want to forecast economic disruption, IMO. Mastery of computer use takes us out of the paradigm of task-specific integrations with AI to a
46.
▲
by
alach11
1y ago
Thus far I didn't have to worry about ChatGPT having bad incentives when giving me advice on product purchases. Now that "Merchants pay a small fee on completed purchases", will the model steer me towards ACP-supported retail
47.
▲
Advice for a young investigator in the last days of the Anthropocene
(docs.google.com)
3 points
by
alach11
1y ago
|
1 comments
48.
▲
by
alach11
1y ago
> using AI is clearly doing more harm than good How do you know this? Wouldn't we expect the benefits of AI in the legal industry to be way less likely to make the front page of HN?
49.
▲
by
alach11
1y ago
I can't fathom why Microsoft has such limited functionality in Copilot for Excel. Some of the ideas in projects like this (and similar) seem so obvious. Are they just being conservative with token cost?
50.
▲
On It, Boss
(github.com)
3 points
by
alach11
1y ago
|
1 comments
51.
▲
by
alach11
1y ago
I'm going to make an unpopular suggestion. Have you considered using a service that will print and ship to you, like CraftCloud? Depending on volume, your total cost would likely be lower. I know you mentioned privacy concerns so this
52.
▲
by
alach11
1y ago
> Florida did drug testing as a condition for welfare benefits... and it cost more than they saved It's more complicated than that. Of the 6352 people who applied for TANF, 2306 dropped out during the process. Then of the 4046 TANF
53.
▲
by
alach11
1y ago
Same with Firefox+Windows 11. I guess they really only care about Chrome...
54.
▲
by
alach11
1y ago
I think you're on the right track here. Most technology pilots fail. As long as risk/investment is managed appropriately, this is healthy. This seems to follow from Surgeon's Law... 90% of everything is crap [0]. [0] https:&
55.
▲
by
alach11
1y ago
OpenAI did it a few weeks earlier when they released text-embedding-3-large, right?
56.
▲
by
alach11
1y ago
Sure. But there are plenty of ways to achieve the same outcome without wasting 8 hours of time for every employee. And once you scale this across all the aspects of company policy/culture you want push, mandatory training classes becom
57.
▲
by
alach11
1y ago
> Everyone had to take an online class (about 8 hours) about effective meetings As soon as I read this line I grimaced. This is a clear sign of an organization that doesn't respect peoples' time. The class should be an email (a
58.
▲
Zuckerberg Expanding His Hawaii Compound. Part of It Sits Atop a Burial Ground
(wired.com)
18 points
by
alach11
1y ago
|
6 comments
59.
▲
by
alach11
1y ago
> it will likely generalize to all kinds of reasoning problems, not just mathematical proofs Big if true. Setting up an RL loop for training on math problems seems significantly easier than many other reasoning domains. Much easier to ve
60.
▲
by
alach11
1y ago
It's very hard for me to imagine the current level of agents serving a useful purpose in my personal life. If I ask this to plan a date night with my wife this weekend, it needs to consult my calendar to pick the best night, pick a bar
More ›