Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eli
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
eli
2mo ago
They're usually leased. So yeah they probably owe Flock new ones.
62.
▲
by
eli
2mo ago
They bill out the ass to lease them
63.
▲
by
eli
2mo ago
I'm not familiar with North Dakota in particular, but at the federal level and in some states ALL works created by the government are exempt from copyright. The public paid for it, so the public owns it.
64.
▲
by
eli
2mo ago
According to a few tasks from my little personal coding benchmark it's very good at coding and kinda bad at web design. (Also excellent at "draw me a picture" one-shot prompts, for whatever that's worth) On a sneaky on
65.
▲
by
eli
2mo ago
To be fair, "people with no skills inflating their own value" is what LinkedIn has always been like. But I guess LLMs are uniquely well positioned for that task.
66.
▲
by
eli
2mo ago
These are the API rates. They don’t train on prompts from the API. In what sense are you the product?
67.
▲
by
eli
2mo ago
I mean, yeah. Alternatively I’ve seen the approach of starter with the cheaper/faster model but give it a tool to handoff or talk to a strong model if it gets stuck.
68.
▲
by
eli
2mo ago
This is the improved iOS parental controls. When they shipped it (and for years afterward) it simply did not work https://www.wsj.com/tech/personal-tech/apples-parental-contr... It really is surprising given App
69.
▲
by
eli
2mo ago
Yup, trying to be really strategic about testing. I didn't end up sticking with it, but I tried requiring test cases to cite a matching clause in the task assignment. But also: these tests are only indicative. Only a human can score a
70.
▲
by
eli
2mo ago
Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.
71.
▲
by
eli
2mo ago
Oh sorry I meant the framework. Though I could see posting a couple of the tasks.
72.
▲
by
eli
2mo ago
I have actually been planning to open source the framework. Only benchmarks I care about are the ones that look like real work I do. So it makes it easy to trawl your own repos looking for benchmark task candidates from real bug fixes or fe
73.
▲
by
eli
2mo ago
In my tests, it averages to much cheaper than Opus 4.8 on real tasks on account of being smarter and more token efficient. I have a benchmark to build a game engine from a set of written instructions. It's a little tricky. Opus 4.8 did
74.
▲
by
eli
2mo ago
It's a new and improved version of an existing model? I don't think it's intentionally befuddling.
75.
▲
by
eli
2mo ago
That's a plausible explanation but I'm not seeing evidence for it. I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.
76.
▲
by
eli
2mo ago
I’m not convinced it wasn’t a PR stunt from the start
77.
▲
by
eli
3mo ago
Right so that's a trade secrets argument. I agree that's probably their best current legal avenue. It isn't related to ToS though.
78.
▲
by
eli
3mo ago
There are a bunch of routers like this eg https://openrouter.ai/openrouter/auto
79.
▲
by
eli
3mo ago
The law was probably fine - that case was just wrongly decided based on the outcome the majority of justices wanted.
80.
▲
by
eli
3mo ago
What do you suppose the damages would be for a ToS violation? The difference between subscription rates and API rates? OpenAI accused Deepseek of misappropriating trade secrets which could have serious penalties but seems like an awfully ha
81.
▲
by
eli
3mo ago
I'm not doubting it's possible to pass such a law, I'm doubting that's it's a practical or worthwhile goal. The terms of service don't even necessarily matter here. OpenAI could cancel your account for almos
82.
▲
by
eli
3mo ago
Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs. But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI sho
83.
▲
by
eli
3mo ago
That's the version that goes directly to the trash.
84.
▲
by
eli
3mo ago
You could just use my.name@gmail.com for the no-alias version. So myname+foo@ works and my.name@ works but myname@ goes directly to the trash.
85.
▲
by
eli
3mo ago
I think a better way to regulate content moderation would be to pass laws that regulate content moderation. Not backdoor regulations through very broad interpretations of product liability. HN surely has some underage users. If they look at
86.
▲
by
eli
3mo ago
These types of lawsuits seem dangerous. But Meta is a pretty awful company and deserves some sort of comeuppance. I worry about what it leads to though. Hard cases make bad laws.
87.
▲
by
eli
3mo ago
I'd suggest first looking into the conditions that enable humans to generate sustained, high quality output.
88.
▲
by
eli
3mo ago
No, you can't just "average" different studies and I'm not sure what "neutral" means in the context of some studies showing a benefit and others not showing a benefit.
89.
▲
by
eli
3mo ago
The placebo effect is not an excuse to allow drug companies to make false claims about the efficacy of the ingredients
90.
▲
by
eli
3mo ago
Obviously there are advantages to not having to do work yourself. But for a benchmark with the goal of picking a model to replace a human on some task? I really think the human should judge which is best. I haven’t gotten very far yet but I
More ›