Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cbg0
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
181.
▲
by
cbg0
5mo ago
They're doubling the five hour limits, but no mention about the weekly limit. So overall it's the same maximum usage, right?
182.
▲
by
cbg0
5mo ago
While I am also a fan of having knobs and switches with that nice mechanical quality, I think we're not really part of a majority. YoY sales dropped "only" 9% for Mercedes passenger cars, so people are not that bothered by th
183.
▲
by
cbg0
5mo ago
The same way they've fought cheaper ICE brands: delivering higher quality materials, a fancy badge and a great driving experience. Currently the Chinese EVs are cheap, but far from Merc levels of refinement.
184.
▲
by
cbg0
5mo ago
The biggest problem isn't the token slot machine refusing to give you the answer, but the fact that multiple refusals can end up flagging your account and getting banned from the service.
185.
▲
by
cbg0
5mo ago
This sounds like a "perfect is the enemy of good" situation. There are certain types of reactors that can reuse uranium to further reduce its half life to around 6000 years so the one million years legal requirement is an unreason
186.
▲
by
cbg0
5mo ago
The real "sleeper" might be https://huggingface.co/ibm-granite/granite-vision-4.1-4b if the benchmarks hold up for such a small model against frontier models for table & semantic k:v extraction.
187.
▲
by
cbg0
5mo ago
So are we saying it's fine that the article is written by an LLM as long as it doesn't have the tell-tale signs of LLMs?
188.
▲
by
cbg0
6mo ago
The quantization for some models can be very detrimental and their quality can drop considerably from the posted benchmarks which are probably at bf16, this is why having considerable RAM can be important.
189.
▲
by
cbg0
6mo ago
Typically if your usage isn't moving it's because you've enabled extra usage and paying with credits.
190.
▲
by
cbg0
6mo ago
You don't need a physical presence to be subject to another country's laws. Disobeying a judicial order would be grounds for issuing a warrant which could easily be expanded to an international warrant for the owners of the platfo
191.
▲
by
cbg0
6mo ago
When it comes to forcing platforms outside of Greece to comply with this, those platforms will just close their service down to Greece. If you want to talk about the concept itself of removing anonymity: on HN the impact would not be huge,
192.
▲
Affordability Still Dominates Americans' Financial Worries
(news.gallup.com)
2 points
by
cbg0
6mo ago
|
1 comments
193.
▲
by
cbg0
6mo ago
Not quite. Large amounts of data going into these models has already been curated, otherwise you would get a tremendous amount of wrong answers for even the most basic questions.
194.
▲
by
cbg0
6mo ago
I think this is more tailored towards enterprise clients that lose money when Github is down, that would probably help with retention.
195.
▲
by
cbg0
6mo ago
Sure there is: contracts, laws and prison time can ensure that doesn't happen.
196.
▲
by
cbg0
6mo ago
> let the community decide Which community are we talking about? The professionals with 10+ years experience using LLMs, the vibe coders that have no experience writing code and everyone in between? If you read some of the online communi
197.
▲
by
cbg0
6mo ago
SWE-bench verified was created in collaboration with OpenAI. It's also an open dataset so prone to contamination, meaning it can be gamed.
198.
▲
by
cbg0
6mo ago
> Who really truly enjoys that and doesn't see it as a chore? This is a whole different discussion, but I just see it as part of the job that I'm getting paid for, I don't need to enjoy it to do it. Functional testing is a
199.
▲
by
cbg0
6mo ago
It should have the same flow as reviewing PRs from humans.
200.
▲
by
cbg0
6mo ago
I was talking about bypassing the ChatGPT safeguards, that's what this bug hunt is about.
201.
▲
by
cbg0
6mo ago
They're probably expecting that it can be done without too much effort so they just want to see all the unique ways people are doing it.
202.
▲
by
cbg0
6mo ago
I've been a fan since the launch of the first Sonnet model and big props for standing up to the government, but you can sure lose that good faith fast when you piss off your paying customers with bad communication, shaky model quality
203.
▲
by
cbg0
6mo ago
The US has over 70 million on Medicare, why would they care about 500K brits?
204.
▲
by
cbg0
6mo ago
The benchmarks say so, but try it out with actual tasks and be the judge.
205.
▲
by
cbg0
6mo ago
Isn't CyberGym an open benchmark so trivial to benchmaxx anyway?
206.
▲
by
cbg0
6mo ago
I downvoted it because it doesn't add anything useful to the conversation, and I don't own any AI stock.
207.
▲
by
cbg0
6mo ago
Doesn't look like it's cheaper, better or uses fewer tokens: https://www.reddit.com/r/Anthropic/comments/1stf6fz/one_week... YMMV, I know.
208.
▲
by
cbg0
6mo ago
If it uses half the tokens to complete a task, then doubling the cost is perfectly fine. But is that actually true?
209.
▲
by
cbg0
6mo ago
For less than 10% bump across the benchmarks? Probably not, but if your employer is paying (which is probably what OAI is counting on) it's all good. It's kind of starting to make sense that they doubled the usage on Pro plans - i
210.
▲
by
cbg0
6mo ago
I often do need in-depth general knowledge in my coding model so that I don't have to explain domain specific logic to it every time and so that it can have some sense of good UX.
More ›