3 ms·
Do you have hard evidence of this assertion?
by dooglius 17d ago
Do you have hard evidence of this assertion?
- simlevesque 17d agoWe can't have hard evidence. It's a SaaS and they own the code and the machine it runs on. So it may be a widespread hallucination. But there's no evidence of that either.
- fragmede 17d agoWe could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.
- marcus_cemes 17d agoYou would need a significant sample size to make any sort of conclusion from such a probabilistic process. Then there's the issue of how you would actually grade/compare.
- ArvidSu 17d agoYou only need to come up with a catchy "SomethingBench" name, post it on reddit/x and now you're an ai sage. Not to disparage the launch/after comparison though, I'd genuinely enjoy a data point like that
- luckydata 17d agosomeone already does that https://aistupidlevel.info/ https://aistupidlevel.info/
- chaimtweiss 17d agoIt's actually a extremely cool site, and fascinating to view the results off the AI bots i use.
- dooglius 17d agoRun a benchmark with a large number of samples, rerun a few days later. Compare results, use statistics to see if there's a statistically significant difference.
- kadoban 16d agoDoing this in a way that doesn't get you noticed or fucked with is potentially going to be quite difficult. If they have a "hey we're being benchmarked" mode, which is not hard to imagine, avoiding tripping it is going to be annoying and difficult to prove.
- dooglius 16d agoYou would want to use your own benchmark, certainly, not something publicly known. But outside of that I don't think "this is a benchmark" is such an easy category to determine -- a benchmark should be similar to a typical scenario anyway.
- kadoban 16d ago> "this is a benchmark" is such an easy category to determine I mean, it's not _that_ hard to determine most likely, and/or it's hard to be sure you didn't get found out by llm-assisted analysis on your traffic. It's not going to be a one-shot request and response it's going to be a whole bunch of them in an artificial way, by nature. And then anything you found is single-use only if you're paranoid because even if they find out later, they have your benchmark now (because you sent it to them to use it even if you don't publish it).
- bradly 17d agoThere is a toot from an Open AI person a couple days ago saying they are "pulling all the levers" because of capacity issues. I have no idea what the heck the person is talking about, but I'm guessing there are consequence for those levers. > "Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very steep growth before. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues."
- nonethewiser 17d agoDepends on the nature of the levers