4 ms·
the 2nd website is not official, just something someone slopped together for some reason.
by sunbum 1mo ago
the 2nd website is not official, just something someone slopped together for some reason.
- Alifatisk 1mo agoI have plenty of these websites, I can’t understand why someone is doing this.
- colesantiago 1mo agoIt is called phishing and grifting. Many people and even software engineers fall for this all the time. Most of these people are from crypto pivoting to AI doing this. AI has made this easier and cheaper and it is going to get a LOT worse. Imagine lots of websites with typosquatting and looking exactly the same as another website, vibe coded and cloned within seconds. The public have no chance.
- Alifatisk 1mo agoWhat is there to phish? These are simple vibe coded websites providing information for a certain topic, nothing else. In this case, that 2nd url is a website with information regarding the new model as well as a broken chat interface to try out.
- colesantiago 1mo agoYou do realise there are hundreds of these types of 'sites'. This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing) Assuming you are technical you are able to discern this, imagine the average person. No chance.
- Alifatisk 1mo ago> You do realise there are hundreds of these types of 'sites'. Yes, and its these sorts of websites I am asking about. > This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing) Again, what is there to phish?
- MrDrMcCoy 1mo agoPhishing implies exploitable data collection. Is that happening here?
- yorwba 1mo agoEven if it weren't slopped together, 65% vs 80% on 10 tasks just isn't a significant difference. For 80% power to distinguish at a significance level of 0.05, you'd need more like 140 samples, if those were the true success probabilities. The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)
- brotchie 1mo agoI'm going steal "slopped together", great quip.