3 ms·
The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for incon
by jesse_dot_id 13d ago
The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.
AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
- moffkalast 13d agoPetition to rename them to the Office of Weights and Biases, haha.
- vatsachak 13d agoTHIS EXACTLY. The only regulation that we need right now is the model that's on tap
- bradleybuda 13d agoAnthropic terms of service: > 12. General terms > Changes to the Services. Our Services are novel and will change. We may sometimes add or remove features, increase or decrease capacity limits, offer new Services, or stop offering certain Services. > Unless we specifically agree otherwise in a separate agreement with you, we reserve the right to modify, suspend, or discontinue the Services or your access to the Services, in whole or in part, at any time without notice to you. Although we will strive to provide you with reasonable advance notice if we stop offering a Service, there may be urgent situations—such as preventing abuse, responding to legal requirements, or addressing security and operability issues—where providing advance notice is not feasible. We will not be liable for any change to or any suspension or discontinuation of the Services or your access to them. You're not buying a gallon of milk or a pound of flour. You're buying hosted software that the host reserves the right to modify.
- alightsoul 13d agoYou are not buying something and expecting it to be what's on the tin? Aka what the benchmarks show?
- Aurornis 13d agoI doubt that would change the perception. Every model release is followed by accusations of nerfing. There are several projects that repeat benchmarks on published models. None has ever found significant fluctations Here's one example https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/claude-code/ Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work. This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.
- ricardobeat 13d agoThis page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'. If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.
- Aurornis 13d ago> This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'. Click the part at the end that says "View historical performance". They wait to collect more data about a new model before adding it to the overall charts. The overall solution rate continues to climb when new models are considered. > If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own. The y-axis is amplified to make differences look larger than they are. Hover over the dots to see the confidence interval. A 1-2% change means nothing.