3 ms·
Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
by mikert89 29d ago
Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
- x3haloed 29d agoYup. Only subjective taste remains.
- GPerson 29d agoNope that will be commodified in short order.
- tedsanders 29d agoDisagree. Examples: - predict a coinflip: easy to verify, hard to learn - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.
- mikert89 29d agothese just need more compute: - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn but we both know these examples go against the spirit of my point
- tedsanders 29d agoPerhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).
- _superposition_ 29d agoYou bring up an interesting point. Isn't the reward itself subjective in many domains?
- mikert89 29d agoyeah but I think you may be underestimating the amount of capital available for compute. if AGI is possible through some 5 trillion of expenditure on computers, there will be money for it. also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye
- ranyume 29d agoDoesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.
- jdthedisciple 29d agoYes, but not necessarily under tight budget constraints.
- mikert89 28d agotheres no budget constraints for AGI
- tomjen3 28d agoThen I propose the tomjen-1 benchmark: prove the N vs NP problem formally undecidable.
- imtringued 28d agoYou mean any repeatable benchmark will be saturated. The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.
- mikert89 28d agothis is a short term problem, over 20 years benchmark gaming will be a blip