3 ms·
GPT 4 was evaluated on multiple reasoning benchmarks in the release paper.[1] The problems in TPTP are not really appropriate for a language model and exceed th
by Sunhold 4y ago
GPT 4 was evaluated on multiple reasoning benchmarks in the release paper.[1] The problems in TPTP are not really appropriate for a language model and exceed the capabilities of most humans working without tools. Clearly humans have some reasoning ability even if they cannot solve those problems.
[1] https://openai.com/research/gpt-4 https://openai.com/research/gpt-4