5 ms·
The various deep research products don't work well for me. For example I asked these tools yesterday, "How many unique NFL players were on the roster for at le
by CSMastermind 1y ago
The various deep research products don't work well for me. For example I asked these tools yesterday, "How many unique NFL players were on the roster for at least one regular season game during the 2024 season? I'd like the specific number not a general estimate."
I as a human know how to find this information. The game day rosters for many NFL teams are available on many sites. It would be tedious but possible for me to find this number. It might take an hour of my time.
But despite this being a relatively easy research task all of the deep research tools I tried (OpenAI, Google, and Perplexity) completely failed and just gave me a general estimate.
Based on this article I tried that search just using o3 without deep research and it still failed miserably.
- simonw 1y agoThat is an excellent prompt to tuck away in your back pocket and try again future iterations of this technology. It's going to be an interesting milestone when or if any of these systems get good enough at comprehensive research to provide a correct answer.
- minraws 1y agoIf you keep the prompt the same at some point the data will appear in training set and we might have answer. So even though today it might be a good check it might not remain as such a good benchmark. I think we need a way to keep updating prompts without increasing complexity in someway to properly verify model improvements. ARC Deep Research anyone?
- ljsprague 1y agoWouldn't somebody need to answer the question below? Or do you mean the discussion of its weakness might somehow make it stronger the next time it's trained?
- minraws 1y agoI think it can be both, what happens if discussing weakness provides more relavent links for the question and help the model that is trained scraped web data to learn somehow. I am not sure if the model will need the exact answer or just the backlinks to site where they can find them is enough. Maybe just documenting how to do it could do the job as well...
- red_trumpet 1y agoWell, to test research capabilities, one could just adopt the year (2024->2025) in the prompt.
- minraws 1y agoI am not sure what happens if some site keeps tracking these metrics and that manages to find its way into the training data. There are some NBA fan sites that do keep track of some of these tournament level final metrics.
- qingcharles 1y agoI had o3 "cheat" yesterday. I tried to demo a Deep Research task to a friend, but o3 managed to find the answer immediately in a Reddit comment I'd made after trying out the same problem previously. I was still impressed though!
- danielmarkbruce 1y agoThis is just a bad match to the capabilities. What you are actually looking for is analysis, similar in nature to what a data scientist may do. The deep research capabilities are much better suited to more qualitative research / aggregation.
- deleted 1y ago[deleted]
- pton_xd 1y ago> The deep research capabilities are much better suited to more qualitative research / aggregation. Unfortunately sentiment analysis like "Tell me how you feel about how many players the NFL has" is just way less useful than: "Tell me how many players the NFL has."
- lucyjojo 1y agoFirst person that makes a good exact aggregation AI will make so much money... Precise aggregation is what so many juniors do in so many fields of work it's not even funny...
- johnnyanmac 1y agoIf AI Can't look up and read a chart, why would I trust it with any real aggregation?
- netghost 1y agoBecause AI is weird and does some things really well, and some things poorly. The terrible/exciting/weird part is figuring out which is which.
- deleted 1y ago[deleted]
- oytis 1y agoSo it's not doing well in things that we can verify/measure, but sure it's doing much better in things we can't measure - except we can't measure them, so we have no idea about how well it is doing actually. The most impressive feature of LLMs stays its ability to impress.
- neom 1y agoIs it accurate that there are 544 rosters? If so, even at 2 minutes a roster isn't that days of work, even if you coded something? How would you go about completing this task in 1 hour as a human? (also chatgpt 4.1 gave me 2,503 and it said it used the NFL 2024 fact book)
- dghlsakjg 1y agoIf the rosters are in some sort of pretty easily parsed or scrapable format from the nfl, as sports stats typically are, this is just a matter of finding every unique name. This is something that I imagine would take less than an hour or two for a very beginner coder, and maybe a second or two for the code to actually run
- krainboltgreene 1y agoFYI for readers: All the major leagues have a stats API, most are public, some are public and "undocumented" with tons of documentation by the community. It's quite a feat!
- CSMastermind 1y ago544 rosters but half as many games (because the teams play each other). Technically I can probably do it in about 10 minutes because I've worked with these kind of stats before and know about packages that will get you this basically instantly (https://pypi.org/project/nfl-data-py/ https://pypi.org/project/nfl-data-py/). It's exactly 4 lines of code to find the correct answer, which is 2,227. Assuming I didn't know about that package though I'd open a site like pro football reference up, middle click on each game to open the page in a new tab, click through the tabs, copy paste the rosters into sublime text, do some regex to get the names one per line, drop the new one per line list into sortmylist or a similar utility, dedupe it, and then paste it back into sublime text to get the line count. That would probably take me about an hour.
- neom 1y agoI see. When you said "game day rosters for many NFL teams are available on many sites" - I thought "that sounds like a lot of hours!!" heh. - I didn't realize it was packaged well, I also know sweet fa about football. Thanks for explaining it more. :)
- paulsutter 1y agoI bet these models could create a python program that does this
- Retric 1y agoMaybe eventually, but I bet it’s not going to work with less than 30 minutes of effort on your part. If “It might take an hour of my time.” to get the correct answer then there’s a lower bond for trying a shortcut that might not work.
- kenjackson 1y agoo3 deep research gave me an answer after I requested an exact answer again (it gave me an estimate first): 2147.
- raybb 1y agoSimilarly, I asked it a rather simple question of giving me a list of AC repair places near me with my numbers. Weirdly, Gemini repeated a bunch of them like 3 or 4 times, gave some completely wrong phone numbers, and found many places hours away but labeled them as in the neighboring city.
- wontonaroo 1y agoI used Google AI Studio instead of Google Gemini App because it provides references to the search results. Google AI Studio gave me an exact answer of 2227 as a possible answer and linked to these comments because there is a comment further down which claims that is the exact answer. The comment was 2 hours old when I did the prompt. It also provided a code example of how to find it using the python nfl data library mentioned in one of the comments here.
- patapong 1y agoSo the time to test data leakage from posting a question and answer to the internet, to LLMs having access to the answer is less than 2h... Does not bode well for the benchmarks of the future!
- gilbetron 1y agoTo avoid "result corruption" I asked a similar question, but for NBA players, and used O4-mini, and got a specific answer: "For the 2023‑24 NBA regular season (which ran from October 24, 2023 to April 14, 2024), a total of 561 distinct players logged at least one game appearance, as indexed by their “Rk” on the Basketball‑Reference “Player Stats: Totals” page (the final rank shown is 561)" Doing a quick search on my own, this number seems like it could be correct.
- deleted 1y ago[deleted]