5 ms·
After reading this I realized I also have an archive of my pocket account (4200 items), so tried the same prompt with o3, gemini 2.5 pro, and opus 4: - chatgpt
by saeedesmaili 1y ago
After reading this I realized I also have an archive of my pocket account (4200 items), so tried the same prompt with o3, gemini 2.5 pro, and opus 4:
- chatgpt UI didn't allow me to submit the input, saying it's too large. Although it was around 80k tokens, less than o3's 200k context size.
- gemini 2.5 pro: worked fine for personality and interest related parts of the profile, but it failed the age range, job role, location, parental status with incorrect perdictions.
- opus 4: nailed it and did a more impressive job, accurately predicted my base city (amsterdam), age range, relationship status, but didn't include anything about if I'm a parent or not.
Both gemini and opus failed in predicting my role, probably understandably. Although I'm a data scientist, I read a lot about software engineering practices because I like writing software and since I don't have the opportunity at work to do this kind of work, I code for personal projects, so I need to learn a lot about system design, etc. Both models thought I'm a software engineer.
Overall it was a nice experiment. Something I noticed is both models mentioned photography as my main hobby, but if they had access to my youtube watch history, they'd confidently say it's tennis. For topics and interests that we usually watch videos rather than reading articles about, would be interesting to combine the youtube watch history with this pocket archive data (although it would be challenging to get that data).
- greenavocado 1y agoYou need to use an iterative refinement pyramid of prompts. Use a cheap model to condense the majority of the raw data in chunks, then increasingly stronger and more expensive models over increasingly larger sets of those chunks until you are able to reach the level of summarization you desire.
- tgtweak 1y agoI think a reasoning/thinking-heavy model would do better at piecing together the various data points than an agentic model. Would be interested to see how o3 does with the context summarized.
- saeedesmaili 1y agoAgreed, that's why I used reasoning models (gemini 2.5 pro and opus 4 with extended thinking enabled).
- tehlike 1y agoYou should take this as a sign, and shoot for SWE jobs - given your interest. What you do at work today doesn't mean you can't switch to a related ladder.
- justusthane 1y agoSometimes it’s nice for hobbies to remain hobbies
- formerphotoj 1y agoExactly this. The need to make money from a thing may well eliminate the value one derives from the thing, and even add negatives such as stress, etc.
- tehlike 1y agoNot really. I do software both as a hobby, and as a career.
- cortesoft 1y agoI believed this, which is what made me avoid computer science in college; I wanted to avoid ruining my favorite hobby. After a few years post graduation, where I wasn't sure what I wanted to do and I floundered to find a career, I decided to give software development a try, and risk ruining my favorite hobby. Definitely the best decision I could have made. Now people pay me a lot of money to do the thing I love to do the most... what's not to love? 20 years later, it I still my favorite hobby, and they keep paying me to do it.
- p1necone 1y agoI think it heavily depends on who you're working for. If they get out of the way and let you do the thing you love how you want to do it you'll get good results for you and them. If they treat you like a cog in a machine and assume they need to carrot and stick you into doing things because you might not really want to be there, you'll be miserable.
- juliendorra 1y agoYou should be able to use Google Takeout to get all of your YouTube data, including your watch history. This article is a nice example of someone using it: > When I downloaded all my YouTube data, I’ve noticed an interesting file included. That file was named watch-history and it contained a list of all the videos I’ve ever watched. https://blog.viktomas.com/posts/youtube-usage/ https://blog.viktomas.com/posts/youtube-usage/ Of course as an European it's a legal obligation for companies to give you access, but I think Google Takeout works worldwide?
- jazzyjackson 1y agoYes I've done this in USA. pretty neat. I have it on my todo list to parse over it and find all the music videos I've watched 3 or more times to archive them.
- toomuchtodo 1y agohttps://archive.zhimingwang.org/blog/2014-11-05-list-youtube-playlist-with-youtube-dl.html https://archive.zhimingwang.org/blog/2014-11-05-list-youtube... might be of use along with https://github.com/yt-dlp/yt-dlp https://github.com/yt-dlp/yt-dlp, might just grab it all and prune later due to rot and availability issues over time within YT.
- viraptor 1y agoIt is available and it can be surprisingly large. I've somehow accumulated multiple GB of data from YT alone. Which feels a bit absurd - there's bound to be lots of waste there.
- yubblegum 1y agoThis can give a false sense of what Google (Alphabet) actually knows about you. That above is Google playing the game of 'ok, here is what we know of your activities on youtube when logged in!' But Google and the rest of the "advertising" (euphemism for surveillance) industry track and create "profiles" based on a basket of data points, from ip/MAC address to the rest of their bag of tricks.
- LoganDark 1y ago> Both models thought I'm a software engineer. You probably still are, even if that's not your career path :)
- larve 1y agore o3: you can zip the file, upload it, and it will use python and grep and the shell to inspect it. I have yet to try using it with a sqlite db, but that's how i do things locally with agents.
- saeedesmaili 1y agoAuthor mentions that by doing that they didn't get a high quality response. Adding the texts into model's context make all the information available for it to use.
- datpuz 1y agoReading 80k tokens requires more than 80k tokens due to overhead
- alexnorton 1y agoI was able to give this a try on every YouTube video I've ever watched by exporting the history from Google Takeout: https://takeout.google.com/settings/takeout/custom/youtube?pli=1 https://takeout.google.com/settings/takeout/custom/youtube?p... And then a combination of pup and jq to parse the video titles from the HTML file: cat watch-history.html \ | pup '.outer-cell .mdl-grid .content-cell:nth-child(2) json{}' \ | jq -r '.[] .children[0] | select(.tag != "br") | select(.text | startswith("https://www.youtube.com/watch?v=") | not) | .text' \ > videos.txt
- UrineSqueegee 1y agoo3 on the webui has a tiny context as do all the models