17 ms·
1,600 days of a failed hobby data science project
- fardo 2y agoThe author’s right about storytelling from day one, but then immediately throws cold water on the idea by saying it would have been a bad fit for this project. This feels in error, as the big value of seeking feedback and results early and often on a project is that it forces you to confront whether you’re going to want or be able to tell stories in the space at all. It also gives you a chance to re-kindle waning interests, get feedback on your project by others, and avoid ratholing into something for about 5 years without having to engage with a public. If a project can’t emotionally bear day one scrutiny, it’s unlikely to fare better five years later when you’ve got a lot of emotions about incompleteness and the feeling your work isn’t relevant anymore tied up in the project.
- rixed 2y agoWould you be able to recommend a project whom author did engage in such public story telling from early on?
- Swizec 2y agoThinking Fast and Slow is a result of some 20 years of regularly publishing and talking about those ideas with others. Most really memorable works fit that same mold if you look carefully. An author spends years, even decades, doing small scale things before one day they put it all together into a big thing. Comedy specials are the same. Develop material in small scale live with an audience, then create the big thing out of individual pieces that survive the process. Hamming also talks about this as door open vs door closed researchers in his famous You And Your Research essay
- rjrdi38dbbdb 2y agoThe title seems misleading. Unless I'm missing something, all he did was scrape a news feed, which should only require a couple days of work to set up. The fact that he left it running for years without finding the time to do anything with the data isn't that interesting.
- amelius 2y agoYes, his #1 advice should be "do something with the data you collected".
- deleted 2y ago[deleted]
- plaidfuji 2y agoI’m not sure I would call this a failure.. more just something you tried out of curiosity and abandoned. Happens to literally everyone. “Failed” to me would imply there was something fundamentally broken about the approach or the dataset, or that there was an actual negative impact to the unrealized result. It’s very hard to finish long-running side projects that aren’t generating income, attention, or driven by some quasi-pathological obsession. The fact you even blogged about it and made HN front page qualifies as a success in my book. > If I would have finished the project, this dataset would then have been released and used for a number of analyses using Python. Nothing stopping you from releasing the raw dataset and calling it a success! > Back then, I would have trained a specialised model (or used a pretrained specialised model) but since LLMs made so much progress during the runtime of this project from 2020-Q1 to 2024-Q4, I would now rather consider a foundational model wrapped as an AI agent instead; for example, I would try to find a foundation model to do the job of for example finding the right link on the Tagesschau website, which was by far the most draining part of the whole project. I actually just started (and subsequently —-abandoned—- paused) my own news analysis side project leveraging LLMs for consolidation/aggregation.. and yeah, the web scraping part is still the worst. And I’ve had the same thought that feeding raw HTML to the LLM might be an easier way of parsing web objects now. The problem is most sites are privy to scraping efforts and it’s not so much a matter of finding the right element but bypassing the weird click-thru screens, tricking the site that you’re on a real browser, etc…
- smcin 2y ago> Nothing stopping you from releasing the raw dataset and calling it a success! Right. OP: release it as a Kaggle Dataset (https://www.kaggle.com/datasets https://www.kaggle.com/datasets) and invite people to collaboratively figure out how to autonate the analyses. (Do you just want to get sentiment on a specific topic (e.g. vaccination, German energy supplies, German govt approval)? or quantitative predictions?) Start with something easy. > for example, I would try to find a foundation model to do the job of for example finding the right link on the Tagesschau website, which was by far the most draining part of the whole project. Huh? To find the specific dates new item corresponding to a given topic? Why not just predict the date-range e.g. "Apr-Aug 2022" > and yeah, the web scraping part is still the worst. Sounds wrong. OP, fix your scraping. (unless it was anti-AI heuristics that kept breaking it, which I doubt since it's Tagesschau). But Tagesschau has RSS feeds, so why are you blocked on scraping? https://www.tagesschau.de/infoservices/rssfeeds https://www.tagesschau.de/infoservices/rssfeeds Compare to: Kaggle Datasets "10k German News Articles for topic classification", Schabus, Skowron Trspp, SIGIR 2017 [https://www.kaggle.com/datasets/abhishek/10k-german-news-articles https://www.kaggle.com/datasets/abhishek/10k-german-news-art...]
- querez 2y agoSome very weird things in this. 1. The title makes it sound like the author spent a lot of time on this project. But really, this mostly consisted of noting down a couple of URLs per day. So maybe 5 min / day = ~130h spent on the project. Let's say 200h to be on the safe side. 2. "Get first analyses results out quickly based on a small dataset and don’t just collect data up front to “analyse it later”" => I think this actually killed the project. Collecting data for several years w/o actually doing anything doesn't with it is not a sound project. 3. "If I would have finished the project, this dataset would then have been released" ==> There is literally nothing stopping OP from still doing this. It costs maybe 2h of work and would potentially give a substantial benefit to others, i.e., turn this project into a win after all. I'm very puzzled why OP didn't do this.
- apwell23 2y agoyep I spent more time on duolingo for 600+ day streak and can barely speak spanish.
- rrr_oh_man 2y agoThat seems to be a pattern
- galleywest200 2y agoIt is because you never really practice talking with Duolingo. I am quite good at reading French now, though.
- pessimizer 2y ago> I am quite good at reading French now, though. If you are, that's actually quite an achievement and good. If you're talking about French outside of Duolingo, that is. I do not normally hear of people getting to reading fluency through Duolingo.
- wizzwizz4 2y ago
- mNovak 2y ago"The data collection process involved a daily ritual of manually visiting the Tagesschau website to capture links" I don't know what to say... I'm amazed they kept this up so long, but this really should never have been the game plan. I also had some data science hobby projects around covid; I got busy, lost interest after 6 months. But the scrapers keep running in the cloud, in case I get motivated again (anyone need structured data on eBay listings for laptops since 2020?), that's the beauty of automation for these sorts of things.
- plaidfuji 2y agoDo you just pay the bill for the resources indefinitely?
- hansvm 2y agoI'm not the person you're asking, but I maintain a number of scraping projects. The bills are negligible for almost everything. A single $3/mo VPS can easily handle 1M QPS (enough for all the small projects put together), and most of these projects only accumulate O(10GB)/yr. Doing something like grabbing hourly updates of the inventory of every item in every Target store is a bit more involved, and you'll rapidly accumulate proxy/IP/storage/... costs, but 99% of these projects have more valuable data at a lesser scale, and it's absolutely worth continuing them on average.
- NavinF 2y agoInbound data is typically free on cloud VMs. CPU/RAM usage is also small unless you use chromedriver and scrape using an entire browser with graphics rendered on CPU. We're taking $5/mo for most scraping projects
- mNovak 2y agoI paying < $0.50 a month, and that's primarily driven by S3. For the scraping itself I'm using lambda, with maybe minutes of runtime per day.
- FrustratedMonky 2y ago"Data Science Project Failing After 1,600 Days" Sounds like my Thesis. How many people have spent 4+ years on a Thesis then just completely gave up, tired, drained, no interest in continuing. The bright eye'd bushy tailed wonder, all gone.
- dankwizard 2y agoI don't speak the language so maybe what you're scraping isn't in this list, but why manual when they seem to have comprehensive RSS feeds? [1] Automating this part should have been day 1. [1] https://www.tagesschau.de/infoservices/rssfeeds https://www.tagesschau.de/infoservices/rssfeeds
- smcin 2y agoThat's what I just concluded. I think the OP was oversold on the idea of using AI to do scraping, NLP and summarization, all in one go.
- smcin 2y agoBest practice (for many reasons) is to separate scraping (and OCR) and store the rawtext or raw HTML/JS, and also the parsed intermediate result (cleaned scraped text or HTML, with all the useless parts/tags removed). This is then the input to the rest of the pipeline. You really want to separate those, both for minimizing costs, and preventing breakage when site format changes, anti-scraping heuristics change, etc. And not exposing garbage tags to AI saves you time/money.
- NeinMiez 2y ago[dead]
- j45 2y agoI don’t know that projects ever fail. Doing them and learning and growing from them is the point. They shed a light on your path and also what you are able to see as possible.
- ddxv 2y agoWhy not open source? I've been slaving away at some possibly pointless data scraping sites that collect app data and the SDKs that apps use. I figure if I at least open source it that data and code is there for others to use.
- kqr 2y agoI see some recommendations about running a small version of the analysis first to see if it's going to work at all. I agree, and the next level up is to also estimate the value of performing the full analysis. I.e. not just whether or not it will work at all, but how much it is allowed to cost and still be useful. You may find, for example, that each unit of uncertainty reduced costs more than the value of the corresponding uncertainty reduction. This is the point at which one needs to either find a new approach, or be content with the level of uncertainty one has.
- brikym 2y agoI know the feeling. I managed 9 months scraping supermarket data before I gave up mostly because a few other people were doing it and I was short on time.
- jfil 2y agoWhat country's data did you scrape? Do you make it available somewhere? I'm on month 10 of scraping Canadian grocer data and make it available publicly at https://jacobfilipp.com/hammer/ https://jacobfilipp.com/hammer/
- barrenko 2y agoPeople relatively new to CS would be wise to be warned about what a colossal time sink it is.
- CRConrad 2y agoYeah, my kid wastes far too much time on CounterStrike.
- wodenokoto 2y ago> Store raw data if possible. This allows you to condense it later. I have some daily scripts reading from an http endpoint, and I can't really decide what to do when it returns html instead of json. Should I store the HTML as it is "raw data" or should I just dismiss it? The API in question has a tendency to return 200 with a webpage saying that the API can't be reached (typically because of a time out)
- IanCal 2y agoI wouldn't store that usually, I'd use that to trigger retries. For you storing the raw data is storing the json that http endpoint returns rather than something like let content = get(url).json() info_i_care_about = content['data']['title'] store(info_i_care_about) as otherwise you'll get stuck when the json response moves the title to data.metadata.title or whatever It's usually less of an issue with structured data, things like html change more often, but keeping that raw data means you can process it in various different ways later. You also decouple errors so your parsing error doesn't stop your write from happening.
- tessierashpool9 2y agothe last thing the world or rather germany needs is a news ticker based on ... the tagesschau LOL
- KeplerBoy 2y agoOh boy, the topic (Covid) alone would have left me exhausted after a few months. I heard enough of it by mid 2021.
- CRConrad 2y agoDid you misspell 2020...?
- rybosworld 2y ago> The data collection process involved a daily ritual of manually visiting the Tagesschau website to capture links to both the COVID and later Ukraine war newstickers. While this manual approach constituted the bulk of the project’s effort, it was necessitated by Tagesschau’s unstructured URL schema, which made automated link collection impractical. > The emphasis on preserving raw HTML proved vital when Tagesschau repeatedly altered their newsticker DOM structure throughout Q2 2020. Another big takeaway is that it's not sustainable to rely on this type of a data source. Your data source should be stable. If the site offers API's, that's almost always better than parsing html. Website developers do not consider scrapers when they make changes. Why would they? So if you are ever trying to collect some unique dataset, it doesn't hurt to reach out to the web devs to see if they can provide a public API.
- abirch 2y agoPlease consider it an early Christmas present to yourself if you can pay a nominal amount for an API instead of spending your time scraping unless you enjoy doing the scraping.
- buddybubble 2y agoI still don't understand what he even tried to do? So he manually collected news articles for a few years without any plan on what to do with them so what? Where is the project? Honestly he could probably just have asked the tagesschau people and they would have given him their archive. The learning from this seems to be: collecting data and never doing anything with it is not a worthy project
- Uptrenda 2y agoI think whether you 'succeed' or 'fail' on a side project they are still valuable. No matter if you can't finish it or it turns out different to how you imagined -- you get to come away as a better version of yourself. A person who is more optimized for a new strategy. And sometimes 'failure' is a worthwhile price for that ability. Who knows, it might be exactly what prepares you for something even bigger in the future.
- fuzzfactor 2y agoI guess the kind of extreme effort that doesn't usually have a promising conclusion is more common in scientific research, or experimentation in general, but sometimes you just have to get accustomed to it. Eventually it doesn't really make any difference if there's no breathtaking milestone because it turned out to be impossible by nature, ran out of runway, or lost interest after a more or less valiant attempt. What can be gained is the strength to overcome the near-impossible next time and all it has to do is be a certain degree less-impossible and you know whether that would take you over the goal line like few others because you've been there. Without even worrying as much about whether you will lose interest or not, that's a lot less stress and pressure when you think about it. This can enable you more realistically to succeed in other areas where peers may find it impossible or not be able to do as well without as big an inconclusive project behind them.
- TheGoodBarn 2y agoWhat I love about projects like this is they are dynamic enough to cover a number of interests all in one. I personally have some side projects that have started as X, transitioned into Y and Z, and then I stole some ideas and built A, which turned to B, which a requirement in my professional job necessitated the Z solution mixed with the B solution and resulted in something else which re-ignited my interest in X and helped me rebuild with a more clear mindset on what I intended in the first place. All that to say, these things are dynamic and a long list of "failed" projects is a historical narrative of learning and interests over time. I love to see it.
- sota_pop 2y agoNice article OP. I and a great many others suffer from the same struggles of bringing personal projects to “completion”, and I’ve gotta respect the resilience in the length of time you hung in there. However, not to be overly pedantic, but I always felt “data science” was an exploratory exercise to discover insights into a given data set. I always personally filed the efforts to create the pipeline and associated automation (i.e. identify, capture, and store a given data set - more commonly referred to as “ETL”) as a “data engineering” task, which these days is considered a different specialty. Perhaps if you scope your problem a little smaller, you may yet be able to capture something demonstrably valuable to others (and something you might consider “finished”). You’d be surprised how simple something that addresses a real issue can be to be able to provide real value for others. Nice work and great effort.
- sshrajesh 2y agoAnyone knows what software is used to create these diagrams: https://lellep.xyz/blog/images/failed_data_science_project/2024-11-01_liveblog_data_format.jpg https://lellep.xyz/blog/images/failed_data_science_project/2...
- regular_trash 2y agoExcalidraw
- tvrg 2y agoLooks like something you could create with excalidraw. It's an awesome tool! https://excalidraw.com/ https://excalidraw.com/
- deleted 2y ago[deleted]
- dowager_dan99 2y agoI for one don't want to start counting everything I lose interest in as a "failure", that would be too depressing. I actually think this is a feature not a flaw. You have very few attention tokens and should be aggressive in getting them back. I think this is very different from the "finishing" decision. That should focus on scope and iterations, while attempting to account for effort vs. reward and avoiding things like sunk cost influences. Combine both and you've got "pragmatic grit": the ability to get valuable shit done.
- ComodoHacker 2y agoTurns out it wasn't nearly as much as 1600 days of labor. So, clickbait headline.