4 ms·
Running code state-of-the-art for federated learning on mobiles: https://github.com/jverbraeken/trustchain-superapp/commits/master https://github.com/jverbraeke
by synctext 5y ago
Running code state-of-the-art for federated learning on mobiles:
https://github.com/jverbraeken/trustchain-superapp/commits/master https://github.com/jverbraeken/trustchain-superapp/commits/m...
(Publication: https://arxiv.org/abs/2110.11006 https://arxiv.org/abs/2110.11006)
People like to code their own work, instead of contributing to others. Incremental work does not count. This is more of a flaw in scientific output counting I believe. Also, our technology stack is completely "zero-server" that makes everything on Android much harder and complex. Disclaimer: I'm the thesis adviser of this student
- mike_hearn 5y ago"People like to code their own work, instead of contributing to others. Incremental work does not count. This is more of a flaw in scientific output counting I believe" Lots of truth here. Actually, maybe the OP could change the title to make it clear you're talking about machine learning models? When I first saw the title I thought it would be about the deeply problematic state of scientific coding and "scientific" modelling in general. The ML world isn't so bad regardless of how it may seem, because it's got deep roots in computer science and industry. The moment you start looking at the code for things like epidemiological models, climate models etc it becomes clear that the incentive structures in science are completely broken. I've not only heard from others but also seen it with my own eyes that people who call themselves "scientists" will happily write models in C that e.g. use pointer values in an equation instead of the dereferenced values, leading to incorrect outputs, and when this is pointed they claim there are no problems, or that they checked the results and the bugs make no difference (i.e. they lie). Or they'll have race conditions that cause unstable outputs and pretend it's because of a (pre seeded) PRNG they used, apparently in the hope that other scientists don't understand pseudo-randomness properly. Or they'll mix up their variables because everything is a single letter and there are thousands of lines of code, or they'll do out-of-bounds reads in a sorting algorithm because they don't know about standard libraries etc. Nobody ever seems to retract papers because of bugs like these, and when the results of their model don't match reality they'll make arguments like "the model was validated against other models, which is a reasonable way to prove validity" or "we don't make predictions we make scenarios". If researchers were more willing to collaborate on shared codebases they'd start to learn programming better, be able to recruit wider and more diverse teams, they'd share infrastructure and best practices and generally things might stand a chance of improving. But, as you say, there are major flaws in how the output of scientists are evaluated (by governments/non-profits), and this has a nasty habit of converting well meaning scientists into what are effectively pseudo-scientists.