3 ms·
Excellent questions, thank you! 1. We could log additional information about the model, such as inference time, number of parameters, memory usage, etc. and ha
by metaphdor 7y ago
Excellent questions, thank you!
1. We could log additional information about the model, such as inference time, number of parameters, memory usage, etc. and have the primary metric be overall efficiency (best NDCG with fewest parameters/fastest runtime/etc).
2. We're experimenting with different kinds of benchmarks, and I am most excited about explicitly collaborative ones. In these there is no contest/prize (hence no incentive to cheat/withhold information); only the shared goal of improving the model and our collective understanding of the problem. I hope we can incentivize information sharing by tracking and acknowledging individual contributions to the eventual best model in the benchmark. We could approximate individual contribution by seeing which scripts, code segments, workflows, architectural changes, writeups, or discussion comments other participants rate as the most helpful or choose to include in their experiments most often as the benchmark evolves. Of course this could only be an estimate--as Shawn says above, any idea could have "actually happened in a hallway conversation". Still, this is much easier to achieve in a logging/visualization platform like W&B than in the current paradigm of "read research papers, clone relevant repos, spend weeks trying to synthesize/reproduce their results, run your own experiments, write them up in a research paper, hope it gets accepted to a conference before other people publish the same idea, try to integrate your changes/publish your own repo, repeat"--and for hundreds of practitioners, ranging from brand new students to PhDs, working on related problems. This cycle is especially challenging for folks who are new to, working outside of, or trying to collaborate across the relatively few established/well-funded academic/industrial teams.
Collaborative benchmarks can be especially impactful for social good projects, where the primary incentive is to figure out and broadly implement the best solution ASAP (e.g. climate change!), not to make money or argue over the credit attribution. So, my long-term goal is for as much sharing of information and collaboration from as many folks as possible--the more inclusive and transparent the field of deep learning research becomes, the safer and better its outcomes. Very open to ideas on how to help make this happen.
~Stacey, deep learning engineer at W&B