5 ms·
I feel like most of this recent Autoresearch trend boils down to reinventing hyper-parameter tuning. Is the SOTA still Bayesian optimization when given a small
by kraddypatties 7mo ago
I feel like most of this recent Autoresearch trend boils down to reinventing hyper-parameter tuning. Is the SOTA still Bayesian optimization when given a small cluster? It was ~3 years ago when I was doing this kind of work, haven't kept up since then.
Also, shoutout SkyPilot! It's been a huge help for going multi-cloud with our training and inference jobs (getting GPUs is still a nightmare...)!
- deleted 7mo ago[deleted]
- ipsum2 7mo agoHyperparam tuning that has better intuition and can incorporate architecture changes automatically. It won't invent something completely new though.
- kraddypatties 7mo agoHm, that's fair. It does feel like there's low hanging fruit in combining "old school" methods for conducting a hyperparameter sweep efficiently _with_ the higher level architecture edit ability of Autoresearch. Probably would cut the number of runs down by a significant number (as far as I can tell it's doing a grid search once it decides to mess with a knob or section of the architecture).
- falcor84 7mo ago> It won't invent something completely new though. I don't necessarily disagree, but am wondering whether you have any particular reason/intuition driving you to claim this. I have seen AI agents be quite creative in other tasks; do you think there's a particular reason why we shouldn't see creativity in architecture research, given enough time and resources?
- karpathy 7mo agoWrong and short-sighted take given that the LLM explores serially learning along the way, and can tool use and change code arbitrarily. It seems to currently default to something resembling hyperparameter tuning in absence of more specific instructions. I briefly considered calling the project “autotune” at first but I think “autoresearch” will prove to be the significantly more appropriate name.
- corndoge 7mo agoWould you say it's fair to describe autoresearch as a form of neural architecture search? I am curious what you think the core differences are between them.
- kraddypatties 7mo agoI can believe that in the long run. Does the agent have access to arxiv (a brief skim of the README didn't have an answer)? If not, it could be that the current approach of relying on the model's weights only is resulting in the perceived local optimum of hyperparameter tuning. Anecdotally, we built a little MCP for arxiv to help with our internal research, noticed a significant boost in the diversity of methods (architecture or otherwise) Claude and friends were able to reference.
- touristtam 7mo agocare to share?
- westurner 7mo agoIs there a cost to converge? And how much does it vary with the random seed? Re: OpenCogPrime:EconomicAttentionAllocation https://news.ycombinator.com/item?id=45518074 https://news.ycombinator.com/item?id=45518074 and something about eWASM (edit) https://news.ycombinator.com/item?id=47171887 https://news.ycombinator.com/item?id=47171887 .. from https://news.ycombinator.com/item?id=46825026 https://news.ycombinator.com/item?id=46825026 re: eWASM and costed opcodes for agent efficiency
- achierius 7mo agoOut of curiosity, what sort of things have you seen it do that better fit 'autoresearch' than 'autotune' thus far? Optimizations it made that wouldn't be been surfaced by an autotune system, I suppose.
- karpathy 7mo agoThe most recent round of autoresearch (round 2) which decreased "time to GPT-2" from 1.8 hours to 1.65 hours had some examples. I adjusted the program.md to "look at modded nanogpt project and draw inspirations from there for things to try" and it came back with a bunch of tuning, but also tried and implemented new architecture changes, some of which actually helped including the smear gate and the backout skip connection. These are not just hyperparameters, they are new PyTorch code. I'm now working on a more general system that can have a queue of ideas that could be sourced from archive papers, github repos, etc.
- wenc 7mo agoI wonder if it's more like "qualitative gradient descent" on a very non-linear non-convex surface. You can try this yourself in a simple fashion -- let's say you have piece of code that you want to speed up. Point your agent to a code profiler (your oracle -- typically your Python profiler) and tell it speed up the code. I've tried it. It works.
- aimarketintel 7mo ago[flagged]