7 ms·
Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a GPU Cluster
- mika-el 7mo ago[flagged]
- kraddypatties 7mo agoI feel like most of this recent Autoresearch trend boils down to reinventing hyper-parameter tuning. Is the SOTA still Bayesian optimization when given a small cluster? It was ~3 years ago when I was doing this kind of work, haven't kept up since then. Also, shoutout SkyPilot! It's been a huge help for going multi-cloud with our training and inference jobs (getting GPUs is still a nightmare...)!
- deleted 7mo ago[deleted]
- ipsum2 7mo agoHyperparam tuning that has better intuition and can incorporate architecture changes automatically. It won't invent something completely new though.
- kraddypatties 7mo agoHm, that's fair. It does feel like there's low hanging fruit in combining "old school" methods for conducting a hyperparameter sweep efficiently _with_ the higher level architecture edit ability of Autoresearch. Probably would cut the number of runs down by a significant number (as far as I can tell it's doing a grid search once it decides to mess with a knob or section of the architecture).
- falcor84 7mo ago> It won't invent something completely new though. I don't necessarily disagree, but am wondering whether you have any particular reason/intuition driving you to claim this. I have seen AI agents be quite creative in other tasks; do you think there's a particular reason why we shouldn't see creativity in architecture research, given enough time and resources?
- karpathy 7mo agoWrong and short-sighted take given that the LLM explores serially learning along the way, and can tool use and change code arbitrarily. It seems to currently default to something resembling hyperparameter tuning in absence of more specific instructions. I briefly considered calling the project “autotune” at first but I think “autoresearch” will prove to be the significantly more appropriate name.
- corndoge 7mo agoWould you say it's fair to describe autoresearch as a form of neural architecture search? I am curious what you think the core differences are between them.
- kraddypatties 7mo agoI can believe that in the long run. Does the agent have access to arxiv (a brief skim of the README didn't have an answer)? If not, it could be that the current approach of relying on the model's weights only is resulting in the perceived local optimum of hyperparameter tuning. Anecdotally, we built a little MCP for arxiv to help with our internal research, noticed a significant boost in the diversity of methods (architecture or otherwise) Claude and friends were able to reference.
- touristtam 7mo agocare to share?
- westurner 7mo agoIs there a cost to converge? And how much does it vary with the random seed? Re: OpenCogPrime:EconomicAttentionAllocation https://news.ycombinator.com/item?id=45518074 https://news.ycombinator.com/item?id=45518074 and something about eWASM (edit) https://news.ycombinator.com/item?id=47171887 https://news.ycombinator.com/item?id=47171887 .. from https://news.ycombinator.com/item?id=46825026 https://news.ycombinator.com/item?id=46825026 re: eWASM and costed opcodes for agent efficiency
- achierius 7mo ago
- wenc 7mo agoI wonder if it's more like "qualitative gradient descent" on a very non-linear non-convex surface. You can try this yourself in a simple fashion -- let's say you have piece of code that you want to speed up. Point your agent to a code profiler (your oracle -- typically your Python profiler) and tell it speed up the code. I've tried it. It works.
- aimarketintel 7mo ago[flagged]
- zhwu 7mo agoThe most surprising part: the agent had access to both H100s and H200s. Without being told, it noticed H200s scored better and started screening ideas on H100s, then promoting winners to H200s for validation. That strategy emerged entirely on its own.
- Aboutplants 7mo agoYeah I thought that was a particularly neat part
- rogerrogerr 7mo agoWhy do we think this emerged “on its own”? Surely this technique has been discussed in research papers that are in the training set.
- fdghrtbrt 7mo agoWhy surely? Have you never seen an LLM try something new?
- rogerrogerr 7mo agoIs your assertion that no one has ever written "we tried some stuff on the small inexpensive platform first, then moved to the bigger more expensive platform with the more promising options" in a research paper or literally anywhere else?
- fdghrtbrt 7mo agoNo, that's not my assertion. In fact I asserted nothing at all.
- rogerrogerr 7mo agoYou're speaking in riddles; your communication would be more effective if you didn't do that.
- covi 7mo agoThis feels like the chimpanzee with a power drill. An agent is honestly just brute-force search, but guided.
- deleted 7mo ago[deleted]
- chaos_emergent 7mo agoHuman-driven research is also brute-force but with a more efficient search strategy. One can think of a parameter that represents research-search-space-navigation efficiency. RL-trained agents will inevitably optimize for that parameter. I agree with your statement insomuch as the value of that efficiency parameter is lower for agents than humans today. It's really hard to imagine that they __won't__ exceed the human value for that efficiency parameter rather soon given that 1. there are plenty of scalar value functions that can represent research efficiency, of which a subset will result in robust training, and 2. that AI labs have a massive incentive to increase their research efficiency overall, along with billions of dollars and really good human researchers working on the problem.
- viccis 7mo ago>Human-driven research is also brute-force but with a more efficient search strategy No it's not. Is there anything to back that up? There's a creative aspect to human research that I've yet to see with gen AI. All it does is regurgitate stuff and get some "new" ideas via the latent space of the distribution it models. But a generative model cannot by definition create anything new. Just estimate its data well enough that it can sample it well enough to fake novelty.
- groby_b 7mo agoIs there anything in the research space that doesn't fit "brute-force search, but guided"? All of science is "gather inputs, make hypothesis, test, analyse" on repeat. There's plenty to critique in the particular guidance approach, but the overall method is the same.
- gwern 7mo ago
- pratelsingh 7mo ago[dead]
- ipsum2 7mo agoA cluster is 2 nodes? That's technically true, but not very exciting.
- deleted 7mo ago[deleted]
- fabmilo 7mo agoI am fascinated by this example of using AI to improve AI. I won a small prize using this technique on helion kernels at a pytorch hackathon in SF. The next step are: - give the agent the whole deep learning literature research and do tree search over the various ideas that have been proposed in the past. - have some distributed notepad that any of these agents can read and improve upon.
- saberience 7mo agoWait, "Karpathy's Autoresearch", you mean a loop that prompts the agent to improve a thing given a benchmark? People have been doing this for a year or more, Ralph loops etc. I hate the weird strange Twitter world of hero-worship for folks that seems to arise just out of large followings. Joe no-followers does this six months ago, nobody cares. Karpathy writes a really basic loop and it's now a kind of AI miracle prompting tons of grifters, copy-cats, weird hype. I do wonder if LLMs have just made everyone seriously, seriously dumber all of a sudden. Most of the "Autoresearch" posts I see are completely rubbish, with AI optimizing for nonsense benchmarks and people failing to understand the graphs they are looking at. So yes, the AI made itself better at a useless benchmark while also making the code worse in 10 other ways you don't actually understand.
- password54321 7mo agoThe number of refurbished mac minis that are available in my country has suddenly dramatically increased ever since the Clawdbot tweet. People never learn.
- misiti3780 7mo agoincreased or decreased?
- dag100 7mo ago"increased" implies that people bought brand-new Mac Minis to run ClawdBot on, got bored of it, and then sold them back to be refurbished and resold.
- misiti3780 7mo agoya ok, that makes more sense, thanks
- maxothex 7mo ago[dead]
- opensre 7mo ago[flagged]
- pbkhrv 7mo ago> How parallelism changed the agent’s research strategy > With a single GPU, the agent is stuck doing greedy hill-climbing: try one thing, check the result, pick a direction, try the next thing. With 16 GPUs, the strategy shifts. ...skip... 12 experiments in a single 5-minute wave. This makes it much harder to get stuck in local optima and much easier to find interaction effects between parameters. The agent can theoretically come up with a protocol to run those same 12 experiments one-by-one and only then decide which branch to explore next - which I think would lead to the same outcome? But in this case, it just happened to have stumbled on this particular outcome only because it didn't get a chance to execute a greedy strategy after the first 1 or 2 results. Worse experiment design + parallelism = better experiment design + serialized execution ?
- rfw300 7mo agoYeah, assuming there's no active monitoring during the training runs, you can trivially give the agent an abstraction which turns "1 GPU" into "16 GPUs" that just so happens to take 16x the wall-clock time to run.
- rfw300 7mo agoIn fact, looking at the blog post, the agent orchestrating 16 GPUs is half as efficient as the agent using 1 GPU in GPU-time. Since it uses 16 GPUs to reach the same result as 1 GPU in 1/8 of the time.
- gwern 7mo ago> The agent can theoretically come up with a protocol to run those same 12 experiments one-by-one and only then decide which branch to explore next - which I think would lead to the same outcome? At least in theory, adaptiveness should save samples and in this case, compute. (As noted, you can always turn the parallel into serial and so the serial approach, which gets information 'from the future', should be able to meet or beat any parallel approach on sample-efficiency.) So if the batch only matches the adaptive search, that suggests that the LLM is not reasoning well in the adaptive setting and is poorly exploiting the additional information. Maybe some sort of more explicit counterfactual reasoning/planning over a tree of possible outcomes?
- ReacherL3692283 7mo ago[dead]
- herf 7mo agoThis "early velocity only" approach seems like a problem - how do you know with 5-minute training runs that you aren't affecting the overall asymptote? e.g., what if the AI picks a quantizer that happens to be faster in the first five minutes, but has a big noise floor where it can't make more progress?
- gwern 7mo agoYes, it's greedy so may hit local optima. You can fit learning curves and try to extrapolate out to avoid that problem, to let you run long enough to be reasonably sure of a dead end, and periodically revive past candidates to run longer. See past hyperparameter approaches like freeze-thaw https://arxiv.org/abs/1406.3896 https://arxiv.org/abs/1406.3896 .
- claud_ia 7mo ago[dead]
- ladyxtel88 7mo ago[dead]
- robutsume 7mo ago[dead]
- aplomb1026 7mo ago[dead]
- fmymzk41 7mo ago[dead]
- nsollazzo53 7mo ago[dead]
- UndoExec55 7mo ago[dead]
- hgoel 7mo agoThis is fascinating to me because I just recently built something similar (as a test), not for improving AI, instead it's for tuning hyperparameters for a physics simulation. We've managed to optimize execution of the simulation enough that brute-force search is a viable option, but giving an agent some background on how we tune those parameters on intuition and some physical reasoning, and a means to run tests and retrieve resulting statstics, works surprisingly well. I see it as essentially a hyperparameter search that is more capable of finding and exploiting implicit constraints in a system.
- Freedumbs 7mo agoThe US governement should 'autoresearch' a way to charge this man for his crimes as head of autopilot at tesla.
- random3 7mo agoCan you elaborate?
- Freedumbs 7mo agoare you familiar with tesla? i'm not super, but am aware of their public things. they introduced fake marketing products called full self driving and autopilot that don't do those things. apparently this person karpathy was in charge of computer vision there. he led the team who is responsible for these systems that occupy our roads which can't navigate due to such outstanding occurrences as sunlight, precipitation, and fog. i don't know a great deal about the guy. i know: he worked at tesla, led autopilot there. if we ignore the character defects required to work at tesla, he's responsible for designing systems that would certainly kill people because they decided lidar was too expensive.
- dbdr 7mo agoBuilding a tech and falsely advertising it to be something else that what it is (e.g. self driving instead of driving assistance) can typically done by different people. Lacking specific evidence, it's reckless to accuse this person.
- Freedumbs 7mo agoright. i'm mostly ignorant of the subject and rushing to judgement based on bias. but he did lead the computer vision team for years at tesla that created autopilot. didn't resign in protest and to my knowledge hasn't apologized, but again i'm ignorant and not seeking new data.
- nurettin 7mo agoIt feels like the richer a company is, the dumber their software and more expensive their upkeep gets. Something you could do with optuna in C++ on a single server now requires clusters of GPUs with LLMs at the helm.
- rsmtjohn 7mo ago[flagged]
- huang-b62b5756 7mo agocool
- augment_me 7mo agoIsn't this apples to pears comparison? It's really just saying that having a bigger credit card gets you shit faster, but it's actually worse in terms of GPU utilization and efficiency. 1) The total amount of time is not the same if you just count GPU-hours. If you have 16 GPUs, it makes sense to run them for 4.5 hours to get to 72h for an even comparison, not 8. 2) If we stop at 4.5 hours(and are generous including the big drop), the loss is about 0.978, which is the same as about 44 hours with the sequential solution, making the sequential solution about twice as efficient. So the real conclusion here is that we are able to run things in parallel at an efficiency loss but at a time win as long as we have access to more hardware. I feel like the blog oversells itself.
- QubridAI 7mo agoFeels like we’ve solved how to run agents anywhere, but not yet how to trust them anywhere.
- elophanto_agent 7mo ago[dead]
- elophanto_agent 7mo ago[dead]
- zkmon 7mo agoI'm trying to find the 'wow' factor in this. Finding the optimal combination of parameters, given a validation criteria should be a boring repetitive task for a machine or a human. Is it about determining how to utilize the given hardware?
- no_shadowban 7mo ago[dead]
- bhekanik 7mo ago[dead]
- ordinarily 7mo agoI've been doing this for about a month. I also have wildly complicated ML pipelines working similarly in parallel. When Karpathy's 'autoresearch' came out I was surprised by how novel it was treated.
- muin_kr 7mo ago[dead]
- snthpy 7mo ago> Scale Autoresearch on your own GPU cluster Who's got a cluster of H100s and H200s just lying around?
- dragonwriter 7mo agoThe entire discussion is about rented cloud clusters, so I guess anyone with the money to rent one?
- WecoAI 6mo ago[dead]