15 ms·
Anthropic's original take home assignment open sourced
- piokoch 9mo agoInteresting... Who would spend hours working for free for some company that promised only that they would invite you for a job interview. Maybe.
- cjrp 9mo agoI guess someone who enjoys solving these kinds of problems anyway, and thinks the potential upside if they do get hired is worth it.
- Aurornis 8mo agoWhen this was being used it was probably given to candidates who had already started the interview loop and been screened. The current e-mail invitation in the README is just another avenue for exceptional people to apply. If someone is already highly qualified from their background and resume they can go through the front door (direct application). For those who have incredible talent but not necessarily the background or resume to unlock the front door yet, this is a fun way to demonstrate it.
- myahio 9mo ago[flagged]
- koolba 9mo agoWhat is the actual assignment here? The README only gives numbers without any information on what you’re supposed to do or how you are rated.
- deleted 9mo ago[deleted]
- glalonde 9mo ago"Optimize the kernel (in KernelBuilder.build_kernel) as much as possible in the available time, as measured by test_kernel_cycles on a frozen separate copy of the simulator." from perf_takehome.py
- vermilingua 9mo agoThink that means you failed :(
- nice_byte 9mo ago+1 being cryptic and poorly specified is part of the assignment just like real code in fact, it's _still_ better documented an self contained than most of the problems you'd usually encounter in the wild. pulling on a thread to end up with a clear picture of what needs to be accomplished is like 90% of the job very often.
- avaer 9mo agoIt's definitely cleaner than what you will see in the real world. Research-quality repositories written in partial Chinese with key dependencies missing are common. IMO the assignment('s purpose) could be improved by making the code significantly worse. Then you're testing the important stuff (dealing with ambiguity) that the AI can't do so well. Probably the reason they didn't do that is because it would make evaluation harder + more costly.
- throwaway81523 9mo agoI didn't see much cryptic except having to click on "perf_takehome.py" without being told to. But, 2 hours didn't seem like much to bring the sample code into some kind of test environment, debug it enough to works out details of its behaviour, read through the reference kernel and get some idea of what the algorithm is doing, read through the simulator to understand the VM instruction set, understand the test harness enough to see how the parallelism works, re-code the algorithm in the VM's machine language while iterating performance tweaks and running simulations, etc. Basically it's a long enough problem that I'd be annoyed at being asked to do it at home for free, if what I wanted from that was a shot at an interview. If I had time on my hands though, it's something I could see trying for fun.
- nice_byte 9mo agoit's "cryptic" for an interview problem. e.g. the fact that you have to actually look at the vm implementation instead of having the full documentation of the instruction set from the get go.
- jackblemming 9mo agoSeems like they’re trying to hire nerds who know a lot about hardware or compiler optimizations. That will only get you so far. I guess hiring for creativity is a lot harder. And before some smart aleck says you can be creative on these types of optimization problems: not in two hours, it’s far too risky vs regurgitating some standard set of tried and true algos.
- rvz 9mo ago> Seems like they’re trying to hire nerds who know a lot about hardware or compiler optimizations. That will only get you so far. I guess hiring for creativity is a lot harder. Good. That should be the minimum requirement. Not another Next.js web app take home project.
- tmule 9mo agoYour comments history suggests you’re rather bitter about “nerds” who are likely a few standard deviations smarter than you (Anthropic OG team, Jeff Dean, proof nerds, Linus, …)
- jackblemming 9mo agoAnd they’re all dumber than John von Neumann, who cares?
- margalabargala 9mo agoTransitively, you haven't thought the most thoughts or cared the most about anything, therefore we should disregard what you think and care about?
- jackblemming 9mo agoThe person replying was trying to turn the conversation into some sort of IQ pissing contest. Not sure why, that seems like their own problem. I was reminding them that there is always someone smarter.
- mips_avatar 9mo agoGoing through the assignment now. Man it’s really hard to pack the vectors right
- avaer 9mo agoIt's pretty interesting how close this assignment looks to demoscene [1] golf [2]. [1] https://en.wikipedia.org/wiki/Demoscene https://en.wikipedia.org/wiki/Demoscene [2] https://en.wikipedia.org/wiki/Code_golf https://en.wikipedia.org/wiki/Code_golf It even uses Chrome tracing tools for profiling, which is pretty cool: https://github.com/anthropics/original_performance_takehome/blob/main/problem.py#L154 https://github.com/anthropics/original_performance_takehome/...
- nice_byte 9mo agoit's designed to select for people who can be trusted to manually write ptx :-)
- KeplerBoy 9mo agoperfetto is pretty widely used for such traces, because building a viewer for your traces is a completely avoidable pain.
- wiz21c 9mo agoI was in the demoscene long ago and that kind of optimisation is definitely in the ballpark of what we did: optimize algorithm down to machine code level (and additionally, cheat like hell to make you believe we ran the algorithm for real :-)). But to be honest, I wonder what algorithm they implement. I have read the code for 2 minutes, and it sound like random forest prediction. Anyone knows what the code does ?
- saagarjha 9mo agoIt’s some useless problem like a random tree walk or something like that, the actual algorithm is not particularly important to the problem
- psb217 9mo agoYeah, I assume it was partly chosen since the problem structure provides some convenient hooks for selectively introducing subtle and less subtle inefficiencies in the baseline algorithm that match common optimization patterns.
- greesil 9mo agoThis is a knowledge test of GPU architecture?
- avaer 9mo agoKind of, but not any particular GPU. The machine is fake and simulated: https://github.com/anthropics/original_performance_takehome/blob/main/problem.py#L66 https://github.com/anthropics/original_performance_takehome/... But presumably similar principles apply.
- benreesman 9mo agoIt's a test of polyhedral layout algebra, what NVIDIA calls CuTe and the forthcoming C++ standard calls std::mdspan. This is the general framework for reasoning about correct memory addressing in the presence of arbitrary constraints like those of hardware.
- zeroCalories 9mo agoIt shocks me that anyone supposedly good enough for anthropic would subject themselves to such a one sided waste of time.
- mips_avatar 9mo agoIt’s kind of an interesting problem.
- browningstreet 9mo agoI’ve been sent the Anthropic interview assignments a few times. I’m not a developer so I don’t bother. At least at the time they didn’t seem to have technical but not-dev screenings. Maybe they do now.
- throwa356262 9mo agoCare to elaborate the first part? Did you apply for a position? Did they send you the assignment without prior discussion?
- sealeck 9mo agoWhy is writing code to execute a program using the fewest instructions possible on a virtual machine a waste of time?
- pclmulqdq 9mo agoI generally have a policy of "over 4 hours and I charge for my time." I did this in the 4-hour window, and it was a lot of fun. Much better than many other take-home assignments.
- sureglymop 9mo agoHaving recently learned more about SIMD, PTX and optimization techniques, this is a nice little challenge to learn even more. As a take home assignment though I would have failed as I would have probably taken 2 hours to just sketch out ideas and more on my tablet while reading the code before even changing it.
- forgotpwd16 9mo agoUnless misread, 2 hours isn't the time limit for the candidate to do this but the time Claude eventually needed to outperform best returned solution. Best candidate could've taken 6h~2d to achieve this result.
- fhd2 9mo agoTheir Readme.md is weirdly obsessed with "2 hours": "before Claude Opus 4.5 started doing better than humans given only 2 hours" "Claude Opus 4.5 in a casual Claude Code session, approximately matching the best human performance in 2 hours" "Claude Opus 4.5 after 2 hours in our test-time compute harness" "Claude Sonnet 4.5 after many more than 2 hours of test-time compute" So that does make one wonder where this comes from. Could just be LLM generated with a talking point of "2 hours", models can fall in love with that kind of stuff. "after many more than 2 hours" is a bit of a tell. Would be quite curious to know though. How I usually design take home assignments is: 1. Candidate has several _days_ to complete (usually around a week). 2. I design the task to only _take_ 2-4 hours, informing the candidate about that, but that doesn't mean they can't take longer. The subsequent interview usually reveals if they went overboard or struggled more than expected. But I can easily picture some places sending a candidate the assignment and asking them to hand in their work within two hours. Similar to good old coding competitions.
- alcasa 9mo agoNo the 2 hours is their time limit for candidates. The thing is that you are allowed to use any non-human help for their take homes (open book), so if AI can solve it in below 2 hours, it's not very good at assessing the human.
- tucnak 9mo agoThe snarky writing of "if you beat our best solution, send us an email and MAYBE we think about interviewing you" is really something, innit?
- sourcegrift 9mo agoPride comes before fall thankfully
- altmanaltman 9mo agoits anthrophic. their entire marketing is just being an pompous ass and AI fear mongering.
- ahussain 9mo agoThey wrote: > If you optimize below 1487 cycles, beating Claude Opus 4.5's best performance at launch, email us at performance-recruiting@anthropic.com with your code (and ideally a resume) so we can be appropriately impressed and perhaps discuss interviewing. That doesn’t seem snarky to me. They said if you beat Opus, not their best solution. Removing “perhaps” (i.e. MAYBE) would be worse since that assumes everyone wants to interview at Anthropic. I guess they could have been friendlier: “if you beat X, we’d love to chat!”
- 0x3f 9mo agoI suppose you could interpret it either way, but having dealt with their interview pipeline I'd choose the snark.
- deleted 9mo ago[deleted]
- dude250711 9mo agoYeah, a nerd bypassed HR and showed their true character. They are swimming in easy money.
- lovich 9mo ago
- pvalue005 9mo agoI suspect this was released by Anthropic as a DDOS attack on other AI companies. I prompted 'how do we solve this challenge?' into gemini cli in a cloned repo and it's been running non-stop for 20 minutes :)
- bird0861 9mo agoWhich Gemini model did you use? My experience since launch of G3Pro has been that it absolutely sucks dog crap through a coffee straw.
- Mashimo 9mo ago> sucks dog crap through a coffee straw. That would be impressive.
- anematode 9mo agoNew LLM benchmark incoming? I bet once it's done, people will still say it's not AGI.
- dotancohen 9mo agoWhen they get the hardware capable of that, a different industry will be threatened by AI. The oldest industry.
- cess11 9mo agoTextile?
- nineteen999 9mo agoThe emperor's (empresses?) new textile.
- darepublic 8mo agoSong of Solomon I guess
- kristianpaul 9mo ago“If you optimize below 1487 cycles, beating Claude Opus 4.5's best performance at launch, email us at performance-recruiting@anthropic.com with your code (and ideally a resume) so we can be appropriately impressed and perhaps discuss interviewing.”
- afro88 9mo ago> at launch Does this confirm they actually do knee cap models after the launch period to save money, without telling users?
- mediaman 9mo agoNo, they later updated the harness for this and it subsequently got better scores.
- sevenzero 9mo agoThe company that wanted to simply get away with the thievery of terabytes of intellectual property, what a great place to work at! Not. Anthropic has no shame.
- NitpickLawyer 9mo agoThe writing was on the wall for about half a year (publicly) now. The oAI 2nd place at the atcoder world championship competition was the first one, and I remember it being dismissed at the time. Sakana also got 1st place in another atcoder competition a few weeks ago. Google also released a blog a few months back on gemini 2.5 netting them 1% reduction in training time on real-world tasks by optimising kernels. If the models get a good feedback loop + easy (cheap) verification, they get to bang their tokens against the wall until they find a better solution.
- lostmsu 9mo ago1% doesn't sound like a lot at all.
- _aavaa_ 8mo agoThat depends on how close to the theoretical max you think they are.
- myahio 8mo agoSakana is a grift from what I understand
- NitpickLawyer 8mo agoEh. I'd call them overly enthusiastic :) I know they publish hype-y stuff, they jumped the gun on a few things, I get that. But their recent result was on a "live" contest, and they did share agent traces, so that's likely a legit result.
- cgearhart 8mo agoI think this is the actual “bitter lesson”—the scalable solution (letting LLMs bang against the problem nonstop) will eventually far outperform human effort. There will come a point—whether sooner or later—where this’ll be the expected norm for handling such problems. I think the only question is whether there is any distinction between problems like this (clearly defined with a verifiable outcome) vs the space of all interesting computer programs. (At the moment I think there’s space between them. TBD.)
- Maro 9mo ago> This repo contains a version of Anthropic's original performance take-home, before Claude Opus 4.5 started doing better than humans given only 2 hours. Was the screening format here that this problem was sent out, and candidates had to reply with a solution within 2 hours? Or, are they just saying that the latest frontier coding models do better in 2 hours than human candidates have done in the past in multiple days?
- dhruv3006 9mo agoI wonder if OpenAI follows suit.
- rvz 9mo agoThey should.
- tayo42 9mo agoI wonder if the Ai is doing anything novel? Or if it's like a brute force search of applying all types of existing optimizations that already exist and have been written about.
- bytesandbits 9mo agoHaving done a bunch of take home for big (and small) AI labs during interviews, this is the 2nd most interesting one I have seen so far.
- petters 9mo agoAnd the answer to the obvious follow-up question is...?
- reader9274 9mo agofries
- deleted 9mo ago[deleted]
- mrklol 9mo agoMilk before cereals
- darkwater 9mo agoMaybe it's under NDA :)
- kevthecoder 9mo ago42
- bytesandbits 8mo ago
- Incipient 9mo ago>so we can be appropriately impressed and perhaps discuss interviewing. Something comes across really badly here for me. Some weird mix of bragging, mocking, with a hint of aloof. I feel these top end companies like the smell of their own farts and would be an insufferable place to work. This does nothing but reinforce it for some reason.
- sponnath 9mo agoI have to agree. It's off-putting to me too. I'm impressed by the performance of their models on this take-home but I'm not impressed at their (perhaps unintentional) derision of human programmers.
- qbane 9mo agoRemember: It is a company that keep saying how much production code can be written by AI in xx years, but at the same time recruiting new engineers.
- yodsanklai 8mo agoThanks for noticing this. I got the same feeling when reading this. It may not sound like much, and it doesn't mean it's an insufferable place to work, but it's a hint it might be. Rant: On a similar note, I recently saw a post on Linkedin from Mistral, where they were bragging to recruit candidates from very specific schools. That sounded very pretentious (and also an HR mistake on several levels IMHO).
- languid-photic 9mo agoNaively tested a set of agents on this task. Each ran the same spec headlessly in their native harness (one shot). Results: Agent Cycles Time ───────────────────────────────────────────── gpt-5-2 2,124 16m claude-opus-4-5-20251101 4,973 1h 2m gpt-5-1-codex-max-xhigh 5,402 34m gpt-5-codex 5,486 7m gpt-5-1-codex 12,453 8m gpt-5-2-codex 12,905 6m gpt-5-1-codex-mini 17,480 7m claude-sonnet-4-5-20250929 21,054 10m claude-haiku-4-5-20251001 147,734 9m gemini-3-pro-preview 147,734 3m gpt-5-2-codex-xhigh 147,734 25m gpt-5-2-xhigh 147,734 34m Clearly none beat Anthropic's target, but gpt-5-2 did slightly better in much less time than "Claude Opus 4 after many hours in the test-time compute harness".
- ponyous 9mo agoVery interesting thanks! I wonder what would happen if you kept running Gemini in a loop for a while. Considering how much faster it ended it seems like there is a lot more potential.
- giancarlostoro 9mo agoI do wonder how Grok would compare, specifically their Claude Code Fast model.
- forgotpwd16 9mo agoCould you make a repo with solutions given by each model inside a dir/branch for comparison?
- lbreakjai 9mo agoI consider myself rather smart and good at what I do. It's nice to have a look at problems like these once in a while, to remind myself of how little I know, and how much closer I am to the average than to the top.
- apsurd 9mo agodisagree. nobody has a monopoly on what metric makes someone good. I don't understand all this leet code optimization. actually i do understand it, but it's a game that will attract game optimizers. the hot take is, there are other games.
- sevenzero 9mo agoAlso leetcode does not really provide insight into ones ability to design business solutions. Whether it be system design, just some small feature implementation or communication skills within a team. Its just optimizers jerking each other off on some cryptic problems 99.999999999% of developers will never see in real life. Maybe it would've been useful like 30 years ago, but all commonly used languages have all these fancy algorithms baked into their stdlib, why would I ever have to implement them myself?
- thorncorona 9mo agoOr more likely, the commonality is how you're applying your software skills? In every other field it's helpful to understand the basics. I don't think software is the exception here.
- sevenzero 9mo agoUnderstanding basics is very different to being able to memorize algorithms. I really dont see why I'd ever have to implement stuff like quicksort myself somewhere. Yes I know what recursion is, yes I know what quick sort is, so if I ever need it I know what to look for. Which was good enough throughout my career.
- 9mo ago
- OhNoNotAgain_99 9mo ago[dead]
- spencerflem 9mo agoOh wow it’s by Tristan Hume, still remember you from EyeLike!
- Graziano_M 8mo agoI recognized the name and dug around too. I played DEFCON CTF with him back in the day!
- tmp-127853716 9mo ago[flagged]
- falloutx 9mo agoWell working under someone who keeps insisting Software engineering is dead sounds like a toxic work environment.
- woof 9mo ago"1) Python is unreadable." Would you prefer C or C++? "2) AI companies are content with slop and do not even bother with clear problem statements." It's a filter. If you don't get the problem, you'll waste their time. "3) LOC and appearance matter, not goals or correctness." The task was goal+correctness. "4) Anthropic must be a horrible place to work at." Depends on what you do. For this position it's probably one of the best companies to work at.
- am17an 8mo ago1) Python is unreadable." Would you prefer C or C++? > Unironically, yes. Unless I never plan to look at that code again
- tap12783487 8mo agoIt is a filter for academics who write horrible Python code and feel smart, yes. I think they also have open positions for stealing other people's code and DDoS-ing other people's websites.
- demirbey05 9mo agoIt's showcase more than being take home assignment. I couldnt understand what the task is ,only performance comparisons between their LLM
- measurablefunc 9mo agoThe task is ill-defined.
- saagarjha 9mo agoYou make it faster
- measurablefunc 9mo agoFewer instructions doesn't mean it's faster. It can be faster but it's not guaranteed in general. Obvious counterexample is single threaded vs multi-threaded code. Single threaded code will have fewer instructions but won't necessarily be faster.
- saagarjha 9mo agoIt does in this case; you can read the assignment to see that it is all single-threaded
- measurablefunc 8mo agoI read it, you're mistaken.
- saagarjha 8mo agoI did the assignment my guy
- measurablefunc 8mo ago
- saagarjha 9mo agoOh, this was fun! If you like performance puzzles you should really do it. Actually I might go back and see if I can improve on it this weekend…
- nottorp 9mo agoIs it "write 20 astroturfing but somewhat believable posts about the merits of "AI" and how it is going to replace humans"?
- svilen_dobrev 9mo agoif anyone is interested to try their agent-fu, here's some more-real-world rabbit-hole i went optimizing in 2024. Note this is now dead project, noone's using it, and probably same for the original. i managed to get it 2x-4x faster than original, took me several days then. btw There are some 10x optimizations possible but they break few edge cases, so not entirely correct. https://github.com/svilendobrev/transit-python3 https://github.com/svilendobrev/transit-python3
- torginus 9mo agoAre you allowed to change the instruction sequence? I see some optimization opportunities - it'd be obviously the correct thing to do an optimizing compiler, but considering the time allotted, Id guess you could hand-optimize it, but that feels like cheating.
- saagarjha 9mo agoYes, in fact this will be one of the first things you will want to do.
- pshirshov 9mo agoYet Claude is the only agent which deadlocks (blocks in GC forever) after an hour of activity.
- abra0 9mo agoThis is a really fun problem! I suggest anyone who likes optimization in a very broad sense to try their hand at it. Might be the most fun I've had while interviewing. I had to spend a week-worth of evenings on it to fully scratch the itch, and I managed to get 1112 cycles. But that was mostly manual, before the current crop of agentic models (clopus 4.5, gpt5.2). I wonder how far you can RalphWiggum it!
- lukah 9mo agoI've never heard AI-assisted coding referred to as "RalphWiggum"ing a problem, and now I will have to use that always. Thank you.
- usgroup 9mo agohttps://awesomeclaude.ai/ralph-wiggum https://awesomeclaude.ai/ralph-wiggum
- clocksmith 8mo agoDid you get an offer?
- kartibbb 9mo ago[flagged]
- kartibbb 9mo ago[flagged]
- karmasimida 9mo agoI am able to beat this 1487 benchmark by switching between LLMs, doesn't seem that hard lol. Albeit, I do not fully understand what the solution is, loll
- lostmsu 9mo agoYeah, GPT 5.2 on high got down to 1293 on the 5th try (about 32mins).
- deleted 9mo ago[deleted]
- alexpadula 9mo agoLooks rather fun!
- FriendlyMike 9mo agoThey should just have you create a problem that can't be solved by an llm in two hours. That's the real problem here
- OisinMoran 8mo ago"You have 1 minute to design a maze that takes 2 minutes to solve"
- ec109685 8mo agoSolvable in more than 2 but not less than 2 would be the real trick.
- potato-peeler 9mo agoWhat does clock cycles mean? Don’t think they are referring to the cpu clock?
- mrdootdoot 9mo ago“In English, Data”
- eisbaw 8mo agoI got to 1364 cycles for now, semi-manually: Using design space exploration organized via backlog.md project, and then recombination from that. 20 agents in parallel. Asked to generate drawio for the winner so I can grok it more easily, then I gave feedback. Edit: 1121 cycles
- eisbaw 8mo ago1023 cycles
- karmasimida 8mo agoSame just make it a survival game
- mannykannot 8mo agoI beat the target by deleting the parts that were causing the cycle count to be too high. /s
- eisbaw 8mo agosubmit and see if Anthropic accepts it
- fabian4 8mo ago[flagged]
- tap12783487 8mo ago[flagged]
- epiccoleman 8mo agoIt definitely bears all the LLM hallmarks we've come to know. emdash, the "this isn't X. it's Y" structure - and then, to cap it off, a single pithy sentence to end it.
- nostrademons 8mo agoAlso bears all the hallmarks of an ordinary post (by someone fairly educated) on the Internet. This would make sense, because LLMs were trained on lots of ordinary posts on the Internet, plus a fair number of textbooks and scientific papers.
- epiccoleman 8mo agoThe — character is the biggest cause of suspicion. It's difficult to type manually so most people - myself included - substitute the easily typed hyphen. I know real people do sometimes use it, but it's a smell.
- nostrademons 8mo agoI think some software will automatically substitute "smart quotes" for regular quotes and an em-dash for a double hyphen -- I know MS Word used to do this. Curious if any browsers do. This comment was typed in Brave, which doesn't appear to, but I didn't check if Chrome or IE or Opera does.
- menaerus 8mo agoThe comment was not wrong though so I am not sure I understand if flagging it for the sole "it was most likely written by the use of AI" reason is completely valid.
- game_the0ry 8mo ago> If you optimize below 1487 cycles, beating Claude Opus 4.5's best performance at launch, email us at performance-recruiting@anthropic.com with your code (and ideally a resume) so we can be appropriately impressed and perhaps discuss interviewing. This is an interesting way to recruit. Much better than standard 2 leetcode medium/hard questions in 45 mins.
- paxys 8mo agoThis is simply to enter the recruiting pipeline. once you're in you will do the same leetcode interviews as everyone else.
- alt227 8mo agoYou would hope that if you manage to beat their engineers best optimisations at launch, then you would leapfrog a certain amount of the initial stages. Then again, this may just be a way to get free ideas at optimising their product from outside the box.
- benlivengood 8mo agoOne could use any number of LLMs on a take-home problem so in-person interviews are a must.
- legel 8mo agoOne could use any number of LLMs on real-world problems. Why are we still interviewing like its 1999?
- game_the0ry 8mo agoOld habits die hard. And engineers are pretty lazy when it comes to interviews, so just throwing the same leetcode problem into coder pad in every interview makes interviews easier for the person doing the interview.
- sublimefire 8mo agoDid a bit of soul searching and manually optimised to 1087 but I give up. What is the number we are chasing here? IMO I would not join a company giving such a vague problem because you can feel really bad afterwards, especially if this does not open a door to the next stage of the interview. As an alternative we could all instead focus on a real kernel and improve it :)
- trishume 8mo agoAuthor of the take-home here: That's quite a good cycle count, substantially better than Claude's, you should email it to performance-recruiting@anthropic.com.
- deleted 8mo ago[deleted]
- NightBlossom 8mo agoI could only cut it down to 41 cycles.
- throwaway0123_5 8mo ago> Claude Opus 4.5 in a casual Claude Code session, approximately matching the best human performance in 2 hours Is this saying that Claude matched the best human performance, where the human had two hours? I think that is the correct reading, but I'm not certain they don't mean that Claude had two hours, and matched the best human performance where the human had an arbitrary amount of time. The former is impressive but the later would be even more so.
- nine_k 8mo agoThis is a kind of task that's best solved by possibly spending more than the allocated 2 hours on it, once any obvious low-hanging fruit is picked. An optimization task is what a machine does best. So the real problem would be to construct a machine that would be able to run the optimization. A right optimization framework that results from the effort could also efficiently solve many more similar problems in the future. I understand that this test is intended to somehow test the raw brianpower, the ability to tackle an unfamiliar and complicated domain, and to work under stress. But I hope it's not representative of the actual working conditions at Anthropic. It's like asking a candidate to play a Quake deathmatch when hiring to a special forces assault squad.
- saagarjha 8mo ago> So the real problem would be to construct a machine that would be able to run the optimization. This is a valid way to solve the problem.
- pickpocket 8mo agoi cleared this one but didn't clear the follow up interview that was way easier than this
- pickpocket 8mo agoI cleared this assignment but did not clear the follow up interview that was way easier than this. So I gave up on tech interviews in general, stayed where I was.
- arsl16 8mo agoI got this but I am an embedded SWE, might not be my cup of tea
- amirhirsch 8mo agoI'm at 1137 with one hour with opus now... Pipelined vectorized hash, speculation, static code for each stage, epilogues and prologues for each stage-to-stage... I think I'm going to get sub 900 since i just realized i can in-parallel compute whether stage 5 of the hash is odd just by looking at bits 16 and 0 of stage 4 with less delay.....
- lalaland1125 8mo agoHow do you avoid the load bottleneck?
- amirhirsch 8mo agotake advantage of index collisions, optimizing round 0 and 11, speculative pre-loading, and the early branch predictor (which now I am doing looking at bits output at stage 3)
- lzhou 8mo agoit's actually pretty funny since opus will suggest both of these with enough prying (though with a single-prompt it might not try it).
- amirhirsch 8mo ago====================================================================== BROADCAST LOAD SCHEDULE ====================================================================== Round | Unique | Load Strategy ------|--------|------------------------------------------ 0 | 1 | 1 broadcast → all 256 items 1 | 2 | 2 broadcasts → groups 2 | 4 | 4 broadcasts → groups 3 | 8 | 8 broadcasts → groups 4 | 16 | 16 broadcasts → groups 5 | 32 | 32 broadcasts → groups 6 | 63 | 63 loads (sparse, use indirection) 7 | 108 | 108 loads (sparse, use indirection) 8 | 159 | 159 loads (sparse, use indirection) 9 | 191 | 191 loads (sparse, use indirection) 10 | 224 | 224 loads (sparse, use indirection) 11 | 1 | 1 broadcast → all 256 items 12 | 2 | 2 broadcasts → groups 13 | 4 | 4 broadcasts → groups 14 | 8 | 8 broadcasts → groups 15 | 16 | 16 broadcasts → groups Total loads with grouping: 839 Total loads naive: 4096 Load reduction: 4.9x
- seamossfet 8mo agoI'm getting flashbacks from my computer engineering curriculum. Probably the first place I'd start is replacing comparison operators on the ALU with binary arithmetic since it's much faster than branch logic. Next would probably be changing the `step` function from brute iterators on the instructions to something closer to a Btree? Then maybe a sparse set for the memory management if we're going to do a lot of iterations over the flat memory like this.
- deleted 8mo ago[deleted]
- htrp 8mo agoIdle side note: surprised that https://github.com/anthropic https://github.com/anthropic is just some random dude in Australia
- LarsKrimi 8mo agoI liked the core challenge. Finding the balance of ALU and VALU, but I think that the problem with the load bandwidth could lead to problems Like optimizing for people who assume the start indices always will be zero. I am close to 100% sure that's required to get below 2096 total loads but it's just not fun If it however had some kind of dynamic vector lane rotate that could have been way more interesting
- NightBlossom 8mo agoI just withdrew my application over this test. It forces an engineering anti-pattern: requiring runtime calculation for static data (effectively banning O(1) pre-computation). When I pointed out this contradiction via email, they ignored me completely and instead silently patched the README to retroactively enforce the rule. It’s not just a bad test; it’s a massive red flag for their engineering culture. They wasted candidates' time on a "guess the hidden artificial constraint" game rather than evaluating real optimization skills.
- hackern3972 8mo agoThis isn't the gotcha moment you think it is. Storing the result on disk is some stupid "erm achkually" type solution that goes against the spirit of the optimization problem. They want to see how you handle low level optimizations, not get tripped over some question semantics.
- NightBlossom 8mo agoYou are missing the point. This isn't "storing result on disk." In high-performance engineering, if the input is static and known at build time, the only correct optimization is pre-computation. I didn't simply "skip" the problem. I implemented a compiler that solves the problem entirely at build time, resulting in O(0) runtime execution. Here is the actual "Theorem" I implemented in my solution. If a test penalizes this approach because it "goes against the spirit," then the test is fundamentally testing for inefficiency. """ Theorem 1 (Null Execution): Let P: M → M be a program with postcondition φ(M). If ∃M' s.t. φ(M') ∧ M ≅ M', then T(P) = 0. Complexity: O(n) compile-time, O(0) runtime """ If they wanted to test runtime loop optimizations, they should have made the inputs dynamic.
- SinghCoder 8mo agowhy is their github handle anthropics and not anthropic :D
- mayankd 8mo agoProblem solving is eternal!
- arsl16 8mo agoFellas should I even attempt it? I got it recently and lets say it brings back memories of computer architecture class.
- yasmineroy33 8mo ago[dead]