8 ms·
> I genuinely believe Groq is a fraud. There is no way their private inference cloud has positive gross margins. > Llama 3.1 405B can currently replace junior
by C-programmer 2y ago
> I genuinely believe Groq is a fraud. There is no way their private inference cloud has positive gross margins.
> Llama 3.1 405B can currently replace junior engineers
I'd like more exposition on these claims.
- refulgentis 2y agoToday, I wrote a full YouTube subtitle downloader in Dart. 52 minutes from starting to google anything about it, to full implementation and tests, custom formatting any of the 5 obscure formats it could be in to my exact whims. Full coverage of any validation errors via mock network responses. I then wrote a web AudioWorklet for playing PCM in 3 minutes, which complied to the same interface as my Mac/iOS/Android versions, ex. Setting sample rate, feedback callback, etc. I have no idea what an AudioWorklet is. Two days ago, I stubbed out my implementation of OpenAI's web socket based realtime API, 1400 LOC over 2 days, mostly by hand while grokking and testing the API. In 32 minutes, I had a brand spanking new batch of code, clean, event-based architecture, 86% test coverage. 1.8 KLOC with tests. In all of these cases, most I needed to do was drop in code files and say, nope wrong a couple times to sonnet, and say "why are you violating my service contract and only providing an example solution" to o1. Not llama 3.1 405B specifically, I haven't gone to the trouble of running it, but things turned some sort of significant corner over the last 3 months, between o1 and Sonnet 3.5. Mistakes are rare. Believable 405B is on that scale, IIRC it went punch for punch with the original 3.5 Sonnet. But I find it hard to believe a Google L3, and third of L4s, (read: new hires, or survived 3 years) are that productive and sending code out for review at a 1/5th of that volume, much less on demand. So insane-sounding? Yes. Out there? Probably, I work for myself now. I don't have to have a complex negotiation with my boss on what I can use and how. And I only saw this starting ~2 weeks ago, with full o1 release. Wrong? Shill? Dilletante? No. I'm still digesting it myself. But it's real.
- nightski 2y agoMost software is not one off little utilities/scripts, greenfield small projects, etc... That's where LLMs excel, when you don't have much context and can regurgitate solutions. It's less to do with junior/senior/etc.. and more to do with the types of problems you are tackling.
- spiderfarmer 2y agoMost software is simple HTML.
- mhh__ 2y agoThis isn't where the leverage is though
- spiderfarmer 2y agoNo, but AI will replace a lot of web developers.
- refulgentis 2y agoThis is a 30KLOC 6 platform flutter app that's, in this user story, doing VOIP audio, running 3 audio models on-device, including in your browser. A near-replica of the Google Assistant audio pipeline, except all on-device. It's a real system, not kindergarten "look at the React Claude Artifacts did, the button makes a POST request!" The 1500 loc websocket / session management code it refactored and tested touches on nearly every part of the system (i.e. persisting messages, placing search requests, placing network requests to run a chat flow) Also, it's worth saying this bit a bit louder: the "just throwing files in" I mention is key. With that, the quality you observed being in reverse is the distinction: with o1 thinking, and whatever Sonnet's magic is, there's a higher payoff from working a larger codebase. For example, here, it knew exactly what to do for the web because it already saw the patterns iOS/Android/macOS shared. The bend I saw in the curve came from being ultra lazy one night and seeing what would happen if it just had all the darn files.
- achierius 2y agoI've definitely noticed the opposite on larger codebases. It's able to do magical things on smaller ones but really starts to fall apart as I scale up.
- noch 2y ago> This is a 30KLOC 6 platform flutter app […] It's a real system, not kindergarten "look at the React Claude Artifacts did, the button makes a POST request!" This is powerful and significant but I think we need to ground ourselves on what a skilled programmer means when he talks about solving problems. That is, honestly ask: What is the level of skill a programmer requires to build what you've described? Mostly, imo, the difficulties in building it are in platform and API details not in any fundamental engineering problem that has to be solved. What makes web and Android programming so annoying is all the abstractions and frameworks and cruft that you end up having to navigate. Once you've navigated it, you haven't really solved anything, you've just dealt with obstacles other programmers have put in your way. The solutions are mostly boilerplate-like and the code I write is glue. I think the definition of "junior engineer" or "simple app" will be defined by what LLMs can produce and so, in a way, unfortunately, the goal posts and skill ceiling will keep shifting. On the other, hand, say we watch a presentation by the Lead Programmer at Naughty Dog, "Parallelizing the Naughty Dog Engine Using Fibers"[^0] and ask the same questions: what level of skill is required to solve the problems he's describing (solutions worth millions of dollars because his product has to sell that much to have good ROI): "I have a million LOC game engine for which I need to make a scheduler with no memory management for multithreaded job synchronization for the PS4." A lot of these guys, if you've talked to them, are often frustrated that LLMs simply can't help them make headway with, or debug, these hard problems where novel hardware-constrained solutions are needed. --- [^0]: https://www.youtube.com/watch?v=HIVBhKj7gQU https://www.youtube.com/watch?v=HIVBhKj7gQU
- xbmcuser 2y agoI agree with you that they are improving not being a programer I can't tell if the code has improved but as user that uses chat gpt or Google gemini to build scripts or trading view indicators. I am seeing some big improvements and many times wording it better detail and restricting it from going of tangent results in working code.
- tonetegeatinst 2y agoYouTube dlp has a subtitle option. To quote the documentation: "--write-sub Write subtitle file --write-auto-sub Write automatically generated subtitle file (YouTube only) --all-subs Download all the available subtitles of the video --list-subs List all available subtitles for the video --sub-format FORMAT Subtitle format, accepts formats preference, for example: "srt" or "ass/srt/best" --sub-lang LANGS Languages of the subtitles to download (optional) separated by commas, use --list-subs for available language tags"
- jiggawatts 2y agoSomething I noticed is that as the threshold for doing something like "write software to do X" decreases, the tendency for people to go search for an existing product and download it tends to zero. There is a point where in some sense it is less effort to just write the thing yourself. This is the argument against micro-libraries as seen in NPM as well, but my point is that the threshold of complexity for "write it yourself" instead of reaching for a premade thing changes over time. As languages, compilers, tab-complete, refactoring, and AI assistance get better and better, eventually we'll reach a point where the human race as a whole will be spitting out code at an unimaginable rate.
- zkry 2y agoThis is a key point and one of the reasons why I think LLMs will fall short of expectation. Take the saying "Code is a liability," and the fact that with LLMs, you are able to create so much more code than you normally would: The logical conclusion is that projects will balloon with code pushing LLMs to their limit, and this massive amount is going to contain more bugs and be more costly to maintain. Anecdotally, supposedly most devs are using some form of AI for writing code, and the software I use isn't magically getting better (I'm not seeing an increased rate of features or less buggy software).
- az226 2y agoShow the code
- JTyQZSnP3cQGa8B 2y agoThat’s why I’m annoyed: they never show the code.
- refulgentis 2y agoYou gotta lower your expectations. :) I didn't see the comment till now, OP made it at 3:30 AM my time. Here you go - https://pastebin.com/8zdMDEnG https://pastebin.com/8zdMDEnG
- IshKebab 2y agoI think they meant the code that the LLM generated.
- refulgentis 2y agoRight: that is it.
- IshKebab 2y agoThis has clearly been heavily edited by humans.
- refulgentis 2y agoI'm happy to provide whatever you ask for. With utmost deference, I'm sure you didn't mean anything by it and were just rushed, but just in case...I'd just ask that you'd engage with charity[^3] and clarity[^2] :) I'd also like to point out[^1] --- meaning, I gave it my original code, in toto. So of course you'll see ex. comments. Not sure what else contributed to your analysis, that's where some clarity could help me, help you. [^1](https://news.ycombinator.com/item?id=42421900 https://news.ycombinator.com/item?id=42421900) "I stubbed out my implementation of OpenAI's web socket based realtime API, 1400 LOC over 2 days, mostly by hand while grokking and testing the API. In 32 minutes, I had a brand spanking new batch of code, clean, event-based architecture, 86% test coverage. 1.8 KLOC with tests." [^2](https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html) "Be kind...Converse curiously; don't cross-examine." [^3](https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html) "Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith."
- lz400 2y agoI don't understand what you guys are doing. For me sonnet is great when I'm starting with a framework or project but as soon as I start doing complicated things it's just wrong all the time. Subtly wrong, which is much worse because it looks correct, but wrong.
- idiocache 2y agoCan you briefly describe your work flow? Are you exchanging information with Sonnet in your IDE?
- deleted 2y ago[deleted]
- Workaccount2 2y agoNot Llama but with Sonnet and O1 I wrote a bespoke android app for my company in about 8 hours of work. Once I polish it a bit (make a prettier UI), I'm pretty sure I could sell it to other companies doing our kind of work. I am not a programmer, and I know C and Python at about a 1 day crash course level (not much at all). However with sonnent I was able to be handheld all the way from downloading android studio to a functional app written in kotlin, that is now being used by employees on the floor. People can keep telling themselves that LLMs are useless or maybe just helpful for quickly spewing boilerplate code, but I would heed the warning that this tech is only going to improve and already helping people forgo SWE's very seriously. Sears thought the internet was a cute party trick, and that obviously print catalogs were there to stay.
- kortilla 2y agoThis is meaningless without talking about the capabilities of the app. I’ve seen examples of this before where non-programmers come up with something using an LLM that could just be a webpage with camera access and some javascript
- deleted 2y ago[deleted]
- throwawaymaths 2y ago> groq i went to a groq event and one of their engineers told me they were running 7 racks!! of compute per (70b?) model. that was last year so my memory could be fuzzy. iirc, groq used to be making resnet-500? chips? the only way such an impressive setup makes any kind of sense (my guess) would be they bought a bunch of resnet chips way back when and now they are trying to square peg in round hole that sunk cost as part of a fake it till you make it phase. they certainly have enough funding to scrap it all and do better... the question is if they will (and why they haven't been able to yet)
- wmf 2y agoYes, Groq requires hundreds or thousands of chips to load an LLM because they didn't predict that LLMs would get as big as they are. The second generation chip can't come soon enough for them.
- arisAlexis 2y agoalso, he believes cerebras is shit but also cerebras runs llama the most efficiently and at top speed. biased ^10
- throwawaymaths 2y agoSo I interacted with people at cerebras at a tradeshow and it seems like you have to have extremely advanced cooling to keep that thing working. IIRC the user agreement says "you can't turn it off or else the warranty is void". With the way their chip is designed, I would be strongly worried that the giant chip has warping issues, for example, when certain cores are dark and the thermal generation is uneven (or, if it gets shut down on accident while in the middle of inferencing an LLM). There may even be chip-to-chip variation depending on which cores got dq'd based on their on-the-spot testing. Already through the gapevine I'm hearing that H100s and B100s have to be replaced more often.... than you'd want? I suspect people are mum about it otherwise they might lose sweetheart discounts from nvidia. I can't imagine that cerebras, even with their extreme engineering of their cooling system, have truly solved cooling in a way that isn't a pain in the ass (otherwise they wouldn't have the clause?) and if I were building a datacenter I would be very worried about having to do annoying and capital intensive replacements.
- arisAlexis 2y agoI have nowhere near the knowledge required to say yes or no to your argument. My point is that the guy that wrote the article is shilling a pre-ipo company whole fuding the competitors which is really surprising to get that many upvotes.
- throwawaymaths 2y agomaybe but it shouldn't be surprising. cerebras's designs were born ~2014 ~pre transformers, and the megachips were initially targetted for hpc workloads. it was definitely "solution looking for a problem" back then and now is drifting into square peg in round hole territory now (see sibling comment about groq). I'm surprised they have gotten their raw perf as high as they have by now.