4 ms·
While I get the MBA-speak of lines-of-code that AI is now able to accomplish, it does make me think about their highly-curated internal codebase that makes them
by fzysingularity 2y ago
While I get the MBA-speak of lines-of-code that AI is now able to accomplish, it does make me think about their highly-curated internal codebase that makes them well placed to potentially get to 50% AI-generated code.
One common misconception is that all LLMs are the same. The models are trained the same, but trained on wildly different datasets. Google, and more specifically the Google codebase is arguably one of the most curated, and iterated on datasets in existence. This is a massive lever for Google to train their internal code-gen models, that realistically could easily replace any entry-level or junior developer.
- Code review is another dimension of the process of maintaining a codebase that we can expect huge improvements with LLMs. The highly-curated commentary on existing code / flawed diff / corrected diff that Google possesses give them an opportunity to build a whole set of new internal tools / infra that's extremely tailored to their own coding standard / culture.
- morkalork 2y agoIs the public gemini code gen LLM trained on their internal repo? I wonder if one could get it to cough up propriety code with the right prompt.
- p1esk 2y agoI’m curious if Microsoft lets OpenAI train on GH private repos.
- happyopossum 2y ago> Is the public gemini code gen LLM trained on their internal repo? Nope
- bqmjjx0kac 2y ago> that realistically could easily replace any entry-level or junior developer. This is a massive, unsubstantiated leap.
- risyachka 2y agoThe issue is it doesn't really replace junior dev. You become one - as you have to babysit it all the time, check every line of code, and beg it to make it work. In many cases it is counterproductive
- throwaway106382 2y agoI’d take pair programming with a junior over a GPT bot any day.
- neaanopri 2y agoI'd take coding by own damn self over either a junior or a gpt bot
- unit149 2y agoPhilosophically, these models are akin to scholars prior to differentiation during their course of study. Throttling data, depending on one's course of study, and this shifting of the period in history step-by-step. Either it's a tit-for-tat manner of exchange that the junior developer is engaged in, when overseeing every edit that an LLM has modified, or I'd assume that there are in-built methods of garbage collection, that another LLM evaluating a hash function partly identifying a block of tokenized internal code would be responsible for.
- jtbetz22 2y ago> Google codebase is arguably one of the most curated, and iterated on datasets in existence I spent 12 years of my career in the Google codebase. This assertion is technically correct in that google3 has been around for 20 years, and all code gets reviewed, but the implication that Google's codebase is a high-quality training set is not consistent with my experience.