13 ms·
Hi, I lead the teams responsible for our internal developer tools, including AI features. We work very closely with Google DeepMind to adapt Gemini models for G
by ntulpule 2y ago
Hi, I lead the teams responsible for our internal developer tools, including AI features. We work very closely with Google DeepMind to adapt Gemini models for Google-scale coding and other Software Engineering usecases. Google has a unique, massive monorepo which poses a lot of fun challenges when it comes to deploying AI capabilities at scale.
1. We take a lot of care to make sure the AI recommendations are safe and have a high quality bar (regular monitoring, code provenance tracking, adversarial testing, and more).
2. We also do regular A/B tests and randomized control trials to ensure these features are improving SWE productivity and throughput.
3. We see similar efficiencies across all programming languages and frameworks used internally at Google and engineers across all tenure and experience cohorts show similar gain in productivity.
You can read more on our approach here:
https://research.google/blog/ai-in-software-engineering-at-google-progress-and-the-path-ahead/ https://research.google/blog/ai-in-software-engineering-at-g...
- reverius42 2y agoTo me the most interesting part of this is the claim that you can accurately and meaningfully measure software engineering productivity.
- valval 2y agoYou can come up with measures for it and then watch them, that’s for sure.
- lr1970 2y agowhen metric becomes the target it ceases to be a good metric. when discovered how it works developers will type the first character immediately after opening the log. edit: typo
- joshuamorton 2y agoOnly if the developer is being judged on the thing. If the tool is being judged on the thing, it's much less relevant. That is, I, personally, am not measured on how much AI generated code I create, and while the number is non-zero, I can't tell you what it is because I don't care and don't have any incentive to care. And I'm someone who is personally fairly bearish on the value of LLM-based codegen/autocomplete.
- valval 2y agoThat was my point, veiled in an attempt to be cute.
- ozim 2y agoYou can - but not on the level of a single developer and you cannot use those measures to manage productivity of a specific dev. For teams you can measure meaningful outcomes and improve team metrics. You shouldn’t really compare teams but it also is possible if you know what teams are doing. If you are some disconnected manager that thinks he can make decisions or improvements reducing things to single numbers - yeah that’s not possible.
- deely3 2y ago> For teams you can measure meaningful outcomes and improve team metrics. How? Which metrics?
- ozim 2y agoThat is what we pay managers -to figure out- for. They should find out which and how by knowing the team, familiarity with domain knowledge, understanding company dynamics, understanding customer, understanding market dynamics.
- seanmcdirmid 2y agoThat's basically a non-answer. Measuring "productivity" is a well known hard problem, and managers haven't really figured it out...
- yorwba 2y agoEconomists are generally fine with defining productivity as the ratio of aggregate outputs to aggregate inputs. Measuring it is not the hard part. The hard part is doing anything about it. If you can't attribute specific outputs to specific inputs, you don't know how to change inputs to maximize outputs. That's what managers need to do, but of course they're often just guessing.
- seanmcdirmid 2y agoMeasuring human productivity is hard since we can't quantify output beyond silly metrics like lines of code written or amount of time speaking during meetings. Maybe if we were hunter/gatherers we could measure it by amount of animals killed.
- UncleMeat 2y agoAt scale you can do this in a bunch of interesting ways. For example, you could measure "amount of time between opening a crash log and writing the first character of a new change" across 10,000s of engineers. Yes, each individual data point is highly messy. Alice might start coding as a means of investigation. Bob might like to think about the crash over dinner. Carol might get a really hard bug while David gets a really easy one. But at scale you can see how changes in the tools change this metric. None of this works to evaluate individuals or even teams. But it can be effective at evaluating tools.
- fwip 2y agoThere's lots of stuff you can measure. It's not clear whether any of it is correlated with productivity. To use your example, a user with an LLM might say "LLM please fix this" as a first line of action, drastically improving this metric, even if it ruins your overall productivity.
- zac23or 2y agoI knew a superstar developer who worked on reports in an SQL tool. In the company metrics, the developer scored 420 points per month, the second developer scored 60 points. “Please learn how to score more points from the leader”, the boss would say. The superstar developer’s secret… he would send blank reports to clients (who would only realize it days later, and someone else would end up redoing the report), and he would score many more points without doing anything. I’ve seen this happen a lot in many different companies. As a friend of mine used to say, “it’s very rare, but it happens all the time.” I have no doubt that AI can help developers, but I don’t trust the metrics of the CEO or people who work on AI, because they are too involved in the subject.
- svieira 2y ago> When people are pressured to meet a target value there are three ways they can proceed: 1) They can work to improve the system 2) They can distort the system 3) Or they can distort the data https://commoncog.com/goodharts-law-not-useful/ https://commoncog.com/goodharts-law-not-useful/
- torginus 2y agoHonestly I doubt he got away with this for long (unless it was a very dysfunctional org). Being the best gets you noticed (in a good way), and screwing people over gets you noticed too (in a bad way), the combination of the two paints a target on your back.
- GeoAtreides 2y ago> Being the best gets you noticed (in a good way), and screwing people over gets you noticed too (in a bad way), ah, to be young again...
- torginus 2y agoI don't know what you're implying - I have had a few instances in my career when I went above and beyond and while I didn't receive too much praise for my efforts directly, after a while I noticed people who had no business knowing who I was, actually did. Now, I was really bad at capitalizing on it, so nothing much came of it, but still, there are some positive things that higher-ups do notice.
- fhdsgbbcaA 2y agoI’ve been thinking a lot lately about how an LLM trained in really high quality code would perform. I’m far from impressed with the output of GPT/Claude, all they’ve done is weight against stack overflow - which is still low quality code relative to Google. What is probability Google makes this a real product, or is it too likely to autocomplete trade secrets?
- gamesetmath 2y ago[flagged]
- hitradostava 2y agoI'm continually surprised by the amount of negativity that accompanies these sort of statements. The direction of travel is very clear - LLM based systems will be writing more and more code at all companies. I don't think this is a bad thing - if this can be accompanied by an increase in software quality, which is possible. Right now its very hit and miss and everyone has examples of LLMs producing buggy or ridiculous code. But once the tooling improves to: 1. align produced code better to existing patterns and architecture 2. fix the feedback loop - with TDD, other LLM agents reviewing code, feeding in compile errors, letting other LLM agents interact with the produced code, etc. Then we will definitely start seeing more and more code produced by LLMs. Don't look at the state of the art not, look at the direction of travel.
- latexr 2y ago> if this can be accompanied by an increase in software quality That’s a huge “if”, and by your own admission not what’s happening now. > other LLM agents reviewing code, feeding in compile errors, letting other LLM agents interact with the produced code, etc. What a stupid future. Machines which make errors being “corrected” by machines which make errors in a death spiral. An unbelievable waste of figurative and literal energy. > Then we will definitely start seeing more and more code produced by LLMs. We’re already there. And there’s a lot of bad code being pumped out. Which will in turn be fed back to the LLMs. > Don't look at the state of the art not, look at the direction of travel. That’s what leads to the eternal “in five years” which eventually sinks everyone’s trust.
- danielmarkbruce 2y ago> What a stupid future. Machines which make errors being “corrected” by machines which make errors in a death spiral. An unbelievable waste of figurative and literal energy. Humans are machines which make errors. Somehow, we got to the moon. The suggestion that errors just mindlessly compound and that there is no way around it, is what's stupid.
- nuancebydefault 2y ago
- pixxel 2y ago[flagged]
- LinuxBender 2y agoIs AI ready to crawl through all open source and find / fix all the potential security bugs or all bugs for that matter? If so will that become a commercial service or a free service? Will AI be able to detect bugs and back doors that require multiple pieces of code working together rather than being in a single piece of code? Humans have a hard time with this. - Hypothetical Example: Authentication bugs in sshd that requires a flaw in systemd which then requires a flaw in udev or nss or PAM or some underlying library ... but looking at each individual library or daemon there are no bugs that a professional penetration testing organization such as the NCC group or Google's Project Zero would find. In other words, will AI soon be able to find more complex bugs in a year than Tavis has found in his career and will they start to compete with one another and start finding all the state sponsored complex bugs and then ultimately be able to create a map that suggests a common set of developers that may need to be notified? Will there be a table that logs where AI found things that professional human penetration testers could not?
- paradox242 2y agoSeems like there is more gain on the adversary side of this equation. Think nation-states like North Korea or China, and commercial entities like Pegasus Group.
- AnimalMuppet 2y agoGoogle's AI would have the advantage of the source code. The adversaries would not. (At least, not without hacking Google's code repository, which isn't impossible...)
- saagarjha 2y agoFWIW: NSO is the group, Pegasus is their product
- 0points 2y agoNo, that would require AGI. Actual reasoning. Adversaries are already detecting issues tho, using proven means such as code review and fuzzing. Google project zero consists of a team of rock star hackers. I don't see LLM even replacing junior devs right now.
- mysterydip 2y agoI assume the amount of monitoring effort is less than the amount of effort that would be required to replicate the AI generated code by humans, but do you have numbers on what that ROI looks like? Is it more like 10% or 200%?
- hshshshshsh 2y agoSeems like everything is working out without any issues. Shouldn't you be a bit suspicious?
- Twirrim 2y ago> We work very closely with Google DeepMind to adapt Gemini models for Google-scale coding and other Software Engineering usecases. Considering how terrible and frequently broken the code that the public facing Gemini produces, I'll have to be honest that that kind of scares me. Gemini frequently fails at some fairly basic stuff, even in popular languages where it would have had a lot of source material to work from; where other public models (even free ones) sail through. To give a fun, fairly recent example, here's a prime factorisation algorithm it produce for python: # Find the prime factorization of n prime_factors = [] while n > 1: p = 2 while n % p == 0: prime_factors.append(p) n //= p p += 1 prime_factors.append(n) Can you spot all the problems?
- senko 2y agoWe collectively deride leetcoding interviews yet ask AI to flawlessly solve leetcode questions. I bet I'd make more errors on my first try at it.
- AnimalMuppet 2y agoWriting a prime-number factorization function is hardly "leetcode".
- atomic128 2y agoEmpirical testing (for example: https://news.ycombinator.com/item?id=33293522 https://news.ycombinator.com/item?id=33293522) has established that the people on Hacker News tend to be junior in their skills. Understanding this fact can help you understand why certain opinions and reactions are more likely here. Surprisingly, the more skilled individuals tend to be found on Reddit (same testing performed there).
- louthy 2y agoI’m not sure that’s evidence; I looked at that and saw it was written in Go and just didn’t bother. As someone with 40 years of coding experience and a fundamental dislike of Go, I didn’t feel the need to even try. So the numbers can easily be skewed, surely.
- bogwog 2y agoIs any of the AI generated code being committed to Google's open source repos, or is it only being used for private/internal stuff?
- wslh 2y agoAs someone working in cybersecurity and actively researching vulnerability scanning in codebases (including with LLMs), I’m struggling to understand what you mean by “safe.” If you’re referring to detecting security vulnerabilities, then you’re either working on a confidential project with unpublished methods, or your approach is likely on par with the current state of the art, which primarily addresses basic vulnerabilities.
- assanineass 2y agoWas this comment cleared by comms
- deleted 2y ago[deleted]
- bcherny 2y agoHow are you measuring productivity? And is the effect you see in A/B tests statistically significant? Both of these were challenging to do at Meta, even with many thousands of engineers —- curious what worked for you.
- nycdatasci 2y agoYou mention safety as #1, but my impression is that Google has taken a uniquely primitive approach to safety with many of their models. Instead of influencing the weights of the core model, they check core model outputs with a tiny and much less competent “safety model”. This approach leads to things like a text-to-image model that refuses to output images when a user asks to generate “a picture of a child playing hopscotch in front of their school, shot with a Sony A1 at 200 mm, f2.8”. Gemini has similar issue: it will stop mid-sentence, erase its entire response and then claim that something is likely offensive and it can’t continue. The whole paradigm should change. If you are indeed responsible for developer tools, I would hope that you’re activity leveraging Claude 3.5 Sonnet and o1-preview.
- deleted 2y ago[deleted]
- ActionHank 2y agoWould you say that the efficiency gain is less than, equal to, or greater than the cost? It's always felt like having AI in the cloud for better autocomplete is a lot for a small gain.