5 ms·
I don't get it? Yes you should require a valid reason before believing something The only objective measures I've seen people attempt to take have at best show
by AstroBen 9mo ago
I don't get it? Yes you should require a valid reason before believing something
The only objective measures I've seen people attempt to take have at best shown no productivity loss:
https://substack.com/home/post/p-172538377 https://substack.com/home/post/p-172538377
https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o...
This matches my own experience using agents, although I'm actually secretly optimistic about learning to use it well
- nfw2 9mo agoWhy do you believe that the sky is blue? What randomized trial with proper statistical controls has shown this to be true?
- AstroBen 9mo agoI can see it, it's independently verifiable by others, and it's measurable
- nfw2 9mo agoThe same is true of AI productivity https://resources.github.com/learn/pathways/copilot/essentials/measuring-the-impact-of-github-copilot/ https://resources.github.com/learn/pathways/copilot/essentia... https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic https://www.anthropic.com/research/how-ai-is-transforming-wo... https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier https://www.mckinsey.com/capabilities/tech-and-ai/our-insigh...
- deleted 9mo ago[deleted]
- davidgerard 9mo agolol those are all self-reports of vibes then they put the vibes on a graph, which presumably transforms them into data
- nfw2 9mo ago"Both GitHub and outside researchers have observed positive impact in controlled experiments and field studies where Copilot has conferred: 55% faster task completion using predictive text Quality improvements across 8 dimensions (e.g. readability, error-free, maintainability) 50% faster time-to-merge" how is time-to-merge a vibe?
- Orygin 9mo agoThe subject is productivity. Time to merge is as useful metric as Lines of Code to determine productivity. I can merge 100s of changes but if they are low quality or incur bugs, then it's not really more productive.
- davidgerard 9mo agothis guy has elsewhere in this thread cited "a16z revenue benchmarks" as evidence of productivity. you know, the sector most famous for setting more money on fire faster than anyone in living memory.
- intended 9mo ago> https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o... Shows that devs overestimate the impact of LLMs on their productivity. They believe they get faster when they take more time. Since Anthropic, GitHub are fair game here’s one from Code Rabbit - https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-gen...
- llmslave2 9mo agoIf you point a spectrometer at the sky during the day in non-cloudy conditions you will observe readings peaking in the roughly 450-495 nanometers range, which crazily enough, is the definition of the colour blue [0]! Then you can research Rayleigh scattering, of which consists of a large body of academic research not just confirming that the sky is blue, but also why. But hey, if you want to claim the sky is red because you feel like it is, go ahead. Most people won't take you seriously just like they don't take similar claims about AI seriously. [0] https://scied.ucar.edu/image/wavelength-blue-and-red-light-image https://scied.ucar.edu/image/wavelength-blue-and-red-light-i...
- admdly 9mo agoI’m not sure why you’d need or want a randomised controlled trial to determine the colour of the sky. There have been empirical studies done to determine the colour and the reasoning for it - https://acp.copernicus.org/articles/23/14829/2023/acp-23-14829-2023.pdf https://acp.copernicus.org/articles/23/14829/2023/acp-23-148... is an interesting read.
- johnfn 9mo agoThe burden you are placing is too high here. Do you demand controlled trials for everything you do or else you refuse to use it or accept that other people might see productivity gains? Do you demand studies showing that static typing is productive? Syntax highlighting? IDEs or Vim? Unit testing? Whatever language you use? Obviously not? It would be absurd to walk into a thread about Rust and say “Rust doesn’t increase your productivity and unless you can produce a study proving it does then your own personal anecdotes are worthless.” Why the increased demand for rigor when it comes to AI specifically?
- llmslave2 9mo agoDo you just believe everything everybody says? No quantifiable data required, as long as someone somewhere says it it must be true? One of the reasons software is in decline is because it's all vibes, nobody has much interest in conducting research to find anything out. It doesn't have to be some double blinded peer reviewed meta analysis, the bar can still be low, it just should be higher than "I feel like"...
- nfw2 9mo agohere are some https://resources.github.com/learn/pathways/copilot/essentials/measuring-the-impact-of-github-copilot/ https://resources.github.com/learn/pathways/copilot/essentia... https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic https://www.anthropic.com/research/how-ai-is-transforming-wo... https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier https://www.mckinsey.com/capabilities/tech-and-ai/our-insigh...
- llmslave2 9mo agoThey're all marketing slop lol. Go look at their methodology. Absolutely shite.
- nfw2 9mo agoThis is what you claimed the bar was "it just should be higher than 'I feel like'" Now you are moving it because your statement is provably false. Your criticism of it is based on vibes. What specifically is wrong with the methodologies? One of them broke randomly developers into two groups, one with access to ai and one without, timed them to complete the same task, and compared the results. That seems fine? Any measurement of performance in a lab environment comes with caveats, but since real world accounts you dismiss as vibes, that seems like the best you can do.