3 ms·
Just to be clear, what I was proposing was a single tool which would, on the basis of a single ~30-minute interaction, purchase a domain name, set up a cloud en
by buu700 10mo ago
Just to be clear, what I was proposing was a single tool which would, on the basis of a single ~30-minute interaction, purchase a domain name, set up a cloud environment, build a full-stack application + cross-platform native apps + useful tests with near-100% coverage, deploy a live test environment, and compile each platform's native app — all entirely autonomously. Are you saying you've used or built something similar to that? That is super interesting if so, even if you're unable to share. A major subset of that could also still be incredibly useful, but the whole solution I described is a very high bar.
I've been very successful building with custom LLM workflows and automation myself, but that's beyond the capabilities of any tooling I've seen, and I wouldn't necessarily expect great results with current models even if current tooling were fully capable of what I described. Even with such tooling, the cost of inference is high enough to deter careless usage without much more rigorous work on the initial spec and/or micromanagement of the development process.
I'm not necessarily advocating for one-shotting in any given context. I'm simply pointing out that there would be huge advantages to LLMs and tooling sufficiently advanced to be fully capable of doing so end-to-end, especially at dramatically lower cost than current models and at superhuman quality. Such an AI could conceivably one-shot any possible project idea, in the same sense that a competent human dev team with nothing but a page of vague requirements and unlimited time could at least eventually produce something functional.
The value of such an AI is that we'd use it in ways that sound ridiculous today. Maybe a chat with some guy at a bar randomly inspires a neat idea, so you quickly whip out your phone and fire off some bullet point notes; by the time you get home, you have 10 different near-production-ready variations to choose from, each with documentation on the various decisions its agent made and why, and each one only cost $5 in account credit. None is quite perfect, but through the process you've learned a lot and substantially refined the idea; you give it a second round of notes and wake up to a new testable batch. One of those has the functional requirements just right, so you make the final decisions on non-functional requirements and let it roll one last time with strict attention to detail on code quality and a bunch of cycles thrown at security review.
That evening, you check back in and find a high-quality final implementation that meets all of your requirements with a performant and scalable architecture, with all infrastructure deployed and apps submitted to all stores/repositories. You subsequently allocate a sales and marketing budget to the AI, and eventually notice that you suddenly have a new source of income. Now imagine that instead of you, this was actually your friend who's never written a line of code and barely knows how to use a computer.
I still agree with you that current models have been "good enough" for some time, in the sense that if LLMs froze today we could spend the next decade collectively building on and with them and it would totally transform the economy. But at the same time, there's definitely latent demand for more and/or better inference. If LLMs were to become radically more efficient, we wouldn't start shuttering data centers; the economy would just become that much more productive.
- nl 10mo agoHave you tried Loveable, Replit, V0 etc? Outside of purchasing the domain and native apps for you they cover a very significant amount of this. If you insist on Native Apps, it's possible Google Jules could do it. With Gemini 2.5 it wasn't strong enough but I think it has Gemini 3 now which can definitely do native apps just fine.
- buu700 10mo agoThanks for the recommendations. Regarding your other comment, Flutter is what I've landed on as well for my next cross-platform app project, and I'm currently in the middle of developing a spec for a fairly complex agentic system that I'm going to try having Codex two-shot (basic project setup + file stubs + exhaustive tests -> manual checkpoint -> TDD the rest). I haven't tried Lovable, V0, or Jules, but I really like Replit for certain things. Having said that, based on my experience, I would characterize it as an amazing tool for rapid frontend iteration with prototype-level backend creation. I'm sure it's gotten better at one-shotting since I tried Agent 2 with Sonnet 3.7 in May, but would still be very (pleasantly) surprised to see that Agent 3 with current models could meet the incredibly high bar of wholly replacing a human dev team. The fact that tools like Replit also include their own hosting environments is definitely neat, but not really what I was getting at as far as deployment. What I had in mind was managing arbitrary cloud platforms, setting up an optimal architecture for your anticipated scale and usage patterns — whether that's a single Hetzner instance with SQLite or horizontally scaled app servers behind an API gateway with Kafka, Valkey, and Spanner or ScyllaDB — and doing all the DevOps to handle that along with things like CI/CD. I'm not downplaying how amazing these capabilities are. Being able to generate high-quality code from natural language feels like magic. But all the parts beyond narrow application code are half of the thing I described: * I'm saying you should be able to send a single off-the-cuff drunk text to an AI and later find a complete production-ready SaaS startup that fully aligns with a reasonable interpretation of your message. * The other half of the whole thing is >=human-level execution. If the AI can't autonomously deliver work comparable to what an experienced CTO would (given the same requirements, an arbitrarily large hiring budget, and a stipulation to never contact you again until the work was done), it's not there yet. Again, none of this is to dunk on agentic coding. My point is that I set an absurdly high bar because I want it to one day be met. Just as a $100 storage budget today is equivalent to $100m a few decades ago, I want to live to see a $100 engineering budget reach equivalency with last decade's $100m.