12 ms·
My frustration with using these models for programming in the past has largely been around their tendency to hallucinate APIs that simply don't exist. The Gemin
by segphault 1y ago
My frustration with using these models for programming in the past has largely been around their tendency to hallucinate APIs that simply don't exist. The Gemini 2.5 models, both pro and flash, seem significantly less susceptible to this than any other model I've tried.
There are still significant limitations, no amount of prompting will get current models to approach abstraction and architecture the way a person does. But I'm finding that these Gemini models are finally able to replace searches and stackoverflow for a lot of my day-to-day programming.
- thefourthchime 1y agoAsk the models that can search to double check their API usage. This can just be part of a pre-prompt.
- ksec 1y agoI have been asking if AI without hallucination, coding or not is possible but so far with no real concrete answer.
- mattlondon 1y agoIt's already much improved on the early days. But I wonder when we'll be happy? Do we expect colleagues friends and family to be 100% laser-accurate 100% of the time? I'd wager we don't. Should we expect that from an artificial intelligence too?
- cinntaile 1y agoIt's tool not a human so I don't know if the comparison even makes sense?
- ziml77 1y agoYes we should expect better from an AI that has a knowledge base much larger than any individual and which can very quickly find and consume documentation. I also expect them to not get stuck trying the same thing they've already been told doesn't work, same as I would expect from a person.
- kweingar 1y agoI expect my calculator to be 100% accurate 100% of the time. I have slightly more tolerance for other software having defects, but not much more.
- asadotzler 1y agoAnd a $2.99 drugstore slim wallet calculator with solar power gets it right 100% of the time while billion dollar LLMs can still get arithmetic wrong on occasion.
- pb7 1y agoMy hammer can't do any arithmetic at all, why does anyone even use them?
- namaria 1y agoDoes it sometimes instead of driving a nail hit random things in the house?
- hn_go_brrrrr 1y agoYes, like my thumb.
- namaria 1y agoLimited blast radii are a great advantage of deterministic tools.
- izacus 1y agoWhat you're being asked is to stop trying to hammer every single thing that comes into your vicinity. Smashing your computer with a hammer won't create code.
- deleted 1y ago[deleted]
- 1y ago
- pohuing 1y agoIt's a tool, not an intelligence, a tool that costs money on every erroneous token. I expect my computer to be more reliable at remembering things than myself, that's one of the primary use cases even. Especially if using it costs money. Of course errors are possible, but rarely do they happen as frequently in any other program I use.
- kortilla 1y agoIf colleagues lie with the certainty that LLMs do, they would get fired for incompetence.
- scarab92 1y agoI wish that were true, but I’ve found that certain types of employees do confidently lie as much as llms, especially when answering “do you understand” type questions
- izacus 1y agoAnd we try to PIP and fire those as well, not turn everyone else into them.
- dmd 1y agoOr elected to high office.
- ChromaticPanic 1y agoHave you worked in an actual workplace. Confidence is king.
- kortilla 1y agoYes. Lying does not work as an engineer.
- deleted 1y ago[deleted]
- ksec 1y agoI dont expect it to be 100% accurate. Software aren't bug free, human aren't perfect. But may be 99.99%? At least given enough time and resources human could fact check it ourselves. And precisely because we know we are not perfect, in accounting and court cases we have due diligence. And it is also not just about the %. It is also about the type of error. Will we reach a point we change our perception and say these are expected non-human error? Or could we have a specific LLM that only checks for these types of error?
- mdp2021 1y agoYes we want people "in the game" to be of sound mind. (The matter there is not about being accurate, but of being trustworthy - substance, not appearance.) And tools in the game, even more so (there's no excuse for the engineered).
- Foreignborn 1y agoTry dropping the entire api docs in the context. If it’s verbose, i usually pull only a subset of pages. Usually I’m using a minimum of 200k tokens to start with gemini 2.5.
- nolist_policy 1y agoThat's more than 222 novel pages: 200k tk = 1/3 200k words = 1/300 1/3 200k pages
- Foreignborn 1y agoIt’s easy to get 500-700k tokens in. I’ll drop research papers, a lot of work docs, get through a bunch of discussion, before writing a PRD-like doc of tasks to work from. That generally seems right to me, given how much we hold in our heads when you’re discussing something with a coworker.
- pizza 1y ago"if it were a fact, it wouldn't be called intelligence" - donald rumsfeld
- codebolt 1y agoI've found they do a decent job searching for bugs now as well. Just yesterday I had a bug report on a component/page I wasn't familiar with in our Angular app. I simply described the issue as well as I could to Claude and asked politely for help figuring out the cause. It found the exact issue correctly on the first try and came up with a few different suggestions for how to fix it. The solutions weren't quite what I needed but it still saved me a bunch of time just figuring out the error.
- M4v3R 1y agoThat’s my experience as well. Many bugs involve typos, syntax issues or other small errors that LLMs are very good at catching.
- ChocolateGod 1y agoI asked today both Claude and ChatGPT to fix a Grafana Loki query I was trying to build, both hallucinated functions that didn't exist, even when telling to use existing functions. To my surprise, Gemini got it spot on first time.
- fwip 1y agoCould be a bit of a "it's always in the last place you look" kind of thing - if Claude or CGPT had gotten it right, you wouldn't have tried Gemini.
- redox99 1y agoMaking LLMs know what they don't know is a hard problem. Many attempts at making them refuse to answer what they don't know caused them to refuse to answer things they did in fact know.
- bezier-curve 1y agoThe best way around this is to dump documentation of the APIs you need them privy to into their context window.
- deleted 1y ago[deleted]
- Volundr 1y ago> Many attempts at making them refuse to answer what they don't know caused them to refuse to answer things they did in fact know. Are we sure they know these things as opposed to being able to consistently guess correctly? With LLMs I'm not sure we even have a clear definition of what it means for it to "know" something.
- redox99 1y agoYes. You could ask for factual information like "Tallest building in X place" and first it would answer it did not know. After pressuring it, it would answer with the correct building and height. But also things where guessing was desirable. For example with a riddle it would tell you it did not know or there wasn't enough information. After pressuring it to answer anyway it would correctly solve the riddle. The official llama 2 finetune was pretty bad with this stuff.
- Volundr 1y ago> After pressuring it, it would answer with the correct building and height. And if you bully it enough on something nonsensical it'll give you a wrong answer. You press it, and it takes a guess even though you told it not to, and gets it right, then you go "see it knew!". There's no database hanging out in ChatGPT/Claude/Gemini's weights with a list of cities and the tallest buildings. There's a whole bunch of opaque stats derived from the content it's been trained on that means that most of the time it'll come up with the same guess. But there's no difference in process between that highly consistent response to you asking the tallest building in New York and the one where it hallucinates a Python method that doesn't exist, or suggests glue to keep the cheese on your pizza. It's all the same process to the LLM.
- pzo 1y agoI feel your pain. Cursor has docs features but many times when I pointed to check @docs and selected one recently indexed one it sometimes still didn't get it. I still have to try contex7 mcp which looks promising: https://github.com/upstash/context7 https://github.com/upstash/context7
- doug_durham 1y agoIf they never get good at abstraction or architecture they will still provide a tremendous amount of value. I have them do the parts of my job that I don't like. I like doing abstraction and architecture.
- mynameisvlad 1y agoSure, but that's not the problem people have with them nor the general criticism. It's that people without the knowledge to do abstraction and architecture don't realize the importance of these things and pretend that "vibe coding" is a reasonable alternative to a well-thought-out project.
- sanderjd 1y agoThe way I see this is that it's just another skill differentiator that you can take advantage of if you can get it right. That is, if it's true that abstraction and architecture are useful for a given product, then people who know how to do those things will succeed in creating that product, and those who don't will fail. I think this is true for essentially all production software, but a lot of software never reaches production. Transitioning or entirely recreating "vibecoded" proofs of concept to production software is another skill that will be valuable. Having a good sense for when to do that transition, or when to start building production software from the start, and especially the ability to influence decision makers to agree with you, is another valuable skill. I do worry about what the careers of entry level people will look like. It isn't obvious to me how they'll naturally develop any of these skills.
- mynameisvlad 1y ago> "vibecoded" proofs of concept The fact that you called it out as a PoC is already many bars above what most vibe coders are doing. Which is considering a barely functioning web app as proof that vibe coding is a viable solution for coding in general. > I do worry about what the careers of entry level people will look like. It isn't obvious to me how they'll naturally develop any of these skills. Exactly. There isn't really a path forward from vibe coding to anything productizable without actual, deep CS knowledge. And LLMs are not providing that.
- Tainnor 1y agoI definitely get more use out of Gemini Pro than other models I've tried, but it's still very prone to bullshitting. I asked it a complicated question about the Scala ZIO framework that involved subtyping, type inference, etc. - something that would definitely be hard to figure out just from reading the docs. The first answer it gave me was very detailed, very convincing and very wrong. Thankfully I noticed it myself and was able to re-prompt it and I got an answer that is probably right. So it was useful in the end, but only because I realised that the first answer was nonsense.
- alex1138 1y agoThe fact that SO much is only discovered after the fact by asking it "Are you sure?" is just insane There has to be some kind of recursive error checking thing, or something
- Tainnor 1y agoI did a bit more than "are you sure", though. I said "I don't think X is right because ..." (after reading the type signature of some function and thinking through what would happen). That seemed to lead it into the right direction.
- gxs 1y agoHuh? Have you ever just told it, that API doesn’t exist, find another solution? Never seen it fumble that around Swear people act like humans themselves don’t ever need to be asked for clarification
- 0x457 1y agoI've noticed that models that can search internet do it a lot less because I guess they can look up documentation? My annoyance now is that it doesn't take version into consideration.
- tough 1y agoYou should give it docs for each of your base dependencies in a mcp/tool whatever so it can just consult. internet also helps. Also having markdown files with the stack etc and any -rules-
- siscia 1y agoThis problem have been solved by LSP (language server protocol), all we need is a small server behind MCP that can communicate LSP information back to the LLM and get the LLM to use by adding to the prompt something like: "check your API usage with the LSP" The unfortunate state of open source funding makes buildings such simple tool a loosing adventure unfortunately.
- satvikpendem 1y agoThis already happens in agent modes in IDEs like Cursor or VSCode with Copilot, it can check for errors with the LSP.
- jug 1y agoI’ve seen benchs on hallucinations and OpenAI has typically performed worse than Google and Anthropic models. Sometimes significantly so. But it doesn’t seem like they have cared much. I’ve suspected that LLM performance is correlated to risking hallucinations? That is, if they’re bolder, this can be beneficial? Which helps in other performance benchmarks. But of course at the risk of hallucinating more…
- mountainriver 1y agoThe hallucinations are a result of RLVR. We reward the model for an answer and then force it to reason about how to get there when the base model may not have that information.
- mdp2021 1y ago> The hallucinations are a result of RLVR Well let us reward them for producing output that is consistent with database accessed selected documentation then, and massacre them for output they cannot justify - like we do with humans.
- johnisgood 1y ago> hallucinate APIs Tell me about it. Thankfully I have not experienced it as much with Claude as I did with GPT. It can get quite annoying. GPT kept telling me to use this and that and none of them were real projects.
- impulser_ 1y agoUse few-shot learning. Build a simple prompt with basic examples of how to use the API and it will do significantly better. LLMs just guess, so you have to give it a cheatsheet to help it guess closer to what you want.
- Jordan-117 1y agoI recently needed to recommend some IAM permissions for an assistant on a hobby project; not complete access but just enough to do what was required. Was rusty with the console and didn't have direct access to it at the time, but figured it was a solid use case for LLMs since AWS is so ubiquitous and well-documented. I actually queried 4o, 3.7 Sonnet, and Gemini 2.5 for recommendations, stripped the list of duplicates, then passed the result to Gemini to vet and format as JSON. The result was perfectly formatted... and still contained a bunch of non-existent permissions. My first time being burned by a hallucination IRL, but just goes to show that even the latest models working in concert on a very well-defined problem space can screw up.
- dotancohen 1y agoAWS docs have (had) an embedded AI model that would do this perfectly. I suppose it had better training data, and the actual spec as a RAG.
- djhn 1y agoBoth AWS and Azure docs’ built in models have been absolutely useless.
- darepublic 1y agoListen I don't blame any mortal being for not grokking the AWS and Google docs. They are a twisting labyrinth of pointers to pointers some of them deprecated though recommended by Google itself.
- perching_aix 1y agoSounds like a vague requirement, so I'd just generally point you towards the AWS managed policies summary [0] instead. Particularly the PowerUserAccess policy sounds fitting here [1] if the description for it doesn't raise any immediate flags. Alternatively, you could browse through the job function oriented policies [2] they have and see if you find a better fit. Can just click it together instead of bothering with the JSON. Though it sounds like you're past this problem by now. [0] https://docs.aws.amazon.com/IAM/latest/UserGuide/access_policies_managed-vs-inline.html#aws-managed-policies https://docs.aws.amazon.com/IAM/latest/UserGuide/access_poli... [1] https://docs.aws.amazon.com/aws-managed-policy/latest/reference/PowerUserAccess.html https://docs.aws.amazon.com/aws-managed-policy/latest/refere... [2] https://docs.aws.amazon.com/IAM/latest/UserGuide/access_policies_job-functions.html https://docs.aws.amazon.com/IAM/latest/UserGuide/access_poli...
- deleted 1y ago[deleted]
- satvikpendem 1y agoIf you use Cursor, you can use @Docs to let it index the documentation for the libraries and languages you use, so no hallucination happens.
- Rudybega 1y agoThe context7 mcp works similarly. It allows you to search a massive constantly updated database of relevant documentation for thousands of projects.
- thr0waway39290 1y agoReplacing stackoverflow is definitely helpful, but the best use case for me is how much it helps in high-level architecture and planning before starting a project.
- yousif_123123 1y agoThe opposite problem is also true. I was using it to edit code I had that was calling the new openai image API, which is slightly different from the dalle API. But Gemini was consistently "fixing" the OpenAI call even when I explained clearly not to do that since I'm using a new API design etc. Claude wasn't having that issue. The models are very impressive. But issues like these still make me feel they are still more pattern matching (although there's also some magic, don't get me wrong) but not fully reasoning over everything correctly like you'd expect of a typical human reasoner.
- toomuchtodo 1y agoIt seems like the fix is straightforward (check the output against a machine readable spec before providing it to the user), but perhaps I am a rube. This is no different than me clicking through a search result to the underlying page to verify the veracity of the search result surfaced.
- disgruntledphd2 1y agoWhy coding agents et al don't make use of the AST through LSP is a question I've been asking myself since the first release of GitHub copilot. I assume that it's trickier than it seems as it hasn't happened yet.
- celeritascelery 1y agoWhat good do you think that would do?
- disgruntledphd2 1y agoI've gotten a bunch of unbalanced parentheses suggestions, as well as loads of non existent variables generated. One could use the LSP errors to remove those completions.
- xmcqdpt2 1y agoMy guess is that it doesn’t work for several reasons. While we have millions of LOCs to train models on, we don’t have that for ASTs. Also, except for LISP and some macro supporting languages, the AST is not usually stable at all (it’s an internal implementation detail). It’s also way too sparse because you need a pile of tokens for even simple operations. The Scala AST for 1 + 2 for example probably looks like this, Apply(Select(scala, Select(math, Select(Int, Select(+)))), New(Literal(1)), Seq(This, New(Literal(2))) etc etc which is way more tokens than 1 + 2. You could possibly use a token per AST operation but then you can’t train on human language anymore and you need a new LLM per PL, and you can’t solve problem X in language Y based on a solution from language Z.
- mannycalavera42 1y agosame, I asked a simple question about javascript fetch api and it started talking about the workspace api. When I asked about that workspace api it replied it was the google workspace API ¯ \ _ (ツ) _ / ¯
- mbesto 1y agoTo date, LLMs can't replace the human element of: - Determining what features to make for users - Forecasting out a roadmap that are aligned to business goals - Translating and prioritizing all of these to a developer (regardless of whether these developers are agentic or human) Coincidentally these are the areas that frequently are the largest contributors to software businesses successes....not wether you use NextJs with a Go and Elixir backend against a multi-geo redundant multi sharded CockroachDB database, or that your code is clean/elegant.
- nearbuy 1y agoWhat does it say when you ask it to?
- mbesto 1y agoIt prompts YOU.
- dist-epoch 1y agoMaybe at elite companies. At half of the companies you can randomly pick those three things and probably improve the situation. Using an AI would be a massive improvement.
- jstummbillig 1y ago> no amount of prompting will get current models to approach abstraction and architecture the way a person does I find this sentiment increasingly worrisome. It's entirely clear that every last human will be beaten on code design in the upcoming years (I am not going to argue if it's 1 or 5 years away, who cares?) I wished people would just stop holding on to what amounts to nothing, and think and talk more about what can be done in a new world. We need good ideas and I think this could be a place to advance them.
- jjice 1y agoI'm confused by your comment. It seems like you didn't really provide a retort to the parent's comment about bad architecture and abstraction from LLMs. FWIW, I think you're probably right that we need to adapt, but there was no explanation as to _why_ you believe that that's the case.
- TuringNYC 1y agoI think they are pointing out that the advantage humans have has been chipped away little by little and computers winning at coding is inevitable on some timeline. They are also suggesting that perhaps the GP is being defensive.
- dml2135 1y agoWhy is it inevitable? Progress towards a goal in the past does not guarantee progress towards that goal in the future. There are plenty of examples of technology moving forward, and then hitting a wall.
- TuringNYC 1y agoI agree with you it isnt guaranteed to be inevitable, and also agree there have been plenty of journeys which were on a trajectory only to fall off. That said, IMHO it is inevitable. My personal (dismal) view is that businesses see engineering as a huge cost center to be broken up and it will play out just like manufacturing -- decimated without regard to the human cost. The profit motive and cost savings are just too great to not try. It is a very specific line item so cost/savings attribution is visible and already tracked. Finally, a good % of the industry has been staffed up with under-trained workers (e.g., express bootcamp) who arent working on abstraction, etc -- they are doing basic CRUD work.
- froh 1y agosearching and ranking existing fragments and recombining them within well known paths is one thing, exploratively combining existing fragments to completely novel solutions quickly runs into combinatorial explosion. so it's a great tool in the hands of a creative architect, but it is not one in and by itself and I don't see yet how it can be. my pet theory is that the human brain can't understand and formalize its creativity because you need a higher order logic to fully capture some other logic. I've been contested that the second Gödel incompleteness theorem "can't be applied like this to the brain" but I stubbornly insist yes, the brain implements _some_ formal system and it can't understand how that system works. tongue in cheek, somewhat, maybe. but back to earth I agree llms are a great tool for a creative human mind.
- dist-epoch 1y ago> Demystifying Gödel's Theorem: What It Actually Says > If you think his theorem limits human knowledge, think again https://www.youtube.com/watch?v=OH-ybecvuEo https://www.youtube.com/watch?v=OH-ybecvuEo
- froh 1y agothanks for the pointer. first, with Neil DeGrasse Tyson I feel in fairly ok company with my little pet peeve fallacy ;-) yah as I said, I both get it and don't ;-) And then the video escapes me saying statements about the brain "being a formal method" can't be made "because" the finite brain can't hold infinity. that's beyond me. although obviously the brain can't enumerate infinite possibilities, we're still fairly well capable of formal thinking, aren't we? And many lovely formal systems nicely fit on fairly finite paper. And formal proofs can be run on finite computers. So somehow the logic in the video is beyond me. My humble point is this: if we build "intelligence" as a formal system, like some silicon running some fancy pants LLM what have you, and we want rigor in it's construction, i.e. if we want to be able to tell "this is how it works", then we need to use a subset of our brain that's capable of formal and consistent thinking. And my claim is that _that subsystem_ can't capture "itself". So we have to use "more" of our brain than that subsystem. so either the "AI" that we understand is "less" than what we need and use to understand it. or we can't understand it. I fully get our brain is capable of more, and this "more" is obviously capable of very inconsistent outputs, HAL 9000 had that problem, too ;-) I'm an old woman. it's late at night. When I sat through Gödel back in the early 1990s in CS and then in contrast listened to the enthusiastic AI lectures it didn't sit right with me. Maybe one of the AI Prof's made that tactical mistake to call our brain "wet biological hardware" in contrast to "dry silicon hardware". but I can't shake of that analogy ;-) I hope I'm wrong :-) "real" AI that we can trust because we can reason about it's inner workings will be fun :-)
- jppittma 1y agoI've had great success by asking it to do project design first, compose the design into an artifact, and then asking it to consult the design artifact as it writes code.
- epaga 1y agoThis is a great idea - do you have a more detailed overview of this approach and/or an example? What types of things do you tell it to put into the "artefact"?
- pdntspa 1y agoI don't know about that, my own adventures with Gemini Pro 2.5 in Roo Code has it outputting code in a style that is very close to my own While far from perfect for large projects, controlling the scope of individual requests (with orchestrator/boomerang mode, for example) seems to do wonders Given the sheer, uh, variety of code I see day to day in an enterprise setting, maybe the problem isn't with Gemini?
- abletonlive 1y agoI feel like there are two realities right now where half the people say LLM doesn't do anything well and there is another half that's just using LLM to the max. Can everybody preface what stack they are using or what exactly they are doing so we can better determine why it's not working for you? Maybe even include what your expectations are? Maybe even tell us what models you're using? How are you prompting the models exactly? I find for 90% of the things I'm doing LLM removes 90% of the starting friction and let me get to the part that I'm actually interested in. Of course I also develop professionally in a python stack and LLMs are 1 shotting a ton of stuff. My work is standard data pipelines and web apps. I'm a tech lead at faang adjacent w/ 11YOE and the systems I work with are responsible for about half a billion dollars a year in transactions directly and growing. You could argue maybe my standards are lower than yours but I think if I was making deadly mistakes the company would have been on my ass by now or my peers would have caught them. Everybody that I work with is getting valuable output from LLMs. We are using all the latest openAI models and have a business relationship with openAI. I don't think I'm even that good at prompting and mostly rely on "vibes". Half of the time I'm pointing the model to an example and telling it "in the style of X do X for me". I feel like comments like these almost seem gaslight-y or maybe there's just a major expectation mismatch between people. Are you expecting LLMs to just do exactly what you say and your entire job is to sit back prompt the LLM? Maybe I'm just use to shit code but I've looked at many code bases and there is a huge variance in quality and the average is pretty poor. The average code that AI pumps out is much better.
- oparin10 1y agoI've had the opposite experience. Despite trying various prompts and models, I'm still searching for that mythical 10x productivity boost others claim. I use it mostly for Golang and Rust, I work building cloud infrastructure automation tools. I'll try to give some examples, they may seem overly specific but it's the first things that popped into my head when thinking about the subject. Personally, I found that LLMs consistently struggle with dependency injection patterns. They'll generate tightly coupled services that directly instantiate dependencies rather than accepting interfaces, making testing nearly impossible. If I ask them to generate code and also their respective unit tests, they'll often just create a bunch of mocks or start importing mock libraries to compensate for their faulty implementation, rather than fixing the underlying architectural issues. They consistently fail to understand architecture patterns, generating code where infrastructure concerns bleed into domain logic. When corrected, they'll make surface level changes while missing the fundamental design principle of accepting interfaces rather than concrete implementations, even when explicitly instructed that it should move things like side-effects to the application edges. Despite tailoring prompts for different models based on guides and personal experience, I often spend 10+ minutes correcting the LLM's output when I could have written the functionality myself in half the time. No, I'm not expecting LLMs to replace my job. I'm expecting them to produce code that follows fundamental design principles without requiring extensive rewriting. There's a vast middle ground between "LLMs do nothing well" and the productivity revolution being claimed. That being said, I'm glad it's working out so well for you, I really wish I had the same experience.
- bboygravity 1y agoThis is hilarious to read if you have actually seen the average (embedded systems) production code written by humans. Either you have no idea how terrible real world commercial software (architecture) is or you're vastly underestimating newer LLMs or both.
- onlyrealcuzzo 1y ago2.5 pro seems like a huge improvement. One area I've still noticed weakness is if you want to use a pretty popular library from one language in another language, it has a tendency to think the function signatures in the popular language match the other. Naively, this seems like a hard problem to solve. I.e. ask it how to use torchlib in Ruby instead of Python.
- viraptor 1y ago> no amount of prompting will get current models to approach abstraction and architecture the way a person does. What do you mean specifically? I found the "let's write a spec, let's make a plan, implement this step by step with testing" results in basically the same approach to design/architecture that I would take.
- nurettin 1y agoJust tell it to cite docs when using functions, works wonders.
- tastysandwich 1y agoRe hallucinating APIs that don't exist - I find this with Golang sometimes. I wonder if it's because the training data doesn't just consist of all the docs and source code, but potentially feature proposals that never made it into the language. Regexes are another area where I can't get much help from LLMs. If it's something common like a phone number, that's fine. But anything novel it seems to have trouble. It will spit out junk very confidently.
- robinei 1y agoSince it's trained on a vast a mount of code (probably all publicly accessible Go code and more), it's seen a vast amount of different bespoke APIs for doing all kinds of things. I'm sure some of that will leak into the output from time to time. And to some extent can generalize, so it may just invent APIs.
- mark_l_watson 1y agoI have a suggestion for you: Create a Gemini Gem for a programming language and put context info for library resources, examples of your coding style, etc. I just dropped version 0.1 of my Gemini book, and I have an example for making a Gem (really simple to do); read online link: https://leanpub.com/solo-ai/read https://leanpub.com/solo-ai/read
- SafeDusk 1y agoI’m having reasonable success specifically with Gemini model using only 7 tools: read, write, diff, browse, command, ask, think. This minimal template might be helpful to you: https://github.com/aperoc/toolkami https://github.com/aperoc/toolkami
- tiahura 1y agoWhy not add the applicable api references as context?
- ookblah 1y agothe future is probably something that looks pretty "inefficient" to us but a non-factor for a machine. i sometimes think a lot of our code structure is just for our own maintenance and conceptualization (DRY, SRP), but if you throw enough compute adn context at a problem im sure none of this even matters (as much). at least for 90% of the CRUD apps out there, you can def abstract away the entire base framework of getting, listing, and updating records. i guess the problem is validating that data for use in other more complex workflows.
- bruce511 1y agoI've spent my career writing code in a language which already abstracts 90% of a CRUD-type app away. Indeed there are a whole subset of users who literally don't write a line of code. We've had this since the very early 90's for DOS. Of course that last 10% does a lot of heavy lifting. Domain expertise, program and database design, sales, support, actually processing the data for more than just simple reports, and so on. And sure, the code is not maximally efficient in all cases, but it is consistent, and deterministic. Which is all I need from my code generator. I see a lot of panic from programmers (outside our space) who worry about their futures. As if programming is the ultimate career goal. When really, writing code is the least interesting, and least valuable part of developing software. Maybe LLMs will code software for you. Maybe they already do. And, yes, despite their mistakes it's very impressive. And yes, it will get better. But they are miles away from replacing developers- unless your skillset is limited to "coding" there's no need to worry.