11 ms·
The barriers to AI engineering are crumbling fast
- croes 2y agoIs that really AI engineering or Software engineering with AI? If a model goes sideways how do you fix that? Could you find and fix flaws in the base model?
- sigmar 2y agoAgree that the use of "AI engineers" is confusing. Think this blog should use the term "engineering software with AI-integration" which is different from "AI engineering" (creating/designing AI models) and different from "engineering with AI" (using AI to assist in engineering)
- crimsoneer 2y agoThe term AI engineer is now pretty well recognised in the field (https://www.latent.space/p/ai-engineer https://www.latent.space/p/ai-engineer), and is very much not the same as an AI researcher (which would be involved in training and building new models). I'd expect an AI engineer to be primarily a software developer, but with an excellent understanding of how to implement, use and evaluate LLMs in a production environment, including skills like evaluation and fine-tuning. This is not some dataset you can just bundle in software developer.
- liampulles 2y agoI wonder if either could be really be called engineering.
- billmalarky 2y agoYou find issues when they surface during your actual use case (and by "smoke testing" around your real-world use case). You can often "fix" issues in the base model with additional training (supervised fine-tuning, reinforcement learning w/ DPO, etc). There's a lot of tooling out there making this accessible to someone with a solid full-stack engineering background. Training an LLM from scratch is a different beast, but that knowledge honestly isn't too practical for everyday engineers given even if you had the knowledge you wouldn't necessarily have the resources necessary to train a competitive model. Of course you could command a high salary working for the orgs who do have these resources! One caveat is there are orgs doing serious post-training even with unsupervised techniques to take a base model and reeaaaaaally bake in domain-specific knowledge/context. Honestly I wonder if even that is unaccessible to pull off. You get a lot of wiggle-room and margin for error when post-training a well-built base model because of transfer learning.
- sharemywin 2y agolooks interesting I'll have to check it out.
- normanthreep 2y agothank you for letting us know. i was wondering if you found it interesting and would check it out
- deleted 2y ago[deleted]
- JSDevOps 2y agoIs anyone instantly suspicious when they introduce themselves these days an "AI Developer"
- noch 2y ago> Is anyone instantly suspicious when they introduce themselves these days an "AI Developer" I'm only suspicious if they don't simultaneously and eagerly show me their Github so that I can see what they've accomplished.
- llm_nerd 2y agoOf the great developers I have worked with in real life, across a large number of projects and workplaces, very few have any Github presence. Most don't even have LinkedIn. They usually don't have any online presence at all: No blog with regular updates. No Twitter presence full of hot takes. Sometimes this industry is a lot like the "finance" industry: People struggling for credibility talk about it constantly, everywhere. They flex and bloviate and look for surrogates for accomplishments wherever they can be found. Peacocking on github, writing yet another tutorial on what tokens are and how embeddings work, etc. That obviously doesn't mean in all cases, and there are loads of stellar talents that have a strong online presence. But by itself it is close to meaningless, and my experience is that it is usually a negative indicator.
- JSDevOps 2y agoIf someone has to tell you either themselves of by proxy they are influential in the industry ... they are not.
- noch 2y ago[flagged]
- llm_nerd 2y ago>Saying "GitHub" is just a way of saying: "Show me what you've accomplished." Do you actually think all development happens in public GitHub repos? Do you even think a majority does? Even a strong minority? Across a number of enormous, well-known projects I've worked on, covering many thousands of contributors, including several very well known names, 0% of it exists in public Github repos. The overwhelming bulk of development is happening in the shadows. If your "field" is "open source software", then sure. But if you're confused into thinking Github -- at least the tiny fraction that you can see -- is "the field" of software development or even just generally providing solutions, I can understand your weird belief about this.
- amelius 2y agoSoon enough we'll have AI that is just integrated into the OS. So individual apps don't need to do anything to have AI.
- michaelmior 2y agoWhat does it mean for an app to "have AI"?
- SteveSmith16384 2y agoIf it makes any kind of decision whatsoever (like an "if" statement), slap the word AI on it.
- amelius 2y agoThink of it as another human having access to the keyboard, mouse and the screen buffer.
- NitpickLawyer 2y agoAt a very minimum I'd say they'll have a way to "chat" with the apps to ask questions / do stuff. Either via APIs that the OS calls or integrated in the app via whatever frameworks will rise to handle this. In 5 to 10 years we'll probably see this. At a very minimum searching docs and "guide" the users through the functionality / do it straight up when asked. Basically what chatgpt did for chatbots, but at app level. There are lots of apps that take a long time to master. But the average joe doesn't need to master them. If I want to lightly edit some photos, I know photoshop can do it, but I have no clue where that specific thing is in the menus, because I haven't used it in 10 years. But it would be cool to type in a chat box "take all the pictures from my sd card, adjust the colors, straighten the ones that need it, and put them in my Pictures folder under "trip to the sea". And then I can go do something else for the 30-60 minutes it would have taken me to google how to do all of that, or script something, etc. The ideea of an assistant that can work like that isn't that far-fetched today, IMO. The apps need to expose some APIs, and the "os" needs an language -> action model capable enough to handle basic stuff for average joes. I'd bet good money sonnet3.5 + proper APIs + a bit of fine-tuning could do it today for 50%+ of average user cases.
- hpen 2y agoDoes AI engineer == API Engineer?
- bormaj 2y agoThe P is silent
- sdesol 2y agoIf the P is silent, how can you "Prompt"?
- Joker_vD 2y agoHow many silent P's are in the word "prompt"? Let me ask ChatGPT...
- tjr 2y agoJust for fun, I did ask ChatGPT: There’s one silent "p" in the word "prompt"—right at the beginning! The “p” isn’t pronounced, but it sneaks into the word anyway.
- dpassens 2y agoNormally, I'd just dismiss this as ChatGPT being trained not to disagree, but as a non-native speaker this has me doubting myself: Is prompt really pronounced rompt? It feels like it can't possible be true, but on the other hand, I'm probably due for having my confidence in my English completely shattered again by learning another weird word's real pronunciation, so maybe this is it.
- Joker_vD 2y agoI am not a native English speaker either, but I am fairly certain it's pronounced "promt", with the second "p", the one between the "m" and the "t", merging into those sounds to the point of being inaudible itself. Also, I too asked ChatGPT and it told me that In the word "prompt," there are no silent P's. All the letters in "prompt" are pronounced, including the P.
- pjmlp 2y agoAssuming an Engineer degree to start with.
- taco_emoji 2y agoWe can all be janitors too, so what?
- sincerecook 2y agoThe only remaining question being, why would you want to?
- aithrowawaycomm 2y agoFunny enough, Helix doesn't know either! They put together a contest hoping that you'll figure it out: https://blog.helix.ml/p/llm-app-challenge-with-helix-10 https://blog.helix.ml/p/llm-app-challenge-with-helix-10 You might say this is about Helix being small and trying to break into a crowded market, but OpenAI and Google offered similar contests / offers that asked users to submit ideas for LLM applications. Considering how many LLM sample apps are either totally useless ("Walter the Bavarian, a chatbot who gives trivia about Oktoberfest!") or could be better solved by classical programming ("a GPT that automatically converts currencies to USD!), it seems AI developers have struggled to find a single marketable use case of LLMs outside of codegen.
- neeleshs 2y agoCodegen, contentgen, and driving existing products in human language (which you can largely bucket into interactive codegen)
- sourcepluck 2y agoI feel like I see this comment fairly often these days, but nonetheless, perhaps we need to keep making it - the AI generated image there is so poor, and so off-putting. Does anyone like them? I am turned off whenever I see someone has used one on a post, with very few exceptions. Is it just me? Why are people using them? I feel like objectively they look like fake garbage, but obviously that must be my subjective biases, because people keep using them.
- RodgerTheGreat 2y agoSome people have no taste, and lack the mental tools to recognize the flaws and shortcomings of GANN output. People who enthuse about the astoundingly enthralling literary skills of LLMs tend to be the kind of person who hasn't read many books. These are sad cases: an undeveloped palate confusing green food coloring and xylitol for a bite of an apple. Some people can recognize these shortcomings and simply don't care. They are fundamentally nihilists for whom quantity itself is the only important quality. Either way, these hero images are a convenient cue to stop reading: nothing of value will be found below.
- nuancebydefault 2y ago> these hero images are a convenient cue to stop reading. If you don't like such content. But I would say don't judge a book by its cover.
- RodgerTheGreat 2y ago"GenAI" slop is a time-vampire. In a world where anyone can ask an LLM to gish-gallop a plausible facsimile of whatever argument they want in seconds, it is simply untenable to give any piece of writing you stumble upon online the benefit of the doubt; you will drown in counterfeit prose. The faintest hint that the author of a piece is a "GenAI" enthusiast (which in this case is already clear from the title) is immediate grounds for dismissing it; "the cover" clearly communicates the quality one might expect in the book. Using slop for hero images tells me that the author doesn't respect my time.
- PreInternet01 2y ago"after years of working in DevOps, MLOps, and now GenAI" You truly know how to align yourself with hype cycles?
- snovv_crash 2y agoThey missed out on drones and blockchain.
- taway1874 2y ago... and Cloud! Don't forget The Klaaooud ...
- kridsdale1 2y agoNah, I see plenty of “my code executed in a CPU on someone else’s computer” in that list.
- stogot 2y agoCloud isn’t really hype. VR or metaverse is the one you’re forgetting
- snovv_crash 2y agoCloud today, edge tomorrow.
- red-iron-pine 2y agodrones may still be a thing, and weren't nearly as hyped as blockchain. my barometer for penetration is how often the non-tech people talk about it, e.g. goofball uncle didn't buy a drone, but he went hard on BTC. if he's still holding he probably made money recently, too.
- snovv_crash 2y agoMy uncle bought a drone but didn't touch BTC. I agree penetration isn't the same, but it's also very different crowds.
- mark_l_watson 2y agoAfter just spending 15 minutes trying to get something useful accomplished, anything useful at all, with latest beta Apple Intelligence with a M1 iPad Pro (16G RAM), this article appealed to me! I have been running the 32B parameters qwen2.5-coder model on my 32G M2 Mac and and it is a huge help with coding. The llama3.3-vision model does a great job processing screen shots. Small models like smollm2:latest can process a lot of text locally, very fast. Open source front ends like Open WebUI are improving rapidly. All the tools are lining up for do it yourself local AI. The only commercial vendor right now that I think is doing a fairly good job at an integrated AI workflow is Google. Last month I had all my email directed to my gmail account, and the Gemini Advanced web app did a really good job integrating email, calendar, and google docs. Job well done. That said, I am back to using ProtonMail and trying to build local AIs for my workflows. I am writing a book on the topic of local, personal, and private AIs.
- mark_l_watson 2y agoAnother thought: OpenAI has done a good enough job productizing ChatGPT with advanced voice mode and now also integrated web search. I don’t know if I would trust OpenAI with access to my Apple iCloud data, Google data, my private GitHub repositories, etc., but given their history of effective productization, they could be a multi-OS/platform contender. Still, I would really prefer everything running under my own control.
- alexander2002 2y agoWho can trust a company whose name contradicts its presence.
- mark_l_watson 2y agoI don’t disagree with you!
- arcanemachiner 2y agoNot me. I learned that lesson after I tried to take a bite out of my Apple Macintosh.
- ein0p 2y agoAn AI engineer with some experience today can easily pull down 700K-1M TC a year at a bigtech. They must be unaware that the "barriers are coming down fast". In reality it's a full time job to just _keep up with research_. And another full time job to try and do something meaningful with it. So yeah, you can all be AI engineers, but don't expect an easy ride.
- actusual 2y agoI run an ML team in fintech, and am currently hiring. If a resumè came across my desk with this "skill set" I'd laugh my ass off. My job and my team's jobs are extremely stressful because we ship models that impact people's finances. If we mess up our customers lose their goddamn minds. Most of the ML candidates I see now are all "working with LLMs". Most of the ML engineers I know in the industry who are actually shipping valuable models, are not. Cool, you made a chatbot that annoys your users. Let me know when you've shipped a fraud model that requires four 9's, 100ms latency, with 50,000 calls an hour, 80% recall and 50% precision.
- btdmaster 2y agoWhat does 50% precision mean in this case? I know 50% accuracy might mean P(fraud_predicted | fraud) = 50%, but I don't understand what you mean by precision?
- chychiu 2y agoPrecision = True Positive / (True Positive + False Positive) = 1 - False Positive Rate On that note, I'm surprised the precision / recall for fin models are 80% / 50%
- disgruntledphd2 2y agoThey are obviously relatively low stakes, otherwise I'd be super worried.
- ein0p 2y ago
- JohnFen 2y agoI don't want to be an "AI engineer" in the way the article means. There's nothing about that sort of job that I find interesting or exciting. I hope there will still be room for devs in the future.
- bayareacommie 2y ago[dead]
- AIFounder 2y ago[dead]
- cess11 2y agoI mean, sure, anyone can cobble together Ollama and a wrapper API and an adjusted system prompt, or go serious with Bumblebee on the BEAM. But that's akin to web devs of old that stitched up some cruft in Perl or PHP and got their databases wiped by someone entering a SQL username. Yes, it kind of works under ideal conditions, but can you fix it when it breaks? Can you hedge against all or most relevant risks? Probably not. Don't put it your toys into production, and don't tell other people you're a professional at it until you know how to fix and hedge and can be transparent about it with the people giving you money.
- fullstackchris 2y agoI still don't see how AI replaces the understanding of what a server is, what DNS is, what HTTP is, or what... I could go on and on. Copy paste is great until you literally dont know where you are copy and pasting
- gabrieledarrigo 2y agoJust...boring.
- lastdong 2y agoLarge Language Models (LLMs) don’t fully grasp logic or mathematics, do they? They generate lines of code that appear to fit together well, which is effective for simple scripts. However, when it comes to larger or more complex languages or projects, they (in my experience) often fall short.
- underwater 2y agoBut humans aren’t either. We have to install programs in people for even basic mathematical and analytical tasks. This takes about 12 years, and is pretty ineffective.
- onlyrealcuzzo 2y agoIt seemingly built the modern world. Ineffective seems harsh.
- krapp 2y agoNo we don't. Humans were capable of mathematics and analytical tasks long before the establishment of modern twelve year education. We don't require "installing" a program to do basic mathematics, as we would never have figured out basic agriculture or developed civilization or seafaring if that were the case. I mean, Eratosthenes worked out the circumference of the Earth to a reasonable degree of accuracy in the third century BC. Even primitive hunter gather societies had concepts of counting, grouping and sequence that are beyond LLMs. If humans were as bad as LLMs at basic math and logic, we would consider them developmentally challenged. Yet this constant insistence that humans are categorically worse than, or at best no better than, LLMs persists. It's a weird, almost religious belief in the superiority of the machine even in spite of obvious evidence to the contrary.
- kmmlng 2y agoI think you both have a point. Clearly, something about the way we build LLMs today makes them inept at these kinds of tasks. It's unclear how fundamental this problem is as far as I'm concerned. Clearly, humans also need training in these things and almost no one figures out how to do basic things like long division by themselves. Some people sometimes figure things out, and more importantly, they do so by building on what came before. The difference between humans and LLMs is that even after being trained and given access to near everything that came before, LLMs are terrible at this stuff.