5 ms·
Anthropic Education the AI Fluency Index
- Kye 8mo agoYou could arrive at the essence of this by just having read and internalized Carl Sagan's The Demon-Haunted World. Especially the Baloney Detection Kit. In my experience good prompting is mostly just good thinking.
- sdf2erf 8mo ago[dead]
- esafak 8mo agoAnd having the experience and judgment to ask the right thing.
- selridge 8mo agoAnd being willing to be wrong and to be misled; finding ways to contain that or build forcing functions against it. In a strange way that's exciting, because it forces me to learn. And sometimes forces me to confront whether stuff I had was domain knowledge or portable as experience.
- bigstrat2003 8mo agoTo the extent that this should be a thing, there are very few people I would want doing it less than the company who has repeatedly been caught lying about its product's achievements. Anthropic should not be taken seriously after their track record.
- deleted 8mo ago[deleted]
- bargainbin 8mo agoI’m not alone in finding this against the claims of the product right? Claude is meant to be so clever it can replace all white collar work in the next n-years, but also “you’re not using it right?” Which one is it?
- dsr_ 8mo agoWhich one will convince you to buy more Claude? Please answer honestly, it's for the sake of profits.
- SpicyLemonZest 8mo agoI'm not quite convinced of the maximalist claims, but these two aren't incompatible. Every time we talk about a company being "mismanaged" by e.g. a private equity buyout, what we mean is that the owners had access to a large volume of high quality white collar work but couldn't figure out how to use it right.
- rsynnott 8mo agoAnthropic in particular seem to be in a weird place where on the one hand they fund some real research, which is often not all roses and sunshine for them, but on the other hand, like all AI companies, they feel the need to make absurdly over-the-top claims about what's coming up Real Soon Now(TM).
- jimbokun 8mo agoAnthropic is a weird company where the CEO almost admits at times they are probably building the Torment Nexus, yet still feel they need to do it anyway…because someone else might do it first?
- sarkarghya 8mo agoHonestly to use llms properly all you need to know is that it’s a next word (or action) prediction model and like all models increased entropy hurts it. Try to reduce entropy to get better results. Rest is just sugarcoated nonsense. To use llms properly you need a physics class.
- Barbing 8mo agoWhich class? Or what subjects
- rishabhaiover 8mo agoAnd then some alignment, prompting structure, and task decomposition.
- arcanemachiner 8mo agoAnd praying that your desired output was embedded into the training data that was used to generate the model.
- dmk 8mo agoSo I guess the key takeaway is basically that the better Claude gets at producing polished output, the less users bother questioning it. They found that artifact conversations have lower rates of fact-checking and reasoning challenges across the board. That's kind of an uncomfortable loop for a company selling increasingly capable models.
- Florin_Andrei 8mo agoI think we're still at the stage where model performance largely depends on: - how many data sources it has access to - the quality of your prompts So, if prompting quality decreases, so does model performance.
- dmk 8mo agoSure, but the study is saying something slightly different, it's not that people write bad prompts for artifacts, they actually write better ones (more specific, more examples, clearer goals,...). They just stop evaluating the result. So the input quality goes up but the quality control goes down.
- candiddevmike 8mo agoWhat does prompting quality even mean, empirically? I feel like the LLM providers could/should provide prompt scoring as some kind of metric and provide hints to users on ways they can improve (possibly including ways the LLM is specifically trained to act for a given prompt).
- dsr_ 8mo agoThat would be a quality metric, and right now they are focused on quantity metrics.
- jimbokun 8mo agoSeems like it’s impossible for output to be good if the prompt is bad. Unless the AI is ignoring the literal instructions and just guessing “what you really want” which would be bad in a different way.
- mlpoknbji 8mo ago> But we know that any person who uses AI is likely to improve at what they do. Do we?
- co_king_5 8mo ago[dead]
- Insanity 8mo agoYah and this seems to be supported by preliminary evidence on the impact of AI on things like retention and cognitive ability.
- rishabhaiover 8mo agoAs a student, I constantly worry about this. But everyone in my class is producing output at a pace I can't compete with without AI assistance.
- Avshalom 8mo agowhat class are you in that "producing output at a [rapid] pace" is relevant to the grade?
- rishabhaiover 8mo agopick any cs class
- Avshalom 8mo agoI have a minor in CS and no -producing the assignment by the deadline is important- grades are not based on quantity of code vs classmates.
- rsynnott 8mo agoI mean, maybe things have changed (I finished college about 20 years ago), but I don't remember producing large volumes of stuff as being a particularly important part of a CS degree.
- kseniamorph 8mo agoI feel like the authors make a logical inconsistency. They present the drop in "identify missing context" behavior in artifact conversations as potentially concerning, like people are thinking less critically. But their own data suggests a simpler explanation: artifact conversations show higher rates of upfront specification (clarifying goals +14.7pp, specifying format +14.5pp, providing examples +13.4pp). It's obvious that when you provide more context upfront, you end up with less missing context later. I'd be more sceptical about such research.
- MarcLore 8mo ago[dead]
- lukev 8mo agoThis is a highly circular method of evaluation. It correlates "fluency behaviors" with longer conversations and more back and forth. What it notably does not correlate any of these these behaviors with is external value or utility. It is entirely possible that those people who are getting the most value out of LLMs are the ones with shorter interactions, and that those who engage in lengthier interactions are distracting themselves, wasting time, or chasing rabbit trails (the equivalent of falling in a wiki-hole, at the most charitable.) I can't prove that either -- but this data doesn't weigh in one way or the other. It only confirms that people who are chatty with their LLMs are chatty with their LLMs. In my own case, I find the longer I "chat" with the LLM the more likely I am to end up with a false belief, a bad strategy, or some other rabbit hole. 90% of the value (in my personal experience) is in the initial prompt, perhaps with 1-2 clarifying follow-ups.
- zahlman 8mo ago> In line with our recent Economic Index, we find that the most common expression of AI fluency is augmentative—treating AI as a thought partner, rather than delegating work entirely. In fact, these conversations exhibit more than double the number of AI fluency behaviors than quick, back-and-forth chats. > But we also find that when AI produces artifacts—including apps, code, documents, or interactive tools—users are less likely to question its reasoning (-3.1 percentage points) or identify missing context (-5.2pp). This aligns with related patterns we observed in our recent study on coding skills. Well, sure. If you're asking the AI to produce artifacts directly, it's likely because you pre-judged yourself less competent to do that kind of analysis.
- rickydroll 8mo agoWhile AI fluency is an important question to ask, affordability is another. Can a low-income person use AI to the same level of fluency as a high-income person? Will fluency become another force for income inequality?