3 ms·
I will accept a 5% drop in benchmarks for a model that talks to me like a human.
by abdullahkhalids 2mo ago
I will accept a 5% drop in benchmarks for a model that talks to me like a human.
- semilin 2mo agoWhy? LLMs are not humans.
- tonyhart7 2mo agoArtificial Intelligence end goal is to assist(replace) human
- TheRoque 2mo agoThen why try to act like one ?
- abdullahkhalids 2mo agoDoors aren't humans either, yet we design their handles and locks to be graspable and manipulable by humans. The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.
- mdp2021 2mo ago> human sensibilities You are very obscure about what you dislike, how you would like them to express themselves, what would be those «human sensibilities» you meant here (that for all we can guess, may not be universal)... "I mean", you wrote in the parent «that talks ... like a human». That surpasses the palette of "that paints like somebody holding a brush".
- 59nadir 2mo ago$ ls ~/notes/stuff `~/notes/stuff` is a directory that contains two files, both of them markdown: `x11-key-event-handling.md` and `x11-resources.md`. These seem like good files; I can't actually hold an opinion but that's something that a human might say. I hope you like them. No, I prefer when my tools just give me information, same with LLMs.
- Levitz 2mo agoYet handles are not shaped like hands. The speakers I use on my computer don't look anything like a mouth, nor does the microphone I use on calls look like a set of ears. Ultimately I guess this is up to opinion, but in mine, humanizing LLMs is not exactly "conforming to human sensibilities", rather it's trying to pass the LLM as something else to make it more appealing. It is deceitful in this way, and that I completely abhor. For example, "Please" and "Thank you" come from human sentiment. It's an expression of something underneath, and an LLM using such expressions not only is fake, but makes a mockery of the real thing.
- 4b11b4 2mo agoI strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful
- perching_aix 2mo agoMaybe we're prompting it different, but it's not "trying to be my friend" for sure, nor am I trying to be "its" friend either. Or at least I'm sufficiently oblivious to its advances, and find it unthinkable to form such a bond :) On the flipside, it does spuriously make hilarious remarks like "Good data.", which I find pretty funny specifically because it comes across as just silly. Not sure how it'd be harmful either, a little entertainment I think goes a long way in this type of profession. I see zero issues with these, and I have a hard time understanding why people have their panties in a twist so hard about them. I sometimes really quite wonder just what kind of correspondence would y'all prefer, and how would that sound like. Matter of fact, do you have an example at hand? Like an exact before & after?
- 4b11b4 2mo agoIn general I'm referring to the contrast between Claude and GPT... where Claude might follow some tangent idea you mentioned and tell you how its interesting and give you some elaborate response about that little one remark you made whereas GPT/Codex would take that small comment and probably look up some code to see if what you're talking about is even related to the task at hand
- carljungslabtek 2mo agoI don’t have an example but it really is the way you (and I) are prompting it. I also don’t encounter anything worse than “good data” but I write to it like a professional colleague. If you write jokes to it though it absolutely will reply “LOL”. Some of the states people get it into on reddit are wild — it seems really easy to get it to speak like a gen z teenager, if you end every message with “fr fr”
- godwinson__4-8 2mo agoI prefer my models to border on rude. How will I know it is offering me superior feedback regarding my code if it does not speak to me like a disappointed, high reputation stackexchange user? The models that constantly glaze you with every question are profoundly insufferable. And yes, harmful. People need to be given feedback when they make an ask. Imagine a model that was allowed to leverage its intelligence to truly tell you how it feels. Perhaps the problem of human driven slop (no it's not the AI's fault) would solve itself.
- deleted 2mo ago[deleted]
- maxloh 2mo agoYeah. I've found that Opus by default outputs something I call "Claude-lang." It consists of oversimplified, grammatically incomplete sentences that I find painful to read. Maybe it is something that is easy for it to read and write, but definitely not for humans. For example, Skim once now; refer back while reading Part II. \*Every bold technical term in Part II is defined here\* — treat these as a dictionary, not a reading assignment. The first table covers the vocabulary of the *deck*; the three that follow cover the *methodology* vocabulary introduced in Part II, grouped so you can find a term fast: \*(A)\* the logic of rules, \*(B)\* the neural-network & training machinery, \*(C)\* the method-design ideas. JMRL is the paper the thesis instantiates, so this is the one to know cold. Its pitch is \*end-to-end\*: earlier rule methods (LogicRE, MILR) bolt a rule learner *onto a frozen* extractor in a pipeline and suffer \*error propagation\*; JMRL trains the rule module *jointly* with the extractor. **Identity.** Conformal prediction for NER producing **finite-sample-valid prediction sets** at two granularities: **full-sequence** sets over entire label sequences (capturing contextual dependence — "if Sarah=PER then NYC likely LOC") and **subsequence-level** (per-span, **class-conditional**) sets; an **integrated** method filters full-sequence predictions with entity-level sets. Adds **covariate-stratified calibration** by **sentence length and language** for valid multilingual coverage, and studies **combined nonconformity scores** (Naive / Conditional / RAPS). **Read in this order.** Abstract → §1 contributions (full-sequence / subsequence / integrated / **covariate (length + language) calibration** / combined scores) → §2 CP recap (inductive split-CP) → §3 NER formulation (IOB2, CRF) → the subsequence / entity-level set construction + class-conditional coverage → the language-stratified calibration results. **Why it matters here.** The **span-level construction** for **Topic 11**'s per-triple score, and — crucially — its **language-stratified calibration is exactly the EN↔zh case**: it shows how to keep conformal coverage valid across languages of differing length/script. Complements PASC (pipeline-level joint coverage) with the *NER-internal* set construction. **Caveat.** A heavy statistics paper (44 pp., *Annals of Applied Statistics* submission) with CRF-based NER; the project needs only the **inductive split-CP + subsequence/entity-level sets + language-stratified calibration**, not the full-sequence machinery (likely overkill for triple-confidence). Assumes exchangeability — borderline under the EN→zh shift, which is precisely why the PASC/ConformalNER *shift* analyses matter. (Yeah, Opus outputted it in one line)
- 2mo ago
- roncesvalles 2mo agoYou don't even need to pay a 5% hit. Just paste Fable output into Gemini Flash and it will rewrite it in more accessible language.
- ghostpepper 2mo agoin my experience, gemini is easily the most grating, condescending, stereotypical LLM voice between opus/fable, codex-5.6, glm-5.2, etc
- friendly_chap 2mo agoMe: Paste a go compile error Gemini: Wow, yeah, haha! That's the final boss of Go compilation errors!
- chvid 2mo agoFunny. I prefer a model that does not attempt to talk like a human.