4 ms·
Article is just a vague summary of https://www.saturnos.com/report/artificial-authority https://www.saturnos.com/report/artificial-authority Anecdotally, curre
by ehe78qhe 7d ago
Article is just a vague summary of https://www.saturnos.com/report/artificial-authority https://www.saturnos.com/report/artificial-authority
Anecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't trust them with direct decision making with actual money due to the training lag time on current tax policy, etc.
- NitpickLawyer 7d ago> certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. Also, models are now good enough that you can give them chapters from "authoritative" books, and they'll integrate that and come up with better answers even if their "vanilla" answers were average. And they'll tailor stuff to your particular situation. It's funny that the "agentic" stuff is only used in coding mostly, while it can and does work in other fields as well. As always, you kinda need to check it (at least spot check) but all in all I'd agree it's better than the average stuff you used to find with a quick google search.
- legostormtroopr 7d agoSo if you know a book that has the information you need, you just need to upload it into the model to get the right answer. Isn't that a bit circular - if you already know the authoritative source, why ask a model?
- Calazon 7d agoBecause it's faster. I've done this on different topics - I know the answer is in a particular eBook/PDF/document, but for whatever reason it's not trivial to look it up. The model can do it a lot more quickly than I can, and then I can still verify the accuracy.
- ImaCake 7d agoDepends on the harness too. The MS copilot 365 chat interface happens to parse PDF and docx files before handing them to the model. Absolutely fantastic if your PDF is 1300 pages of concatenated reports and you want to find a single detail in it that you can't easily ctrl+f for.
- lconnell962 7d agoSome of the more widely spread and tolerated LLM outputs seem to be AI slop replacing Journalism/Blog slop. Places where people complained about quality already, but tolerated it if important enough. So to name some of the more common ones Translation, Summarization, and Reiteration of a source material. Humans put spin on things, how much you trust a source might not reflect the source's factual accuracy. It might just mean you liked reading it better from one source than another.
- theshrike79 7d agoIt's basically a fancy context aware ctrl-f to the book. I use this regularly with RPG manuals. I _know_ the stuff, but don't remember every detail by heart. And just ctrl-f:ing through a Mörk/Pirate Borg -style PDF isn't really productive (they're "artistically" laid out). But I can just ask an AI bot that has the pdf indexed like "how does the medical kit work?" and it'll give me a summary along with the relevant rolls within seconds.
- LorenPechtel 7d agoIs there any local version of an AI that can do this sort of thing?
- tygon 7d agoWhy use a bulldozer to push down a mound of dirt when you can do the same with a shovel? It is faster, easier, and less prone to giving you back pain. Even assuming you are only reading the section of the book related to your issue, a computer works a fair bit faster.
- FearNotDaniel 7d agoImportant to note is that what is being measured here is the ability of the models not of the chat tools themselves, which combine model completions with other tools that the models can call upon. The mainstream labs already know this about models, it's no secret, and in fact training materials from e.g. Anthropic are at pains to point out that users, or analysts designing workflows, have the reponsibility to ensure the correct tools are used and that human verification takes place at appropriate stages depending on the risk/consequences of the task at hand. Of course a language-completion model with a training cutoff date won't have up-to-date information on tax rules or the ability to carry out correct numerical calculations, but when you combine that with (in Claude terminology) web search and code execution tools invoked by the chat agent, you immediately have much more reliable results.
- ehe78qhe 7d agoI keep finding that the current harnesses, when encountering syntax that was invalid at training time but is now valid due to new language versions or custom extensions, don't correctly figure out why and assume something is wrong with the codebase or toolchain. I would hate to have that happen with my taxes.
- htrp 7d agoI guess the question becomes, how much of this is harness versus model?
- AnimalMuppet 7d agoWell, lag time against current tax law is definitely model.
- johnnienaked 7d agoIt has very little to do with lag time on tax policy