7 ms·
Mixing instructions and data is never a good idea. And I thought people understood that.
by fg137 2mo ago
Mixing instructions and data is never a good idea.
And I thought people understood that.
- bossyTeacher 2mo agoIsn't React, the most popular JS library, an example of that? Clearly people don't understand that
- an0malous 2mo agoNo it’s not an example of that. Do you store components in your component state?
- cj 2mo agoHe probably means JSX mixing HTML with Javascript... function Greeting({ name }) { return <h1>Hello, {name}</h1>; }
- TeMPOraL 2mo agoYou just did that in a HN comment, yet nothing happened :). Could it be that the whole idea is silly misunderstanding of fundamental tenets of reality in the first place?
- cj 2mo agoWell duh. You can also post a random AI malware prompt, and I can assure you nothing will happen. What's your point?
- TeMPOraL 2mo agoMy point is that "mixing instructions and data" is a red herring, and that reality has no such distinction; what is code and what is data isn't just context-dependent but question-dependent, which you aptly demonstrated by posting "code" that in this context is just "data".
- cj 2mo agoThat's a lot of words to argue that we shouldn't be separating prompts from the (often untrusted) data the prompts are acting on. I still don't see your point.
- dev_l1x_be 2mo agoThere are so many better alternatives but it seems many people really like Word for some weird reason. The last time I cared I had to look up how to make a document starting the page numbering on the 2nd page. It turns out there are totally different ways between different versions of Word. shrug.jpg
- volkl48 2mo agoSuch as? Word hits the sweet spot of having support for all the complexity the average person may encounter/want to create. Libre, Apple Pages, and Google Docs all seem like clearly worse tools in most aspects in my experience. LaTeX is extremely powerful, but also way too complicated for the average non-HN person/person who doesn't live in complicated documents.
- gus_massa 2mo agoI almost agree. Have you tried to add an image in LaTeX that does not wander to a random page? \begin{figure}[HERE!!!!!!] or something like that. And in the old compiler, I remember a problem with bounding boxes, and keeping a eps and pdf version of each image to get a correct dvi and pdf. I think this part is fixed now.
- winddude 2mo agoid say markdown for LLMs
- IshKebab 2mo agoWhat are the "so many better alternatives"? Google Docs is pretty decent but a fair bit more basic. Proper technical authoring systems like Typst, LyX and LaTeX are way too hard for the average person. LibreOffice is much worse than MS Word.
- dev_l1x_be 2mo agoI mean workflows. Gdocs is disastrous in many scenarios. I think document management should be more like software version management with a nice ui. Editing the content (what it has) vs editing the template (how it looks). Markdown + a typst template gets you far. I have exactly that for my docs and it works very well. Not sure yet about how a decent size company could adopt anything like that.
- TeMPOraL 2mo agoSeparation of instructions and data is artificial. Reality has no such separation. A general purpose system needs not to have them either; it's a design feature, not a bug. People get too hung up on this fundamentally wrong idea, and the space of security, instead of progressing, is just running in circles like a headless chicken, making a mess of everything.
- yoz-y 2mo agoOnly in systems that need to be themselves super generalist. Which is almost never the case.
- jclulow 2mo agoLiterally all of software is artificial? Being explicit and reasoned about how you choose to allow or deny a particular computation is, surely, at the heart of a lot of computer security?
- TeMPOraL 2mo agoCode/data separation is at the heart of computer security in the same way slapstick comedy is at the heart of humor. There's an endless supply of people who think they know what is Code and what is Data, and they're always arguing with others who also think that, and neither realize that Code/Data classification is an opinion, a perspective. It doesn't hold in general. Having a separation like this makes sense for super narrow systems, where you can define the allowed and disallowed use cases, enforce the distinction (because it's not real - therefore you have to enforce it mechanistically within your system), and willing to accept that some useful operations will be denied by your system.
- WJW 2mo agoSecurity minded programmers understand that. "People" as a whole have not even heard about mixing instructions and data, and certainly not the reasons why it is not a good idea. And AI chatbots are very much targeted at the second group, not the first.
- NegativeLatency 2mo agoEven engineers like doing it sometimes. The old telephone system was so hackable because of in band signaling.
- dylan604 2mo agoSome physical constraints do not allow for the best security, and those physical constraints will always win out in the real world. When presented with the pick two of three options of fast, cheap, secure/done right, fast and cheap will always win out.
- kayodelycaon 2mo agoAnd sometimes security can’t be a high priority. When it comes to building things: fire, building codes, inspectors, public, utilities, lenders, and insurance all go before security. All those dictate whether or not you can build in the first place. Real world constraints are everywhere. :)
- TeMPOraL 2mo agoAlso more commonly, security is at direct odds with utility. Not in the least here.
- dylan604 2mo agoNothing more secure than a system the users cannot use
- DANmode 2mo agoThe “new” phone system (SS7) still relies on implicit trust and a lack of security.
- wongarsu 2mo agoPeople understand that. They just don't know how to implement that with LLMs In the GPT-2 era LLMs were just data. Instructions did not exist, and if you added them to your data they would not be followed. Then around 2022 we figured out how to patch in instruction following with a bit of fine tuning, leading to the current AI bubble. That's an ugly hack that leads to all these issues. But it's what this entire AI bubble is founded on. And nobody seems to have found a better way (or at least one that actually scales and doesn't make unreasonable sacrifices)
- TeMPOraL 2mo agoSure they would be. But for those old models, you'd have to prompt it in a framing of a screenplay or something. You're forgetting that LLMs just output a stream of tokens - the interpreter that acts on those is a piece of classical code, and sits outside of the model.
- xienze 2mo ago> the interpreter that acts on those is a piece of classical code, and sits outside of the model. Correct, but it's an LLM that's reasoning about what stream of interpretable tokens should be emitted. The interpreter can certainly apply some security measures around what's being asked of it (like ask for confirmation), but that can only go so far. Is the human in the loop always capable of understanding what's safe to execute? If not, should we pass it through another fallible LLM to help make that judgement call? Some security measures can be handled in a purely deterministic manner. But not all of them, and that's the problem.
- TeMPOraL 2mo agoThe mistake is in treating the LLM as just another deterministic, narrow computer program it can reason about. It's not. It's a "DWIM" system, and unless you can always precisely express what you mean - which you can't (not the very least because often people only realize what they meant after they get a result that's not it) - you have to treat the LLM as a human-like component. It's what it was designed to be anyway.
- catlifeonmars 2mo agoPeople who use machines based on the von Neumann architecture?
- JKCalhoun 2mo agoWhen working on PDFKit for MacOS, one short-coming our implementation had was the lack of support for Javascript in PDF's. Oops. (I mean, I'm one engineer and I was not going to try and hoist a JS runtime in my little PDFKit framework. And besides, the sample PDF's we were running into with JS were rare—usually tax-like forms that would add numbers from A and B and display the result in C. It seemed like a huge effort for such a small gain . Oh, and a security vulnerability.)
- DANmode 2mo agoArguably the biggest vuln of the filetype.
- cwmoore 2mo agoCode is data is symbolic reality. I don’t think people’s understanding changes this.
- anthk 2mo agoTell that to middle bosses and CEOs and MS Office VBA bootlickers and Excel workshippers. Meanwhile, CSV files parsed with custom reviewed AWK scripts can be 100% safe with charts made from Gnuplot. Heck, even some notebook like Ipython with a CSV module would be far more desirable than a spreadsheet. Any of them. Just look at the Genomics Disaster on Excel because of shitty parsing.
- teamonkey 2mo ago‘(Lisp would like a word)
- reaperducer 2mo agoAnd I thought people understood that. The graybeards know it. But they only know it through experience. It's blue/red/pink box phone phreaking all over again. The technology changes, but the mistakes remain the same.