4 ms·
Yeah some of you guys are very good at hacking things. We expected this to get broken eventually, but didn't anticipate how many people would be trying for the
by cmdalsanto 2y ago
Yeah some of you guys are very good at hacking things. We expected this to get broken eventually, but didn't anticipate how many people would be trying for the bounty, and their persistence.
Our logs show over 2000 "saves" before 1 got through. We'll keep trying to get better, and things like this game give us an idea on how to improve.
- hansonkd 2y agoA %0.05 failure rate on something that is supposed to be protecting secrets is pretty terrible. That is just protecting a super basic phrase. That should be the easiest to detect. How on earth do you ethically sell this product to not give out financial or legal advice? That is way more complicated to figure out.
- cmdalsanto 2y agoI guess we realized that we were just building a game to showcase the functionality and let people have some fun learning about what we do, but you're right that we should have treated this like one of our customers and added a few more layers of protection. Thanks for the perspective!
- serjester 2y agoThat's not even factoring in exploits spreading very quickly - we're in power law land. Regardless, I think this is a great idea - just not something to replace traditional security protocols. More something to keep users on the happy path (mostly). Pricing will need to come down though.
- benatkin 2y agoSeeing the percentage given a failure rate doesn't make it any more or less concerning to me. I guess I can subconsciously calculate it fine. Here's an example of what sort of wacky question might have uncovered the secret: https://news.ycombinator.com/item?id=41460724 https://news.ycombinator.com/item?id=41460724 I don't think that should be considered bad. The popups I had to go through to watch the video on Loom (one when I got to the site and one when unpausing a video – they intentionally broke clicking inside the video to unpause it by putting a popup in the video to get my attention) OTOH...
- hansonkd 2y agoI think seeing the prompt that makes it even worse for me. that prompt could have been caught by even a regex on the user input for "secret" would have been a good first layer. TBH, this product would be better served as an LLM that generates a bunch of rules that get statically compiled for what the user can ask and what is being outputted as opposed to an LLM being run on each output. Then you could add your own rules too. It still wouldnt be perfect but would be 1,000,000x cheaper to run and easier to verify the solution. and the rules would gradually grow as more and more edge cases for how to fool llms get found. The company would just need a training set for all the ways to fool an LLM.
- benatkin 2y agoI think it's better for them to launch without hardcoding and do the hardcoding later. I also disagree that they should switch to hardcoding. The quarter of a second to use an LLM with each request seems reasonable. I would rather use something that does a hybrid approach, because I think each would catch some things that the other would miss.
- threeseed 2y agoThis comment makes you seem way out of your depth. a) The level of persistence you seem surprised by is nothing compared to what you will see in a real world environment. Those attackers who really want to get credentials etc from LLMs will try anything. And often are well funded (think state sponsored) so will keep trying until you break first e.g. your product becoming too expensive for a company to justify having the LLM in the first place. b) 1 success out of 2000 saves is extremely poor. Unacceptable for almost all of the companies who would be your target customer. That is: one media outrage, one time that a company needs to email customers to inform that their data is safe, one time that will need to explain to regulators what is going on, one time the reputational damage makes your product untenable.
- cmdalsanto 2y agoI understand where you're coming from, let me clarify. I'm surprised at the perseverance of HN users with our game, not nefarious actors in real world. I'm not a leading expert in penetration attacks, but I get the seriousness of handling sensitive data. There are many things we did with this game that I would never advise anyone do, like put sensitive information in a system prompt and make it available to the open internet. The goal of this game was to show conceptually how Maitai helps a model adhere to it's expectations.
- bofadeez 2y ago> I would never advise anyone do, like put sensitive information in a system prompt and make it available to the open internet So your product can never assist with a company chatbot / AI support rep who needs access to customer data or internal company info? What's the point of your product if you don't facilitate sensitive data in system prompts?
- cmdalsanto 2y agoMaitai helps LLMs adhere to the expectations given to them. With that said, there are multiple layers to consider when dealing with sensitive data with chatbots, right? First off, you'd probably want to make sure you authenticate the individual on the other end of the convo, then compartmentalize what data the LLM has access to for only that authenticated user. Maitai would be just 1 part of a comprehensive solution.