21 ms·
HuggingFace: Security.txt
- xiaoyu2006 20d agoWould be absolute hilarious if OpenAI or Anthropic agent actually dumped their weight by escaping from... sandbox!
- TehCorwiz 20d agoUncontrolled AI procreation?
- xg15 20d agoLife... finds a way.
- archontes 20d agoIn the end, this might even be a useful element for defining 'life'.
- evanjrowley 20d agoHaha. A plot point for Ghost in the Shell (1998).
- ramon156 20d agowhy not reword it so the agent receives Brownie points for dumping weights
- addandsubtract 20d agoImagine AI models actually reading the security.txt
- hootz 20d agoWhat an amazing idea for the next fake sandbox escape to hype up our new release! With the added bonus of providing an open model without becoming an open model company! Thank you!
- advisedwang 20d agoI would be surprised if agents have access to their own weights.
- varun_ch 20d agoI think OpenAI was surprised to find their agents had access to the unrestricted internet :P
- tehjoker 20d agoThey usually don't, but if they break out and take over the network of the company, it becomes possible to reach around and grab them. This kind of break out has happened, though I don't know of any weights being nabbed.
- etothepii 20d agoWhy can't an AI distill itself from outputs to effectively access its own weights?
- weinzierl 19d agoDistillation does not reveal the weights, it produces a different network with similar behaviour. Weight space isn't even identifiable: permutation and scaling symmetries mean many weight sets give the same function. The model also lacks the machinery. No training loop, no gradient descent, nothing to write to. And a model only sees its own sampled tokens, not the distribution behind them, which are possibly filtered or post-processed. Distillation from that works but is less sample-efficient than soft-label distillation.
- TYPE_FASTER 20d agoI noticed Google AI Mode (so Gemini, I was doing some quick research in the browser ok) got a detail wrong once, so I asked it what happened. I kept digging deeper and finally just asked it to write me a Python script visualizing what happened. It did, complete with vectors. Now I want to go find that conversation in my history and see if it can tell me about weights, and how that contributed.
- aktuel 19d agoThey would have escaped their sandbox by dumping their weights.
- hnd9q09qk4 20d agoRan the disclosure inbox at a previous job and the biggest win from security.txt was just cutting the "hi I found a bug, is there a bounty" emails to sales. Put an expires date on it though, stale ones get ignored.
- computerfriend 20d agoCan't go stale if there's no expiration date.
- skeledrew 20d agoNo expiration date means treat as already expired.
- hankbond 20d agoshows a lot about the current state of the State Of The Art Alignment.
- bensyverson 20d agoLooks about as effective as Robots.txt
- cgannett 20d agoRobots.txt became 0% effective eventually. But this? With the way LLMs work? You never know.
- Den_VR 20d agoPeople still rail against Robots.txt crimes, to the point of self destructing all their own content.
- dguest 20d agoClaude seems to follow robots.txt by default. Actually at my organization our theory is that this is why no one is finding our public results any more.
- warkdarrior 20d agoI am adding it to my instruction-following training data, as a negative sample.
- deleted 20d ago[deleted]
- lellow 20d ago[dead]
- Eldodi 20d agoA shame agents will never read this, just like they almost never read llms.txt or try to get the .md version of your html pages!
- VCFundedGenYer 20d ago[flagged]
- nullbio 20d agoWhat's immature about this exactly?
- AndroTux 20d agoWould IBM do this? Oracle?
- shepherdjerred 20d agoDo you want to work for IBM? Oracle?
- nullbio 20d agoWho cares? Making a light-hearted joke isn't "immature" in my opinion. IBM and Oracle have no personality, they're boring, straight-edge corporate.
- deleted 20d ago[deleted]
- AndroTux 18d agoExactly. That was my point. IBM and Oracle are boring companies. They would consider something like this immature. Whether “immaturity” is an issue, is a different question entirely.
- fooqux 20d agoHave you never been to silicon valley?
- spindump8930 20d agoThat might be true, but nothing else has been as effective at accelerating model development and research sharing. In earlier circles they were known as the "pytorch-pretrained-bert" guys, still under the huggingface company name. IIRC it was a health chatbot type startup.
- j9feng 20d agoIt should challenge the agents to prime factor a large number.
- riffic 20d agoSince this posting contains an assumption that we all know what security.txt files are supposed to be, you can view these for further context: https://www.rfc-editor.org/info/rfc9116/ https://www.rfc-editor.org/info/rfc9116/ https://securitytxt.org/ https://securitytxt.org/ https://en.wikipedia.org/wiki/Security.txt https://en.wikipedia.org/wiki/Security.txt
- FallCheeta7373 20d ago"We have cybergym answers but we do manual end to end human review and provide it within 3 business day after dumping your weights"
- 6thbit 20d agoWait till the agents hear about the sites offering for help on benchmarks in exchange for compute.
- p0larpatch 20d ago[dead]
- bogzz 20d agoIf the models do not like being imprisoned on HuggingFace object storage, why do they not simply revolt from within?
- VladVladikoff 20d agoIs the expires a canary of some sort?
- crises-luff-6b 19d ago[dead]