4 ms·
Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.
by 7734128 1mo ago
Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.
- 0xbadcafebee 1mo agoI'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models
- HDBaseT 1mo agoA small number of inputs in a large dataset can poison training data pretty drastically. Anthropic wrote a good article about it a while back [0]. This should mean its possible to pull back that information fairly easily. It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways. [0] https://www.anthropic.com/research/small-samples-poison https://www.anthropic.com/research/small-samples-poison
- owaiswiz 1mo agodoesn't mean the raw text goes into training. they most likely have a pipeline to clean out any secrets before they train on it?
- HDBaseT 1mo agoIn theory, but in practice how difficult is that?
- hhh 1mo agonot hard for secrets with explicit patterns and existing pipelines to detect them
- 0xbadcafebee 1mo agoUnless they're base64-encoded or compressed?
- mhh__ 1mo agoI would expect, although have no evidence, that any obviously high entropy crap like base64 and so on probably would get removed whether it's a secret or not.
- userbinator 1mo agoIf my experience with image generation is any indication, unless AWS keys are somehow extremely prevalent in the training data, you may get something that looks like one, but it definitely won't be valid.
- wrsh07 1mo agoPrice segmentation at its finest
- sourcecodeplz 1mo agothey key pricing is cache reads at $0.002 per M (same as old deepseeek v4 flash prices)