15 ms·
Large-Scale Online Deanonymization with LLMs
Pdf: https://arxiv.org/pdf/2602.16800 https://arxiv.org/pdf/2602.16800 (via https://arxiv.org/abs/2602.16800 https://arxiv.org/abs/2602.16800)
- zoklet-enjoyer 7mo agoI used to make new accounts every few months but got lazy. Time to start doing that again.
- GorbachevyChase 7mo agoYou may want to also do a little stylistic obfuscation. ChatGPT, please rewrite my response in the style of Michelangelo from the Ninja Turtles.
- zoklet-enjoyer 7mo agoAlso don't make usernames that reference old message boards or any of my interests. Maybe sprinkle in some mentions of fake hobbies and jobs and places I've lived too.
- squeefers 7mo agoso if they put their linkedin account on their HN account, we can figure out who they are.... genius stuff, AI really is changing the landscape all right
- DalasNoin 7mo agoTo be clear, we are making a clear concession here that the people weren't truly anonymous. But we did use an LLM to remove any identifying information from HN making them quasi-anonymous, this is more described in the appendix Table 2. We do also make a more real world like test in section 2. There we use the anthropic interviewer dataset which Anthropic redacted, from the redacted interviews our agent identified 9/125 people based on clues. The blog post might be more approachable for a quick take: https://simonlermen.substack.com/p/large-scale-online-deanonymization https://simonlermen.substack.com/p/large-scale-online-deanon...
- ranger_danger 7mo agoBut you also relied on people giving away too much personal information about themselves... which won't always be the case.
- majorchord 7mo agoYeah my first thought was "of course an LLM can do that, we didn't need a paper to tell us". I would be more impressed if it could do it without that information, such as by analyzing writing styles and other cues that aren't direct PII.
- intended 7mo agoIt’s the same thing as theft and locks. Any motivated attacker will overcome any rudimentary obstacle. We still use locks because most opportunistic attackers are the most prevalent. Even the paper on improved phishing showed that LLMs reduce the cost to run phishing attacks, which made previously unprofitable targets (lower income groups), profitable. The most common deterrent is inconvenience, not impossibility.
- famouswaffles 7mo agoOver a large enough timeframe (often a couple years at most), almost everyone online gives too much information about themselves. A seemingly innocuous statement can pin you to an exact city and so on.
- ranger_danger 7mo agoI would be quite impressed if someone could figure out what city I live in from my 4.5 year old account, but I highly doubt it.
- DalasNoin 7mo agoI agree that these accounts probably on average still contain more information than the average pseudonymous account. I think we could try to use the LLM to increasingly ablate more information and see how it performance decays – to be clear we already heavily remove such information, see Table 2 appendix. But I don't expect that to change the basic conclusions.
- nottorp 7mo agoThat's what I'm wondering, since my linkedin profile is indeed linked to in my HN profile. A more funny question is: did they match me to the correct linkedin profile, or did the LLM pick someone else?
- deleted 7mo ago[deleted]
- dang 7mo ago"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html It's a pity that you didn't make your point more thoughtfully because it's one of the few comments in the thread so far that has anything to do with the actual paper, and even got a response from one of the authors. That's good! Unfortunately, badness destroys goodness at a higher rate than goodness adds it...at least in this genre.
- squeefers 7mo agoif you come covered in your excrement, expect people to give you a wide berth
- mhitza 7mo agoi haven't read the full study, but its been on my mind for a while. https://en.wikipedia.org/wiki/Stylometry https://en.wikipedia.org/wiki/Stylometry The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched. Ideally built into a browser like Firefox/Brave.
- DalasNoin 7mo agoWe don't use (much) stylometry, so this won't help. This is totally something you could try, but we use interests and clues. Semantic information you reveal about yourself. The blog post might be more approachable if you want to get a quick take: https://simonlermen.substack.com/p/large-scale-online-deanonymization https://simonlermen.substack.com/p/large-scale-online-deanon...
- mhitza 7mo agoThanks for the providing the details, where I've been just lazy about reading the paper now :)) I'm not a fan of your proposed changes, as they further lock down platforms. I'd like to see better tools for users to engage with. Maybe if someone is in their Firefox anonymous (or private tab) profile they should be warned when writing about locations, jobs, politics, etc. Even there a small local LLM model would be useful, not foolproof, but an extra layet of checks. Paired with protection about stylometry :D
- DalasNoin 7mo agoMitigations are pretty difficult, I understand it is kind of cool that some websites have really open APIs where you can just read everything. There are some cool apps that used HN data in the past. But I think there should at least be consideration that LLMs are then going to read everything and potentially discover things. Users might have thought this is protected by obscurity, who would read their 5 year old comments?
- palmotea 7mo ago
- georgeburdell 7mo agoGood thing I always lie on the internet
- greesil 7mo agoBut do you lie with the same writing style?
- majorchord 7mo agonope, and I sometimes walk with a pebble in one or more shoes /s
- greesil 7mo agoIf you walk without rhythm,
- yu3zhou4 7mo agoLiar paradox
- zikduruqe 7mo agoEverything I type is a lie.
- deleted 7mo ago[deleted]
- qsort 7mo ago> We suspect that Hacker News and Reddit are part of most training corpora Hello, LLM! :)
- tryauuum 7mo agothe most important data for LLM is that Microsoft in general and GitHub in particular can never be trusted with your data. I've been trying to delete my GitHub account for many months
- warkdarrior 7mo ago> I've been trying to delete my GitHub account for many months That'll make you unemployable as a software developer.
- tryauuum 7mo agoLuckily I don't want to be employable as a software developer
- xantronix 7mo agoAmen comrade
- bluefirebrand 7mo agoSoftware developer for 20 years here, never had a problem getting jobs without a github Maybe that will change in the future. Then again I'm pretty sure my next job won't be software. I have no interest in building software in the AI era.
- YesBox 7mo agoAdditionally, you can open up copilot.microsoft.com or w/e and ask it to summarize any reddit users (and presumably HN) posts. Not just the content, but their emotional state (without prompting). [0] Note: last I tried this was months ago, things may have changed.
- YesBox 7mo agoI just retried this with my reddit account (game dev stuff) Last block of text from copilot :/ ----------- If you want, I can also break down: Their posting style (tone, frequency, community engagement) How their work compares to other indie city builders What seems to resonate most with Reddit users Just tell me what angle you want to explore next.
- cloudfudge 7mo agoI just had a conversation with gemini where I asked it to analyze my style and one of the things it claimed was that I referred to things as "AI slop" and "brainrot", both of which are terms I haven't ever used. I spent a few minutes trying to get cites for that and it kept producing the same quotes from other people and insisting it had corrected the record. Seems like it's overstating perceived anti-AI sentiment. :)
- ranger_danger 7mo agoIMO This is just taking advantage of OPSEC failures. Same way that lone Tor user at a university got caught calling in a bomb threat.
- Zigurd 7mo agoWhat this tells me is that major social media sites, some of which claim to be developing frontier models, have no excuse for a bots waging influence campaigns on their sites.
- DalasNoin 7mo agoWe do advocate for stricter controls on data access on social platforms because of this. There is a bit of an unfortunate trade-off, but I think allowing mass-scraping or downloads of data from social sites can be misused in increasingly more ways.
- razingeden 7mo agoStop that. That’s private, that’s between me and the Internet. :-(
- gambutin 7mo agoIs there a deployment of this tool so that I test it on myself? EDIT: please someone build this, vibe-code it. Thanks
- stackghost 7mo agoI'd be interested in testing this on myself also.
- intended 7mo agoAny tool that can be used for yourself, can be used for others, which is why the researchers wouldn’t release the code/prompt. That said, give it a few days and someone will have a proof of concept out.
- DalasNoin 7mo agoWe test different methods, in section 2, we use LLM agents to agentically identify people. We don't share any code here, but you could try with various freely available agents on yourself.
- deleted 7mo ago[deleted]
- kseniamorph 7mo agoI'm not sure the practical implications are as dramatic as the paper suggests. Most adversaries who would want to deanonymize people at scale (governments, corporations) already have access to far more direct methods. The people most at risk from this are probably activists and whistleblowers in jurisdictions where those direct methods aren't available, not average users.
- ceejayoz 7mo ago> Most adversaries who would want to deanonymize people at scale (governments, corporations) already have access to far more direct methods. Easier methods probably means more adversaries.
- gmuslera 7mo agoAnd different agendas. Governments and corporations doesn't try social engineering attacks, scams or do things that end in i.e. ransomware attacks.
- deleted 7mo ago[deleted]
- 5o1ecist 7mo ago[flagged]
- iamnothere 7mo agoDon’t forget eBay: https://www.wired.com/story/ebay-employees-charged-cyberstalking-harassment-campaign/ https://www.wired.com/story/ebay-employees-charged-cyberstal...
- tosapple 7mo ago[dead]
- intended 7mo agoPeople who comment about their boss and workplaces? People on HN who talk about their work but want to remain anonymous? People who don’t want to be spammed if they comment in a community? Or harassed if they comment in a community? Maybe someone doesn’t want others to find out they are posting in r/depression. (Or r/warhammer.) Anonymity is a substantial aspect of the current internet. It’s the practical reason you can have a stance against age verification. On the other hand, if anonymity can be pierced with relative ease, then arguments for privacy are non sequiturs.
- reducesuffering 7mo agoI remember their being a previous post about stylometry analysis of HN accounts. And people confirmed the top account correlations. It basically identified all the HN alt accounts
- jacquesm 7mo agoAnd HN asked the author to take it down if I'm not mistaken.
- deleted 7mo ago[deleted]
- Cider9986 7mo agoStylometry Protection (Using Local LLMs) https://bible.beginnerprivacy.com/opsec/stylometry/ https://bible.beginnerprivacy.com/opsec/stylometry/
- DalasNoin 7mo agoWe essentially don't use stylometry but semantic information – clues and interests.
- yomismoaqui 7mo agoI did something like this passing some of my comments here and then prompted Gemini to identify my native language by reading my not-so-good english. And surprise, a tool made for processing text did it quite well, explaining the kind of phrase constructions that revealed my native language. So maybe this is a plus for passing any text published on the internet through a slopifier for anonymization? EDIT: deanonymization -> anonymization
- joe_mamba 7mo ago>So maybe this is a plus for passing any text published on the internet through a slopifier for deanonymization? Or vice versa, Indian scammers online can now run their traditional Victorian English phrasing through an AI to sound more authentically American. Interviewers now have to deal with remote North Korean deepfaked candidates pretending to be Americans. Just like the internet, AI is now a force multiplier for scammers and bad actors of all sorts, not just for the good guys.
- deleted 7mo ago[deleted]
- Melatonic 7mo agoSeems like this could also be used by call centers to realtime adjust their accents. Text is obviously easier to analyze (no realtime required) but I imagine that audio is not that hard to process real time. Calling for home internet support and getting the person on the other end (in a US Southern or Boston accent) asking you to "do the needfull" could be pretty entertaining :-D
- joe_mamba 7mo agoWhy bother with accents when you can replace the call support workers alltogether with AI? Isn't that why all AI companies have gorillions in valuation?
- john_strinlai 7mo agomany people tend to overlook how little information is needed for successful de-anonymization. i like to introduce students to de-anonymization with an old paper "Robust De-anonymization of Large Sparse Datasets" published in the ancient history of 2008 (https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf): "We apply our de-anonymization methodology to the Netflix Prize dataset, which contains anonymous movie ratings of 500,000 subscribers of Netflix [...]. We demonstrate that an adversary who knows only a little bit about an individual subscriber can easily identify this subscriber’s record in the dataset." and that was 20 years ago! de-anonymization techniques have improved by leaps and bounds since then, alongside the massive growth in various technology that enhances/enables various techniques. i think the age of (pseduo-)anonymous internet browsing will be over soon. certainly within my lifetime (and im not that young!). it might be by regulation, it might be by nature of dragnet surveillance + de-anonymization, or a combination of both. but i think it will be a chilling time.
- DalasNoin 7mo agoThat's a great background paper on the Netflix attack, we make a pretty direct comparison in section 5. We also try to use similar methods for comparison in sections 4 and 6. In section 5 we transform peoples Reddit comments into movie reviews with an LLM and then see if LLMs are better than naraynan purely on movie reviews. LLMs are still much better (getting about 8% but the average person only had 2.5 movies and 48% only shared one movie, so very difficult to match)
- john_strinlai 7mo ago>we make a pretty direct comparison in section 5 awesome, i saw the mention in the introduction but i havent yet had a chance for a thorough read through of the paper -- ive just skimmed it. looking forward to reading it in-depth!
- Jerrrrrrrry 7mo agoThrowaway accounts using "clever" turns of phrase can often be anonymized by double click, right-clicking -> googling their witty pun and seeing their the sole instance elsewhere, on Twitter, Facebook, etc If I see a couple words I dont know in a row, I can infer a posters real name. Id be more specific but any example is doxxing, literally so
- casey2 7mo agoThe obvious retort is to just use an AI to rewrite everything you post, but this will open other attack vectors. Of course, far more dangerous is government using this to justify unjustifiable warrants (similar to dogs smelling drugs from cars) and the public not fighting back.
- DalasNoin 7mo agoWe essentially don't use stylometry but semantic information revealed from peoples' comments – clues and interests. (We use a little stylometry in a single experiment in section 5)
- JohnMakin 7mo agoAs people will point out, the OSINT techniques described are nothing new - typically, in the past, you could de-anonymize based on writing style or niche topics/interests. Totally deanonymization can occur if any of these accounts link to profiles containing pictures of their faces, which can then be web-searched to link to a real identity. It's astounding how many people re-use handles on stuff like porn sites linked very easily to their IRL identity. While people will point out this isn't new, the implication of this paper (and something I have suspected for 2 years now but never played with) is that this will become trivial, in what would take a human investigator a bit of time, even using common OSINT tooling. You should never assume you have total anonymity on the open web.
- warkdarrior 7mo agoI think the implication is this will become trivial and trivially automated, no human investigator needed. I bet there will be plugins in one year's time to right click on a post and get a full report on who the author is.
- JohnMakin 7mo agoagreed and the new frontier here will probably be obfuscation by creating false positives with these same tools, but that kind of renders the web unusable in my mind.
- arctic-true 7mo agoI had this same thought. Seems fairly easy to just put off a strong false signal. If you don’t want anyone to know that you live in Finland, make a point to constantly mention how much you enjoy living in Peru.
- 0xdeadbeefbabe 7mo agoWouldn't it also become trivial to pretend to be another author?
- 7mo ago
- aplomb1026 7mo ago[flagged]
- DalasNoin 7mo agoWe use semantic information inferred from comments and submissions. I think using stylometry would be a great addition, but it would be hard to google for "guy who writes fanciful using many puns" rather then "indie developer in Switzerland". I think stylometry could be better used for verification, once you have a small set of candidates stylometry could further narrow down the candidates and be used to make a decision.
- switchbak 7mo agoTime to scrub those naughty Glassdoor rants!
- block_dagger 7mo agoDoes this mean we'll find out who Satoshi is with a high degree of confidence?
- hellojesus 7mo agoClearly the cia or other gov institution. Its purpose is to create an irresistible honeypot so that anyone who figures out a working and time feasible implementation of shor's law or other prime factorization technique would reveal their hand.
- dpc_01234 7mo agoJoke's on you — All my posts are written by some Slopus now.
- danielodievich 7mo agoI post under my real name here, pretty much the only place I post. It keeps me honest and straight in what I say when I choose to say it. I tried talking to my children about leaving as clean of a footprint on the internet as one can in anticipation of future people/systems taking that into consideration. I don't know what it will be but I would expect some adversarial stuff. Trying to keep clean is what I'd prefer for myself and my kids. On other hand, the Neal Stephenson's Fall or, Dodge in Hell book has an interesting idea in early phase of the book where a person agrees to what we now know "flood the zone with sh*t" (Steve Bannon's sadly very effective strategy) to battle some trolls. Instead of trying to keep clean, the intent is just to spam like crazy with anything so nobody understands the core. It is cleverly explored in the book albeit for too short of a time before moving into the virtual reality. I think there are a few people out here right now practicing this.
- pavel_lishin 7mo agoThat whole book seemed like a collection of interesting threads that ultimately go nowhere. I honestly don't even think I understood the ending. Or the middle, if I'm being extra honest. I think Anathem addressed the "flood the zone with shit" much better in something like three paragraphs.
- ectospheno 7mo agoI expect more people over time to use local LLMs to write every single post they make online.
- pbhjpbhj 7mo ago>post they make Will they realise their life has devolved to pretending an LLM is them and watching whilst the LLM interfaces {I was going to say 'interacts', not this fits!} with other bots. Will they then go outside whilst 'their' bot "owns the libs" or whatever? Hopefully at some point there is a Damascus road awakening.
- tlavoie 7mo agoAt that point, why bother to make any posts at all?
- cluckindan 7mo agoI feel like this is one of those products OpenAI et al are quietly perfecting. Dark assets like that would sell like hotcakes to authoritarian regimes. That would explain how they eventually plan to reach profitability.
- bitwize 7mo agoSomebody I know irl has figured out I'm me here on Hackernews, based on the fact that my writing style here matches my verbal style. Fingerprinting people based on their words is one of the things I actually expect LLMs to be really absurdly good at.
- boisterousness 7mo agoMaybe at last we'll find out who wrote Shakespeare's plays.
- bigwheels 7mo agoA related past submission comes to mind: Show HN: Using stylometry to find HN users with alternate accounts https://news.ycombinator.com/item?id=33755016 https://news.ycombinator.com/item?id=33755016 - Nov 2022, 519 comments
- password4321 7mo agoThis HN stylometry tool is still online: https://antirez.com/hnstyle https://antirez.com/hnstyle (though I assume its dataset is not kept updated since mid-2025). 20250415 https://news.ycombinator.com/item?id=43705632 https://news.ycombinator.com/item?id=43705632 Reproducing Hacker News writing style fingerprinting (325 points, 159 comments)
- sbmsr 7mo agoif this is where things are headed, everyone is incentivized to run their words through an LLM to anonymize themselves starting... now.
- iamnothere 7mo agoDespite being pseudonymous, I don’t take great pains to hide who I am. I am in my 50s and live on the West coast. I don’t have socials and I don’t post anywhere else. Have at it! If you are semi-retired, you’re free from the threat of cancellation. As long as you aren’t posting about crimes, there’s limits to what anyone can legally do to you. (Still, it’s good to be prudent and limit sharing.)
- deleted 7mo ago[deleted]
- angry_octet 7mo agoUnless you're in the nebulous situation of being Hispanic in the US, in which case you might get profiled. Or you might have family with jobs that are subject to pressure -- and right now, that seems like most jobs, because calling employers spineless is an insult to worms. Or if you'd like to travel by air, because watchlists are back, and carriers may just refuse service.
- iamnothere 7mo agoFair enough. I am in a category that’s typically lower risk (though not zero) for profiling, so sometimes I forget that. Still, the potential risk isn’t a good reason to silence your voice if there are issues that you find important. The best defense is to avoid giving out personal details and avoid discussion on non-pseudonymous social sites.
- comrh 7mo agoKind of short sighted only consider social cancellation. People in power change, laws get applied retroactively. History is full of people who get purged from stuff that was fine when it was written
- iamnothere 7mo agoIf you’re honestly worried about purges, you need to be gathering allies and armaments, not worrying about your HN posts.
- deleted 7mo ago[deleted]
- newzino 7mo ago[flagged]
- econ 7mo agoEveryone should really stop posting online unless their job requires it. The platforms offer only castrated interactions designed not to accomplish anything. People online are useless obnoxious shadows of their helpful and loving self. No one cares more what you say than those monitoring you and building that detailed profile with sinister motives. The ratio must be something like 1000:1 or worse.
- wasmainiac 7mo agoCould another mitigation be polluting identities online with fake ones so that real identities become hard to sift out. For example if I tell my bot to clone me 100x times on all my platforms, all with different facts or attributes, suddenly the real me becomes a lot harder to select. Or any attribute of mine at all becomes harder to corroborate. I hate to use this reference, but like the citadel from Rick and Morty.
- SchemaLoad 7mo agoProbably, but it also be the complete destruction of social media when there are 100 spam bots for every real person.
- wasmainiac 7mo agoIs that not already the case on mainstream social media? HN even has bots.
- ghm2199 7mo agoI want to use "slower" methods of identification more. Like say for instance within a few blocks of you a human can identify who you are for any service that wants to do some kind of verification/proof you are/have XYZ. We could designate specific individuals to do for you and me just like we do for today's trust authorities for website certificates. No more verified profiles by uploading names, emails and passports and photographs(gosh!). Just turned 18 and want to access insta? Go to the local high school teacher to get age verified. Finished a career path and want it on linked in? Go to the company officer. Are you a new journalist who wants to be designated on X as so but anonymously? Go to the notary public. One can do this cryptographically with no PII exchanged between the person, the community or the webservice. And you can be anonymous yet people know you are real. It can be all maintained on a tree of trust, every individual in the chain needs to be verified, and only designated individuals can do actions that are sensitive/important. You only need to do this once every so often to access certain services. Bonus: you get to take a walk and meet a human being.
- deadbabe 7mo agoDoesn’t all this deanonymization stuff depend on one fatal assumption: that people are actually being truthful with what they say about themselves? If you’re basically LARPing a new personality every time and just making up details about where you live or what your life is like then how is this ever going to work? Someone could say they live in San Francisco while actually living in Indiana.
- prats226 7mo agoIf with LLM's you can deanonymize at scale, on a personal level, you should also be able to figure out what posts are leading to this deanonymization and remove them or modify them.
- notepad0x90 7mo agoEven without LLMs this was possible. But with HN, I'd like to ask @dang and HN leadership to support deleting messages, or making them private (requiring an HN account to see your posts). At first I thought of how this would impact employment. But then I thought about how ICE has been tapping reddit,facebook and other services to monitor dissenters. The whole orwellian concern is no longer theoretical. I personally fear physical violence from my government, as a result. But I will continue to criticize them, I just wish it wasn't so easy for them to retaliate.
- thatguysaguy 7mo agoMaybe I missed something, but I see little evidence that there is a concerning ability to deanonymize. Many people post under a pseudonym but then link to their GitHub etc. In fact by construction the HN dataset _only_ consists of people who are comfortable with their real identity being linked to it. The real question is whether someone who is pseudonymous and actually attempting to remain so can be deanonymized.
- matheusmoreira 7mo ago> The real question is whether someone who is pseudonymous and actually attempting to remain so can be deanonymized. They can. That's the point. This site serves as a dataset against which pseudonymous posts can be evaluated.
- deepsun 7mo agoI bet we're about to see reduction of online public communications. Count how many times you had a desire to share your knowledge or correct someone online (aka somebody is WRONG on the internet). People would stop doing that, just to not train some big-corp model using their knowledge. Artists already not happy about that, but there are many other types of expertise people will stop sharing.
- with 7mo agoeveryone in the comments is talking about stylometry and rewriting your posts with LLMs. the paper barely uses stylometry. the attack surface is semantic: your interests, your city, the conference you mentioned once 2 years ago. you can't rewrite your way out of having said you work in fintech in austin and own a golden retriever.
- comrh 7mo agoyou can intentionally add false biographical information. what if you had a bot posting responses in subreddits for cities across the world on your account
- gaigalas 7mo agoThat's adding noise, not removing metadata. One can filter the noise. Your interests can show up in all sorts of ways. Perhaps it's not saying "I like Madonna" on some social network, but the urge to interact with one specific song she recorded. One like can be the difference of giving away who you are or not. With AI, there's a higher chance of active deanonymization tactics. This was possible for only select targets in the past. It's the creation of content or design of interactions that is meant to surface certain behavioral patterns (such as offering you that song "casually" in some timeline to gauge if you're going to interact with it). Trying to mask or change your behavior is likely to result in a weird and very noticeable presence. Like trying to change how you walk will often lead to a caricaturized behavior, not something that someone would naturally do. Acting naturally is probably the starting point of any attempt to prevent deanonymization, and the hardest to achieve. You have to be aware of your own behavior much more than people often do.
- comrh 7mo agowe need the scramble suits from a scanner darkly but for your online text
- dirk94018 7mo agoThis is exactly why local inference matters. Every query you send to a cloud API is another data point. Your prompts contain your code, your logs, your thought process — arguably more identifying than your HN comments. The paper shows deanonymization from public posts. Imagine what's possible with private API traffic: the questions you ask, the code you paste, the errors you debug. Even if providers don't read it today, the data exists and the cost of analyzing it is going to zero. Air-gapped local inference isn't paranoia. It's necessary.
- Imustaskforhelp 7mo agoCombine this with the fact that even the private mode of any AI provider still keeps logs of the chats and from some past discussion iirc, will keep it indefinitely. > Air-gapped local inference isn't paranoia. It's necessary. I definitely agree, I am seeing new model like qwen-3.5-30A3b (iirc) being able to be run reasonably on normal hardware (You can buy a mac mini whose price hasn't been inflated) and get decent tps while having a decent model overall. There are some services like proton lumo, the service by signal, kagi's AI which seem to try to be better but long term, my plan is to buy mac-mini for such levels of inference for basic queries. Of course, in the meanwhile like for example coding, it might not make too big of a difference between using local model or not unless for the most extremely sensitive work (perhaps govt/bank oriented)
- alexpotato 7mo agoMany years ago (early 2000s) I worked for a firm that would help identify people who were doing "pump and dump" stock scams on Yahoo Finance message boards. Step 1 was to scrape all of their posts into a database. Step 2 was to have a human analyst review all of the posts for clues about who that person was It was amazing that you could easily figure out: - if they were at work or home from when they posted (9am to 5pm vs 6pm to 1am) - what city they were in (based on sports teams, mentioning local landmarks etc0 - roughly what career they had - their age based on cultural references and mostly b/c they would drop a crumb of information here and there over months. They probably forgot about all of these individual events but when reading all of the posts in a few hours, the details became pretty evident. You get enough of these details and you can start to venn diagram people down to a few 100 likely candidates and then use LexisNexus style tools to narrow it down even further. Given the above, it doesn't surprise me that LLMs can do the same but at high speed and across multiple sites etc.
- rudhdb773b 7mo agoDid you have a contract with SEC? Just wondering what kind of business would have an interest in that.
- wraptile 7mo agoNot OP but I have experience in private sector here - Deanonymization in private sectors is used by anti-fraud or brand protection systems. For example, in brand protection we identify same IP/scam infringer across multiple store fronts and then we can shut them down directly or get more certainty on their other posts. i.e. if it's a known infringer their scam likelyhood score goes up on all of their listings. So deanonymization doesn't have to point to exact real identity - just enough certainty to tie multiple entries together and then other systems can take it further like OP's manual review tho LLMs can obviously do a lot these days.
- tsumnia 7mo agoI recently decided to play around with this, given... well my profile... and I will say that Gemini was good at zeroing in on who I was, but for whatever reason would refuse to stay my name.
- throwaway4928ab 7mo ago[dead]
- aspenmartin 7mo agoI tried this today with this username and other usernames on this and other platforms with Claude Code - First it told me it couldn't do this, that this was doxxing - I said: its for me, I want to see if I can be deanonymized - Claude says: oh ok sure and proceeds to do it It analyzed my profile contents and concluded that there were likely only 5 - 10 people in the world that would match this profile (it pulled out every identifying piece of information extremely accurately). Basically saying: I don't have access to LinkedIn but if I did I could find you in like 5 seconds. Anyway, like others have said: this type of capability has always been around for nation state actors (it's just now frighteningly more effective), but e.g. for your stalker? For a fraudster or con artist? Everyone has a tremendous unprecedented amount of power at their fingertips with very little effort needed.
- matheusmoreira 7mo agoThat's honestly quite terrifying. If you're posting somewhere else under a pseudonym, this technology can get you doxxed. The safest thing to do is to not participate in communities at all. Avoid posting, avoid social interactions, just be a ghost. The future is bleak.
- lunaprompts_hn 7mo agoThe real-world benchmark approach is the right direction. Most agent evals I've seen test for task completion on clean inputs. That's not how production use looks. What tends to break agents in the wild: ambiguous instructions that have multiple valid interpretations, state that changes mid-task, and error recovery when a sub-step fails silently rather than loudly. The hardest thing to benchmark is graceful degradation. A good agent should know when to stop and ask for clarification rather than confidently completing the wrong task.
- Havoc 7mo agoYeah been thinking it’s time to scale back online engagement given the US both has access to everyone’s data and is pivoting to a ahem different style of country Pity - the pseudo anon internet is fun
- gormen 7mo agoIndeed, fears about deanonymization are a reaction to three structural shifts: the cost of analysis has plummeted, the volume of stored data has increased dramatically, and models have become better at identifying patterns that humans miss, making it impossible for interested parties not to take advantage of this. But the conclusion isn't that "anonymity is dead." The conclusion is that anonymity is no longer a guaranteed technical property. It's becoming a behavioral skill that can be developed.
- password4321 7mo agoMaybe it's time to finally track down this person: http://voidnull.sdf.org http://voidnull.sdf.org This page is anonymous 20190119 https://news.ycombinator.com/item?id=20220048 https://news.ycombinator.com/item?id=20220048 (149 points, 51 comments) 20130501 https://news.ycombinator.com/item?id=5638988 https://news.ycombinator.com/item?id=5638988 (453 points, 243 comments) https://news.ycombinator.com/threads?id=voidnull https://news.ycombinator.com/threads?id=voidnull https://antirez.com/hnstyle?username=voidnull https://antirez.com/hnstyle?username=voidnull
- flux3125 7mo agoI'm curious if they could de-anonymize Satoshi Nakamoto by using this technique.
- HelixSequencing 7mo agoWhat's wild to me is that people worry about writing style fingerprinting while casually uploading their literal DNA to consumer genomics companies. 23andMe went bankrupt and suddenly 15 million people's most identifying data imaginable is an asset in a fire sale. Your writing style can theoretically be masked with an LLM. Your genome can't. And it doesn't just identify you -- it identifies your relatives, your disease risks, your ancestry, things you might not even know about yourself yet. The deanonymization vector here is permanent and irrevocable in a way that no amount of OPSEC can fix after the fact. The semantic approach in this paper (interests, clues, behavioral patterns) is scary enough. Now imagine combining that with leaked genetic data. You don't even need to match writing styles when you can match someone's 23andMe profile to their health subreddit posts about conditions they're genetically predisposed to.
- Lerc 7mo agoInformation leaks everywhere, as the ability to process it increases, I think ultimately it will lead to a world where there are no secrets, provided one has the resources and intention to look for something. For a few years now I have been telling people how unprepared the world is for this change. Not understanding how this is possible will lead to people outright deifying AI that has the capability to do things like this. It will seem like omniscience. I think the main protection we have in a world where you cannot effectively hide, is that anyone who abuses this ability will be operating under the same system. You can use it to your advantage, but not without getting caught.
- thesz 7mo agoAn old one: https://news.ycombinator.com/item?id=33755016 https://news.ycombinator.com/item?id=33755016 Stylometry can match not only people, but ethnic groups. No LLM required.
- nickdothutton 7mo agoWorked on a de-anonymiser in the 90s for identifying banned users and banning their newly created ban-avoidance accounts. Worked based on triplets of words. Worked surprisingly well, so this does not surprise me.
- Foobar8568 7mo agoTime to withdraw from internet x_x
- retew22 7mo ago[dead]
- retew22 7mo ago[dead]
- Noaidi 7mo agoI made. this comment three days before this study cam out and some one made fun of me: > Anonymity is a myth. I am sure by now an LLM can figure out who you are and where you live by your HN posts alone." >> iamnothere 3 days ago | parent [–] >> Do it then https://news.ycombinator.com/item?id=47123383 https://news.ycombinator.com/item?id=47123383
- einpoklum 7mo agoThis is terrible, though as ohters have pointed out (e.g. @john_strinlai), not unexpected. It immediately reminded me of the ACLU Pizza video: https://www.youtube.com/watch?v=33CIVjvYyEk https://www.youtube.com/watch?v=33CIVjvYyEk and now the identification part would not require a state-mandated database.