4 ms·
Whoa what. This a game breaker. I’ve been using this tool for some personal side project stuff and had no idea that this was happening. Whoa. Not good. To edit
by soultrees 3y ago
Whoa what. This a game breaker. I’ve been using this tool for some personal side project stuff and had no idea that this was happening. Whoa. Not good.
To edit: I was never expecting full privacy as I assumed everything being fed into your service was being padded to your data set to make phind a competently fine tuned agent, but open web is not good.
In the end, the amount of times I’ve sworn at your bot with some foul language after it forgets everything on the 3rd chain and I have to go back and redo everything, will probably harm your seo if those swears get published. Just saying.
- rushingcreek 3y agoThe cached link is currently shareable but won’t get indexed anywhere unless you yourself share it somewhere like Twitter. But we’ve heard the community’s concern about security through obscurity and we will address this shortly.
- _3ruj 3y agoAre the 'cached' pages eventually deleted? I've seen some shared `https://www.phind.com/search?cache= https://www.phind.com/search?cache=*` URLs which redirect to the Phind landing page. Wiping these pages would not be good from the preservation standpoint, especially since the URL doesn't seem human-readable.
- hanselot 3y agoSo how hard would it be to brute-force to some share links then? I have been using this for a lot of searching.Should probably have read the ToS. But to be fair, it feels way worse to use now than when it came out, so I guess I won't lose much by switching away at this point. I hate this thing assuming it knows what I want. Previously it had predictable behaviour, and acted like a slow LLM powered search engine. Now I do a search and it starts asking me shit. I already tabbed away to three other places to do similar searches and come back to some half baked shit that has nothing to do with what I searched for. Plus I have to wait for it to go and do inference again before I can get to a link? Why forget all the things that search engines got right over the years, but keep all the shit they got wrong? https://www.phind.com/agent?cache=cll1bg5np0005l008molmt7d2 https://www.phind.com/agent?cache=cll1bg5np0005l008molmt7d2 import random import string import requests # Static prefix prefix = "cll1" # An empty set to store the cache links we've already tried tried_links = set() # Website URL url = "http://target.com/agent?cache=" # Dictionary to store cachelink and its response text cache_link_dict = {} while True: # Generate a random 21 character base 36 encoded string suffix = ''.join(random.choices(string.ascii_lowercase + string.digits, k=21)) # Combine the prefix and suffix to create the cache link cache_link = prefix + suffix # Continue with a new loop iteration if we've already tried this link if cache_link in tried_links: continue # Add the cache link to our set of tried links tried_links.add(cache_link) # Make a request to the website with the cache link response = requests.get(url + cache_link) # Store the cachelink and its response text in the dictionary cache_link_dict[cache_link] = response.text # Add a delay between requests to avoid overwhelming the server time.sleep(1) So it took 17 minutes to use your website to create the code needed to start enumerating these links. Now given, I am pretty sure a large number of these will return nothing, but I am also willing to bet there are people that put GDPR data into these searches, perhaps a lookup of a phone number, or an address, names? It would be quite trivial adding in the code to run these requests through a pool of proxies so they don't trigger anything too suspicious on your WAF. I don't think I'm particularly skilled either, so if I can do this, I'm sure someone already has.
- rushingcreek 3y agoYou can still use the basic search mode if you prefer maximum control.
- hanselot 3y agoI actually loved the original version when it just came out. It beat ChatGPT every time simply because it unlocked a portion of ChatGPT locked down by the knowledge cutoff. It was also quite speedy, and even the way it resists going into infinite loops was much better than ChatGPT at the time. I assumed my data might be used to train AI in an obtuse obfuscated kinda way, but I never would have imagined I could just brute-force cache links.
- Sunhold 3y agoWhy do you believe you can brute-force cache links? There could be some vulnerability but that string looks far too long to brute-force.
- hanselot 3y agoYou are probably right, I was just sleep deprived and angered when I wrote that reply. I am now only sleep deprived. I still use phind, and just came back here to say sorry to whoever, since it really does help me with productivity. Also, it will require more time to reverse engineer the cache links, and I don't want to spend more time. First all the requests return 403 so think there will need to be a selenium component (user-agent trickery is not sufficient).
- Frotag 3y agoTheres ~10^31 possible links. Assuming theres 100 billion valid links and someones scanning 1mil links per second, they'd still be averaging 1 discovered link per million years.
- jbellis 3y agofwiw I do like how easy it is to share conversations now. if you change it, please make the friction to sharing minimal.
- rushingcreek 3y agoI think that we'll made it private by default but with an easy "share" button you can click that opens it up.