3 ms·
I was hoping there would be some machine learning in here. This just seems to be cross referencing a couple of different data sources.
by languagehacker 10y ago
I was hoping there would be some machine learning in here. This just seems to be cross referencing a couple of different data sources.
- jgalt212 10y agofair enough, but people cross reference because the total is often more than the sum of its parts.
- donalhunt 10y agoagreed. I think the OP needs to break the problem down into 1) access from bots that identify themselves (e.g. googlebot, bingbot, etc) 2) access from bots that masquerade as humans and 3) activity that is from humans. #1 should be trivial by looking at the user-agents. #2 and #3 could utilise machine learning to categorise the behaviour of the connecting party, etc It's not particularly clear why the OP wants to detect bots... I suspect it's to get a clearer signal of what assets are being accessed by humans.
- yeukhon 10y ago#2 and #3 is a moot point. Is my selenium / python script bot when I am reproducing human behaviors for my testing? I guess you can called it a bot or an automata. So we need a definition of what is considered a robot. The best detections for human vs non-human so far have been (1) introducing captcha and (2) based on view time and interaction (hotspot). With both we can have a reasonable criteria to build a somewhat simple model. If based on server access log, then you need to group resources together (css, js, images, html) or ignore most of the resources, and calculate view time. Based on some reasonable expectation, a user who views multiple pages at the same time or within +X seconds would likely be a robot than a human despite we humans are used to opening up several tabs once we are used to how to use the website. As many tests do put a sleep in between (since there's always lag) so for the non-intrusive robot, we will have a difficult time distinguishing within some reasonable confidence. If we enabled tracing / tracking, then this becomes slightly easier as well since we can learn the behavior of "new" and "veteran" users. This is why tracking is such as a privacy issue for many.
- gav 10y agoYou could feed in a whole bunch of data to determine if a user is a bot or not: * Did they request robots.txt? * User-agent * Logged in vs. guest account (if applicable) * Number of requests/time period * Patterns of requests The real value of this kind of analysis (to me) is to bucket the types of visitors: * bot * well-behaved (e.g. googlebot, etc.) * valuable * not-valuable * badly behaved * visitor The valuable vs. non-valuable traffic is interesting to me, I've seen sites where the bot traffic was 25-35% of the total, and that some of the bots, though well-behaved, didn't really bring any business.