3 ms·
Hi folks: So we're monitoring most major 2020 Presidential Candidates' sites for visual + HTML/DOM + network + extracted text changes. (You can see all detected
by bluepeter 7y ago
Hi folks: So we're monitoring most major 2020 Presidential Candidates' sites for visual + HTML/DOM + network + extracted text changes. (You can see all detected changed at the above link.) There's a lot of noise! So we're using ML to identify significant changes. (You can see these findings so far at the top of the page.)
We've trained our model using detected changes from corporate sites and some earlier political sites. Each change for our model was human-rated in terms of relevance/importance, and we also feed in other descriptive attributes about each change, such as DOM location, immediate parent tag, and several other attributes.
- bhl 7y agoTo compute the text differences, are you using the HTML/DOM and rendered webpage to extract the text? Or are the method for each diff separately determined? I'm curious to what the input and output of the ML model consists of.
- bluepeter 7y agoEach diff is done slightly differently depending on what it is... we use headless Chrome + Puppeteer on Fargate for crawling. We use Puppeteer itself to take screenshots and output HTML. We then use a separate Lambda function to extract the text. From that, we feed the results into "diff" Lambda functions to compute image, text, HTML, and network diffs. We treat the text diff as the primary diff type, and so only if we have a text diff do we do the other comparison types. From these diffs, we then feed the text + some DOM info into ML. There, we use added text, deleted text, and, for each, the shortest unique CSS selector, the immediate parent tag, and some other items that I may be missing (possibly some approximation of where in the main text the change appears... e.g., top 10%, top 20% IIRC). Hopefully this answers your question?
- social_quotient 7y agoCongrats on what you’ve made! What sort of noise are you running in to? (Curious) In an old project we actually did something a bit more based on visual changes. Basically detected visual diffs and got coordinates of offsets to the change. We then found the smallest html container that encompassed the diff and highlighted it. Using visual diff you can fuzz things a bit to handle artifact and small movements. A good lib we have some miles on https://github.com/mapbox/pixelmatch https://github.com/mapbox/pixelmatch And this write up on niffy is pretty good https://segment.com/blog/perceptual-diffing-with-niffy/ https://segment.com/blog/perceptual-diffing-with-niffy/
- bluepeter 7y agoIn terms of overall noise, we're running into a lot. (This is what led us to ML as a way to hopefully reduce it.) Comparing the pure DOM reveals almost constant change due to various inserted Javascript from Google/Facebook/etc modifying the DOM. Looking at text/images also result in a fair amount of expected noise, but it's mostly in the form of interstitial marketing banners, fundraising targets, etc. We already have various options to "filter" out certain DOM areas. (As an example, remove all footers, headers, or any other CSS selector.) These work really well... but they require a fair bit of setup. Thanks for the pixelmatch GH repo link... I have it starred so must have taken a look in the past. We need to evolve our image diff, so we may end up using this!