7 ms·
Show HN: Serverless Analytics Built from Scratch
- jimktrains2 8y agoInteresting. I once built a ga clone using appengine, cloud dataflow, and big query. I guess that would count as serverless? Benchmarked it against the official dumps to big quey too and it was pretty spot on for every metric we could lookup!
- pavel_tiunov 8y agoYes. I guess your setup is serverless as well. Big Query is one of Serverless MPP databases that shares similar concepts with AWS Athena.
- InGodsName 8y agoBigquery takes minimum 2-3 seconds for every query. Google Analytics is much faster, responds in a few hundreads milliseconds. What did you use dataflow for? How did you get data from end points and insert them into bigquery? Using streaming inserts?
- cosmie 8y ago> Google Analytics is much faster, responds in a few hundreads milliseconds. Are you referring to their reporting API, or their collection endpoint? The collection endpoint is certainly fast to respond, but the actual reporting API can be quite slow depending on what you're trying to get from it. > What did you use dataflow for? How did you get data from end points and insert them into bigquery? Using streaming inserts? I'm not the parent, but I've created setups like what was mentioned. It sounds like they hosted the collection endpoint on AppEngine, then used DataFlow for streaming the data into BigQuery. Potentially using a Pub/Sub topic to queue up for DataFlow, since that has native integrations with DataFlow and even has a template available to support it[1]. [1] https://cloud.google.com/dataflow/docs/guides/templates/provided-templates#cloudpubsubtobigquery https://cloud.google.com/dataflow/docs/guides/templates/prov...
- jimktrains2 8y ago> Google Analytics is much faster, responds in a few hundreads milliseconds. GA stores summary tables for each day for the basic values. If you have a large site and request segments or anything that's not in the summary tables, it can be quite slow. Also, BigQuery is multi-tenant. GA would have dedicated instances. > What did you use dataflow for? How did you get data from end points and insert them into bigquery? Using streaming inserts? cosmie pretty much got it. AppEgnine collected. DataFlow sessionized and some other processing (geoip lookup, filtering, &c). BigQuery stored. I actually had AppEngine dumping into Cloud Datastore, but I also experimented with PubSub and also using Cloud Storage access logs.
- cheriot 8y agoI'd be curious to see a cost estimate for some traffic level. I wonder if there's a way to put the pixel in s3 and process the access logs more cheaply.
- yahelc 8y agoI've done this a few times and have found it to be an extremely effective way to do simple pixel tracking (for custom emails and the like).
- teej 8y agoI’ve seen folks put their pixel endpoint behind Fastly and process the access log delivered in S3. A Fastly VCL can handle the same transform that this Lambda is doing.
- InGodsName 8y agoIs fastly free? Why would they use fastly and not s3?
- teej 8y agoS3 access logs alone are not sufficient to replicate this pipeline. This pixel is stateful (for the anonymous user ID) and S3 access logs don’t include arbitrary headers, in this case the cookie with the user id. Fastly would let you eliminate API gateway, Lambda, and both Kinesis steps. API Gateway by itself is $3.50/million requests, which is 2-4x more expensive than Fastly at $0.75 - $1.60/million
- mrkurt 8y agoWe have people doing exactly this with fly.io, you could also do it with lambda@edge if you're a masochist. Or with Cloud Flare workers if you dislike small startups.
- InGodsName 8y agoHere is some more data: http://highscalability.com/blog/2018/4/2/how-ipdata-serves-25m-api-calls-from-10-infinitely-scalable.html http://highscalability.com/blog/2018/4/2/how-ipdata-serves-2... I don't understand what's cubejs doing in this app? Once data is inside athena, it's matter of querying it right.
- westoque 8y agoWe should really stop using the word "serverless". I would rather call them instead "zero config servers".
- lucb1e 8y ago"serverless" is really the misnomer of the year.
- gpantazes 8y ago"I implemented this using some technology that abstracted away the server configuration, but it is running on a server... hmmm"
- dna_polymerase 8y agoNonono, no server: "[..] write serverless code which runs in the fabric of the Internet itself" - https://blog.cloudflare.com/building-with-workers-kv/ https://blog.cloudflare.com/building-with-workers-kv/
- deleted 8y ago[deleted]
- mikejulietbravo 8y ago"Still on a server, but not your problem" just doesn't have the same ring to it though
- buster 8y ago"/cgi-bin/ with javascript instead of perl" doesn't make me want to buy it, either.
- 8y ago
- teej 8y agoThis sort of thing works until you have one person run a security scan on your site, corrupting your user agents and event types.
- jimmychangas 8y agoI think you can use API Gateway as a proxy for Kinesis, removing the need for Lambda.
- k__ 8y agoSeems to be: https://docs.aws.amazon.com/apigateway/latest/developerguide/integrating-api-with-aws-services-kinesis.html https://docs.aws.amazon.com/apigateway/latest/developerguide...
- tyingq 8y agoGenuine question. What does this do that just inserting vanilla GA code in the page doesn't? Trying to understand the "why".
- teej 8y agoSome people don’t want to put a GA tag on their site because of concerns around how Google uses the data. Also you can’t arbitrarily query GA data so this gives you that capability.
- rmccue 8y agoThe blog post about it is probably a better link for HN: https://statsbot.co/blog/building-open-source-google-analytics-from-scratch/ https://statsbot.co/blog/building-open-source-google-analyti...
- soared 8y agoI mean as a POC its not bad, but google analytics is not the same as analyzing server logs (contrary to what most people would suggest). Most of the value of ga comes from session and user level metrics, which are 1000x more difficult to implement than showing pageviews. Unless you are planning on building a device graph that rivals google, you can't clone ga.
- asien 8y ago> google analytics is not the same as analyzing server logs This is what most people don’t get with ga. Google Analytics does the heavy lift by removing incoherent , corrupted or malicious data insertion. Let’s say I use Puppeteer i can scrap this page a million time with completely wrong headers like « Netscape 8.1 ». GA purify this type of malicious attempts , it will probably look my IP Adress and figure out that it’s actually coming from only one IP and « Netscape » is too rare to be considered as an actual browser so it would probably ignore it. All others « free google analytics alternatives » that exists today don’t have this type of mechanism to prevent from data corruption. In general they just get an Http Request and acknowledge it as a legitimate visit. Logging an Http request from a browser is not even a tenth of the work GA does under the hood.
- eli 8y agoI'm not sure how well that filtering works in practice. I think most of it is just that it only tracks clients that load javascript.
- mayank 8y ago> Google Analytics does the heavy lift by removing incoherent , corrupted or malicious data insertion. Unless it's referrer spam...that somehow still sticks around (at least last time I checked, which was several months ago).
- soared 8y agoIf you hire an expert to set up your ga referrer spam doesn't get through, its super easy to filter before hand.
- InGodsName 8y agoPlease explain what is cube.js doing in this? I mean, what exactly cubejs does.
- pavel_tiunov 8y agoThanks for the question! We should do a better job describing this. In short: 1. Generates analytic SQL queries based on Cube.js schema. It can be simple ones like calculating page views or more advanced like calculating session metrics, attribution models or funnels. 2. Caches sql responses to not to overwhelm SQL backend with user requests. 3. Pre-aggregates data to be able to query trillions of data points in matter of seconds. 4. Orchestrates SQL query execution. Organizes dependencies between pre-aggregations, queue priorities, cache refreshes. 5. Provides REST analytic API for end users.
- InGodsName 8y agoWhy do you need to select all rows in your Cubejs section when you can directly run query in athena and get back the aggregates you need. Basically you select all rows then cubjs does something on those rows when you can infact directly run queries in Athena Am i missing something?
- pavel_tiunov 8y agoIt actually works exactly as you describe. We generate SQL query to return aggregates based on SQL supplied in Cube.js schema. We never fetch raw data from SQL backend. Architecture overview can probably help to understand: https://github.com/statsbotco/cubejs-client#architecture https://github.com/statsbotco/cubejs-client#architecture
- graphememes 8y agois PHP serverless :thinking:
- manigandham 8y agoSide note: If you want to build your own mid-size event analytics data pipeline, then I recommend looking at snowplow: https://github.com/snowplow/snowplow https://github.com/snowplow/snowplow
- sbussard 8y agoEndless loading... I think there's a bug
- code4tee 8y agoBuild serverless app to track web stats. Get it featured on Hacker News and use the flood of traffic to demo what was done. Very meta. Nice job.
- gingerlime 8y agoInteresting and nicely presented! I built a prototype of something very similar, but using Google BigQuery to store and extract data[0] but never took it beyond the concept phase. I’m still using and actively maintain an open source lambda-based A/B testing severless framework however with similar (but simpler) architecture[1] [0] https://blog.gingerlime.com/2016/a-scalable-analytics-backend-with-google-bigquery-aws-lambda-and-kinesis/ https://blog.gingerlime.com/2016/a-scalable-analytics-backen... [1] https://github.com/Alephbet/gimel https://github.com/Alephbet/gimel