3 ms·
Here is an example Pipeline which supports loading data from an unbounded source to BigQuery in batches using load jobs (evading BigQuery's Streaming Insert cos
by xstartup 8y ago
Here is an example Pipeline which supports loading data from an unbounded source to BigQuery in batches using load jobs (evading BigQuery's Streaming Insert cost)
See: https://zero-master.github.io/posts/pub-sub-bigquery-beam/ https://zero-master.github.io/posts/pub-sub-bigquery-beam/
- cobookman 8y agoFYI dataflow's bigqueryio does not use stream inserts. Instead it batches the data into many load jobs
- oomkiller 8y agoThat is configurable, but for unbounded PCollections it appears the default is to use streaming inserts https://beam.apache.org/documentation/sdks/javadoc/2.4.0/org/apache/beam/sdk/io/gcp/bigquery/BigQueryIO.Write.Method.html https://beam.apache.org/documentation/sdks/javadoc/2.4.0/org...
- AWebOfBrown 8y agoAre you the author? OT but I'm amazed the author charged merely 100 euro for implementing that solution for the subject startup, even if they're cash-strapped. I'm not familiar with BigQuery, but I'm curious what a normal rate for solving that issue would look like.
- mentat 8y agoThat's shockingly cheap for the value.
- manigandham 8y agoYou can also skip pub/sub and/or use it to write files to cloud storage, then load from there by using a cloud function that will trigger the load job when a new object is created.