Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alexott
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
alexott
3y ago
New Python SDK covers all public Databricks REST API operations, so you can use it for all things around automation on Databricks...
32.
▲
Official Python SDK for Databricks
(github.com)
9 points
by
alexott
3y ago
|
4 comments
33.
▲
by
alexott
3y ago
Not directly related, but https://github.com/nosqlbench/nosqlbench is very flexible benchmark tool for Cassandra and other distributed systems
34.
▲
by
alexott
3y ago
I stumbled this week on few answers on a topic that I’m following, and although they were really nicely looking, correctly formatted, etc. - they were absolutely wrong - ChatGPT was hallucinating, showing calls to functions that aren’t in s
35.
▲
by
alexott
3y ago
Yes from my previous experience there were common language pairs, especially on the short texts - Spanish / Catalan, often German / Dutch, Bulgarian / Ukrainian, Dutch / Afrikaans, and few more…
36.
▲
by
alexott
3y ago
Looking into generated ngrams, I’m not sure that it’s good idea of getting rid of spaces between words, and not having markers for begin/end of word. It would be interesting to check on the dataset linked to this blog post from 6 years
37.
▲
by
alexott
3y ago
Grokking series is from Manning and do they sell PDF and ePub/mobi directly
38.
▲
by
alexott
3y ago
Who said that people are rational? Especially in green party?
39.
▲
by
alexott
3y ago
Did you reinstall OS, or it works as upgrade?
40.
▲
Apache Spark 3.4.0 has been released
(spark.apache.org)
3 points
by
alexott
3y ago
|
1 comments
41.
▲
by
alexott
3y ago
It includes Python client for Spark Connect, more ANSI SQL compliance, very good improvements in Structured Streaming, new built-in functions, and much more. There is a blog post on Databricks with more details for most notable features: h
42.
▲
by
alexott
4y ago
But why such complexity? Is it easier to maintain than terraform code?
43.
▲
by
alexott
4y ago
it was before, but it was relying on the "normal clusters" with longer startup times, limited scalability, etc. Now it's serverless with fast startup times, scaling to 0, etc.
44.
▲
by
alexott
4y ago
10 years ago I’ve attended a talk from one of German car producer - they talked about 25+ years of software maintenance. You invest into tools 5 years before release to market, and they should work 20 years after release. So companies need
45.
▲
by
alexott
4y ago
I have an Integration of Planet Clojure directly using API (direct posting allows attribution to specific person). So, 8.5k followers will be cut from that news source…
46.
▲
by
alexott
4y ago
Im struggling with previous behavior all the time. It would be much better with this
47.
▲
by
alexott
4y ago
Even for data transformation logic, SQL isn’t the best choice. How would you handle the case when you need to apply the same transformations to few dozens or hundreds columns?
48.
▲
by
alexott
4y ago
there is Pandas on Spark, included into Spark itself (originally Koalas) - the switch to it is very easy, and you get parallelization.
49.
▲
by
alexott
4y ago
yes, but it's a best thing of the cloud - your cluster doesn't run when you don't need it, plus you can advantages of spot instances, autoscaling, etc.. And you won't do TPC test each every half an hour.
50.
▲
by
alexott
4y ago
That number should be "The total 3-year price of the entire Priced Configuration must be reported, including: hardware, software, and maintenance charges", so they just took the cost of the hardware used for benchmark, and extende
51.
▲
by
alexott
4y ago
Same for me - sometimes UI changes aren’t good, but usually they were addressed
52.
▲
by
alexott
4y ago
There are some implementation details about Databricks’ Photon engine in this paper: https://cs.stanford.edu/~matei/papers/2022/sigmod_photon.pdf
53.
▲
by
alexott
4y ago
If you are using IDE you can look to dbx Python package that has the “dbx sync” command that allows to sync local code with Repos and then you can quickly test your code in notebooks
54.
▲
by
alexott
4y ago
You can use Databricks Repos ( https://docs.databricks.com/repos/index.html ) specifically files in repos ( https://docs.databricks.com/repos/work-with-notebooks-other-... ) functionality that allows
55.
▲
by
alexott
4y ago
As I remember it was introduced in 2018th: https://www.elastic.co/blog/an-introduction-to-elasticsearch... , although there was an open source extension for that… Although I wasn’t very hard to implement it for a subset
56.
▲
by
alexott
4y ago
It may work, until you start to perform maintenance things, like repair, or new node bootstrap, etc. Then it may fail with high probability
57.
▲
by
alexott
4y ago
But then you need to push these segments into partitions, and big partitions are really bad, especially for old versions of Cassandra… Although I met customers with partitions of size of 100Gbs…
58.
▲
by
alexott
4y ago
https://www.gnu.org/software/emacs/manual/html_node/emacs/Ke... ?
59.
▲
by
alexott
4y ago
Have you seen Database Internals? https://www.databass.dev/
60.
▲
by
alexott
4y ago
Yes, 100%. I’m trying to use registration information for cybersecurity stuff, and it’s a mess. Some TLDs just doesn’t provide that information or provide it only to registered accounts or only inside their country. Parsing is a mess. Many
More ›