Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
closed
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
closed
5y ago
Author here. My concern in the article isn't that you can't do groupby fast, but that the approach above is not composable. * If f() converts grouped data to something ungrouped, then you can't use a similar function f2(f(gro
32.
▲
by
closed
5y ago
I've been working on using github projects beta for the past couple months. I'm super excited for it, but right now it's hard to use because because it lacks basic features. Things projects have that beta projects don't
33.
▲
by
closed
5y ago
> What I’m willing to bet, however, is that you’ve heard of deliberate practice in the context of Malcolm Gladwell’s ‘10,000 hour rule’ — the mistaken notion that 10,000 hours of practice would turn anyone, at any age, for any skill, int
34.
▲
by
closed
5y ago
I think the quote flouting the rule of 3, and the discussion in the article is analogous to a similar phenomenon with face attractiveness. If you average across many faces, you get an attractive face. It's a safe bet for an attractive
35.
▲
by
closed
5y ago
On my laptop each run of boto3.client("s3") takes about 4 milliseconds. I'm guessing this is okay for most crud apps..! Edit: especially compared to any interactions with s3
36.
▲
by
closed
5y ago
Weirdly enough, Momento is a big hit with a good chunk of memory researchers, and had been used as a stimulus in memory experiments! https://www.aalto.fi/en/news/film-memento-helped-uncover-how...
37.
▲
by
closed
5y ago
The article mentions Brier score is just mean squared error, so it's connected to binomial through that (e.g. where correct prediction is 1, incorrect is 0, it is the mean of the binomial).
38.
▲
by
closed
5y ago
> Bad code doesn't work. > It's very hard, and maybe impossible, to determine if a novel "works". I wonder if there is a an assumption here about what it means to "work" vs to be "bad". Psycholo
39.
▲
by
closed
5y ago
For what it's worth, I maintain a library called siuba that lets you generate SQL code from pandas methods. It's crazy to me how people use SELECT * -> pandas, but also how people in SQL type a ton of code over and over. https
40.
▲
by
closed
5y ago
In case you're interested in what's missing--I maintain a port of dplyr from R to python called siuba, and gave a talk recently on why pandas might be hard to use: https://www.rstudio.com/resources/rstudioglob
41.
▲
by
closed
5y ago
Is there a specific aspect of pandas research you're interested in? There are a lot of useful guides around table-based workflows that might be helpful :). I would start w/ different strategies on how to model data in tables. One
42.
▲
by
closed
6y ago
What you described is pretty similar to my experience. I wonder if part of the author's sentiment about it being more widely deployed can be explained in part by stack overflow trends data. Basically, MySQL used to make up a much large
43.
▲
by
closed
6y ago
I'm surprised every time someone says: 1: This field requires a lot skill. 2: I am not skilled in this field. 3: It's obvious to me <observation that might require skill in domain>. And then there's no appeal to an expe
44.
▲
by
closed
6y ago
It seems like the bulk of OSS developers I know do not get paid, but are obsessed with a particular problem domain (or have essentially merged with their tool and become a finely tuned cyborg).
45.
▲
by
closed
6y ago
Have you tried the python port of ggplot, plotnine? I screencast doing live data analyses with jupyter notebooks, virtualenv, and plotnine. I'm definitely much quicker in R, but it's not too bad! https://youtu.be/z
46.
▲
by
closed
6y ago
I've been building data analysis tools on top of SQLAlchemy's declarative system over the past couple years. It's got to be the most well documented, carefully designed library I've ever interacted it :). It looks like m
47.
▲
by
closed
6y ago
I'm seeing a lot of comments on why SELECT * is fine or not, but it seems like the bigger issue is that (in general) SQL has two very limited ways of letting your select columns: explicitly naming each column, or getting all columns. T
48.
▲
by
closed
6y ago
It's worth noting that the general models ELO approximates, called item response theory (IRT) models, could handle larger than 1v1. Its similar to how you model situations where multiple latent skills contribute to performance (multidi
49.
▲
by
closed
6y ago
I usually reach for a friend, or someone I've met before, since using the first version of a doc is asking a lot! (And they're often part of the target audience).
50.
▲
by
closed
6y ago
I love architecture docs, but find they're often written using a funny process: 1. Spend a long time writing the doc. 2. Wait for a person to chance upon it. 3. Hope you anticipated their questions. It seems like the most im
51.
▲
by
closed
6y ago
I'm familiar with both Rmarkdown and Jupyter Book. Rmarkdown also uses pandoc. Both are very flexible.
52.
▲
by
closed
6y ago
AFAIK most prepackeged UPS devices are a battery and inverter. My hot take is that if you need to run this device continuously it will be a substantial (& custom?) build. The biggest factor is how portable it needs to be. Overview: Most
53.
▲
by
closed
6y ago
One critical piece methods miss is they can't decentralize contribution. For example, the gganimate package in R gives user new ggplot functions. With a `+` users can use functions from any package, so the gganimate approach works. Wit
54.
▲
by
closed
6y ago
In case anyone is wondering what the big deal is with ggplot / plotnine, I record myself doing hour long data analyses in python with it! I've noticed that a lot of bootcamp grads can use matplotlib to do very simple plots, but wh
55.
▲
by
closed
6y ago
> Nor do I find charming the belligerent lack of any magical syntactic sugar for `self`. Does Python force you to pass it as an argument to make some kind of clever point? On the class, you can call the method like a normal function (pas
56.
▲
by
closed
6y ago
I'm still debating chaining vs piping, but you can do.. from siuba import _ from siuba.data import mtcars # mtcars is a pandas DataFrame mtcars \ .groupby("cyl") \ .siu_summarize(avg_hp=_.hp.mean())
57.
▲
by
closed
6y ago
Using the experimental fast grouped pandas functions, it should run at the speed of optimized pandas code! Since siuba functions just run on pandas DataFrames, you can always hand tune for performance, but imo most of the time pandas code r
58.
▲
by
closed
6y ago
Hey, creator of siuba here. I think siuba's big advantage is that it can generate SQL code. The architecture necessary to pull off executing either pandas or SQL also makes it very extensible (e.g. to spark or dask in the future :). h
59.
▲
by
closed
6y ago
siuba does the SQL translation :). pandas is used for local data, since it does a lot of optimization in c++, similar to dplyr's low level code. Thanks for bringing this up--the docs could be clearer here
60.
▲
by
closed
6y ago
Siuba uses type hints to dispatch the appropriate versions of custom functions! For example, siuba allows users to create custom functions using a thin wrapper around functools.singledispatch. When deciding how to run... df >> f
More ›