Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dandermotj
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
dandermotj
10y ago
That was the joke I was making...
62.
▲
by
dandermotj
10y ago
The "big data vs. small data" argument isn't the issue for data scientists. It's the total lack of decent statistical libraries, visualisation tools, programming capabilities and on and on. Don't think anything I&#x
63.
▲
by
dandermotj
10y ago
So is Hadley Wickham... https://twitter.com/hadleywickham/status/748392441154248704
64.
▲
by
dandermotj
10y ago
It is absolutely the interface to excel that causes errors. I can take your excel spreadsheet, format cells, add things, change a reference and generally fuck it up and hand it back to you and you would never know. Whereas, if I change a te
65.
▲
by
dandermotj
10y ago
So just to give you an opposing perspective of what you just said: "we have __one__ person that does all...". So apart from the fact he can't access a database appropriately, he's the _only_ person who does this work and
66.
▲
by
dandermotj
10y ago
This guy has three options: 1) short 2) go long 3) do nothing. Given he knows about this vulnerability, what is the "ethical" thing to do. But the stock? Don't engage? I think, in fact, he has a duty to short the stock.
67.
▲
by
dandermotj
10y ago
Everyone is missing this. They used simple random forests and survival models with basic k-fold cross validation. There's no breakthrough in ML here!
68.
▲
by
dandermotj
10y ago
I genuinely think these type of thoughts come from having extensive experience as a programmer, that can consider building out systems that might reach the performance ceiling of R. For 99% of people multicore/distributed architecture
69.
▲
by
dandermotj
10y ago
Very good, just look at sparklyr and the likes. People seem to think there is some hang up here. R has some of the best developed packages for working out of memory.
70.
▲
by
dandermotj
10y ago
Multidimensional scaling is designed to preserve the most variance (or some variation of that idea) among points in a lower dimensional space. So for the most part, the cluster memberships assigned by clustering in an n-dimensional space ar
71.
▲
by
dandermotj
10y ago
Go one step further with Wickham's purrr too. Offers well developed functional programming tools!
72.
▲
by
dandermotj
10y ago
The straight forward answer you'll get is to study a pure/applied science in university up to MS or PhD level. This is probably half of the answer. The other half is being voraciously curious in many many areas and actively self t
73.
▲
by
dandermotj
10y ago
The US is one of the only countries on earth that follow this practice.
74.
▲
by
dandermotj
10y ago
But in the real world AML isn't going away, and it continues to be expensive.
75.
▲
by
dandermotj
10y ago
There seems to be some bizarre skepticism by some HNers of financial institutions' AML practices. But the thing is, AML is really difficult and expensive for banks, hiring floors of consultants dedicated to AML. This is a billion dolla
76.
▲
by
dandermotj
10y ago
From Ireland, so I know what you're talking about. I suppose my point was really that Lidl and Aldi have swept into an market over the last 10 years, where customers are hard fought for, seemingly with ease. The Tescos of the world def
77.
▲
by
dandermotj
10y ago
Aldi and Lidl are blasting through other super market chains with their low cost, high quality goods. If there were ever perfect case studies for usurping incumbents in a sharply competitive market Aldi and Lidl are it.
78.
▲
by
dandermotj
10y ago
The only cost tied to these issues is the salary costs for a statistician/data scientist. The best tools of the trade are open source. You won't find lowering costs until more people are educated or trained up in these fields.
79.
▲
by
dandermotj
10y ago
Next step is have it output the LaTeX code. Wouldn't that be nice!
80.
▲
by
dandermotj
10y ago
I agree and I think it's clear TPOT and similar tools, are the first generation. Genetic algorithms might be slow/costly but the concept is there with an integrated, implementable solution. If it's as useful as I think it cou
81.
▲
TPOT: Automatically Optimize ML Pipelines with Genetic Algorithms
(github.com)
4 points
by
dandermotj
10y ago
|
0 comments
82.
▲
by
dandermotj
10y ago
There's a key statement here that most will either miss or never reach. The second last paragraph > Machine Learning Automation Tuning models is a pain, and if you don't understand the model and all its parameters (like most so
83.
▲
by
dandermotj
10y ago
This is awesome
84.
▲
by
dandermotj
10y ago
The entire scene - a gigantic white cuboid shape delicately placed between mountains and desert - is science fiction brought to life.
85.
▲
by
dandermotj
10y ago
This article's attempt to contribute Leicester's success this season to "statistics" is just painful. As if every other team in the Premier League and three tiers down don't use extensive analysis in every aspect of
86.
▲
by
dandermotj
10y ago
I understand where you're coming from, but these two lines are exactly 'taking a list of libraries for import'. packages is a character vector of package names, lapply is by definition 'list apply'. We're takin
87.
▲
by
dandermotj
10y ago
In terms of managing packages in session, readability and keystrokes, I think it definitely wins out over successive library calls. It's common to have >5 packages in any one script, especially if you avoid base R like many do.
88.
▲
by
dandermotj
10y ago
This is a really interesting post. I have to say the data science team at Stitchfix are clearly doing really good, applied work that is central to their business. It's so cool to see. Here's a tip for any R users who read thro
89.
▲
by
dandermotj
10y ago
This book is extremely practical - I would definitely recommend it for doing actual time series analysis. That said, it's not a learn time series book . It's a do basic time series in R book.
90.
▲
by
dandermotj
10y ago
Another open source Bayesian book/course all in python: http://www.greenteapress.com/thinkbayes/t
More ›