Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sergey-obukhov
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
sergey-obukhov
10y ago
you're absolutely right!
2.
▲
by
sergey-obukhov
10y ago
While the research and prototyping were done in R the service itself is in python. The implementation is so simple that it doesn't require any specific ML libs. More complex tasks will probably require using scikit (if desired to stay
3.
▲
by
sergey-obukhov
10y ago
There are many approaches https://en.wikipedia.org/wiki/Cluster_analysis . In the research k-means clustering was used. It's probably not the best clustering algo for this task and a much better results would be ac
4.
▲
by
sergey-obukhov
10y ago
What part? :)
5.
▲
by
sergey-obukhov
10y ago
Re-pasting my comment here where it belongs as a reply: It could be a great idea - to plot the data on 2 axes (x - html length or html size, y - processing time, if I understand this correctly). It's simple and elegant. I'll try t
6.
▲
by
sergey-obukhov
10y ago
Html size and html tags count was a natural choice. If it didn't work out the next step would be to try something else. You're right that it's a very naive example and that in a way it was solved before any ml was applied. Th
7.
▲
by
sergey-obukhov
10y ago
It could be a great idea - to plot the data on 2 axes (x - html length or html size, y - processing time, if I understand this correctly). It's simple and elegant. I'll try that. It could be though that the chart will get messy wi