Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vishvananda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
vishvananda
6y ago
Less years, but I concur.
32.
▲
by
vishvananda
6y ago
There is also a rust implementation that I wrote in my time at Oracle. Unfortunately they no longer maintain it, but there is a fork with some more recent updates: https://github.com/drahnr/railcar
33.
▲
by
vishvananda
6y ago
Unfortunately I don't remember the exact numbers, but I think it was a couple percentage points worse than we were able to get with the large models.
34.
▲
by
vishvananda
6y ago
we did a lot of our early experimentation with small networks. I don't think we went any smaller than 5 layers of 64 filters as we mentioned here: https://medium.com/oracledevs/lessons-from-alpha-zero-part-5...
35.
▲
by
vishvananda
6y ago
The lc0 group has switched the result prediction to predict win, loss, and draw probabilities instead of just win/loss. Some information can be found in https://lczero.org/blog/2020/04/wdl-head/
36.
▲
by
vishvananda
6y ago
Nice work on this! I was behind the implementation at oracle which you referenced in the tutorial. I still keep tabs on the lc0 crowd which seems to be pushing into new ideas. Did you pull anything else from the leela crowd besides prior-te
37.
▲
by
vishvananda
6y ago
My team did an implementation of alpha zero connect four a couple of years ago. Our findings are in a series of blog posts starting at https://medium.com/oracledevs/lessons-from-implementing-alph... . We didn't man
38.
▲
by
vishvananda
7y ago
also the overhead of calling out to c can actually be quite high: https://www.cockroachlabs.com/blog/the-cost-and-complexity-o...
39.
▲
by
vishvananda
7y ago
the parent mentions in parentheses that the dir doesn't have to be in dropbox because it appears to monitor all files.
40.
▲
by
vishvananda
7y ago
Prior to my life in technology I had similar experiences to the author. I was involved with multiple spiritual groups that could be classified as cults. One mistake that I have seen people without first-hand experience make is assuming that
41.
▲
by
vishvananda
7y ago
Creator of https://github.com/vishvananda/netlink here. Would be happy to have support added.
42.
▲
by
vishvananda
8y ago
I think this is primarily due to the immense effort that NVIDIA has put into CUDA. It works very well and it is extremely fast. The alternatives for AMD are OpenCL and ROCm which have seriously lagged behind CUDA in every respect. EDIT: lot
43.
▲
by
vishvananda
8y ago
This is actually an annoying challenge of reproducible builds. In many cases it is actually useful to have a build timestamp, git sha, or build number available for debug output from the program. I've often gone as far as embedding a s
44.
▲
by
vishvananda
8y ago
Unfortunately, nix does not produce fully reproducible builds. The build environment is portable and produced in a way that it can be repeated, but due to the limitations of the software that is being built, the builds are not binary reprod
45.
▲
by
vishvananda
8y ago
I also found the tone pretty condescending, but I wasn't sure if that was a result of being translated from another language. I agree that both paths are viable.
46.
▲
by
vishvananda
8y ago
Isn't this mistaking the path for the goal? I think it is more accurate to say that the knowledge of how OSes work, basic discrete math, etc. is important regardless of how that knowledge is gained. A university education is only one w
47.
▲
by
vishvananda
8y ago
Contractors have their own hassles when it comes to open source. Pre-openstack, there was a small contracting company called Anso Labs that consisted of a few of us that were creating a private cloud for NASA. After a few months of banging
48.
▲
by
vishvananda
8y ago
Is it just me or does solving the moral dilemma of prioritization based on attributes of the person seem like a worthless endeavor? In practice, I can't come up with a case where it would matter. For example, a more reasonable metric f
49.
▲
by
vishvananda
8y ago
My evidence is mostly anecdotal, unfortunately. Most of my experimentation was KVM in KVM from about 2012 and i saw frequent long pauses in the kernel, lockups, and kernel panics. At the time, the attitude from the qemu-kvm community was th
50.
▲
by
vishvananda
8y ago
I believe this is only because ec2 does not allow nested virtualization. In my experience, nested virtualization is still buggy and can suffer from major performance issues. So although it is possible to run it in a vm, I'm not sure I
51.
▲
by
vishvananda
8y ago
I used go in machine learning contexts extensively while writing graphpipe[1]. Go is a fantastic language for servers, and distributed communication. Unfortunately, the lack of generics and dependence on interfaces and reflection makes writ
52.
▲
by
vishvananda
8y ago
I think you have this correct, as I understand it. Just to explicitly state the differences: 1. You can modify GPL code as much as you want, and as long as you don't distribute the software, you do not need to make the modifications av
53.
▲
by
vishvananda
8y ago
Yes I totally understand. In general model serving doesn't seem to be something people are thinking about at the moment. I expect it gets more attention in the next couple of years.
54.
▲
by
vishvananda
8y ago
I like how you made this ultra-simple. It makes it very easy to understand and use. There are times when a python+json server isn't performant enough. To help with this, my team recently built a protocol for efficient model serving cal
55.
▲
Show HN: Connect Four Powered by the AlphaZero Algorithm
(azfour.com)
1 points
by
vishvananda
8y ago
|
0 comments
56.
▲
Connect Four Powered by the AlphaZero Algorithm
(azfour.com)
1 points
by
vishvananda
8y ago
|
0 comments
57.
▲
GraphPipe – Dead Simple ML Model Serving via a Standard Protocol
(oracle.github.io)
1 points
by
vishvananda
8y ago
|
0 comments
58.
▲
by
vishvananda
8y ago
not op, but I assume it stands for Get Sh*t Done
59.
▲
Lessons from AlphaZero: Improving the Training Target
(medium.com)
4 points
by
vishvananda
8y ago
|
0 comments
60.
▲
by
vishvananda
8y ago
You might find http://www.fast.ai/ useful. Depending on your learning style, their courses can either be amazing or somewhat annoying. Their library includes jupyter notebooks so that you can work through the examples.
More ›