Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tmurray
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
tmurray
14y ago
Look at part three of this series: http://roguelikedeveloper.blogspot.com/2008/01/death-of-leve... Considering this article is four years old, a lot of this didn't pan out. Spore failed to come remotely close to its hype, Borderlands' gun
62.
▲
by
tmurray
14y ago
"You know some things to be impossible. Most things that were impossible or impractical years ago became possible or will become possible some time later. Your experience might tell you that something you want to do can’t be done. Other peo
63.
▲
by
tmurray
14y ago
Your API documentation (or specification if you have one) is not your specification. Your API implementation is your specification, especially if you are the only provider of the API. I realize how trite that sounds, but it's true. Defensiv
64.
▲
by
tmurray
14y ago
This post is very accurate. I build APIs for a living (CUDA), and this lines up pretty well with my experience. Writing APIs is very tough, you will get a lot of things wrong, and the fixes available to you after you realize your mistake ar
65.
▲
by
tmurray
15y ago
I took the class that this book was written for when I was at CMU, and it's an excellent book. If you're a programmer with some understanding of C (read: pointers don't fill you with terror), this book will give you a solid overview of syst
66.
▲
by
tmurray
15y ago
It sounds like you're basically proposing Plan 9--lots of very small, single-purpose components with one standardized way of passing data between them used across the entire system. Also, you can ask a lot of people to give up binary compat
67.
▲
by
tmurray
15y ago
"If we're really moving into a many cores future as people suggest, then we should have architectures that give each application a dedicated core. No context switching." You probably don't want that either because the cost of moving data be
68.
▲
by
tmurray
15y ago
$85-100k seems to be about entry level for software engineers with a BS in CS from a good school at established companies in this area. My impression is that the Google/Facebook talent war has significantly increased this over the past few
69.
▲
by
tmurray
15y ago
this is a good description of the individual concerns, but it's missing a really important consideration that most API designers seem to ignore: what's the ideal way (or ways) for someone to use your API? once you have some answers to that,
70.
▲
by
tmurray
15y ago
Does anyone know what this improvement actually is? The article makes it sound like their card is doing DMA to pinned user allocations, but that doesn't sound too revolutionary. I know next to nothing about storage controllers, but see the
71.
▲
by
tmurray
15y ago
I wonder if they have a patent on using an automated natural language processing system to fight patent trolls.
72.
▲
by
tmurray
15y ago
There's another side to being optimistic that's extremely important when you're proposing an idea. I don't mean assuming every new idea is great or that every proposal will work out exactly as planned; that's obviously a recipe for disaster
73.
▲
by
tmurray
15y ago
Yeah, this article left me speechless. Why would you use jwz to justify startup culture? jwz left the tech industry completely in 2000 and bought DNA Lounge.
74.
▲
by
tmurray
15y ago
It depends entirely on where your bottlenecks are. If the bottleneck is entirely within your node, then this isn't going to be compelling. If you're doing something that's very light on the resources within your node (serving static content
75.
▲
by
tmurray
15y ago
Each CPU is a separate node in this configuration--separate DRAM, IO, etc.
76.
▲
by
tmurray
15y ago
You assume the application is compute limited and that the extra performance on the Xeon translates into extra performance on a given application. That's probably not a good assumption for this kind of workload.
77.
▲
by
tmurray
15y ago
From an OS POV both of those tests are doing almost exactly the same thing: make a system call, have the kernel spawn a new unit of execution (whether it's a thread or a process), wait for the child to be scheduled, have it terminate, wait
78.
▲
by
tmurray
15y ago
As the paper mentions, latency of this magnitude is already available to HPC using RDMA on top of InfiniBand or 10GigE. I don't think 10GigE will be a special purpose interconnect that much longer and IP over OFED (OpenFabrics, the very low
79.
▲
by
tmurray
16y ago
CUDA and BrookGPU share the same creator, but there's not really anything else in common. The programming models are fundamentally different (multi-level memory hierarchy with sharing of local data with fast synchronization versus pure stre
80.
▲
by
tmurray
16y ago
Global memory.
81.
▲
by
tmurray
16y ago
I think there's some preallocated memory for dynamic allocations from the device. As far as I know virtual functions work as well, and the usual thinking about warp divergence does apply.
82.
▲
by
tmurray
16y ago
You can take an arbitrary pointer value in a function, look up its location, and do the copy if necessary, instead of requiring multiple entry points for CPU and GPU pointers. This also simplifies multi-GPU programming a lot--copy from one
83.
▲
by
tmurray
16y ago
It's not a true unified address space in the way I think you're getting at. We carve out the address space, but you can't just directly dereference a GPU pointer from CPU code or vice-versa. What we guarantee is that you can determine the l
84.
▲
by
tmurray
16y ago
PCIe can be a limitation, but there are a lot of ways to amortize the latency of copying data to/from the GPU. Overlapping transfers with kernels, direct load/store from system memory to a kernel, multiple kernels running on the chip at the
85.
▲
by
tmurray
16y ago
Been working on the API improvements for this for quite a while now. If there are any questions I can answer, I'll tell you whatever I can. (edit: in case it's not obvious, I work for NVIDIA on CUDA)
86.
▲
by
tmurray
16y ago
That's exactly what I did, and I work at NVIDIA. What a company says at career fairs is not necessarily applicable to everyone.
87.
▲
by
tmurray
16y ago
Work on interesting projects, be active in the appropriate communities--that gets you noticed. Grades aren't everything.
88.
▲
by
tmurray
16y ago
The Remote Desktop issue is actually fixed as of CUDA 3.2 if you have a Tesla card that is only doing compute (e.g., you're not running a Tesla C2050 as a display card). full disclosure: I work on the CUDA driver stack and the dedicated com
89.
▲
by
tmurray
16y ago
No, shared memory is only for intra-block communication. __syncthreads() only ensures that every thread in a block is at a particular point rather than every block in a grid. Take Fermi, for example--you can potentially have 128 blocks runn
90.
▲
by
tmurray
16y ago
It's not really suited to CUDA/OpenCL; for the purposes of this discussion, we can treat the two as the same because the execution model of actual kernels is almost identical. The problems are - the size of the grid is fixed at kernel launc
More ›