Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jhj
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
jhj
3y ago
This is product quantization (a vector is chopped up into sub-vectors where each sub-vector is quantized using vector quantization (VQ)), not scalar quantization (which is what you're comparing it to here). Also most scalar quantizatio
32.
▲
by
jhj
3y ago
k-D trees and BSPs don't work well in high dimensions (say >20 to 30 dims), since your true nearest neighbor is highly likely to lie on either side of any dividing hyperplane. If each dividing hyperplane separates a set of vectors i
33.
▲
by
jhj
3y ago
Faiss GPU author here. Vector binarization via LSH for dense vector, high-dimensional (say 30 - 2000ish dimensions) tends to have poor recall and quality of results. As the Annoy author implied here as well, there's a lot of literature
34.
▲
by
jhj
3y ago
Yes but coordinating the computationally expensive parts usually needs to be done across multiple CPU threads. Typically one should feed each GPU from a different CPU thread, even in native C/C++ land, as in many cases the kernels bein
35.
▲
by
jhj
3y ago
It's an important feedstock for producing a wide variety of chemicals, and will likely continue being so.
36.
▲
by
jhj
3y ago
For performance, it's always better to explicitly manage GPU memory and host/device copies for performance than to depend upon the unified memory paging mechanism, if it's possible to go the extra effort. My feeling is that u
37.
▲
by
jhj
3y ago
Not speaking to their implementation, but prefix sums/scans are simply a very useful primitive tool for parallelizing many otherwise sequential operations. For instance, appending a variable number of items per worker to a shared coale
38.
▲
by
jhj
3y ago
Cool, this definitely seems like a good enumeration of techniques, nice to see that they discuss stuff like kernel fission as well. Having a good understanding of loop nest optimization transformations (tiling, fission, fusion, strip mini
39.
▲
by
jhj
3y ago
This seems impractical, it's likely that the data is highly redundant and you'd probably do just as well by just picking a random projection to a much smaller subspace (or simply just perform a random subsampling of the dimensions
40.
▲
by
jhj
3y ago
This feels very disconnected from the realities of hardware to the point of impracticality. More energy is typically burned on RAMs/flops (storing bits and shuffling them around) than the combinational logic portion (adders/multip
41.
▲
by
jhj
3y ago
I would suspect this is more likely than not (at least in the US) due to the problem of legal liability and shielding the company from lawsuits alleging unfair hiring practices if detailed feedback to the interviewee were given.
42.
▲
by
jhj
3y ago
Non-engineering roles at Meta have different salary bands and RSU comp per level than engineering ones (usually less and sometimes a lot less).
43.
▲
by
jhj
3y ago
A (16, 1) posit gives a maximum of 2 extra fractional bits of precision over float16, and (16, 2) gives you 1 extra bit. But also while the numerical distribution a posit assumes is slightly better but still not well geared to the distribut
44.
▲
by
jhj
3y ago
PQ is just a vector compression/quantization method, not an approximate k-NN search algorithm (like cell-probe or bucketing based methods such as IVF or LSH, or graph-based methods such as HNSW). HNSW can be used together with PQ encod
45.
▲
by
jhj
3y ago
Regarding reduced precision, depending upon what you are trying to do, I think it doesn't work quite as well in similarity search as it does for, say, neural networks. If you are concerned about recall of the true nearest neighbor (k=1
46.
▲
by
jhj
3y ago
SCaNN (the paper) is roughly two different things: 1. a SIMD-optimized form of product quantization (PQ), where code to distance lookup can be performed in SIMD registers 2. anisotropic quantization to bias the database towards returning be
47.
▲
by
jhj
3y ago
> Pinecone has zero moat and quite a few free alternatives (Faiss, Weviate, pg-vector) Faiss is a collection of algorithms for in-memory exact and approximate high-dimensional (e.g., > ~30 dimensional) dense vector k-nearest neighbor,
48.
▲
by
jhj
4y ago
> Yeah, it's an "M0" role. Not true, TLMs can be D1+ level too at FB/Meta. Higher level TLMs more often than not tend to occur in more research-oriented organizations though (where the TLM is effectively a "princ
49.
▲
by
jhj
4y ago
ANS is super fast and trivially parallelizable, faster than Huffman or especially arithmetic encoding, and is superior to Huffman at least as far as compression ratios are concerned (symbol probabilities are not restricted to powers of 2, b
50.
▲
by
jhj
4y ago
ML-assisted combinatorial optimization and searching in high-dimensional discrete spaces (MILP/SAT/SMT solvers/theorem provers, etc) ML-assisted continuous regime optimization (e.g., in solid/fluid mechanics problems, ch
51.
▲
by
jhj
4y ago
In CUDA, some transfers involving pageable host memory are completely synchronous from the perspective of the host, even if you use `cudaMemcpyAsync`: https://docs.nvidia.com/cuda/cuda-runtime-api/api-sync-behav..
52.
▲
by
jhj
4y ago
GDP does not represent wealth, it's a measure of current economic production. Furthermore, GDP is skewed for some countries where GNP might be a better figure (e.g., Ireland, where a fair chunk of the domestic production accrues to for
53.
▲
by
jhj
5y ago
The energy requirement for a quire is very high unless you are talking less than 12 bit posits or so (in which case the strategy actually becomes superior from what I've seen due to the lack of needing to convert back to a float/p
54.
▲
by
jhj
5y ago
Having investigated posits by running RTL implementations through synthesis on 7 and 28 nm nodes, I don't buy the claim that a posit FPU is smaller than a IEEE float FPU. The implication is probably more around that one could use a 32
55.
▲
DietGPU: Fast ANS compression codec for Nvidia GPUs
(github.com)
28 points
by
jhj
5y ago
|
1 comments
56.
▲
by
jhj
5y ago
1 vector against 1 million vectors in 768 dims at k = 10 takes 259 ms for me using Faiss CPU IndexFlatL2 with Intel MKL: https://gist.github.com/wickedfoo/165b69075cfcceba872aec1c46...
57.
▲
by
jhj
5y ago
The cycle threshold (Ct) is the quantitative value from PCR. This is mapped to counts per volume based on results in https://pubmed.ncbi.nlm.nih.gov/32607521/ I believe.
58.
▲
by
jhj
5y ago
That's an estimated gross sales figure, not Apple's take of that or profit.
59.
▲
by
jhj
5y ago
Despite the historical enmity, it seems that the public viewpoint towards Japan is shifting in SK though: > Political elites here are usually careful not to antagonize China, the country’s largest trading partner. But Mr. Yoon’s blunt rh
60.
▲
by
jhj
6y ago
500 billion is substantially off, that figure is not for San Francisco the city, it's for the combined SF-Oakland-Hayward, CA metropolitan statistical area, which has a total population of around 4.7 million people. https://
More ›