Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
varunkmohan
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
by
varunkmohan
4y ago
This is one of the next IDEs we will be supporting! We're keeping folks on our discord ( https://discord.com/invite/3XFf78nAx5 ) updated on this.
62.
▲
by
varunkmohan
4y ago
Yes, good point. Interestingly, our VSCode extension also works on VSCodium since we uploaded it to the OpenVSX registry!
63.
▲
by
varunkmohan
4y ago
Would be curious to see your benchmarks. Btw, Nvidia will be providing support for fp8 in a future release of CUDA - https://github.com/NVIDIA/TransformerEngine/issues/15 I think TMA may not matter as much fo
64.
▲
by
varunkmohan
4y ago
I'm not sure any of this is accurate. 8 bit inference on a 4090 can do 660 Tflops and on an H100 can do 2 Pflops. Not to mention, there is no native support for FP8 (which are significantly better for deep learning) on existing CPUs. T
65.
▲
Show HN: Codeium: Free Copilot Alternative for Vim / Neovim
(github.com)
94 points
by
varunkmohan
4y ago
|
80 comments
66.
▲
by
varunkmohan
4y ago
Find it hard to believe developers would pay $100/month unless the product embeds deeply into their workflow. I do think it's an interesting idea to place some limitations and still have a free tier.
67.
▲
by
varunkmohan
4y ago
I think it's likely the other way around, a product like Copilot is actually much cheaper than if you were to build off of OpenAI's API. ChatGPT, the product, will be much cheaper than if you actually ran the API yourself. Some co
68.
▲
Ask HN: How much will a ChatGPT monthly subscription be?
9 points
by
varunkmohan
4y ago
|
14 comments
69.
▲
by
varunkmohan
4y ago
The model runs on a remote GPU as most users don't have powerful enough GPUs to run run the model with reasonable latency.
70.
▲
by
varunkmohan
4y ago
Posted this on a comment above but systems like Whatsapp likely sent an insane amount of data as well but used only 16 servers over 1.5 billion users at time of acquisition. Modern NICs can handle millions of requests a second - I still fee
71.
▲
by
varunkmohan
4y ago
Genuinely asking, why do you think Twitter needs 24 million vcpus to run? This is not apples to apples but Whatsapp is a product that entirely ran on 16 servers at the time of acquisition (1.5 billion users). It really begs the question why
72.
▲
by
varunkmohan
4y ago
Good analysis. Obviously, this doesn't handle cases like redundancy and doesn't handle some of other critical workloads the company has. However, it does show how much real compute bloat these companies actually have - https:
73.
▲
by
varunkmohan
4y ago
This feels like what a company that hasn't accepted the future would say. Yes, ChatGPT doesn't do exactly what Google does. Can it be augmented by well understood search retrieval engines to generate a much better response? I thin
74.
▲
Show HN: Codeium: Free Copilot alternative that works in Jupyter notebooks
(codeium.com)
11 points
by
varunkmohan
4y ago
|
3 comments
75.
▲
by
varunkmohan
4y ago
Any plans to be able to directly run this in the browser similar to DuckDB. Nice to see more options in the space.
76.
▲
by
varunkmohan
4y ago
Posts like these are always awesome to look at how much we can push consumer hardware. It's hard not to really appreciate some of the devices we have today. For instance, an RTX 4090 is capable of 660 TFlops of FP8 (MSRP 1600). Would n
77.
▲
by
varunkmohan
4y ago
Wow, very smooth to use! Would be extra neat if this could support plugins as well. We're building out a free AI code completion tool, Codeium that will support Neovim as an IDE and would love to hook into this as well.
78.
▲
by
varunkmohan
4y ago
Interesting, I guess we implemented this op properly :) It is strange that cub does not guarantee determinism here - it's not too hard an op to implement in CUDA.
79.
▲
by
varunkmohan
4y ago
It's very strange that they don't guarantee determinism. We trained a model at Codeium that does have determinism at inference time - it's not super hard if you don't just use random seeds everywhere. This makes it a lot
80.
▲
by
varunkmohan
4y ago
Really thorough post! It seems hard to prevent these prompt injections without some RLHF / finetuning to explicitly prevent this behavior. This might be quite challenging given that even ChatGPT suffered from prompt injections.
81.
▲
by
varunkmohan
4y ago
There are a bunch of really good ideas used to train this model - multi query attention, infilling, near deduplication and dataset cleaning. I do wish that the demo was a little more interactive (not needing to click buttons to create a gen
82.
▲
by
varunkmohan
4y ago
That's not the issue here, it's just saying any odd number is prime, which is false
83.
▲
by
varunkmohan
4y ago
The magical number for performance is actually memory bandwidth which is actually lower for TPUs compared to A100s. They have more aggregate compute, but it's not trivial to use that to get very low latency on a per request basis.
84.
▲
by
varunkmohan
4y ago
The neovim plugin mostly actually communicates with a node-js service seen here ( https://github.com/github/copilot.vim/tree/release/copilot/d... ). This is why they require you to install node for us
85.
▲
by
varunkmohan
4y ago
Very true, I think the issue though is unless that output is very likely to be 100% correct, a user would always prefer something that is incomplete but quicker to iterate on. It would be interesting to see if we can get to a paradigm like
86.
▲
by
varunkmohan
4y ago
Curious, what functionality did you find the most useful? Sometimes on edits, I find it adding more than expected (potentially entirely new files) which causes me to not accept the suggestion. Explain does work well sometimes though!
87.
▲
by
varunkmohan
4y ago
This is a pretty cool idea even for just engineering the prompt! It's a complicated tradeoff to decide what should go into the context and what should be selected from other files (2000 tokens is a lot but sometimes not long enough for
88.
▲
by
varunkmohan
4y ago
Most likely latency and cost reasons. A model that's 10x as big requires 10x the hardware to serve at the same latency. Since most generations are not too long, a smaller finetuned model should work well enough.
89.
▲
by
varunkmohan
4y ago
A 1T model would be capable of much more than what the current version of Copilot in terms of autocompletion and even code correction. However, at that point, even with a lot of model parallelism to speedup inference, it's likely to be
90.
▲
by
varunkmohan
4y ago
It's an embedding model so it generates vector embeddings not text generations. That's to be expected.
More ›