Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
helloericsf
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
helloericsf
2y ago
True. More benchmark metrics here: https://x.com/deepseek_ai/status/1872242657348710721/photo/2
32.
▲
by
helloericsf
2y ago
HF link: https://huggingface.co/deepseek-ai/DeepSeek-V3 Aider link: https://aider.chat/docs/leaderboards/ Pricing($0.14/$0.28 per 1M tokens) reference: https://x.com/xingy
33.
▲
DeepSeek v3 beats Claude sonnet 3.5 and way cheaper
(huggingface.co)
48 points
by
helloericsf
2y ago
|
9 comments
34.
▲
by
helloericsf
2y ago
Tweets from Chinese Scholars call for investigation: - https://x.com/xiaoyongzhu001/status/1867724288369004716 - https://x.com/tydsh/status/1867965300706255060 - https://x.com
35.
▲
by
helloericsf
2y ago
NeurIPS: - https://x.com/NeurIPSConf/status/1867759121023336464 - https://neurips.cc/Conferences/2024/StatementOnInclusivity Picard: - https://www.media.mit.edu/posts&#
36.
▲
NeurIPS and Dr. Picard released statement for singling out Chinese scholars
(twitter.com)
2 points
by
helloericsf
2y ago
|
2 comments
37.
▲
by
helloericsf
2y ago
Call Jensen and Lisa!lol
38.
▲
Tencent Hunyuan-Large
(github.com)
148 points
by
helloericsf
2y ago
|
103 comments
39.
▲
by
helloericsf
2y ago
- 389 billion parameters and 52 billion activation parameters, capable of handling up to 256K tokens. - outperforms LLama3.1-70B and exhibits comparable performance when compared to the significantly larger LLama3.1-405B model.
40.
▲
by
helloericsf
2y ago
Thank you! Great insight.
41.
▲
by
helloericsf
2y ago
Congrats on the launch! Curious to know, which OSS models you see works best at the moment?
42.
▲
by
helloericsf
2y ago
Credit to https://x.com/calebfahlgren https://huggingface.co/spaces/cfahlgren1/model-release-heatm...
43.
▲
Chinese AI Community: open-source Heatmap
(huggingface.co)
1 points
by
helloericsf
2y ago
|
1 comments
44.
▲
by
helloericsf
2y ago
OpenCL was discussed more frequently in classes about a decade ago. However, I haven't heard it mentioned in the last five years or so.
45.
▲
by
helloericsf
2y ago
How many coding co-pilot we have on the market?
46.
▲
Poolside is raising $400M+ at a $2B valuation to build a coding co-pilot
(techcrunch.com)
3 points
by
helloericsf
2y ago
|
1 comments
47.
▲
by
helloericsf
2y ago
I don't think github is using cloudflare.
48.
▲
by
helloericsf
2y ago
4-bit quantization tends to come at the cost of output quality losses. https://github.com/ggerganov/llama.cpp/issues/9
49.
▲
by
helloericsf
2y ago
Personally, I never seen onnx used for LLM.
50.
▲
by
helloericsf
2y ago
Seems interesting! https://github.com/turboderp/exllama "A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights."
51.
▲
Is LMDeploy the Ultimate Solution? Why It Outshines VLLM, TRT-LLM, TGI, and MLC
(bentoml.com)
16 points
by
helloericsf
2y ago
|
8 comments
52.
▲
by
helloericsf
2y ago
Seems like is not another model company. https://x.com/ssi/status/1803472825476587910
53.
▲
by
helloericsf
2y ago
Agreed. custom models could be a hit or a miss.
54.
▲
by
helloericsf
2y ago
Truly mind-blowing to see Mixtral 47B running on a smartphone.
55.
▲
by
helloericsf
2y ago
Project Github link: https://github.com/SJTU-IPADS/PowerInfer
56.
▲
21.2× faster than llama.cpp? plus 40% memory usage reduction
(arxiv.org)
43 points
by
helloericsf
2y ago
|
14 comments
57.
▲
by
helloericsf
2y ago
Also, their AI/ML side of the house is a mess. :(
58.
▲
by
helloericsf
2y ago
Snowflake unveiled Polaris yesterday, Databricks acquires Tabular today. What's coming tmr?
59.
▲
Databricks acquires Tabular, Snowflake fork Iceberg?
(datagravity.dev)
2 points
by
helloericsf
2y ago
|
3 comments
60.
▲
New Yi 1.5 models under Apache 2.0
(huggingface.co)
2 points
by
helloericsf
2y ago
|
0 comments
More ›