Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
convexstrictly
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
convexstrictly
3y ago
Pricing input: $8/1M tokens output: $24/1M tokens https://docs.mistral.ai/platform/pricing/
32.
▲
by
convexstrictly
3y ago
Jeremy makes compelling arguments. Here are some more mundane corollaries: It is only a matter of time before you and your company are affected by the pending regulations. In the future, almost all software products will be using AI model
33.
▲
by
convexstrictly
3y ago
Some information here. https://www.ntia.gov/federal-register-notice/2024/dual-use-f...
34.
▲
by
convexstrictly
3y ago
The comments will inform the drafting of regulations on open weight models under the Biden executive order on AI using his powers under the Defense Production Act. Fact Sheet: https://www.whitehouse.gov/briefing-room/
35.
▲
(US Dept of Commerce) NTIA Solicits Comments on Open-Weight AI Models
(commerce.gov)
1 points
by
convexstrictly
3y ago
|
0 comments
36.
▲
by
convexstrictly
3y ago
"By enabling the use of a single high-precision base model accompanied by multiple 1-bit deltas, BitDelta dramatically reduces GPU memory requirements by more than 10x, which can also be translated to enhanced generation latency in mul
37.
▲
by
convexstrictly
3y ago
https://github.com/FasterDecoding/BitDelta
38.
▲
BitDelta: Your Fine-Tune May Only Be Worth One Bit
(arxiv.org)
2 points
by
convexstrictly
3y ago
|
2 comments
39.
▲
by
convexstrictly
3y ago
Twitter summary: https://twitter.com/ssgrn/status/1738256456250470853 Github: https://github.com/KaiNylund/lm-weights-encode-time
40.
▲
Time is encoded in the weights of finetuned language models
(arxiv.org)
124 points
by
convexstrictly
3y ago
|
55 comments
41.
▲
by
convexstrictly
3y ago
Everything you say makes sense. Training is definitely more compute intensive than inference. Training is both memory throughput and compute constrained. Much research in speeding up training goes into optimizing HBM to SRAM communication
42.
▲
by
convexstrictly
3y ago
Could you explain the blockers to getting back-propagation working well on your chips?
43.
▲
by
convexstrictly
3y ago
Research suggesting that much of the power of the transformer architecture comes from associative recall over long sequences that does not require scaling model dimensions. They design state space models that narrow the gap. Overview http
44.
▲
Zoology 1: Measuring and Improving Recall in Efficient Language Models
(hazyresearch.stanford.edu)
2 points
by
convexstrictly
3y ago
|
1 comments
45.
▲
by
convexstrictly
3y ago
"... we find that a duo of a 1.3B generation model and a 1.3B verifier model can achieve 81.5% accuracy, outperforming existing models that are orders of magnitude larger."
46.
▲
TinyGSM: Achieving >80% on GSM8k with small language models
(arxiv.org)
2 points
by
convexstrictly
3y ago
|
1 comments
47.
▲
by
convexstrictly
3y ago
The paper claims it builds upon the concepts in HashGraph, an efficient CUDA hashtable implementation. HashGraph (2019) https://arxiv.org/abs/1907.02900 Anyone know what the most performant CUDA hash table implementati
48.
▲
by
convexstrictly
3y ago
Github repo https://github.com/harp-lab/gdlog
49.
▲
by
convexstrictly
3y ago
Twitter thread with video introduction https://twitter.com/1x_tech/status/1730610445541638378
50.
▲
Androids built to meet the labor demands
(1x.tech)
1 points
by
convexstrictly
3y ago
|
1 comments
51.
▲
by
convexstrictly
3y ago
Emily Chang from Bloomberg reports: Satya Nadella was "blindsided" and is furious. Details in paywalled article: https://www.bloomberg.com/news/articles/2023-11-18/openai-al... Chang does not make i
52.
▲
Sam Altman likely to start company with researchers from OpenAI: Bloomberg
(twitter.com)
4 points
by
convexstrictly
3y ago
|
6 comments
53.
▲
by
convexstrictly
3y ago
If you were logged into Bing, those prompts may be in your history. They can be viewed using the Edge browser. A few weeks ago, I had spotty service with Bing Chat where it would keep resetting the conversation which I assumed was due to l
54.
▲
by
convexstrictly
3y ago
Github Copilot Chat is in beta and is marketed as a separate product. I suspect it is using a later generation (and better) underlying model. I used it many months ago, and I agree it is much better than vanilla Github Copilot. I haven&#x
55.
▲
by
convexstrictly
3y ago
Yes, I have been using it the last few days. I haven't noticed the failure case of Bing Chat not answering the question at all. Are you sure you had the modes I recommended turned on? Try reporting the problems to Mikhail Parakhin,
56.
▲
by
convexstrictly
3y ago
After the recent updates to the UI, the old 50 messages every 3 hours warning went away. Though I haven't really tried pushing it past that limit to tell if that restriction is still there.
57.
▲
by
convexstrictly
3y ago
Bing Chat (Choose "Creative Mode" or "Use GPT-4" depending on UI) is GPT-4 with retrieval augmented generation using the Bing Search engine. It seems to be tuned a little differently, but I haven't found it any bet
58.
▲
by
convexstrictly
3y ago
"Greg Brockman works 60 to 100 hours per week, and spends around 80% of the time coding. Former colleagues have described him as the hardest-working person at OpenAI." https://time.com/collection/time100-ai&#x
59.
▲
by
convexstrictly
3y ago
"I am currently on leave from MIT and spending it at OpenAI." https://madry.mit.edu/
60.
▲
by
convexstrictly
3y ago
Jakub Pachocki and Szymon Sidor have worked on mu-parametrization/tensor programs and Dota 2. https://www.semanticscholar.org/author/J.-Pachocki/2713380?s... As @eachro pointed out, Aleksander Madry is on lea
More ›