Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rasbt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
rasbt
3y ago
Yes, totally agree. These implementations are meant for educational purposes. You could in theory use them to train a model though (GPT-2 also had a from-scratch implementation if I recall correctly). In practice, you probably want to use F
62.
▲
Coding Self-Attention, Multi-Head Attention, Cross-Attention, Causal-Attention
(magazine.sebastianraschka.com)
142 points
by
rasbt
3y ago
|
11 comments
63.
▲
Noteworthy AI Research Papers of 2023
(magazine.sebastianraschka.com)
3 points
by
rasbt
3y ago
|
0 comments
64.
▲
AI Research Papers in November 2023: hallucinations and reasoning capabilities
(magazine.sebastianraschka.com)
5 points
by
rasbt
3y ago
|
0 comments
65.
▲
Practical Tips for Finetuning LLMs Using LoRA (Low-Rank Adaptation)
(magazine.sebastianraschka.com)
342 points
by
rasbt
3y ago
|
27 comments
66.
▲
AI Research Papers (October 2023)
(magazine.sebastianraschka.com)
5 points
by
rasbt
3y ago
|
0 comments
67.
▲
AI and Open Source in 2023: A Review of the Year's Highs and Lows
(magazine.sebastianraschka.com)
2 points
by
rasbt
3y ago
|
0 comments
68.
▲
Takeaways from hundreds of LLM finetuning experiments with LoRA
(lightning.ai)
258 points
by
rasbt
3y ago
|
39 comments
69.
▲
by
rasbt
3y ago
It's fascinating that it's a 1B business. I thought it was just uploading screen recordings to the cloud (basically UI around uploading. Like macOS QuickTime + YouTube private video upload)
70.
▲
Atlassian to Acquire Loom in $975M Deal
(marketwatch.com)
4 points
by
rasbt
3y ago
|
2 comments
71.
▲
OpenAI plans to unveil major developer updates on November 6 like memory storage
(reuters.com)
3 points
by
rasbt
3y ago
|
0 comments
72.
▲
AMD to acquire AI software startup as it seeks to catch up with Nvidia
(cnbc.com)
2 points
by
rasbt
3y ago
|
0 comments
73.
▲
AI chips, acquisitions, new "small" open-source LLMs, and new LoRA techniques
(magazine.sebastianraschka.com)
5 points
by
rasbt
3y ago
|
0 comments
74.
▲
AI news editorial from custom AI chips to new "small" LLMs like phi and Mistral
(magazine.sebastianraschka.com)
1 points
by
rasbt
3y ago
|
0 comments
75.
▲
OpenAI is exploring making its own AI chips
(reuters.com)
116 points
by
rasbt
3y ago
|
94 comments
76.
▲
Meta pitches EU to charge a $10 subscription for ad-free Facebook and Instagram
(wsj.com)
41 points
by
rasbt
3y ago
|
70 comments
77.
▲
by
rasbt
3y ago
> When Yaccarino was first asked about user metrics during the interview, she seemingly wanted to move away from that particular conversation, saying that X had between 200 and 250 daily active users. I hope they mean 200 "million&q
78.
▲
Q&A with AMD CEO on Competing with Nvidia's H100 and PyTorch
(theverge.com)
1 points
by
rasbt
3y ago
|
1 comments
79.
▲
Intel begins chip production using extreme ultraviolet lithography machines
(reuters.com)
5 points
by
rasbt
3y ago
|
1 comments
80.
▲
AI research papers summaries and highlights (Aug to Sep)
(magazine.sebastianraschka.com)
3 points
by
rasbt
3y ago
|
0 comments
81.
▲
by
rasbt
3y ago
I think it could potentially make the model smarter, but it's up to how you collect the data to train the reward models. Currently, companies & papers that use RLHF focus on "safety" rankings, for example. But you could p
82.
▲
by
rasbt
3y ago
RLHF is a popular candidate, but the focus is more on "helpfulness" and "safety" -- I don't think it necessarily improves LLMs on reasoning benchmarks
83.
▲
Training and aligning LLMs with RLHF and RLHF alternatives
(magazine.sebastianraschka.com)
102 points
by
rasbt
3y ago
|
14 comments
84.
▲
by
rasbt
3y ago
Yes, when I remember correctly, they said they didn't release the 34B Llama 2 model yet because they haven't had a chance for "red teaming" that one, where with "red teaming" they mean something along the lines
85.
▲
by
rasbt
3y ago
Haven't seen that one, yet. Thanks for sharing!
86.
▲
by
rasbt
3y ago
Good catch. Above that paragraph, I wrote that the Code Llama models were initialized with the Llama 2 weights, which makes this contradictory, indeed. What I meant to say here was 500B domain-specific tokens. Maybe domain-specific is not t
87.
▲
by
rasbt
3y ago
Interesting, I thought GPT-3.5 was considered GPT-3 + InstructGPT-style RLHF on a large scale, whereas GPT-4 is considered to be an MoE model.
88.
▲
by
rasbt
3y ago
I think so too. But in general, it could also be due to other reasons: faster hardware, lower timeout for batched inference, optimizations like flash attention and flash attention 2, quantization, ... I'd say that it's probably a
89.
▲
Understanding Llama 2 and the New Code Llama LLMs
(magazine.sebastianraschka.com)
170 points
by
rasbt
3y ago
|
34 comments
90.
▲
Llama 2, CodeLlama, and GPT-4 performance: recent LLM developments and research
(magazine.sebastianraschka.com)
1 points
by
rasbt
3y ago
|
0 comments
More ›