Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rasbt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
26 ms
·
91.
▲
AI Research Highlights in 3 Sentences or Less (July-August 2023)
(magazine.sebastianraschka.com)
1 points
by
rasbt
3y ago
|
0 comments
92.
▲
NeurIPS 2023 LLM Efficiency Challenge Starter Guide
(lightning.ai)
2 points
by
rasbt
3y ago
|
0 comments
93.
▲
Does it beat LLMs? NN+Gzip method reimplemented and explained step-by-step
(magazine.sebastianraschka.com)
3 points
by
rasbt
3y ago
|
0 comments
94.
▲
AI and DL paper highlights June-July 2023
(magazine.sebastianraschka.com)
1 points
by
rasbt
3y ago
|
0 comments
95.
▲
State of Computer Vision 2023
(magazine.sebastianraschka.com)
2 points
by
rasbt
3y ago
|
0 comments
96.
▲
Optimizing Memory Usage for Training LLMs and Vision Transformers in PyTorch
(lightning.ai)
1 points
by
rasbt
3y ago
|
0 comments
97.
▲
Accelerating PyTorch Model Training 10x (With Mixed-Precision and FSDP)
(magazine.sebastianraschka.com)
2 points
by
rasbt
3y ago
|
0 comments
98.
▲
Understanding Encoder and Decoder LLMs
(magazine.sebastianraschka.com)
5 points
by
rasbt
3y ago
|
0 comments
99.
▲
New book on more advanced concepts in machine learning, deep learning, and AI
(leanpub.com)
2 points
by
rasbt
3y ago
|
0 comments
100.
▲
AI Research Highlights in 3 Sentences or Less (May-June 2023)
(magazine.sebastianraschka.com)
2 points
by
rasbt
3y ago
|
0 comments
101.
▲
Recapping recent LLM research concerning tuning strategies and data efficiency
(magazine.sebastianraschka.com)
2 points
by
rasbt
3y ago
|
0 comments
102.
▲
by
rasbt
3y ago
Agreed, compared to other architectures, transformers are actually quite straight-forward. The complicated part comes more from training it in distributed setups, making the data loading and tensor parallelism work due to the large size etc
103.
▲
by
rasbt
3y ago
> Misspelling "Attention is All Your Need" twice in one paragraph makes for a rough start to the linked post. 100%! LOL. I was traveling and typing this on a mobile device. Must have been some weird autocorrect/autocomplet
104.
▲
by
rasbt
3y ago
So weird, I posted it with almost the original title (only slightly abbreviated to make it fit: "Why the Original Transformer Figure Is Wrong, and Some Interesting Tidbits About LLMs". Not sure what happened there. Someone must ha
105.
▲
Why the original transformer figure is wrong, and some other tidbits about LLMs
(magazine.sebastianraschka.com)
237 points
by
rasbt
3y ago
|
49 comments
106.
▲
Finetuning LLMs Efficiently with Adapters
(old.reddit.com)
2 points
by
rasbt
3y ago
|
1 comments
107.
▲
AI Research Highlights in 3 Sentences or Less (April-May 2023)
(magazine.sebastianraschka.com)
1 points
by
rasbt
3y ago
|
0 comments
108.
▲
Understanding Parameter-Efficient Finetuning of Large Language Models
(sebastianraschka.com)
2 points
by
rasbt
3y ago
|
0 comments
109.
▲
by
rasbt
4y ago
I think some businesses and people are worried about using GPL code in their code bases because that's incompatible with their own licenses.
110.
▲
by
rasbt
4y ago
Not sure, but I think the point was that if you have something in GPL license (like the code in this case) it's open source, but that doesn't mean you can use that for your business application. That's because GPL requires yo
111.
▲
by
rasbt
4y ago
I guess that means time to fire up a few GPUs later today and get some weights! We should have a weight exchange platform for that maybe, haha.
112.
▲
Keeping Up with AI Research and News
(sebastianraschka.com)
1 points
by
rasbt
4y ago
|
0 comments
113.
▲
Latest research on reinforcement learning w human feedback for language models
(magazine.sebastianraschka.com)
1 points
by
rasbt
4y ago
|
0 comments
114.
▲
Show HN: Some Techniques to Make Your PyTorch Models Train Faster
(sebastianraschka.com)
1 points
by
rasbt
4y ago
|
0 comments
115.
▲
Understanding the Self-Attention Mechanism of Large Language Models from Scratch
(sebastianraschka.com)
2 points
by
rasbt
4y ago
|
0 comments
116.
▲
Understanding Large Language Models – A Transformative Reading List
(sebastianraschka.com)
2 points
by
rasbt
4y ago
|
0 comments
117.
▲
What Are the Different Approaches for Detecting AI-Generated Content?
(sebastianraschka.com)
2 points
by
rasbt
4y ago
|
0 comments
118.
▲
The new ChatGPT release still fails at basic math
(twitter.com)
1 points
by
rasbt
4y ago
|
0 comments
119.
▲
Comparing Different Automatic Image Augmentation Methods in PyTorch
(sebastianraschka.com)
1 points
by
rasbt
4y ago
|
0 comments
120.
▲
Machine Learning Q and AI – A human vs. ChatGPT explaining ML concepts
(leanpub.com)
2 points
by
rasbt
4y ago
|
0 comments
More ›