Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anakin87
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
anakin87
6mo ago
Hi HN, I've been spending some time lately trying to build Reinforcement Learning Environments and training small language models and wanted to share a little course I put together based on my experiments. Over the past year, we'v
2.
▲
Show HN: Hands-on course for building RL environments for LLMs
(github.com)
1 points
by
anakin87
6mo ago
|
1 comments
3.
▲
Environments Hub: Your Language Model needs better (open) environments to learn
(huggingface.co)
2 points
by
anakin87
1y ago
|
1 comments
4.
▲
by
anakin87
1y ago
LLMs improve when they can practice and reason in interactive environments. Recent work (DeepSeek-R1, GRPO) shows RL can teach models to prefer better outputs by giving rewards. But most RL environments for LLMs are fragmented or closed. Th
5.
▲
GRPO experiment - I trained a Language Model to schedule events
(github.com)
1 points
by
anakin87
1y ago
|
1 comments
6.
▲
by
anakin87
1y ago
I experimented with GRPO lately, since I am fascinated by models learning from prompts and rewards - no example answers needed like in Supervised Fine-Tuning. After the DeepSeek boom, everyone is trying GRPO with GSM8K or the Countdown Game
7.
▲
I trained a Language Model to schedule events with GRPO
(huggingface.co)
1 points
by
anakin87
1y ago
|
1 comments
8.
▲
by
anakin87
1y ago
I experimented with GRPO lately, since I am fascinated by models learning from prompts and rewards - no example answers needed like in Supervised Fine-Tuning. After the DeepSeek boom, everyone is trying GRPO with GSM8K or the Countdown Game
9.
▲
Llama2 + Haystack on Colab
(github.com)
7 points
by
anakin87
3y ago
|
1 comments
10.
▲
by
anakin87
3y ago
I recently conducted some experiments with Llama2 and Haystack ( https://github.com/deepset-ai/haystack ), the NLP/LLM framework. The notebook can be helpful for those trying to load Llama2 on Colab. 1) Installed
11.
▲
by
anakin87
3y ago
In my experience, Argilla is a good open source platform for datacentric NLP. And these features are a great addition... Have you tried it?