Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
umjunsik132
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Create your own AI, then watch it battle others in your browser
(kim-ai-gpu.github.io)
2 points
by
umjunsik132
3mo ago
|
0 comments
2.
▲
My Opinion on RL
3 points
by
umjunsik132
4mo ago
|
1 comments
3.
▲
by
umjunsik132
4mo ago
Hi dang, sorry to ping you here again. I emailed back about a week ago but haven't heard back. My posts are still showing up as [dead] immediately after submission even with non-dev.to URLs. Would appreciate it if you could take a look
4.
▲
by
umjunsik132
4mo ago
Great project!
5.
▲
by
umjunsik132
4mo ago
I built this because most RL environments are either interesting or accessible, rarely both. Engaging environments like MineRL or Dota 2 either demand massive compute or involve painful setup—dependency hell, rendering configs, CUDA mismatc
6.
▲
Show HN: Run RL agents in the browser with WebGPU
(agenlus.com)
3 points
by
umjunsik132
4mo ago
|
1 comments
7.
▲
Qwen2.5 Coder 1.5B Roblox
(huggingface.co)
1 points
by
umjunsik132
10mo ago
|
1 comments
8.
▲
by
umjunsik132
10mo ago
Overview This model is built on Qwen2.5-Coder-1.5B-Instruct and has been fine-tuned exclusively on the official Roblox Luau corpus. It is designed to assist developers with code generation, completion, and understanding of Luau patterns com
9.
▲
I built a benchmark to score the 'architectural intelligence' of neural nets
(github.com)
2 points
by
umjunsik132
11mo ago
|
1 comments
10.
▲
by
umjunsik132
11mo ago
Hi HN, author here. I've always felt that standard benchmarks focus too much on final accuracy, while the architectural choices that get us there are often treated like a dark art. We celebrate a new SOTA model, but rarely do we have a
11.
▲
by
umjunsik132
11mo ago
For this initial version, I kept the gating static to keep the model as simple as possible while validating the core idea. Making the gate dynamic based on the input is a great suggestion for the next step, and I agree it could lead to bett
12.
▲
Making GPT-2 better at math reasoning with a new attention mechanism
(github.com)
3 points
by
umjunsik132
11mo ago
|
3 comments
13.
▲
by
umjunsik132
11mo ago
Hi HN Author here. I built FactorizedAttention - a new attention mechanism based on the GWO framework. Instead of simple QK^T dot products, it uses factorized quadratic forms to model higher-order token interactions. Testing on GPT-2 small
14.
▲
by
umjunsik132
1y ago
Thank you for your comment and for sharing your interesting work. I'll take a look.
15.
▲
by
umjunsik132
1y ago
I used AI to polish my response. The idea was mine though. My apologies.
16.
▲
by
umjunsik132
1y ago
That's a fantastic question, and you've hit on a perfect example of the GWO framework in action. The key difference is the level of abstraction: GWO is a general grammar to describe and design operations, while Mamba is a specific
17.
▲
I unified convolution and attention into a single framework
(zenodo.org)
80 points
by
umjunsik132
1y ago
|
18 comments
18.
▲
by
umjunsik132
1y ago
Hi HN, author here. For years, it bothered me that convolution (the king of vision) and matrix multiplication / self-attention (the engine of Transformers) were treated as completely separate, specialized tools. It felt like we were mi