Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
suryabhupa
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
suryabhupa
2y ago
In practice, and at scale, that's exactly what having <bos> and <eos> tokens allow you to easily and programmatically do.
2.
▲
by
suryabhupa
2y ago
The announcements are live on Twitter! See this for example: https://x.com/suryabhupa/status/1806342617191379167
3.
▲
by
suryabhupa
2y ago
Surya here from the core Gemma team -- we can think of a distillation loss as learning to model the entire distribution of tokens that are likely to follow the prefix thus far, instead of only the token in the training example. If you do so
4.
▲
by
suryabhupa
5y ago
This is really remarkable! How hard do you think it will be to support new models, i.e. does the tooling you’ve built generalize to you being able to serve other large scale models easily?
5.
▲
by
suryabhupa
8y ago
Hi everyone! One of the creators of DFL here. In an attempt to more deeply understand fundamental concepts in machine learning, we designed Depth First Learning. It's a pedagogy for diving deep into machine learning by carefully tailor
6.
▲
Depth First Learning Fellowship: $4000 grants to build ML curricula
(fellowship.depthfirstlearning.com)
4 points
by
suryabhupa
8y ago
|
1 comments
7.
▲
Hitchhiker’s Guide to Organizing an Academic Workshop
(medium.com)
3 points
by
suryabhupa
8y ago
|
0 comments
8.
▲
by
suryabhupa
9y ago
Many machine learning and reinforcement learning models are susceptible to adversarial attacks; it's not unique to deep learning. However, because so many systems that are currently deployed in applications use deep learning, it's
9.
▲
Vicarious: General Game Playing with Schema Networks
(vicarious.com)
2 points
by
suryabhupa
9y ago
|
0 comments
10.
▲
by
suryabhupa
9y ago
There are talks of incorporating this into Excel at some point in the future, but it may take a _while_ before it can be fully productionized.
11.
▲
by
suryabhupa
9y ago
That would be pretty cool to see what it learns, but I don't think we've tried that :P
12.
▲
by
suryabhupa
9y ago
That's one manifestation of this kind of research being used in real life by programmers around the world. :)
13.
▲
by
suryabhupa
9y ago
I'm not too familiar with evolutionary computation methods, but I imagine the approaches may be similar in nature.
14.
▲
by
suryabhupa
9y ago
Theorem solving is very closely related to program induction (we just change the grammar). Just as with Python, the underlying search space would be incredibly large, and while in theory, we could simply change the DSL and it should work, i
15.
▲
by
suryabhupa
9y ago
It turns out the full grammar of Python (and almost all real programming languages) is quite large; this is very early and new work in neural program synthesis, and so we chose a pretty limited DSL to make sure that we could at least solv
16.
▲
by
suryabhupa
9y ago
Eventually, yes.
17.
▲
by
suryabhupa
9y ago
One of the authors here -- would love to answer any questions about the work! :)
18.
▲
Show HN: Deep Learning for Program Synthesis
(microsoft.com)
93 points
by
suryabhupa
9y ago
|
23 comments
19.
▲
by
suryabhupa
10y ago
How do you imagine going about this?
20.
▲
by
suryabhupa
11y ago
The idea is that even there's a policy network that is able to decide at some point what the best possible move is, the tree search is done to refine this choice and to "evaluate" it. This is why a value network is derived fr