Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
phillypham
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
phillypham
5y ago
Typically, these libraries are unusable for this purpose because you don't want to ship a Python interpreter with your video game. I usually prefer to to rewrite my training step as a pure function, so the model weights are just inputs
32.
▲
by
phillypham
5y ago
> If you can make it into a FAANG/MAMAA/whatever flavor of the year acronym, you can make it into most tech companies (and still make a salary that's more than enough to live a dignified life and support yourself, and fami
33.
▲
by
phillypham
5y ago
At Seattle and NYC farmers markets, I regularly pay $12 a dozen. The eggs are often sold out, so others are, too.
34.
▲
by
phillypham
5y ago
Serious question. What's wrong with living next to basketball hoops? I live across the street from one in NYC and consider it a big positive to have a park nearby.
35.
▲
by
phillypham
5y ago
Those few hours you spent are probably worth at least 2 years of YouTube Premium if you're a SWE. And you'd be supporting creators.
36.
▲
by
phillypham
5y ago
It wouldn't surprise me if at least one factor contributing to the higher yield of Asian American students is that they are being discriminated against at Ivy+ schools and have to "settle" for UCs that don't practice aff
37.
▲
by
phillypham
5y ago
It was used in https://ai.googleblog.com/2020/01/reformer-efficient-transfo... for faster and more memory-efficient attention.
38.
▲
by
phillypham
6y ago
Traffic prediction https://deepmind.com/blog/article/traffic-prediction-with-ad...
39.
▲
by
phillypham
6y ago
Edit: This is not meant to criticize the Tensorflow, TPU, or XLA team. They responded quickly to our bugs and made the work possible. I just meant there was some extra organizational overhead.
40.
▲
by
phillypham
6y ago
I think the insight is not incredibly original and follows naturally from OpenAI's Sparse Transformer. The idea is similar to Longformer. Two teams at Google had a similar insight, hence the high number of authors. The original impleme
41.
▲
by
phillypham
6y ago
This is something we want to explore. BigBird just replaces the attention mechanism in BERT. We believe something like BigBird can be complementary to GPT-3. GPT-3 is still limited to 2048 tokens. We'd like to think that we could gener
42.
▲
by
phillypham
6y ago
Not really. The proofs are more of a curiosity really. I think the strong performance on QA takes that require multihop reasoning give some evidence that the model is capable of complex reasoning.
43.
▲
by
phillypham
6y ago
Right now, BigBird is encoder only so it doesn't generate text. Causal attention with the global memory is a bit weird, but we could probably do it. GPT-3 is only using a sequence length of 2048. In most of our paper, we use 4096, but
44.
▲
by
phillypham
6y ago
I'm one of the authors. I may be able to answer some questions or concerns.
45.
▲
by
phillypham
6y ago
GPT-2 fine tuned on the works of Robin DiAngelo.
46.
▲
RobotDiAngelo: GPT-2 learning to be antiracist
(twitter.com)
4 points
by
phillypham
6y ago
|
1 comments
47.
▲
by
phillypham
7y ago
I know it isn't perfect, but I personally am a huge fan of Google's interview process. It's a boon to people switching careers because data structures and algorithms can be studied without any professional experience. Behavio