Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jerpint
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
17 ms
·
181.
▲
by
jerpint
3y ago
I agree that these benchmarks don’t mean as much anymore because it’s highly likely they were already present in the training set, but also believe it’s likely these tools will be significantly better in a few research cycles
182.
▲
by
jerpint
3y ago
Cunningham's Law states "the best way to get the right answer on the internet is not to ask a question; it's to post the wrong answer." https://meta.wikimedia.org/wiki/Cunningham%27s_Law
183.
▲
by
jerpint
3y ago
Companies have been syphoning this data for free for years to train their models. Why would the proprietor of the data not be able to sell it? Assuming of course their terms of use plainly and clearly state that the data belongs to them
184.
▲
by
jerpint
3y ago
He won a Nobel prize for his works so not sure how much of it would be refuted
185.
▲
by
jerpint
3y ago
I wonder if future versions of PyTorch will automatically apply mixed precision if not specified, the article makes it seem like a no brainer to use them by default
186.
▲
by
jerpint
3y ago
Thanks for saving me a click
187.
▲
by
jerpint
3y ago
The core professors involved in that startup are pioneers in shrinking model sizes while maintaining performance, they may have bought them out for that kind of ability
188.
▲
by
jerpint
3y ago
It’s also not meant to be a generative model - only to be used as an encoder model (they list retrieval as a potential use case )
189.
▲
by
jerpint
3y ago
Interesting that it’s not vision based, I suspect you will get much better performance once vision is incorporated, using e.g LLaVa style models
190.
▲
by
jerpint
3y ago
The architecture as a whole is referred to as a transformer (autoregressive decoder-only architecture in the case of gpt style LLMs). Note that there can be other types of LLMs too that are not necessarily transformer based (SSM, RNN, etc)
191.
▲
by
jerpint
3y ago
This is why open source can be net beneficial for companies too, helping them spot bugs easily and improve their own tools
192.
▲
by
jerpint
3y ago
It sounds like you are simply not familiar with how Colab works, this has nothing to do with the original work
193.
▲
by
jerpint
3y ago
Some editors display permanent indentation lines which help visually resolve the indentation
194.
▲
by
jerpint
3y ago
Why would curly braces stop you from making the same logic error?
195.
▲
by
jerpint
3y ago
I’m also not sure with this method if tje LLM can exactly cite its source, which is another great benefit of RAG
196.
▲
by
jerpint
3y ago
This + ohmytmux is my go to
197.
▲
by
jerpint
3y ago
A new kind of search engine
198.
▲
by
jerpint
3y ago
Interesting prompt hack, but not sure it required a whole article about it, this will probably be patched in the coming days
199.
▲
by
jerpint
3y ago
Perhaps signing any kind of upload or content with a GPG key that proves your identity?
200.
▲
by
jerpint
3y ago
I use seaborn to plot directly with pandas, does Altair have any extra advantages? Or is it a similar style?
201.
▲
by
jerpint
3y ago
It seems unikely to me that today’s generation would be less willing to share details about their sex lives compared to 20 years ago
202.
▲
by
jerpint
3y ago
The only OSS model I’ve been wowed by so far is CC mixtral which from limited usage gave me a vibe closer to gpt3.5 turbo
203.
▲
by
jerpint
3y ago
> Meta and Google are releasing these models arguably to kneecap any possible next OpenAI. They want to basically set the market value of anything below state of the art at $0.00, ensuring that there is no breathing room below the $2T co
204.
▲
by
jerpint
3y ago
I would really be surprised if just adding noise would give you convergence
205.
▲
by
jerpint
3y ago
A big Streamlit competitor right now is gradio, which is especially popular for machine learning demos and makes prototyping very easy, they have a very active community
206.
▲
by
jerpint
3y ago
Streamlit reloads and runs the entire app on every change which can make it very slow for data intensive tasks and is hard to work with for complex UIs in my experience
207.
▲
by
jerpint
3y ago
I like the concept of not having to collect UI elements ins data structure; very elegant and I’m now very curious :)
208.
▲
by
jerpint
3y ago
Probably, the openAI api got a lot better since I made that post, though if you stream audio at 2x speed you have to expect a drop in quality since on average most clips whisper is trained on are not at 2x
209.
▲
by
jerpint
3y ago
If you’re interested I did a YouTube video and short blog post about it https://www.jerpint.io/blog/yougptube/ https://www.youtube.com/watch?v=WtMrp2hp94E
210.
▲
by
jerpint
3y ago
It feels weird to me to use the hyper parameters as the variables to iterate on, and also wasteful. Surely there must be a family of models that give fractal like behaviour ?
More ›