Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jayalammar
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
jayalammar
4y ago
What do you mean by naturally biased? That people seem to favor them?
32.
▲
by
jayalammar
4y ago
So it really depends on what you use for clustering. In this case, I'm clustering by the original embeddings so the UMAP results are different. I've also seen: 1- Clustering by UMAP. Here the plot would show clean separation of to
33.
▲
by
jayalammar
4y ago
[1] https://txt.cohere.ai/combing-for-insight-in-10-000-hacker-n... [2] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_... [3] https://assets.cohere.ai/blog/text-
34.
▲
Show HN: Analyzing top HN posts with language models
117 points
by
jayalammar
4y ago
|
43 comments
35.
▲
by
jayalammar
4y ago
These are good explanations: - https://www.assemblyai.com/blog/how-dall-e-2-actually-works/ - https://www.youtube.com/watch?v=F1X4fHzF4mQ
36.
▲
by
jayalammar
4y ago
Yes. This is already done in research and commercially ( https://openai.com/blog/instruction-following/ ).
37.
▲
by
jayalammar
4y ago
Agreed that people should pay attention to cherry-picking of model outputs. For this one in particular, here are a few more results for Battlestar and The Office: https://twitter.com/Miles_Brundage/status/153247388
38.
▲
by
jayalammar
4y ago
Not Dalle2 specifically, it's proprietary and there's a waitlist. Older open source alternatives include https://huggingface.co/dalle-mini
39.
▲
DALL-E 2 generates images of Kermit The Frog in various films
(twitter.com)
390 points
by
jayalammar
4y ago
|
192 comments
40.
▲
by
jayalammar
4y ago
Oh wow. Yeah that's way off.
41.
▲
How the Vanderbilt heirs spent $300B in 50 years
(collaborativefund.com)
2 points
by
jayalammar
4y ago
|
2 comments
42.
▲
Enjoy the Silence – GridSearch Is Not Enough: Part 7
(koaning.io)
2 points
by
jayalammar
4y ago
|
0 comments
43.
▲
by
jayalammar
5y ago
Sorry if I've missed your connection request. Please email me at alammar at gmail if I can be of service.
44.
▲
by
jayalammar
5y ago
Yeah, mentioned it earlier, but added a note to this thread too.
45.
▲
by
jayalammar
5y ago
Cohere (language understanding and generation with large language models) is very different from Sagemaker (general ML platform). Cohere abstracts training and deploying language models for developers and companies that don't have an a
46.
▲
by
jayalammar
5y ago
Content moderation extends to plenty of sub-problems other than spam. A lot of use cases need detection of different types of online harm, for example (bullying, hate speech...etc). A lot of these cases can be improved by training classifie
47.
▲
by
jayalammar
5y ago
Thanks, sdoering. Cohere engineer here. Happy to provide some context. You can can find performance benchmarks here: https://txt.cohere.ai/launch-larger-embed-models#model-compa... Cohere provides an API to access and finet
48.
▲
Default Now
(nateliason.com)
1 points
by
jayalammar
5y ago
|
0 comments
49.
▲
Mle-monitor: A lightweight experiment and resource monitoring tool (2021)
(roberttlange.github.io)
15 points
by
jayalammar
5y ago
|
0 comments
50.
▲
The Data Science Venn Diagram (2010)
(drewconway.com)
2 points
by
jayalammar
5y ago
|
0 comments
51.
▲
by
jayalammar
5y ago
It always amazes me how important and useful open source software is often supported by small team doing heroic efforts. I found it very interesting to learn about QuantStack and what they do to support Jupyter data science communities. I&#
52.
▲
QuantStack (small company that carries Jupyter): 2021 in review
(medium.com)
1 points
by
jayalammar
5y ago
|
1 comments
53.
▲
by
jayalammar
5y ago
Hi HN, This is a gentle guide to building simple semantic search features that go beyond keyword search. Uses sentence embeddings and Annoy to build a "similar questions" feature. Semantic search is one of the most powerful ideas
54.
▲
Intro to Basic Semantic Search
(docs.cohere.ai)
2 points
by
jayalammar
5y ago
|
1 comments
55.
▲
by
jayalammar
5y ago
I'm intrigued by them and wanted to dig deeper into Switch Transformer. So hopefully yeah if they continue to show promise. Thank you!
56.
▲
by
jayalammar
5y ago
Back of the envelop: assuming 3 characters per token, 3 bytes per character (unicode is 1-4), that's 18 trillion bytes. So about 18 TB? Reasonable for disk size, unreasonable to be loaded in GPU memory. Compute: Building the database r
57.
▲
by
jayalammar
5y ago
The size of the database, the training set, the details of the architecture, as well as results on benchmark tasks should all be considered in the comparison. I'm also a fan of Behavioral Testing [1]. Parameter count is not very accura
58.
▲
by
jayalammar
5y ago
Hi HN, Summary: The latest batch of language models can be much smaller yet achieve GPT-3 like performance by being able to query a database or search the web for information. A key indication is that building larger and larger models is no
59.
▲
The Illustrated Retrieval Transformer
(jalammar.github.io)
75 points
by
jayalammar
5y ago
|
13 comments
60.
▲
Show HN: Language model analysis and visualization toolkit
(github.com)
4 points
by
jayalammar
5y ago
|
0 comments
More ›