Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bluecoconut
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
bluecoconut
2y ago
Curious for those that are reading comments here - 1. Are you users of google dataset search? 2. What other dataset searches are you using? At my company, we recently (last year) did a crawl of over 600 million tables (TabLib) and released
32.
▲
by
bluecoconut
3y ago
Those curves of "embedding displacement" are very interesting! quickly scanning the blog led to this notebook which shows how they're computed and shows other examples too with similar behavior. https://github.com
33.
▲
by
bluecoconut
3y ago
I haven't published it nor have I seen it published. I can copy paste some of my raw notes / outputs from poking around with a small model (Phi-1.5) into a gist though: https://gist.github.com/bluecoconut/6a08
34.
▲
by
bluecoconut
3y ago
Directly answering your question requires making some assumptions about what you mean and also what "class" of models you are asking about. Unfortunately I don't think it's just a yes or no, since I can answer in both di
35.
▲
by
bluecoconut
3y ago
We've made a lot of data tooling things based on LLMs, and are in the process of rebranding and launching our main product. 1. sketch (in notebook, ai for pandas) https://github.com/approximatelabs/sketch 2. datad
36.
▲
Computational Complexity Theory
(en.wikipedia.org)
3 points
by
bluecoconut
3y ago
|
0 comments
37.
▲
Units of Textile Measurement
(en.wikipedia.org)
2 points
by
bluecoconut
3y ago
|
0 comments
38.
▲
by
bluecoconut
3y ago
I started Approximate Labs to tackle this problem! We just recently addressed one of the biggest gaps we believe exists in the DL research community for this modality, dataset; we openly released the largest dataset of tables with annotatio
39.
▲
TabLib – Repo of 627M tables for training Large Data Models
(approximatelabs.com)
4 points
by
bluecoconut
3y ago
|
0 comments
40.
▲
Cayley Graphs and Pretty Things
(juliapoo.github.io)
7 points
by
bluecoconut
3y ago
|
1 comments
41.
▲
by
bluecoconut
3y ago
Nice update. I think the key question they added that clarifies a lot is #3 (quoted below) Can I input an extensive text, like a book, into StreamingLLM for summarization? While you can input a lengthy text, the model will only r
42.
▲
by
bluecoconut
3y ago
EDIT: the authors have updated the readme to add a clarified FAQ section that directly addresses this: https://github.com/mit-han-lab/streaming-llm#faq Just tested it - this definitely doesn't seem to be giving en
43.
▲
by
bluecoconut
3y ago
I think people are misreading this work, and assuming this is equivalent to full dense-attention. This is just saying its an efficiency gain over sliding window re-computation, where instead of computing the L^2 cost over and over (T times)
44.
▲
by
bluecoconut
3y ago
Pretty cool, I made something similar (lambdaprompt[1]), with the same ideal of functions being the best interface for LLMs. Also, here's some discussion about this style of prompting and ways of working with LLMs from a while ago [2].
45.
▲
by
bluecoconut
3y ago
We're doing this at https://www.approximatelabs.com
46.
▲
by
bluecoconut
3y ago
From making a few variations on data chatbots in the past year, I found that my favorite / most fun to use ones seem to be more "chain-of-thought" and conversational rather than "retrieval-augmented" style. Less abo
47.
▲
Value of Information
(en.wikipedia.org)
1 points
by
bluecoconut
3y ago
|
0 comments
48.
▲
by
bluecoconut
3y ago
I love this question, because this is exactly what I'm currently focused on doing! I founded approximatelabs to essentially chase this out and show how much rich fruit there is in this space. See approximatelabs.com It is not tradition
49.
▲
by
bluecoconut
3y ago
We've been using windmill for our internal tooling and dashboards and its been great! Genuinely excited to see GPT-4 integration, we'll definitely give it a go. Some things we've done with windmill so far: * Slack bots
50.
▲
by
bluecoconut
3y ago
To be honest, this is a hunch thing more than a I can teach and explain it thing. To me it's just "a lot of the same stories" that are told about cuprates, and less of a "I can explain the mechanism". Roughly, one o
51.
▲
by
bluecoconut
3y ago
I'd probably approach 90% confidence with ~2-3 different measurements from different labs showing similar stories. Select from the grab bag of possible experiments that could all collaborate and provide consistency in their stories: sp
52.
▲
by
bluecoconut
3y ago
I got my PhD studying band structure of high-tc superconductors (experimentalist, ARPES). These Cu d-d interactions right at the fermi energy give me huge hope. Feels very familiar to other superconductors (re: all the cuprates). (Note: I s
53.
▲
by
bluecoconut
3y ago
I believe that the comment about CAP theorem violation / treating the problem as a technically unsolved thing isn't true. Eg. See the dataflow paper that sets up more clear tradeoffs for latency and correctness in large scale data
54.
▲
by
bluecoconut
3y ago
Pretty good examples and simple explanations. I didn't realize Claude 2 was so good at working with PDFs natively. I wonder if they're doing anything special? Is this just due to larger context length they have? Also, biased opini
55.
▲
by
bluecoconut
3y ago
I wonder if market manipulation is statistically higher during highs of hype cycles? or is it just always happening? > Faking volume allowed the site to climb in the rankings on these sites, which exposed it to a greater audience of peop
56.
▲
Introducing LakehouseIQ (Databricks)
(databricks.com)
2 points
by
bluecoconut
3y ago
|
0 comments
57.
▲
Micromort
(en.wikipedia.org)
5 points
by
bluecoconut
3y ago
|
1 comments
58.
▲
by
bluecoconut
3y ago
Really appreciate and like how he calls out agents as difficult. They're super fascinating, but just not there yet. Agents, like self-driving, are easy to imagine, easy to demo, but to do right will "take a decade". "If
59.
▲
The Dataflow Model: Practical Approach to Balancing Correctness Latency and Cost
(research.google)
2 points
by
bluecoconut
3y ago
|
0 comments
60.
▲
by
bluecoconut
3y ago
They seemed to be pretty mindful of this contamination, and call out that they agressively pruned some training dataset and still observed strong performance. That said, I agree, I really want to try it out myself and see how it feels, and
More ›