Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
refibrillator
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
61.
▲
by
refibrillator
2y ago
One confounding factor here is the proliferation of autocorrect and grammar “advisors” in popular apps like gmail etc. One algorithm tweak could change a lot of writing at that scale. While the word frequency stats are damning, there doesn’
62.
▲
Colmap-Free 3D Gaussian Splatting
(oasisyang.github.io)
3 points
by
refibrillator
2y ago
|
0 comments
63.
▲
by
refibrillator
2y ago
Congrats on the paper! Any chance the code will be released? Also I’d be curious to hear, what are you excited about in terms of future research ideas? Personally I’m excited by the trend of eliminating the need for traditional SfM preproce
64.
▲
by
refibrillator
2y ago
Do you have any perspectives to share on Ryan's observation of a potential scaling law for these tasks and his comment that "ARC-AGI will be one benchmark among many that just gets solved by scale"?
65.
▲
by
refibrillator
2y ago
> # TODO: Add Claude support There are no cases for Claude models yet. I wonder if anyone has run a bunch of messages through Anthropic's API and used the returned token count to approximate the tokenizer?
66.
▲
by
refibrillator
2y ago
There are some neat videos and images in the repo: https://github.com/ccli3896/RLWorms Basically they took a worm and hijacked the genetic machinery such that a subset of neurons could be activated or inhibited via a l
67.
▲
by
refibrillator
2y ago
> project management risks: “Lead developer hit by bus” > software engineering risks: “The server may not scale to 1000 users” > You should distinguish them because engineering techniques rarely solve management risks, and v
68.
▲
by
refibrillator
2y ago
The problem is an ANN or LLM has no physical location or “point of view”, which has led ML practitioners toward the concept of embodiment: where algorithms and agents no longer learn from datasets of images, videos or text curated primarily
69.
▲
by
refibrillator
2y ago
The samples of leafy vegetables were from Israel, Italy, Spain, and Switzerland. I have a bad feeling the numbers are much worse in urban and suburban United States, where car culture utterly dominates in close proximity to many farms and h
70.
▲
by
refibrillator
2y ago
There’s a few subtle misconceptions being spread here: 1) Hallucination rate is not inversely proportional to number of samples, unless you assume statistical independence. As you’re sampling from the same generative process each time, any
71.
▲
by
refibrillator
2y ago
Can Wuffs provide stronger safety guarantees than techniques like WasmBoxC? My understanding is that compiling unsafe C to WASM and back would also guarantee safety with respect to buffer overflows, integer arithmetic overflows and null poi
72.
▲
by
refibrillator
2y ago
GP meant deterministically add jitter. Long ago I was responsible for implementing a “rate limiting algorithm”, but not for HTTP requests. It was for an ML pipeline, with human technicians in a lab preparing reports for doctors and in dire
73.
▲
by
refibrillator
2y ago
Why do you think language is so special? There's an extensive body of literature across numerous domains that demonstrates the benefits of Multi-Task Learning (MTL). Actually I have a whole folder of research papers on this topic, here
74.
▲
by
refibrillator
2y ago
You don’t need a DB, I would avoid that for a one time job (I’ve used pgvector a lot). Since your data fits in memory (18 GB @ FP32), I would start with a simple python script that does naive exhaustive search, which is O(n^2). You can do a
75.
▲
by
refibrillator
2y ago
Props for mentioning BirdNET as a potentially more accessible starting point for less technical folks. There are a couple relative advantages of your approach that I feel are notable though: Squeezed wav2vec2 (SEW) architecture leverages Tr
76.
▲
by
refibrillator
2y ago
> The center for Medicare and Medicaid is the primary source of graduate medical education (residency) funding. Per the Graham Center interactive GME data tool Mt. Sinai received around $175k a year per resident, and pays them a salary
77.
▲
by
refibrillator
3y ago
Even if you choose a more sophisticated similarity measure as suggested by another commenter, you’ll still need to set a threshold on that metric to perform binary classification. In my experience there are two paths forward, the one I reco
78.
▲
by
refibrillator
3y ago
It’s really surprising how well this works. Intuitively this illustrates how over-parameterized many LLMs are, or conversely how under-trained they might be. The drop and rescale method outlined in the paper makes the latent space increasin
79.
▲
by
refibrillator
3y ago
Biggest model is 30b MoE trained on 100b tokens, max sequence length 4096. A bit underwhelming compared to recent announcements like the open source Large World Model [1]. Absolutely no benchmarks against GPT4 present in the paper. Notably
80.
▲
by
refibrillator
3y ago
FLOPs utilization is arguably the industry standard metric for efficiency right now and it should be a good first approximation of how much performance is left on the table. But if you mean the reported utilization in nvtop is misleading I
81.
▲
by
refibrillator
3y ago
I’ve seen this before! Indeed it has nothing to do with Python or the GIL. The OS scheduler has to execute M threads on N cpu cores, while also balancing competing priorities like latency and power usage. Because each separate process uses
82.
▲
by
refibrillator
3y ago
Speculating? I think branch predictors would eat this for lunch
83.
▲
by
refibrillator
3y ago
If regression on wavelet coefficients solves your task then sure DL is unnecessary. I spent half a decade on a data science team developing algorithms for ambulatory ECG monitoring at scale. 100+ human years of 256 Hz multichannel ECG colle
84.
▲
by
refibrillator
3y ago
It's not mentioned in the paper but this month OpenChat 3.5 released the first 7b model that achieves results comparable to ChatGPT in March 2023 [1]. Only 8k context window, but personally I've been very impressed with it so far.
85.
▲
by
refibrillator
3y ago
PSA: This is a read-modify-write pattern, thus it is not safe under concurrency unless a transaction isolation level of SERIALIZABLE is specified, or some locking mechanism is used (select for update etc).
86.
▲
by
refibrillator
3y ago
> The benzene content of typical gasoline is 0.76% by mass (gasoline composition). A spill of 10 gallons of gasoline (only 0.1% of the 10,000 gallon tank, a quantity undetectable by manual gauging and inventory control) contains about 23
87.
▲
by
refibrillator
3y ago
I've noticed the skeptics here like to raise this argument on every article related to psychedelic research. As if they've found a gaping hole that somehow casts "doubts" on the methodology of decades worth of clinical t
88.
▲
by
refibrillator
3y ago
This is linked in the article but easy to miss, it has helpful visualizations of the “time programmable frequency comb”: https://www.nist.gov/news-events/news/2022/10/break-new-grou... Some notable numbe
89.
▲
by
refibrillator
3y ago
I’ve only used DD-WRT (for many years on my routers), and awhile back I perused the version control history which led me to believe that it was maintained by only a few key individuals. Perhaps I’m wrong, or it’s simply a very mature projec
90.
▲
by
refibrillator
3y ago
This seems to happen a lot here lately, particularly on earth and space science related topics. After you read the comments on a paper in your field, it becomes depressingly obvious that a lot of the most upvoted commentary is armchair skep
More ›