Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
smaddox
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
smaddox
2y ago
Transformers have quadratic computational complexity in sequence length, i.e. O(N^2) where N is the sequence length. RNNs, Linformer, Mamba, etc. have linear or quasi-linear computational complexity in sequence length, which often bottlenec
32.
▲
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
(transformer-circuits.pub)
1 points
by
smaddox
2y ago
|
1 comments
33.
▲
by
smaddox
2y ago
There's https://hero.handmade.network/forums , which is sort of a part of https://handmade.network/forums , which has a jobs board. And there's an associated discord chat, and loosely associated Ha
34.
▲
by
smaddox
2y ago
> with unprecedented nano-fabricated coatings with sub-wavelength structure on optical glasses. Also known as an anti-reflection coating. Definitely not unprecedented. Cool project, though.
35.
▲
by
smaddox
2y ago
The image quality looks quite good, and the variable depth of focus is something that I've been eagerly awaiting. I personally find the 3D effect wears off after a bit with existing stereographic systems, likely due to the fixed depth
36.
▲
The Matrix: A Bayesian learning model for LLMs
(arxiv.org)
3 points
by
smaddox
2y ago
|
0 comments
37.
▲
by
smaddox
2y ago
Very interesting! Two-photon pumping of traditional lasers with only 2 states would require absurd energy densities, but I guess the combination of a strong quadrupole coupling and long state decay times makes it feasible for a nuclear lase
38.
▲
by
smaddox
2y ago
> Retrying makes sense because LLMs aren’t deterministic even at temperature zero. This is news to me. I'm trying to think where non-determinism would come in at temperature zero, but coming up with nothing. What am I missing?
39.
▲
by
smaddox
2y ago
It was a major issue with aluminum interconnects. It was primarily addressed by switching to a different metal that exhibits less electromigration, namely copper. But you have to be careful putting copper in a silicon fab, because if it get
40.
▲
by
smaddox
2y ago
To build a laser, you would need at least three energy levels, and ideally four, with particular constraints in the transition probabilties so that you can create population inversion. And you would need to pump it with a higher-energy (sho
41.
▲
by
smaddox
2y ago
For free atoms, yes. For atoms in a crystal lattice (or other solid), it's quite common for electrons to decay through phonon interactions, i.e. by emitting vibrations (i.e. heat) to the lattice.
42.
▲
by
smaddox
3y ago
Bag of words models use a context that is a "bag" (i.e. an unorder map from elements to their counts) of words/tokens. GPT's use a context that is a sequence (i.e. an ordered list) of words/tokens.
43.
▲
by
smaddox
3y ago
AFAIK, the part that's been debunked is that there's a complete separation of concerns between the two hemispheres. From studies on split-brain patients, there does appear to be some specialization, but it's much fuzzier than
44.
▲
by
smaddox
3y ago
I feel your pain. I've spent more time playing Dungeon Crawl Stone Soup than I care to admit. The fact that this can be played on Android makes it even more dangerous.
45.
▲
by
smaddox
3y ago
Anyone care to comment on why this concise response to the parent's question, supported with a link to the relevant data, was down voted? I'm confused.
46.
▲
by
smaddox
3y ago
It doesn't help that most of Europe is facing significant population decline: https://zeihan.com/demographics-part-4-the-european-breakdow...
47.
▲
by
smaddox
3y ago
They don't disclose the embedding dimension for gpt-3.5, but based on table 4, comparing the Size and # Queries columns, gpt-3.5-turbo presumably has an embedding dimension of roughly 20,000? Interesting...
48.
▲
by
smaddox
3y ago
Agree, that will be the proof. But tokomaks have yet to do the same despite decades of investment. We know that helion is able to recover 95% of the energy of every pulse. And they've measured the scaling laws. There doesn't seem
49.
▲
by
smaddox
3y ago
It's relevant when it doesn't work after all that time. As I argue further down the thread, the trajectory of FRCs looks much more promising.
50.
▲
by
smaddox
3y ago
If tokomaks worked as well as Rockets, then I would agree. But the age of a technology does become relevant when you're discussing the history and development trajectory. The trajectory for FRC's looks far more promising. Direct e
51.
▲
by
smaddox
3y ago
Tokamaks are 1960's technology. The future of economical fusion appears much more likely to be based on the field-reversed configuration (FRC). Helion expects to produce net positive energy production from a reactor designed primarily
52.
▲
by
smaddox
3y ago
Damn. Well, I guess I better hurry up and write and publish a paper on the Ternary Neural Network research that I've been doing (part-time) for the last several months, before it all gets scooped.
53.
▲
by
smaddox
3y ago
I read about this explanation years ago. It's not novel. I don't remember where, though... It might have been from Feynman.
54.
▲
by
smaddox
3y ago
That's just HTML that looks like a PDF, though. Incredible feat, but not really what I want from PDF turned to HTML. I want something mobile friendly.
55.
▲
by
smaddox
3y ago
CPU-local could work, but would give up the possibility of static type verification. Agree passing around context can be tedious, but Rust currently doesn't have a way to implicitly pass context (although there have been some proposals
56.
▲
by
smaddox
3y ago
That isn't a very good example. The vectors for each token are randomly initialized with each element taken from the normal distribution. After training, similar words will have some cosine similarity, but almost never as much cosine s
57.
▲
by
smaddox
3y ago
You would probably need to pass around a context type that encodes information about the current context and which interrupts are possible. You would then acquire the lock via that context, which would handle disabling those interrupts.
58.
▲
by
smaddox
3y ago
I don't understand what this has to do with AI. It seems like Heroku for Fast API?
59.
▲
by
smaddox
3y ago
> If replicated, it's arguably the first linear-time architecture with Transformer-quality performance!! RetNet was published in August (although I only learned of it last week): https://arxiv.org/abs/2307.08621
60.
▲
by
smaddox
3y ago
Yeah, seems like they would need to use gold or platinum electrodes to rule out copper (or whatever) leaching. But even if it's copper leaching that boosts growth, that seems interesting in and of itself---just not really due to electr
More ›