Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cpldcpu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
Misguided Attention – Prompts to challenge the reasoning ability of LLMs
(github.com)
2 points
by
cpldcpu
2y ago
|
0 comments
92.
▲
by
cpldcpu
2y ago
Which LLM is Amazon Q based on?
93.
▲
Multi-Layer Perceptron Visualization
(cpldcpu.github.io)
2 points
by
cpldcpu
2y ago
|
0 comments
94.
▲
Claude-Artifacts-Starter: Deploy Claude Artifacts to GitHub Pages
(github.com)
2 points
by
cpldcpu
2y ago
|
0 comments
95.
▲
by
cpldcpu
2y ago
>That _should_ only make a difference for memory usage if your C compiler isn’t perfect Considering that the PMC150 has an accumulator based 8 bit architecture which is almost hostile to C, it is safe to assume that the compiler is not p
96.
▲
by
cpldcpu
2y ago
Yes, you could implement it in a way where the first layer is streamed and accumulate on output activations in parallel in the memory. This would limit the memory requirements for the input activations, but would increase execution time, as
97.
▲
by
cpldcpu
2y ago
There are no performance profiling mechanisms on these small devices, and the timers are rather coarse. But it is easily possible to estimate the execute time: - mulacc of one weight takes 11 clock cycles. - There are 1696 weights in the mo
98.
▲
by
cpldcpu
2y ago
Yes, as far as i remember the limit was somewhere around 1kbyte total parameters size.
99.
▲
by
cpldcpu
2y ago
Well, you can also build microprocessors out of them: https://hackaday.io/project/182915-555enabled-microprocessor
100.
▲
Implementing neural networks on the "3 cent" 8-bit microcontroller
(cpldcpu.wordpress.com)
159 points
by
cpldcpu
2y ago
|
20 comments
101.
▲
by
cpldcpu
2y ago
The tokens are immediately transformed into embeddings (very large vectors), so the 17 bit values are not used for any computation.
102.
▲
by
cpldcpu
2y ago
>We introduce a new Autoencoder (AE) that aggressively increases the scaling factor to 32. Compared with AE-F8, our AE-F32 outputs 16× fewer latent tokens, Basically they compress/decompress the images more, which means they need le
103.
▲
by
cpldcpu
2y ago
Bill Dally from nvidia introduced a log representation that basically allows to replace a multiplication with an add, without loss of accuracy (in contract to proposal above) https://youtu.be/gofI47kfD28?t=2248
104.
▲
by
cpldcpu
2y ago
I have to disagree. Nvidia spent a lot of effort on researching improved numerical representations. You can see a summary in this talk: https://www.youtube.com/watch?v=gofI47kfD28 A lot of their work was published but went
105.
▲
by
cpldcpu
2y ago
It puzzles me that there does not seem to be a proper derivation and discussion of the error term in the paper. It's all treated indirectly way inference results.
106.
▲
by
cpldcpu
2y ago
Well, it's basically the technical implementation of Moore's law, since Moore's law is just an empirical observation. (And maybe also a self-fulfilling prophecy)
107.
▲
by
cpldcpu
2y ago
what are you referencing?
108.
▲
by
cpldcpu
2y ago
So we will also see an increate in obesity?
109.
▲
by
cpldcpu
2y ago
the performance is still a bit degraded though.
110.
▲
by
cpldcpu
2y ago
Can any of these tools do anything that the Github copilot cannot do? (Apart from using other models?). I tried Continue.dev and cursor.ai, but it was not immediately obvious to me. Maybe I am missing something workflow specific?
111.
▲
by
cpldcpu
2y ago
There probably are not too many at this point. What would be awesome is to have an open and standardized NN accelerator to go with RISC-V, but that is a dream.
112.
▲
by
cpldcpu
2y ago
What does ESP32 have to do with RPI? The equivalent of to the ESP32 would be the Rasperry Pi Pico 1 / 2 / W. They start at $4, which is a fair price.
113.
▲
by
cpldcpu
2y ago
>whoever runs this site has been engaging in You are suggesting that BFL is using "organic marketing" to push their product?? It may be worth mentioning that the BFL team actually consists of the people who invented Latent Diff
114.
▲
by
cpldcpu
2y ago
They mentioned somewhere that they could not designate the actual area increase since it is part of the synthesized digital area. The area increase to include the RISC-V cores was negligible, apparently.
115.
▲
by
cpldcpu
2y ago
It's also an interesting juxtaposition, because it directly allows to benchmark the architectures in the same system environment. In the RP2350 it is possible to either use the RISC-V cores, the CM33 cores or even use one of each.
116.
▲
by
cpldcpu
2y ago
My understanding is that they do not predict the target of the next branch but of the one after the next (2-ahead). This is probably much harder than next-branch prediction but does allows to initiate code fetch much earlier to feed even de
117.
▲
by
cpldcpu
2y ago
At least they put all the relevant information into the title, so that it is not necessary to actually read the article.
118.
▲
by
cpldcpu
2y ago
Claude-3.5-Sonnet is SCHNITZEL mit BRATKARTOFFELN. Waiting for 3.5-Opus or OpenAUs response...
119.
▲
by
cpldcpu
2y ago
Searching for the string on google yields hundreds of hits. Likely that it appears in recent webscrape-data.
120.
▲
by
cpldcpu
2y ago
That is why the ML/AI community is usually shortcutting this process by publishing preprints on Arxiv. A publication based on an LLM that was state-of-the art only until March of 2023 cannot be justified by long review times. Edit: To
More ›