Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dkhudia
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
dkhudia
2y ago
This is true but the instruction already existed and it doesn't support uint16_t accumulation. For the reason you mention, activations are uint8_t and weights are int8_t so it worked out well for neural networks.
2.
▲
by
dkhudia
2y ago
> It's quite common in machine learning operations to multiply a matrix of unsigned byte by a matrix of signed byte. Don't ask me why, but that's the case. Overflow is the reason. Intel's vpmaddubsw takes int8_t and u
3.
▲
by
dkhudia
3y ago
@tome for the deterministic system, what if the timing for one chip/part is off due to manufacturing/environmental factors (e.g., temperature) ? How does the system handle this?
4.
▲
LLM Inference Performance Engineering: Best Practices
(databricks.com)
2 points
by
dkhudia
3y ago
|
0 comments
5.
▲
LLM Inference Performance Engineering: Best Practices
(databricks.com)
3 points
by
dkhudia
3y ago
|
0 comments
6.
▲
by
dkhudia
3y ago
MosaicML/Databricks | San Francisco Bay Area or New York | Machine Learning Engineer - Performance Optimization | Full-time Founded in late 2020 by a small group of machine learning engineers and researchers, MosaicML enables companies
7.
▲
by
dkhudia
3y ago
Disclaimer: I work for MosaicML (MosaicML is the creator of the training platform used by Replit). Training these models from scratch on your domain specific data is not as expensive as one might think. We have provided some cost estimates