3 ms·
The connection between type systems and neural net structure is underexplored in practice. One thing I'd add: when you're dealing with multi-modal inputs in pro
by Daffrin 5mo ago
The connection between type systems and neural net structure is underexplored in practice. One thing I'd add: when you're dealing with multi-modal inputs in production — say, mixed structured and unstructured content — the type-safety problem compounds. You end up with implicit contracts at inference boundaries that are very hard to enforce.
Has the author written anything on how this applies to transformer architectures specifically? The attention mechanism seems like a place where a richer type theory would be genuinely useful.
- bgavran 5mo agoThere's been some exciting work generalising transformers to data structures that aren't just pure arrays: https://glaive-research.org/2025/02/11/Generalized-Transformers-from-Applicative-Functors.html https://glaive-research.org/2025/02/11/Generalized-Transform... I've implemented these in Idris 2: https://github.com/bgavran/TensorType/blob/main/src/NN/Architectures/Transformer/Attention.idr https://github.com/bgavran/TensorType/blob/main/src/NN/Archi...