4 ms·
Check the article. They have a learned preprocessing step that translates time slices containing multiple data points into tokens, so the transformer is actuall
by jsenn 3y ago
Check the article. They have a learned preprocessing step that translates time slices containing multiple data points into tokens, so the transformer is actually predicting larger chunks of time rather than individual time points.
- daxfohl 3y agoOh, right. I missed the point of that step on first read.