4 ms·
Agreed and yes would be awesome... But currently this does not infer large meta 7b model efficiently - like 1 token or lower per seconds. But the small toy stor
by AMICABoard 3y ago
Agreed and yes would be awesome... But currently this does not infer large meta 7b model efficiently - like 1 token or lower per seconds. But the small toy story (not so useful) models are fast.
If the above mentioned API / python binding is ready, I'll make a streamlit interface demo. The streamlit demo should be simple. But I have to figure out python binding.