3 ms·
benchmarks: we provide plenty in the over 100 page tech report here https://github.com/swiss-ai/apertus-tech-report/blob/main/Apertus_Tech_Report.pdf https://gi
by lllllm 1y ago
benchmarks: we provide plenty in the over 100 page tech report here
https://github.com/swiss-ai/apertus-tech-report/blob/main/Apertus_Tech_Report.pdf https://github.com/swiss-ai/apertus-tech-report/blob/main/Ap...
quantizations: available now in MLX https://github.com/ml-explore/mlx-lm https://github.com/ml-explore/mlx-lm (gguf coming soon, not trivial due to new architecture)
model sizes: still many good dense models today lie in the range between our small and large chosen sizes
- dcreater 1y agoThank you! Why are the comparisons to llama3.1 era models?
- lllllm 1y agowe compared to GPT-OSS-20B, Llama 4, Qwen 3, among many others. Which models do you think are missing, among open weights and fully-open models? Note that we have a specific focus on multilinguality (over 1000 languages supported), not only on english