3 ms·Actually NVIDIA made one earlier this year, check out their Fast-dLLM paperby nathan-barry 11mo agoActually NVIDIA made one earlier this year, check out their Fast-dLLM papergdiamos 11mo agoThanks I’ll check it out!gdiamos 11mo agoDid I miss something? https://github.com/NVlabs/Fast-dLLM/blob/main/llada/chat.py https://github.com/NVlabs/Fast-dLLM/blob/main/llada/chat.py That’s inference code, but where is the high perf web server?