3 ms·Cutting LLM inference costs by 36% with prompt caching2 points by lizakatz 1mo agoTokenLat 29d ago[flagged]