2 ms·
Hi Troy, I skimmed through the outline. Will take a look at the individual videos when I'm on PC. But I have been through this cost-saving phase. I didn't see
by bhagyeshsp 5mo ago
Hi Troy,
I skimmed through the outline. Will take a look at the individual videos when I'm on PC.
But I have been through this cost-saving phase. I didn't see "prompt distillation" as one of the techniques in your outline.
The idea is to reduce your fixed prompt token size such as "system prompt" by removing semantic words completely and other methods. I saw a whopping 60% decrease in my fixed prompt token budget.
Pls note, my scale is small but the technique works nonetheless.
Edit: so, have you tried it? Or if you have tried it, how did it go?
- troymagennis 5mo agoThat's a great one. THANKS. I'll do some more research
- bhagyeshsp 5mo agoYou're welcome. While you research, this is the link to my article I had written a few weeks ago on my practical experience with the whole thing. I'm going to post it on HN today. http://sisyphusconsulting.org/case-studies/2026/04/01/scaling-llms-at-the-edge http://sisyphusconsulting.org/case-studies/2026/04/01/scalin...