5 ms·
llama.cpp uses a hard cutoff. The agent then does "something" that is specific to the agent's implementation and configuration. It might summarize and then "f
by ttkciar 2mo ago
llama.cpp uses a hard cutoff. The agent then does "something" that is specific to the agent's implementation and configuration. It might summarize and then "finish the thought" with a different model, and then resubmit the prompt to the llama.cpp API endpoint with <think>..</think> prefilled. The primary model then infers the remainder of the reply.
- cyanydeez 2mo agollama.cpp does a hard cut off on budget; it can set a reasoning-message as default but the client _can_ set a per message reasoning-message, so it's possible a smart harness could inspect the cut of thoughts and trim and do whatever.