3 ms·
Wow, those responses are better than I expected. Part of me was expecting terrible responses since Phi-3 was amazing on paper too but terrible in practice.
by syntaxing 2y ago
Wow, those responses are better than I expected. Part of me was expecting terrible responses since Phi-3 was amazing on paper too but terrible in practice.
- refulgentis 2y agoOne of the funniest tech subplots in recent memory. TL;DR it was nigh-impossible to get it to emit the proper "end of message" token. (IMHO the chat training was too rushed). So all the local LLM apps tried silently hacking around it. The funny thing to me was no one would say it out loud. Field isn't very consumer friendly, yet.
- TeMPOraL 2y agoSpeaking of, I wonder if and how many of the existing frontends, interfaces and support packages that generalize over multiple LLMs, and include Anthropic, actually know how to prompt it correctly. Seems like most developers missed the memo on https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/use-xml-tags https://docs.anthropic.com/en/docs/build-with-claude/prompt-..., and I regularly end up in situation in which I wish they gave more minute control on how the request is assembled (proprietary), and/or am considering gutting the app/library myself (OSS; looking at you, Aider), just to have file uploads, or tools, or whatever other smarts the app/library does, encoded in a way that uses Claude to its full potential. I sometimes wonder how many other model or vendor-specific improvements there are, that are missed by third-party tools despite being well-documented by the vendors.
- refulgentis 2y agoHah, good call out: there was such a backlash and quick turnaround on Claude requiring XML tool calls, I think people just sort of forgot about it altogether. You might be interested in Telosnex, been working on it for ~year and it's in good shape and is more or less designed for this sort of flexibility / allowing user input into requests. Pick any* provider, write up your own canned scripts, with incremental complexity: ex. your average user would just perceive it as "that AI app with the little picker for search vs. chat vs. art" * OpenAI, Claude, Mistral, Groq Llama 3.x, and one I'm forgetting....Google! And .gguf
- regularfry 2y agoIn a field like this the self-doubt of "surely it wouldn't be this broken, I must just be holding it wrong" is strong.