4 ms·
I've come to the conclusion that anything that "abstracts" the openai complete/chat complete API call is just bad practice and to stay away from the entire fram
by fswd 3y ago
I've come to the conclusion that anything that "abstracts" the openai complete/chat complete API call is just bad practice and to stay away from the entire framework, with the exception of microsoft guidance. Just because you can, doesn't mean you should. And if you do abstract the completion API, then it must either reduce friction or increase capabilities over just calling openai with http fetch/axios. Which microsoft guidance does this.
- jawerty 3y agoYea learning how to use the core API directly should be the focus for any engineer. Lots of frameworks built on top of LLMs are being made very quickly each with their own philosophy. It's a good time to stay with the fundamentals as much as possible until the dust settles. Learning langchain will take you less than a day if you have fundamentals don't worry about not staying up to date. Now's the time to learn how LLMs work from the ground up not being a framework chaser. (watch Karpathy's GPT from scratch video and read through huggingface's LLM documentation from RLHF to PEFT fine-tuning)
- victor106 3y ago> with the exception of microsoft guidance. why?
- benjreinhart 3y agoYes, we largely agree with you on that. The APIs are high-level enough that wrapping really doesn't add much value in many circumstances. We chose to do this for our first module to take a stab at integrating RAG pipelines in a coherent manner, but we don't plan on following this pattern in all modules within our framework. There is possibly one exception here, which is that an interface that allows composable middleware for things like logging, error handling, or redirecting of requests may justify wrapping in some places. The next steps for us involve lower-level functionality. One need we see again and again is more robust data extraction and processing. Most people we talk to who use other community projects (e.g., langchain or llama) find that data loading and chunking are among the most valuable parts of those libraries. We agree, but would like more robust functionality for these tasks, so this is one thing we're working towards next. Beyond that, we're working on infrastructure. Easy model serving from Node (for OSS or proprietary models), monitoring, and pipelines for fine-tuning based on production inference results.