3 ms·
Yeah, we were playing around with doing some semantic chunking. Works okay for some use cases. We have some ideas to go further on that. Generally we have foun
by ddematheu 3y ago
Yeah, we were playing around with doing some semantic chunking. Works okay for some use cases. We have some ideas to go further on that.
Generally we have found that recursive chunking and character chunking tend to be short sighted.
- hrpnk 3y agoDon't you find it dangerous to just run the code w/o any sanitizing? Why not capture a few strategies that the LLM returns as code that can be properly audited (and ran locally improving the overall performance)?
- ddematheu 3y agoIt is dangerous, part of the reason that we haven't productized that further. One of the ideas we had to productize the capabilities further was to leverage edge / lambda functions to compartmentalize the code generated. (Plus it becomes a general extensibility for folks that are not using semantic code generation and simply want to write their own code.) The idea of auditing the strategy is interesting. The flow that we have used for the semantic chunkers up to date has been along these lines where we : 1) Use the utility to generate the code snippets (and do some manual inspection) 2) Test the code snippets against some sample text 3) Validate the results
- treprinum 3y agoWhy not use Stanford Stanza?