3 ms·
This interface needs to have a better relationship with streaming, there is always a lag in response and a lot of people are going to want to stream the respons
by kyledrake 2y ago
This interface needs to have a better relationship with streaming, there is always a lag in response and a lot of people are going to want to stream the response in non blocking threads instead of hanging the process waiting for the response. Its possible this is just a documentation issue, but either way streaming is a first class citizen on anything that takes more than a couple seconds to finish and uses IO.
Aside from that the DSL is quite excellent.
- joevandyk 2y agoFrom https://rubyllm.com/#have-great-conversations https://rubyllm.com/#have-great-conversations # Stream responses in real-time chat.ask "Tell me a story about a Ruby programmer" do |chunk| print chunk.content end
- jupp0r 2y agoThis will synchronously block until ‘chat.ask’ returns though. Be prepared to be paying for the memory of your whole app tens/low hundreds of MB of memory being held alive doing nothing (other than handling new chunks) until whatever streaming API this is using under the hood is finished streaming.
- andrewmutz 2y agoThreads?
- jupp0r 2y agoRails is a hot ball of global mutable state. Good luck with threads.
- andrewmutz 2y agoThe default rails application server is puma and it uses threads
- jupp0r 2y agoYes, it does. Ruby has a global interpreter lock (GIL) that prevents multiple threads to be executed by the interpreter at the same time, so Puma does have threads, they just can’t run Ruby code at the same time. They can hide IO though.
- andrewmutz 2y agoThe GIL is released during common IO operations like the HTTP requests that power LLM communication
- jupp0r 2y agoThe Rails documentation has lots of info about this: https://guides.rubyonrails.org/tuning_performance_for_deployment.html https://guides.rubyonrails.org/tuning_performance_for_deploy... Concurrency support is missing from the language syntax and this particular library as a concept. This is by design, to not distract from beautiful code. Your request will make zero progress and take up memory while waiting for the LLM answer. Other threads might make progress on other requests, but in real world deployments this will be a handful (<10). This server will get 10s of requests per second when something written in JS or Go will get many 1000s. It’s amazing how the Ruby community argues against their own docs and doesn’t acknowledge the design choices their language creators have made.
- kyledrake 2y agoThat looks good, I didn't see that earlier.
- bradgessler 2y agoThere’s a whole world of async IO in Ruby that doesn’t get enough attention. Checkout the async gem, including async-http, async-websockets, and the Falcon web server. https://github.com/socketry/falcon https://github.com/socketry/falcon
- earcar 2y agoThank you for your kind words! Valid point. I'm actually already working on testing better streaming using async-http-faraday, which configures the default adapter to use async_http with falcon and async-job instead of thread-based approaches like puma and SolidQueue. This should significantly improve resource efficiency for AI workloads in Ruby - something I'm not aware is implemented by other major Ruby LLM libraries. The current approach with blocks is idiomatic Ruby, but the upcoming async support will make the library even better for production use cases. Stay tuned!