3 ms·
> a lazy sequence that does sensible prefetching and asks the server for more data as it’s iterated through Isn't this an accurate description of cursor-based
by giaour 4y ago
> a lazy sequence that does sensible prefetching and asks the server for more data as it’s iterated through
Isn't this an accurate description of cursor-based pagination? You can have the server hold the cursor and associate a request for the next page based on the client ID (like in the PostgreSQL wire protocol), or you can make the flow stateless (at the network level) by handing the client a ticket for the next page.
- jayd16 4y agoPagination implies many requests where as a streaming API would keep to a single request, no? No worry about multiple servers or juggling state.
- giaour 4y agoIt's still multiple requests on a streaming API if you're asking the client to notify you when it's ready for the next page. Most DB engines use TCP connection affinity to identify which cursor is in play, but the client still needs to send a packet asking for the next page. Both client and server need to keep the TCP connection open, which is state for both parties. I see you mentioned gRPC streams in a sibling comment, which are a great alternative to REST-based pagination, but they are highly stateful! Streams are built into the protocol, so a generic gRPC client will handle most of the state juggling that would have been your responsibility with a REST client.
- jayd16 4y agoThis is a bit inaccurate I think. You can make a RESTful api over gRPC. The protocol is irrelevant. An open connection is not a violation of statelessness. If anything, a streaming API where you can close the cursor upon completion and treat new requests as fully new is a lot less stateful. We're also talking about APIs in general, not just http REST. >It's still multiple requests on a streaming API Perhaps, but the requests can be pipelined. You don't need to wait for the response to complete before asking for more.
- giaour 4y agoThe original article is talking about pagination over stateless (at the protocol level) HTTP requests. (Referring to APIs that are offered over stateless HTTP as "REST APIs" is technically incorrect but reflects common usage and is a convenient shorthand.) gRPC is able to offer powerful streaming abstractions because it utilizes HTTP/2 streams, which are cursor based rather than strictly connection oriented. The state is still there; it's just a protocol-level abstraction rather than an application-level abstraction. > Perhaps, but the requests can be pipelined. You don't need to wait for the response to complete before asking for more. That sort of defeats the purpose of grpc flow control, doesn't it?
- jayd16 4y ago>That sort of defeats the purpose of grpc flow control, doesn't it? Why do you say that? I don't know the full implementation details themselves but generally there's no reason you can't safely ask for even more after having asked and consumed some. If you have a buffer of M bytes and ask for M, then consume N, you could immediately ask for N more without waiting to receive all of the original M. Although, I wasn't speaking about gRPC in that case though. I'm not sure how exactly gRPC achieves back pressure. I was only speaking abstractly about a pipelined call vs many pages. You seemed to claim that multiple requests was a requirement. But fine, ignoring gRPC, you could possibly tune your stack such that normal http calls achieve back pressure from network stacks and packet loss. Http is built on top of TCP streams, after all. That doesn't make it inherently stateful does it? Going all the way back, I still think its fair to say that a streaming response with backpressure can be stateless and without multiple requests. If you want to argue that multiple packets or TCP signals are needed then perhaps so, but I think that's a far cry from the many separate requests a paginated call requires and I dont think its accurate enough to conflate them.
- giaour 4y ago> I don't know the full implementation details themselves but generally there's no reason you can't safely ask for even more after having asked and consumed some I think we're saying the same thing but using different formulations. If you send an HTTP request for a list with 20 items, then get back a response with 10 items and a link for the next page, that is essentially the same as cosuming 20 items over a stream with flow control. The point of returning a cursor in the response and having the client send a new request for the next page is to support a stateful stream over a stateless protocol. In neither case are you waiting for the response to be complete before processing items, since your message indicates that "complete" here means the full result set has been sent to the client. > But fine, ignoring gRPC, you could possibly tune your stack such that normal http calls achieve back pressure from network stacks and packet loss. Http is built on top of TCP streams, after all. That doesn't make it inherently stateful does it? That's pretty much how streaming responses are implemented in TCP-based protocols (like the SQL query APIs exposed by Postgres or MySQL). TCP connections can be terminated for unrelated reasons, which is why you don't see this pattern very often in recent protocols. When a TCP connection is dropped and restarted due to, say, network congestion, you have to restart the paginated operation from the head of the list. H2 (and, vicariously, gRPC) streams are built to be more resilient to network noise. But to answer your question, yes, that pattern is inherently stateful. Pagination has to be, since the server has to know what the client has seen in order to prepare the next batch of results. You can manage this state at the protocol level (with streams), or you can push it to the application level (with cursor tokens embedded in the messages exchanged). The streaming approach requires a stateful server and client, whereas the application-level approach only requires a stateful client.