23 ms·
Pagination with rel=“next” and rel=“prev” (2011)
- kevindeasis 10y agoHow do you guys implement / What's the best practice for paginating your db using sql and nosql? I know there are a bunch of gotcha's to look out for
- ldjb 10y agoI feel there are essentially two approaches. You could store the entire article text as a single chunk of data. The text would be marked-up to indicate where the page breaks are. Each time a page of the article is requested, the entire article is fetched from the database, but the website software extracts only the necessary extract. This is rather inefficient in principle, although it is possible to optimise, and it does keep things relatively simple. On the other hand, you could have an "articles" table and a "pages" table with the relationship: articles 1 <--> 1..* pages Each page's text is stored as a separate row in the "pages" table, along with the article ID and a page number. When a page is requested, only the page that matches the requested article's ID and the appropriate page number. The database schema here is more complex, but is likely to be more efficient.
- err4nt 10y agoI'm no DB programmer, but despite the inefficiencies of the first approach (fetch entire article each time and display a different chunk) it would allow for users (or admins) to change page length after the article is written a LOT more easily than if each page was stored separately. Is this a potential concern of using the second method where each chunk is in its own field?
- ldjb 10y agoAssuming there is an actual user interface for composing and publishing articles, that interface doesn't have to be dictated by the way the data is stored. Whichever method you choose, the software will need to be able to transform the data between two formats: one suitable for the user interface, and one suitable for the database. However you designed the UI, it would be possible to use either approach for the database. For example, suppose the UI used a single textarea and had a special notation for delimiting pages. For the first approach, the contents of the textarea can simply be stored in its own field. For the second approach, the software would take the contents of the textarea, split it into its individual pages, then store each page separately. If the UI had separate textareas for each page, then for the first approach, the software combines the contents of the textareas and puts in delimiters. For the second approach, each textarea is mapped to their own database field. I suppose some of combinations are easier for the programmer than others, but so long as the UI is designed to be easy to use, the user need never know how the article is actually represented in the database.
- fbonetti 10y agoIt's really easy to paginate data with SQL - just use LIMIT and OFFSET. limit = page_size offset = (current_page - 1) * limit For example, let's say you have a table with 1000 products in it, and you want to display page 3 with a page size of 50. Here's how you would write it: SELECT * FROM products LIMIT 50 OFFSET 100
- pmarreck 10y agoAnd if the data frequently changes during pagination?
- objclxt 10y agoThat's why it's often better to use a cursor in your client-facing API (or a pseudo-cursor that maps to something you have in your database), rather than passing the limit/offset through.
- danneu 10y agoThe downside is that OFFSET is slow if you have a lot of items, like on a forum. At which point you could assign each post in a topic an indexed `position` column that always increases from 0 within that topic, which lets you jump to arbitrary pages and parameterize perPage. fromPos = (page - 1) * perPage SELECT * FROM posts WHERE topic_id = $1 AND position >= fromPos AND position < fromPos + perPage ORDER BY id For simpler needs, your "Next" button could load items that come after the last ID on the current page but with a LIMIT. <ul> {% for item in items %} <li>{{ item }}</li> {% endfor %} </ul> <a href="/items?afterId={{ items[items.length - 1].id }}">Next</a> SELECT * FROM items WHERE id > $afterId LIMIT $perPage That's fast and really easy.
- lox 10y agoThis sounds like a really premature optimization to me. Let the DB do it's job unless you have an explain trace showing it's slow for your workload. Trying to do things like sequence columns adds complexity, when this is often handled by the engine.
- tantalor 10y agoOne great approach is cursors. Basically, you keep track of the value of the sort function (e.g., relevance) of the last record on the page, and add that to your filter (e.g., WHERE relevance > $LAST_RELEVANCE) to get the next page. This is very efficient when your table is indexed by that field; you can seek directly to the next record. Some database APIs do this for you! I wrote up my experience implementing this on App Engine here, http://johntantalo.com/blog/paginating-with-bookmarks-in-app-engine/ http://johntantalo.com/blog/paginating-with-bookmarks-in-app...
- xrstf 10y agoAs fbonetti already said, LIMIT is the canonical way to do it with SQL databases (I would just add to their comment that you should ensure a positive page_size and offset, because "LIMIT -5" is invalid in at least MySQL, and to limit the page size to some sensible values). For NoSQL, specifically (only?) CouchDB, you start iterating from a given key (in SQL, this would effectively be "WHERE id >= $start ORDER BY id ASC LIMIT n"). This difference is why I would recommend, especially for paginated REST APIs, to provide links to the next/prev pages, so API clients can follow these links without thinking about how pagination is accomplished exactly. To prevent people from guessing the inner workings and working around your links, you can encode the parameters in a "cursor" (so instead of ?pagesize=20&page=10 you could have ?cursor=base64encode("20,10") (SQL) or ?cursor=base64encode("$startKey") (NoSQL)). IIRC Facebook does it this way. Encrypting or authenticating the cursor is IMHO overblown here, given cursor values can be easily validated and rejected if fiddled with. Using an opaque cursor also gives you the ability to change how pagination works without breaking existing API clients. If you plan for lots of data and many pages, think twice before outputting links to all pages (on a website) or outputting the total number of elements (in an API). COUNT(*) on InnoDB is slow and I heard it's not the quickest thing in NoSQL databases either.
- yeukhon 10y agoHow do you keep the cursor / know where you left off with the database? For example, I don't want to repeat what I saw in page 1 in page 2 because someone else just added more articles.
- HappyTypist 10y agoUse continuations.
- spottedquoll 10y agoWhy would you break content into "pages"? If the requestor wants the data, give it to them Forcing artificial page-breaks is abusive.
- gnaritas 10y agoSo if someone requires the entire table you think a single API call should just send it? Does it not occur to you that people used page data to limit response sizes and manage the performance of their API's?
- detaro 10y agoSo HN should be a single page with all submissions ever? ;) Yes, don't unnecessarily break up content. But most pages are going to include lists of some sort that you want to be paginated in some way. If you look at HNs /new queue, there are submission IDs in the next links, so you can click "more" multiple times without seeing content multiple times, even if new submissions have pushed them down in the global list by now. Useful if there are many updates and new entries come in at the "top" of the list.
- deathanatos 10y agohttp://use-the-index-luke.com/blog/2013-07/pagination-done-the-postgresql-way http://use-the-index-luke.com/blog/2013-07/pagination-done-t... Essentially, once your rows are ordered, your next page is "rows > the last row on the previous page, limit <number of items to return>" This has the fantastic benefit of being index-friendly, which OFFSET is not, and if you structure the URLs for the pages well, obvious what you're getting back. (e.g., if your records are by date, you might get /daily-reports?since=2015-01-04&limit=50 for the page that starts at 2015-01-04) It's from an SQL site, but the concept is not specific to SQL. I'm working on an implementation of it for data stored in Mongo, at the moment.
- emmelaich 10y agoUse the seek method: http://use-the-index-luke.com/sql/partial-results/fetch-next-page http://use-the-index-luke.com/sql/partial-results/fetch-next... edit: more on this via a slideshare presentation or pdf: http://use-the-index-luke.com/blog/2013-07/pagination-done-the-postgresql-way http://use-the-index-luke.com/blog/2013-07/pagination-done-t...
- deleted 10y ago[deleted]
- CiPHPerCoder 10y agoThis is a blog post from 2011. The current HN submission title does not indicate this article's age.
- DiabloD3 10y agoI agree, but Google still seems to recommend this usage.
- kmfrk 10y agoLiterally one of the greatest things about the Opera browser was that you could browse an entire forum or whatever (longform article etc) with the Space key, because the browser would automatically go to the `next` page on hitting the bottom of the page. They also did some cool stuff with swiping to the next page on mobile. Opera seemed like the only browser vendor that truly championed next/prev. I only do it now as a habit because of Opera - and, since switching to Chrome after Opera's nadir, for accessibility, regardless of whether screenreaders respect it. But I probably wouldn't know about it if not for Opera. It's both wistful and aspirational to use them, because this is the way browsers were supposed to work - as was generally the case with most things Opera, of course.
- landr0id 10y agoI saw recently that Safari's reader view actually pieces together paginated articles into one long view. I haven't had a chance to try it myself though.
- Jerry2 10y agoSafari's Reader view is one of the primary reasons why Safari's my default browser. I still use Chrome for development but 99% of browsing & reading, I do in Safari. Low power usage and the speed of Reader view is just unbeatable. It works so well and concatenates most of the paginated articles as well. If you couple it with Page One extension (it has a bunch of rules that automatically re-directs you to a single-page view of articles), 99% of the sites will be a single page in Reader view.
- knight17 10y agoClearly extension on Firefox also grabs the text from 'next' page(s) content into the readable view. The built-in reader in Firefox can not do that. It is a very useful feature to have. edit: While looking for the addon at addons.mozilla, I find that it has been discontinued by Evernote: 'This add-on has been removed by its author.' https://addons.mozilla.org/en-GB/firefox/addon/clearly/ https://addons.mozilla.org/en-GB/firefox/addon/clearly/ Anyone know of another addon that works similarly?
- deleted 10y ago[deleted]
- Roberto_ua 10y agoWhat are the recommendations for infinite scroll?
- kachnuv_ocasek 10y agoDon't use it. It's an antipattern and an awful meme.
- ajkjk 10y agoWhy's that?
- D-Coder 10y agoInfinite scroll breaks text search. :-(
- currysausage 10y agoPages become so complex that browsers start lagging. If you click a link on the infinitely scrolling page and then come back, your position is lost and you have to scroll down the whole way again (see Facebook, Twitter).
- kmfrk 10y agoIt is extremely rare that it doesn't bring my browser to its knees after a certain amount of scrolling. (This is with 12 GB on Chrome.) You will also end up losing your reading position after a browser restart or on iOS where Mobile Safari reloads a page to save memory. It's a dumb, pointless feature that at the very least shouldn't be the default option. I'm sure there are hypothetical scenarios that may or may not warrant it, but it is a pain and a half for the user. Tumblr blogs have it in spades, and it's a miserable experience to browse.
- Guest91283 10y agoI don't understand why more sites don't just use pagination with larger result sets. Return pages with 100 or 200 results each. It takes a little more bandwidth, but it's a good middle ground between pagination and infinite scroll.
- draw_down 10y agoOne of the nice things this does is give the browser's "reader mode" a hint where the next page is. And it's great. Only problem is, publishers seem to have no interest in helping the reader mode. On the contrary, many seem hell-bent on defeating its content detection algorithm, etc, to make it useless so we all look at their stupid ads.
- dumbguy 10y agoAlso known as a good way to break many a SPA.
- ycombobreaker 10y agoSingle page application? Can you explain why the existence of next/prev relationships would break these? I would expect such a site to simply not use this. It's definitely nice for linearly-linked content on supporting browsers (e.g. old opera, as mentioned elwewhere in this thread).
- dumbguy 10y agoWell, look at it this way, when creating a SPA, you're running everything in a single place. A single HTML page. In some instances, it makes sense to push the state to the query string, but often it doesn't, and more often still are the occasions developers haven't thought about it. So every time you hit the back button, and weren't where you thought you would be, that's the same issue. SPA frameworks's are getting better at that, but this still happens to me on the daily.
- detaro 10y agoI wonder why they recommend putting <link rel="..."> tags into the head, instead of marking up existing <a href> tags (which you'll have in most cases for prev/next) with rel-attributes. Actual difference in parsing? Just because <a> tags might not be on every site?
- ldjb 10y agoWell, for one thing, it's in the HTML standard [1]. But also consider websites such as forums that allow users to post custom HTML to a page. Although there is often some sanitisation, that often doesn't stretch to removing attributes from <a> tags. In those cases, it might be possible for a user to hijack the sequential links and make them point to some other website. [1] https://www.w3.org/TR/html/links.html#sequential-link-types https://www.w3.org/TR/html/links.html#sequential-link-types
- detaro 10y agoyour link says: The next keyword may be used with link, a, and area elements of course it is allowed to put it in <link>. And if your badly sanitized HTML (which hopefully did strip the malicious onClick handler ;)) includes an <a rel="" href="">, are we sure that parsers will prefer the <link> in the header? Or will Google and others actively ignore the <a>, despite the standard allowing it on both? That would be useful to know. For simple bots putting it in an easy-to-discover link in <head> might be good, but a search engine and a browser both have to parse the entire page anyways...
- tyingq 10y agoI believe it's to unambiguously tie the "next" to the page. Where the page content might have, for example, an <a> that moves images around in a carousel.
- ibejoeb 10y agoThere is only one next in the sequence, but there may be multiple navigational elements that bring the user there. For example, links in both a header and a footer.
- 10y ago
- SilasX 10y agoOne of their examples is a single sentence split over three pages of an article. I'm not sure if that's just a toy example, or also a jab at the way clickbaiters game the view counts.
- mitchtbaum 10y agoI would guess that whoever wrote it subtly wants to say, if web servers\apps send partial responses to requests without a Range header, then something has broken - probably somewhere between our keyboards and chairs.
- deleted 10y ago[deleted]
- peterburkimsher 10y agoThe iPod Notes format used this type of pagination! It was essential back then because each note could only have a maximum of 1000 characters.
- wyuenho 10y agoBesides making the big G happy, these semantic relationships also helps libraries authors to easily traverse a series purely on the front end. Years ago, I looked at Github's pagination API [1] and discovered this approach also makes infinite paging much easier, just send me a bunch of next pointers until there aren't any more. So I decided to support this format by default. [1]: https://developer.github.com/guides/traversing-with-pagination/ https://developer.github.com/guides/traversing-with-paginati... [2]: https://github.com/backbone-paginator/backbone.paginator https://github.com/backbone-paginator/backbone.paginator
- captainmuon 10y agoAh, this might be the cause of the infuriating behavior where you search for a term, and it appears in one page in a huge forum thread, but after clicking on a search result you end up somewhere completely else in the thread. I don't know whether Google sends you there, or whether the site redirects you from the "view all" page to the first page, but the result is pretty annoying.
- deleted 10y ago[deleted]