5 ms·
The import/export is built-in to WordPress. It's under Tools in the menu and you can read more about it here: https://learn.wordpress.org/tutorial/tools-import-
by desas 3y ago
The import/export is built-in to WordPress. It's under Tools in the menu and you can read more about it here: https://learn.wordpress.org/tutorial/tools-import-and-export/ https://learn.wordpress.org/tutorial/tools-import-and-export...
The file format is called wxr, if you search for [platform name] wxr import/export or [platform name] wordpress import/export you'll find a write-up.
I went from wordpress to Jekyll to Hugo and back to WordPress, importing and exporting my data along the way.
- thunderbong 3y agoThanks! Didn't know about this.
- chrismorgan 3y agoI handled a migration from WordPress last year, and was unimpressed with its export format. Well, with WordPress as a whole, frankly. The export format was very clearly a half-baked implementation for migrating from WordPress to WordPress. It didn’t include all content (and as far as I could easily tell, the missing content even seemed to be WordPress-shaped rather than from a plugin that stored things in a separate database table, which I could more understand missing though it should still have hooks plugins can implement to avoid that happening, no idea if it does or not). You had to fetch all the media files completely separately. And as far as non-WordPress interactions are concerned, the actual content markup format was something abominable that mixed old-style almost-HTML-but-line-breaks-are-awful-magic and new-style Gutenberg blocks, including within individual pages, and I believe the site wasn’t even three years old. Making the content suitable for importing to something else required, to begin with, copying/porting/applying WordPress’s wpautop function (which has the innocuous description “Replaces double line breaks with paragraph elements.”, but it’s way worse than that because it’s trying to avoid damaging HTML and you end up with a monstrosity that might almost be worse than Markdown in its fairly arbitrary interactions with HTML, and it’s doing it all with regular expressions, like even the new-style blocks, and it does mangle and destroy content, and not doing this stuff will mangle and destroy content, beyond just missing paragraph breaks). And then there are the URLs… what should be the URL, since it’s extending RSS, actually isn’t the post’s URL. Sometimes it is, by coincidence. Other times, it isn’t, but that URL still works under WordPress, by coincidence, and maybe is in use. (For that matter, other URLs may be in use for linking to the content—so unless you duplicate a lot more of WordPress, you will hazard breaking links.) Other times, it isn’t and doesn’t work at all. Still other times, it should be, but WordPress matches some other piece of content at that URL first, and so the page is completely inaccessible (though if you know the post ID, which is normally not present when served, /?p=… may work), and all you can learn about it is the tantalising excerpt on a blog post list (because when you follow the link, it serves you a different page—and I’ve seen this on at least two other WordPress installations). The actual answer in the export is, as far as I can tell, completely undocumented and difficult to get at, requiring mixing a few unrelated things. But half of this stuff is really more flaws in WordPress itself than its export model. Because its export format exposes a lot of the awfulness of WordPress, which is apparently a mountain of poor design decisions and technical debt. If I didn’t already detest it for its atrocious security model (I’ve had to help fix more than a few hacked WordPress sites over the years), this would have made me scorn WordPress utterly.