8 ms·
Local PDF Tools – Powered by WebAssembly
- djrogers 6y agoThere was an error merging PDFs. Not very helpful, can you tell me what the error was or how to avoid it?
- georgeutsin 6y agoLooking forward to using this tool! Are there plans to make this open source?
- kc0bfv 6y agoThere are a few versions of tools like this, or similar, available. Here's mine: https://kc0bfv.github.io/WASM-PDF-Combiner/ https://kc0bfv.github.io/WASM-PDF-Combiner/ I used existing wasm compiles of PDF tools. This use of wasm is pretty awesome to me - I often end up working on very restricted desktop clients with little customization possible, but they always let me run a browser.
- kickbeak 6y agoLooks cool
- redman25 6y agoSeems to be stuck at “Loading” for me on iOS safari.
- lights0123 6y agoThat website loads WASM by embedding base64 in the HTML, which is good for saving it as a single file but horrible for WebKit support ("The operation is insecure", it complains), transfer size, and speed.
- kc0bfv 6y agoThanks for that - I haven't tried it there... Yup - single file operation was a design goal, but I bet there's a better way to get it working than the hack I used.
- danvk 6y agoI’d love a tool (that’s not Acrobat) to manage comments on PDFs.
- franga2000 6y agoOkular has support for comments, although I don't know if they're compatible with Acrobat's.
- yomansat 6y agoCan one easily install such apps as a Chrome app/PWA, and deactivate access to the internet since it doesn't need it and one can merge personal PDFs?
- simonmales 6y agoI have this same idea on my to-do list. Great that people are experimenting with webapps that don't send any data!
- redman25 6y agoJavaScript apis for browsers can do a lot now. It’s great! I built something similar recently with Mozilla’s PDF library. It’s for diffing PDFs but everything happens locally. https://parepdf.com https://parepdf.com
- simonmales 6y agoSweet, when I discovered the Mozilla PDF.js I thought client side manipulation of PDFs would be a breeze. I built a tool that required to count the number of pages of a PDF (ca. 2014-2015). At the time server side counting was the 'sure' way in my brief research.
- pabs3 6y agoHmm, I think I would just compile the pdfcpu Go source to native code, that might be faster than WebAssembly?
- whoisjohnkid 6y agoDefinitely
- codetrotter 6y agoConvenient if you are on a machine where you can’t install software. (Corporate computer, school computer, library computer etc.) For Linux and macOS computers that you are allowed to install software on I recommend the pdftk command line tool. Ubuntu family: sudo apt install pdftk macOS with Homebrew: brew install pdftk-java
- stephenr 6y ago> For Linux and macOS computers that you are allowed to install software on I recommend the pdftk command line tool. If you're on a Mac, the built in Preview tool has had the ability to merge and manipulate PDF documents for years.
- codetrotter 6y agoYea, true. But pdftk is really handy if you are assembling very many pages. And it can also help you do stuff involving all odd-numbered pages or all even-numbered pages for example. So pdftk goes quite a far beyond what you can do with Preview.
- naedish 6y agoThis is helpful. Generally if I need to do any pdf manipulation when I'm away from my own machine I use an android app - PDF Utils [1]. [1] https://play.google.com/store/apps/details?id=pdf.shash.com.pdfutility https://play.google.com/store/apps/details?id=pdf.shash.com....
- travis729 6y agoFor something like this, how do we know that the files are not sent to a server? Am I just trusting the web app? Is there any way to be sure other than having and reading the source?
- fireattack 6y agoOpen dev tool and monitor network?
- simonmales 6y agoFor someone who knows that these tools exists, yes this is a way. For an ordinary user, the only 'easy' way a user can verify the claimed behaviour is to literally go offline. Browsers do not currently have badge to verify that the app is not sending any data. I'm thinking how we were brought to trust the padlock icon browsers display for TLS supporting sites.
- kickbeak 6y agoSomething like that would be great, seeing that you can do more and more with wasm locally it would be really useful
- boustrophedon 6y agoYou can load the website and then disable the network, either by turning off the connection in your OS or via File -> Work Offline in Firefox. I just did this and it worked.
- travis729 6y agoGood tip. Though it would be nice to be able to disable network activity per tab.
- grok22 6y agoDisconnect from the network and try it? Or disconnect from the network always when using this?
- kickbeak 6y agoHey, Thanks for Posting it here, i built this tool, hope you like it, feel free to look at my source and contribute. https://github.com/jufabeck2202/localpdfmerger https://github.com/jufabeck2202/localpdfmerger
- Abishek_Muthian 6y agoIs there a .pdf tool which allows compression to a defined file size? Tools like ghostscript can compress a .pdf to different levels of quality by using different setting but not a defined file size; I understand that this has to do with the compression algorithm itself and that data could be compressed only to a certain limit, but what if the file size limit is within that limit? I'm asking this because an user of my problem validation platform wanted a solution for this[1], because websites requiring document upload have a file size limit and often the compressed file is either above or below the prescribed file size limit thereby loosing out on quality unnecessarily. [1]'Reduce document file size to specific size' (I have added the link to it on my profile, since it's my own platform).
- hnick 6y agoIt's a bit more complicated than it sounds, text streams are generally just compressed as good as they can be using whatever available scheme, the bulk of the space usage is often fonts and images. A PDF itself is not compressed so much as each part of it is compressed individually. There's not much to do with fonts except don't embed them unless you need to, and don't have duplicate/overlapping subsets, if you do have these it is very tricky to untangle. I'm not aware of any good tool to do it automatically. For images, it depends on the format. If your PDF has JPEG (DCTDecode) images then it will have to resample to the JPEG spec, if it's TIFF likewise, you can change the number of colour bits, or you can downsample the DPI which is a simple gs command line option, or you can change the compression scheme within the JPEG itself then replace it. There are so many avenues to approach this that I'm not sure it's something easily achieved while still obtaining a good result across all possible PDFs. Within a problem domain though, like PDFs that are just pages of scanned images, you could probably iterate and downsample until you hit your target size.
- Abishek_Muthian 6y agoI appreciate your detailed comment. >Within a problem domain though, like PDFs that are just pages of scanned images, you could probably iterate and downsample until you hit your target size. I presumed any solution for this problem would involve multiple passes through the compression routine to hit the target size. Having to deal with text, font, images separately as you said does make that complex. Right now, I'm just waiting to see if there's really a need gap for this. It's usually the Govt. websites which has very low limit for document upload size like 100KB, I personally take the image out of .pdf and compress it to minimum jpeg quality; but documents with multiple pages make it tricky.
- desmap 6y agoSomething I still miss is a free and easy PDF tool which lets you delete, reorder and add pages from multiple PDFs. On Windows there is just Xodo but its UX is unfortunately subpar and on macOS you have Preview where the UI is better but once you have multiple PDFs from where you get the pages it can get confusing.
- c-st 6y agoOn Linux there is https://github.com/pdfarranger/pdfarranger https://github.com/pdfarranger/pdfarranger It has a nice UI and is pretty easy to use.
- karthickgururaj 6y agoPdfTk does this. I use the CLI version, but I think there is one with a GUI as well.
- hnick 6y agopdftk is by far my favourite for this too. It's quite fast but it does run into some file size problems when merging files as it doesn't deduplicate resources, and can crash outright on some bigger files.
- number6 6y agoOn Windows there ist: https://www.pdf24.org/ https://www.pdf24.org/
- svat 6y agoLast year I wrote a couple of similar "local" PDF tools that run in the browser with no network requests. Each is just a single HTML file that will work offline: - https://shreevatsa.net/pdf-pages/ https://shreevatsa.net/pdf-pages/ is for extracting pages, inserting blank pages, duplicating or reversing pages, etc. - https://shreevatsa.net/pdf-unspread/ https://shreevatsa.net/pdf-unspread/ is for splitting a PDF's "wide" pages (consisting of two-page spreads) in the middle. - https://shreevatsa.net/mobius-print/ https://shreevatsa.net/mobius-print/ is the earliest of these, and written for a niche use-case: "Möbius printing" of pages, which is printing out an article/paper two-sided in a really interesting order. (I've tried it and love it.) These don't use WebAssembly, but just use the excellent "pdf-lib" JS library. To keep the file self-contained, I put the whole minified source into a <script> tag at the bottom of the (otherwise hand-written) HTML file.
- skrebbel 6y agoI hope to never again print more than 2 pages in a go in my life, but if I do, I'm definitely going to use your Möbius printing tool, it's genius.
- hackerbrother 6y agoVery cool!! Thanks for sharing :)
- justsayinghi 6y agoVery useful pdf tools
- sigvef 6y agoLooks like this thread is all about sharing our own related local browser-based PDF tools. Here’s mine: https://pdftotext.github.io https://pdftotext.github.io
- mgm__ 6y agoI created a PDF table extractor tool last year with the same idea that it should be local only. Try it here: https://pdftableutil.possiblenull.com/app/ https://pdftableutil.possiblenull.com/app/ Also as a Google Docs addon (still local only) https://workspace.google.com/marketplace/app/pdf_table_importer/646940040599 https://workspace.google.com/marketplace/app/pdf_table_impor... I had a bad case of scope creep, so the tool can also extract tables from scanned/image PDFs using OpenCV.js and tesseract OCR wasm build!
- redman25 6y agoThis is interesting. How accurate would you say it is?
- mgm__ 6y agoI haven't seen anything better. It started as a PoC and I decided not to include table detection on the page and require the user to draw box around the table. I use Tabula under the hood for the cell/row detection and it is really good given the correct mode is selected for the type of table. The modes are stream (find cells by spacing) or lattice (find cells by ruling lines). The OCR/OpenCV seemed to be fine as well as long as the text isn't too blurry. Here is a GIF of the OCR/OpenCV running on an example Image PDF: https://lh3.googleusercontent.com/-OobUBBtnydg/X6Vn_Ls3juI/AAAAAAAAEAg/q__NHhsbhb0WLl1KYuSPBt16b0OyY3stwCLcBGAsYHQ/s640-w640-h400/Kvgkz4NcqO.gif https://lh3.googleusercontent.com/-OobUBBtnydg/X6Vn_Ls3juI/A...
- kickbeak 6y agoWow That looks awesome, what did you use to display the PDF in the Browser? feels all really responsive!
- mgm__ 6y agoI used Mozilla's PDF.js https://mozilla.github.io/pdf.js/ https://mozilla.github.io/pdf.js/ It is what firefox uses on desktop to show PDFs!
- not_knuth 6y agoDoes anyone know some good tutorials/explanations for understanding the PDF format at the byte level?
- maest 6y agoThis: https://www.adobe.com/content/dam/acom/en/devnet/pdf/pdfs/pdf_reference_archives/PDFReference.pdf https://www.adobe.com/content/dam/acom/en/devnet/pdf/pdfs/pd... or a newer version.
- hnick 6y agoThe spec is really well written, and will take you far if you just want the basics (i.e. not the scripting stuff they added after 1.4 which is around where you should usually stop if you just care about printing). My only issue with it is when it comes to fonts or images, you'll have to break out to additional specs to understand those formats since PDF is more of a container. I do like this minimal example as a way to get started and see how a very basic PDF is built. https://brendanzagaeski.appspot.com/0004.html https://brendanzagaeski.appspot.com/0004.html
- gpvos 6y agoThe official PDF Reference is very readable, see maest's link. Just start at the bit about the file structure and data types, then the basics about the commands in the Contents stream, the graphics and text states, then whatever takes your fancy. Later versions added some extra complications such as compression of the xref table and object streams, you don't need them unless you encounter them. Don't delve too deep into the bit about fonts unless you have to, it might bend your brain. A tool like qpdf to deconstruct a PDF to an uncompressed form is very handy.
- hnick 6y ago> A tool like qpdf to deconstruct a PDF to an uncompressed form is very handy. Forgot to say this in my reply, likewise pdftk's uncompress option is my #1 stop for learning about how a PDF is built when we have issues. You can poke around and hex edit for learning too. It is easy in an uncompressed PDF to comment out a line with % to remove objects that I think are causing issues without messing up the xref offsets. Using a fully fledged tool to edit may have side effects in how things are rewritten which ruins the point of the investigation. Acrobat Pro has some very useful features too, but not everyone has access to that.
- Ciantic 6y agoIt's great to see more PDF tools. Many times I just want to clip white margins from PDFs so that it is easier to view on tablets or phones. Most viewers don't have a way to force the clipping of pages, so when you change page the zoom is lost and suddenly all the content is squished to center. Last time I found cli programs to do it aprox. five years ago it was really difficult to find good tools to edit PDFs like that. It's actually not trivial task, as sometimes pages have different margins, e.g. odd and even pages has different margins on folding side of page.
- ithkuil 6y agoAnybody knows a simple tool I can use to turn an academic two-column paper into a single column pdf (so I can read it easily on e-paper like a remarkable)? (Ideally I'd like to be able to run such a tool from browser/phone)
- jerry_tk 6y agoMaybe try k2pdfopt (https://www.willus.com/k2pdfopt/ https://www.willus.com/k2pdfopt/) It is a Windows/Linux application though.
- brailsafe 6y agoI'd certainly be curious why wasm ends up being 15x slower than native binary in this case, but it's not insurmountable. All of the major commercial PDF editing suites use wasm + their own C++ based pdf engine to great effect. The article that this is based on is here, and a good read. It seems like it's at least non-trivial to get it working, and I'd wonder how the process looks for other compiled binaries, having not tried to do that implementation from scratch. https://dev.to/wcchoi/browser-side-pdf-processing-with-go-and-webassembly-13hn https://dev.to/wcchoi/browser-side-pdf-processing-with-go-an...
- horst_vie 6y agoThis is a nice usecase for pdfcpu. If you are pdftk user give the pdfcpu CLI a spin. It is multi platform and has some nice features baked in. https://pdfcpu.io/ https://pdfcpu.io/